Task Compass Skill: deterministic, auditable task routing for AI agents and OpenClaw
#benchmark
30 posts
Timestamps: 00:00 - Intro 02:05 - 3D Model Test Overview 04:02 - 3D Model Test 10:26 - 3D Model Results 13:10 - AI Magazine Test Overview 1...
Workbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more
[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
Private, local-first AI assistant for Windows - use your own model, keep durable memory, and approve every sensitive tool.
Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, lay...
Open benchmark for AI coding agents on SWE-bench Verified. Compare resolution rates, cost, and unique wins.
Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Portable memory layer for AI agents -*pre-release
Real-time trustworthiness evaluation and safety interception for AI agents. Semantic analysis, safe alternative suggestions, multi-step atta...
A comprehensive security benchmark for evaluating infrastructure-layer defenses in MCP-based AI agent systems