[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
#benchmark
39 posts
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
Private, local-first AI assistant for Windows - use your own model, keep durable memory, and approve every sensitive tool.
Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, lay...
Open benchmark for AI coding agents on SWE-bench Verified. Compare resolution rates, cost, and unique wins.
Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.
Portable memory layer for AI agents -*pre-release
Real-time trustworthiness evaluation and safety interception for AI agents. Semantic analysis, safe alternative suggestions, multi-step atta...
A comprehensive security benchmark for evaluating infrastructure-layer defenses in MCP-based AI agent systems
Multi-tool LLM agent benchmark: 496 end-to-end MCP tasks, deterministic evaluation, Russian localization
DevOps MCP server, Flight recorder for AI infrastructure agents. The prescribe/report protocol captures intent before execution and outcome...
Assay — a canary-oracle benchmark for MCP security: a frozen task set scored by a recomputable HMAC-canary oracle (structural-zero false pos...