← All tags

#benchmark

39 posts

Claude Skill 83

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

Python MIT Updated 1mo ago
MCP Server 35

MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols

Python MIT Updated 6mos ago
MCP Server 8

Private, local-first AI assistant for Windows - use your own model, keep durable memory, and approve every sensitive tool.

C# Apache-2.0 Updated 1mo ago
MCP Server 6

Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, lay...

Python MIT Updated 4mos ago
MCP Server 11

Open benchmark for AI coding agents on SWE-bench Verified. Compare resolution rates, cost, and unique wins.

Shell MIT Updated 4mos ago
Claude Skill 9

Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.

Elixir NOASSERTION Updated 1mo ago
MCP Server 4

Portable memory layer for AI agents -*pre-release

Rust Apache-2.0 Updated 1mo ago
MCP Server 24

Real-time trustworthiness evaluation and safety interception for AI agents. Semantic analysis, safe alternative suggestions, multi-step atta...

Python NOASSERTION Updated 2mos ago
MCP Server 14

A comprehensive security benchmark for evaluating infrastructure-layer defenses in MCP-based AI agent systems

Python Apache-2.0 Updated 4mos ago
MCP Server 2

Multi-tool LLM agent benchmark: 496 end-to-end MCP tasks, deterministic evaluation, Russian localization

Python Apache-2.0 Updated 1mo ago
MCP Server 13

DevOps MCP server, Flight recorder for AI infrastructure agents. The prescribe/report protocol captures intent before execution and outcome...

Go Apache-2.0 Updated 2mos ago
MCP Server 6

Assay — a canary-oracle benchmark for MCP security: a frozen task set scored by a recomputable HMAC-canary oracle (structural-zero false pos...

Python MIT Updated 2mos ago