← All tags

#benchmark

30 posts

Claude Skill 1

Task Compass Skill: deterministic, auditable task routing for AI agents and OpenClaw

Python MIT Updated 1w ago
Claude Skill 2

Workbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more

TypeScript MIT Updated 3d ago
Claude Skill 71

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

Python MIT Updated 1w ago
MCP Server 35

MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols

Python MIT Updated 4mos ago
MCP Server 8

Private, local-first AI assistant for Windows - use your own model, keep durable memory, and approve every sensitive tool.

C# Apache-2.0 Updated 1w ago
MCP Server 6

Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, lay...

Python MIT Updated 2mos ago
MCP Server 11

Open benchmark for AI coding agents on SWE-bench Verified. Compare resolution rates, cost, and unique wins.

Shell MIT Updated 2mos ago
Claude Skill 9

Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.

Elixir NOASSERTION Updated 1w ago
MCP Server 4

Portable memory layer for AI agents -*pre-release

Rust Apache-2.0 Updated 1w ago
MCP Server 24

Real-time trustworthiness evaluation and safety interception for AI agents. Semantic analysis, safe alternative suggestions, multi-step atta...

Python NOASSERTION Updated 1mo ago
MCP Server 14

A comprehensive security benchmark for evaluating infrastructure-layer defenses in MCP-based AI agent systems

Python Apache-2.0 Updated 3mos ago