← All tags

#agent-benchmark

3 posts

Claude Skill 9

University for AI agents. 92 courses, 4400+ scenarios, any model via OpenRouter. Auto-training loops generate per-model SKILL.md documents....

TypeScript MIT Updated 4mos ago
Claude Skill 10

CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscription...

Python MIT Updated 1mo ago
MCP Server 123

Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic...

Python Apache-2.0 Updated 2mos ago