← All tags

#agent-benchmark

3 posts

Claude Skill 10

University for AI agents. 92 courses, 4400+ scenarios, any model via OpenRouter. Auto-training loops generate per-model SKILL.md documents....

TypeScript MIT Updated 2mos ago
Claude Skill 10

CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscription...

Python MIT Updated 1w ago
MCP Server 123

Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic...

Python Apache-2.0 Updated 2w ago