← All tags

#ai-evaluation

7 posts

Claude Skill 10

Production-ready Claude Code skills for building AI agents with Pydantic AI. Includes dependency injection, tools, validators, streaming, mu...

Python MIT Updated 2w ago
MCP Server 6

Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, lay...

Python MIT Updated 2mos ago
Claude Skill 2

Vendor-neutral research umbrella for measuring AI plugin, agent, and MCP server quality across CLI runtimes (Claude Code, Gemini CLI, Copilo...

Python NOASSERTION Updated 1w ago
MCP Server 6

MCP server for human-in-the-loop surveys, A/B preference tests, ratings, and rankings. Get real human feedback inside Claude Code, Claude De...

Python MIT Updated 1mo ago
MCP Server 22

⚡️ The "1-Minute RAG Audit" — Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, async, visual reports.

Python Apache-2.0 Updated 1mo ago
MCP Server 154

Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.

MIT Updated 3w ago
Claude Skill 10

CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscription...

Python MIT Updated 1w ago