Skill instalable para Claude Code que mide si tu agente cumple las reglas establecidas. Lee el historial de sesiones de tu disco y devuelve...
#ai-evaluation
10 posts
Software for HackerRank Orchestrate: an evaluator, an engineering-memory CLI that remembers what you measured and rejected, a mentor that co...
Today we'll be comparing Claude Opus 5 and Fable 5 to GPT 5.6 Sol Pro and GPT 5.6 Sol Ultra (through ChatGPT Work) on a research maths prob...
Production-ready Claude Code skills for building AI agents with Pydantic AI. Includes dependency injection, tools, validators, streaming, mu...
Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, lay...
Vendor-neutral research umbrella for measuring AI plugin, agent, and MCP server quality across CLI runtimes (Claude Code, Gemini CLI, Copilo...
MCP server for human-in-the-loop surveys, A/B preference tests, ratings, and rankings. Get real human feedback inside Claude Code, Claude De...
⚡️ The "1-Minute RAG Audit" — Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, async, visual reports.
Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.
CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscription...