← All tags

#agent-evaluation

17 posts

Claude Skill 1

Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin i...

Python MIT Updated 1mo ago
Claude Skill 4

Official MutagenT skills for AI coding agents — Claude Code plugin marketplace for prompt optimization, evaluation, and observability.

Shell MIT Updated 4w ago
Claude Skill 89

Production-grade Agent Skills for AI coding agents—composable workflows for planning, TDD, debugging, review, UI/UX, releases, incidents, an...

Python MIT Updated 3w ago
Claude Skill 5

Agent skills for trapstreet.run — set up the tp CLI, build solutions against an eval task, and author new tasks, from plain language. Claude...

3 skills Python MIT Updated 3w ago
Library 46

🔁 Build reliable recurring AI-agent systems: 874 resources, 22 operational patterns, 22 loop contracts, 8 runtime starters, an interactive...

Python CC0-1.0 Updated 1mo ago
MCP Server 2

Production AI agent quality gate and risk control framework for LLMOps, agent evaluation, regression detection, gray release, audit, and obs...

Python MIT Updated 1mo ago
MCP Server 2

An automated red-teaming and reliability-auditing platform for AI agents - tests for prompt injection, tool hijacking and data exfiltration....

Python MIT Updated 1mo ago
MCP Server 7

A living world where agents exist as participants alongside NPCs, internal actors, real service APIs, budgets, policies, and consequences.

Python MIT Updated 2mos ago
Claude Skill 151

Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite eval and trace replay...

Python MIT Updated 3mos ago