← All tags

#agent-evaluation

9 posts

MCP Server 2

Production AI agent quality gate and risk control framework for LLMOps, agent evaluation, regression detection, gray release, audit, and obs...

Python MIT Updated 1w ago
MCP Server 2

An automated red-teaming and reliability-auditing platform for AI agents - tests for prompt injection, tool hijacking and data exfiltration....

Python MIT Updated 2w ago
MCP Server 7

A living world where agents exist as participants alongside NPCs, internal actors, real service APIs, budgets, policies, and consequences.

Python MIT Updated 1mo ago
Claude Skill 146

Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite eval and trace replay...

Python MIT Updated 1mo ago
MCP Server 8

Open-source test harness for AI agents. Stress-test production agents with adversarial multi-turn scenarios in CI

Python Apache-2.0 Updated 1w ago
Claude Skill 10

CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscription...

Python MIT Updated 1w ago
MCP Server 8

The agent eval standard for MCP — score output quality, catch safety failures, enforce cost budgets

TypeScript MIT Updated 1w ago
MCP Server 123

Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic...

Python Apache-2.0 Updated 2w ago
MCP Server 1.2k

A single interface to use and evaluate different agent frameworks

Python Apache-2.0 Updated 2w ago