Production AI agent quality gate and risk control framework for LLMOps, agent evaluation, regression detection, gray release, audit, and obs...
#agent-evaluation
9 posts
An automated red-teaming and reliability-auditing platform for AI agents - tests for prompt injection, tool hijacking and data exfiltration....
A living world where agents exist as participants alongside NPCs, internal actors, real service APIs, budgets, policies, and consequences.
Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite eval and trace replay...
Open-source test harness for AI agents. Stress-test production agents with adversarial multi-turn scenarios in CI
CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscription...
The agent eval standard for MCP — score output quality, catch safety failures, enforce cost budgets
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic...
A single interface to use and evaluate different agent frameworks