#agent-evaluation
17 posts
Human-gated Content OS for Chinese social content, with Agent Skills and an explicit production-readiness CI audit.
Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin i...
Official MutagenT skills for AI coding agents — Claude Code plugin marketplace for prompt optimization, evaluation, and observability.
Production-grade Agent Skills for AI coding agents—composable workflows for planning, TDD, debugging, review, UI/UX, releases, incidents, an...
Agent skills for trapstreet.run — set up the tp CLI, build solutions against an eval task, and author new tasks, from plain language. Claude...
🔁 Build reliable recurring AI-agent systems: 874 resources, 22 operational patterns, 22 loop contracts, 8 runtime starters, an interactive...
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A...
Production AI agent quality gate and risk control framework for LLMOps, agent evaluation, regression detection, gray release, audit, and obs...
An automated red-teaming and reliability-auditing platform for AI agents - tests for prompt injection, tool hijacking and data exfiltration....
A living world where agents exist as participants alongside NPCs, internal actors, real service APIs, budgets, policies, and consequences.
Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite eval and trace replay...