ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.
#benchmark
39 posts
Turn "does this AI feature work?" into a scorecard someone can argue with. Rubrics, failure taxonomies, gates and confidence intervals — ren...
让 AI 像资深测试工程师一样工作:全生命周期 QA Agent Skills 框架——方法论 + 10 Skills + 可复现 Benchmark(Claude Code 等 Agent 可用)
A permission-aware Agent Skill for reproducible, auditable batch episode collection across VLA, policy, world-model, and embodied benchmark...
Agent Skill for ASD-STE100 Simplified Technical English. Makes LLMs write text that survives one read. Benchmarked on 12 models: 95.5% fewer...
Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published
Agent skills for XORCISE — benchmark AI models on hands-on cyber missions and get an eval-card report.
Every AI slide-deck skill worth knowing, on one page — both the HTML and native-PPTX routes, compared on what actually decides the choice ra...
Anti-slop writing skill that preserves facts, voice, and density, with public benchmarks.
Task Compass Skill: deterministic, auditable task routing for AI agents and OpenClaw
Timestamps: 00:00 - Intro 02:05 - 3D Model Test Overview 04:02 - 3D Model Test 10:26 - 3D Model Results 13:10 - AI Magazine Test Overview 1...
Workbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more