← All tags

#benchmark

39 posts

Claude Skill 3

Turn "does this AI feature work?" into a scorecard someone can argue with. Rubrics, failure taxonomies, gates and confidence intervals — ren...

Python MIT Updated 3w ago
Claude Skill 6

让 AI 像资深测试工程师一样工作:全生命周期 QA Agent Skills 框架——方法论 + 10 Skills + 可复现 Benchmark(Claude Code 等 Agent 可用)

Python MIT Updated 2w ago
Claude Skill 4

A permission-aware Agent Skill for reproducible, auditable batch episode collection across VLA, policy, world-model, and embodied benchmark...

Python MIT Updated 1mo ago
Claude Skill 4

Agent Skill for ASD-STE100 Simplified Technical English. Makes LLMs write text that survives one read. Benchmarked on 12 models: 95.5% fewer...

Python MIT Updated 3w ago
Claude Skill 1

Claude Code skills where every entry ships receipts — accuracy-gated benchmarks against baseline and placebo, rejects published

JavaScript MIT Updated 1mo ago
Claude Skill 2

Agent skills for XORCISE — benchmark AI models on hands-on cyber missions and get an eval-card report.

Python NOASSERTION Updated 1mo ago
Claude Skill 26

Every AI slide-deck skill worth knowing, on one page — both the HTML and native-PPTX routes, compared on what actually decides the choice ra...

Python NOASSERTION Updated 3w ago
Claude Skill 1

Task Compass Skill: deterministic, auditable task routing for AI agents and OpenClaw

Python MIT Updated 1mo ago
Claude Skill 2

Workbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more

TypeScript MIT Updated 1mo ago