Claude Skills · Code Review & Testing
Agent Eval
ericrisco/rsc-harnessUse when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is `building-agents`).
At a glance
This skill is for Code Review & Testing and helps you evaluate llm model performance, score agent trajectories, and assess rag system quality.
git clone --depth 1 https://github.com/ericrisco/rsc-harness
cp -r rsc-harness/skills/agent-eval ~/.claude/skills/agent-eval
Setup, runtime and requirements describe ericrisco/rsc-harness, the repo this skill ships in.
llm-evaluationGolden SetsRag Scoringagent-testingFaithfulnessbenchmark
Also in ericrisco/rsc-harness
View the repoUse when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CU...
Use when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA...
Use when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platfo...
Use when bounding an LLM agent that already runs — scoping its task domain, gating tools to least privilege, defending against prompt inject...
Use when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-vid...
Use when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PI...
Use when constitution, spec, plan and tasks all exist and you want them cross-read against each other before any code is written — the rsc-s...
Use when building, refactoring, or debugging Angular (v20/21+): standalone components, signals, zoneless change detection, @if/@for/@defer c...
Use when writing a client for someone else's REST or GraphQL API: auth flow choice and token refresh, pagination to exhaustion, retry-with-j...
Use when settling the contract of an API you expose, before implementation: resources/URLs, REST vs GraphQL, versioning, one RFC 9457 error...
Use when writing one long-form article end to end — answer-first lede, question-shaped headings, plus its on-page surface (title, meta, slug...
Use when building a content-driven or marketing site with Astro 6: static-first pages, islands and partial hydration, content collections, s...
Other Code Review & Testing skills
A relentless interview to sharpen a plan or design, which also creates docs (ADR's and glossary) as we go.
Use when you need to resolve an in-progress git merge/rebase conflict.
Create exercise directory structures with sections, problems, solutions, and explainers that pass linting. Use when user wants to scaffold e...
Migrate test files from `as` type assertions to @total-typescript/shoehorn. Use when user mentions shoehorn, wants to replace `as` in tests,...
Move issues and external PRs through a state machine of triage roles — categorise, verify, grill if needed, and write agent-ready briefs.
A relentless interview that asks every frontier question at once, round by round.