← All tags

#llm-evaluation

3 posts

Claude Skill 5

Portable Agent Skills for disciplined software delivery, learning, editing, and LLM evaluation

Python NOASSERTION Updated 3w ago
Claude Skill 2

Evidence-first evaluation and creation for Agent Skills.

Python MIT Updated 1mo ago
MCP Server 7

Production-grade RAG pipelines with evaluation baked in

Python MIT Updated 4mos ago