← All tags

#evaluation

11 posts

Claude Skill 2.1k

YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

1 skills Python MIT Updated 2w ago
MCP Server 9

The Full-Stack LLM Engineering Playbook. Architectural patterns for Agents (MCP) & RAG, coupled with advanced Post-Training recipes (SFT, DP...

Updated 2w ago
MCP Server 2

Multi-tool LLM agent benchmark: 496 end-to-end MCP tasks, deterministic evaluation, Russian localization

Python Apache-2.0 Updated 1w ago
MCP Server 6

PrometheusLLM is a unique transformer architecture inspired by dignity and recursion. This project aims to explore new frontiers in AI resea...

Python GPL-3.0 Updated 1w ago
kakz 0 0
MCP Server 4

A judgment engine: register a thesis, steelman each claim, measure it against the evidence that could refute it, and refine the weakest axis...

Python NOASSERTION Updated 1w ago
MCP Server 22

⚡️ The "1-Minute RAG Audit" — Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, async, visual reports.

Python Apache-2.0 Updated 1mo ago
MCP Server 4

MCPLab - Test and evaluate MCP servers with LLMs

TypeScript MIT Updated 3w ago
MCP Server 8

Field-fit agents to real work. Turn tacit know-how into reusable capabilities.

Python Apache-2.0 Updated 4w ago
MCP Server 13

The open testing standard for voice AI agents. Deterministic + semantic + RAG augmented evaluation. Local first. Zero telemetry.

Python NOASSERTION Updated 4w ago