← All tags

#benchmarking

9 posts

Claude Skill 33

Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

Python MIT Updated 1w ago
MCP Server 168

Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)

Python MIT Updated 4mos ago
MCP Server 146

Goku is an HTTP load testing application written in Rust

Rust MIT Updated 3w ago
MCP Server 19

Auditable context capsules for LLM handoffs, coding agents, and OpenCode MCP workflows.

Python MIT Updated 1mo ago
Claude Skill 14

A benchmarking harness for coding agents.

Python MIT Updated 2w ago
MCP Server 40

Iterative agent harness improvement: run a coding agent on a hard task, generate the reusable tooling it was missing, qualify it, and replay...

Python Apache-2.0 Updated 2w ago
MCP Server 13

The open testing standard for voice AI agents. Deterministic + semantic + RAG augmented evaluation. Local first. Zero telemetry.

Python NOASSERTION Updated 4w ago