← All tags

#benchmarking

12 posts

Claude Skill 5

Codex plugin for reproducible Blender modeling, validation, animation, MCP tooling, and agent benchmarking

TypeScript MIT Updated 2w ago
Claude Skill 32

Cross-platform .NET performance engineering skill for coding agents, covering CPU, memory, GC, benchmarking, concurrency, startup, native pr...

Python MIT Updated 1mo ago
Claude Skill 37

Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

Python MIT Updated 4w ago
MCP Server 168

Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)

Python MIT Updated 5mos ago
MCP Server 146

Goku is an HTTP load testing application written in Rust

Rust MIT Updated 2mos ago
MCP Server 19

Auditable context capsules for LLM handoffs, coding agents, and OpenCode MCP workflows.

Python MIT Updated 2mos ago
Claude Skill 14

A benchmarking harness for coding agents.

Python MIT Updated 2mos ago
MCP Server 41

Iterative agent harness improvement: run a coding agent on a hard task, generate the reusable tooling it was missing, qualify it, and replay...

Python Apache-2.0 Updated 2mos ago
MCP Server 13

The open testing standard for voice AI agents. Deterministic + semantic + RAG augmented evaluation. Local first. Zero telemetry.

Python NOASSERTION Updated 2mos ago