← All tags

#benchmark

39 posts

MCP Server 11

Multi-engine LLM benchmark & monitoring CLI for Apple Silicon

Python Apache-2.0 Updated 1mo ago
Claude Skill 6

Race a baseline vs a skill or MCP on real tasks. Hard checks show if it got better, faster, or cheaper. Numbers, not vibes.

Python Updated 2mos ago
Claude Skill 71

Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs

TypeScript MIT Updated 3mos ago
MCP Server 2

Measure whether an MCP server or skill actually improves your coding agent. A/B arms, deterministic gates, your own models. Negative results...

JavaScript MIT Updated 2mos ago
MCP Server 7

Multi-hop cross-prompt injection benchmark for multi-agent AI systems. 250 attack cases, 8 taxonomy categories, 4 defenses evaluated. Watch:...

Python MIT Updated 2mos ago
MCP Server 146

Goku is an HTTP load testing application written in Rust

Rust MIT Updated 2mos ago
MCP Server 21

C0

An external memory for LLMs: a bi-temporal knowledge graph with hybrid (keyword + vector) retrieval and a self-improving reflection loop. Be...

Rust MIT Updated 1mo ago
MCP Server 4

Self-hosted engineering memory for coding agents with CLI, MCP, proof gates and benchmark artifacts.

TypeScript MIT Updated 2mos ago
Claude Skill 14

A benchmarking harness for coding agents.

Python MIT Updated 2mos ago
MCP Server 6

An Agent Infra Benchmark Suite for Agentic Tool Serving (AgentCore, Lambda, GCP, AgentRun, E2b etc.)

Python Apache-2.0 Updated 2mos ago
MCP Server 119

An MCP server that lets LLM agents play Civilization VI.

Python MIT Updated 2mos ago