← All tags

#benchmark

30 posts

MCP Server 2

Multi-tool LLM agent benchmark: 496 end-to-end MCP tasks, deterministic evaluation, Russian localization

Python Apache-2.0 Updated 1w ago
MCP Server 13

DevOps MCP server, Flight recorder for AI infrastructure agents. The prescribe/report protocol captures intent before execution and outcome...

Go Apache-2.0 Updated 3w ago
MCP Server 6

Assay — a canary-oracle benchmark for MCP security: a frozen task set scored by a recomputable HMAC-canary oracle (structural-zero false pos...

Python MIT Updated 1mo ago
MCP Server 11

Multi-engine LLM benchmark & monitoring CLI for Apple Silicon

Python Apache-2.0 Updated 1w ago
Claude Skill 6

Race a baseline vs a skill or MCP on real tasks. Hard checks show if it got better, faster, or cheaper. Numbers, not vibes.

Python Updated 1mo ago
Claude Skill 71

Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs

TypeScript MIT Updated 1mo ago
MCP Server 2

Measure whether an MCP server or skill actually improves your coding agent. A/B arms, deterministic gates, your own models. Negative results...

JavaScript MIT Updated 3w ago
MCP Server 7

Multi-hop cross-prompt injection benchmark for multi-agent AI systems. 250 attack cases, 8 taxonomy categories, 4 defenses evaluated. Watch:...

Python MIT Updated 1mo ago
MCP Server 146

Goku is an HTTP load testing application written in Rust

Rust MIT Updated 3w ago
MCP Server 21

C0

An external memory for LLMs: a bi-temporal knowledge graph with hybrid (keyword + vector) retrieval and a self-improving reflection loop. Be...

Rust MIT Updated 1w ago
MCP Server 4

Self-hosted engineering memory for coding agents with CLI, MCP, proof gates and benchmark artifacts.

TypeScript MIT Updated 1mo ago