Multi-tool LLM agent benchmark: 496 end-to-end MCP tasks, deterministic evaluation, Russian localization
#benchmark
30 posts
DevOps MCP server, Flight recorder for AI infrastructure agents. The prescribe/report protocol captures intent before execution and outcome...
Assay — a canary-oracle benchmark for MCP security: a frozen task set scored by a recomputable HMAC-canary oracle (structural-zero false pos...
Tool-neutral attack corpus for AI agent egress security
Multi-engine LLM benchmark & monitoring CLI for Apple Silicon
Race a baseline vs a skill or MCP on real tasks. Hard checks show if it got better, faster, or cheaper. Numbers, not vibes.
Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs
Measure whether an MCP server or skill actually improves your coding agent. A/B arms, deterministic gates, your own models. Negative results...
Multi-hop cross-prompt injection benchmark for multi-agent AI systems. 250 attack cases, 8 taxonomy categories, 4 defenses evaluated. Watch:...
Goku is an HTTP load testing application written in Rust
An external memory for LLMs: a bi-temporal knowledge graph with hybrid (keyword + vector) retrieval and a self-improving reflection loop. Be...
Self-hosted engineering memory for coding agents with CLI, MCP, proof gates and benchmark artifacts.