Tool-neutral attack corpus for AI agent egress security
#benchmark
39 posts
Multi-engine LLM benchmark & monitoring CLI for Apple Silicon
Race a baseline vs a skill or MCP on real tasks. Hard checks show if it got better, faster, or cheaper. Numbers, not vibes.
Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs
Measure whether an MCP server or skill actually improves your coding agent. A/B arms, deterministic gates, your own models. Negative results...
Multi-hop cross-prompt injection benchmark for multi-agent AI systems. 250 attack cases, 8 taxonomy categories, 4 defenses evaluated. Watch:...
Goku is an HTTP load testing application written in Rust
An external memory for LLMs: a bi-temporal knowledge graph with hybrid (keyword + vector) retrieval and a self-improving reflection loop. Be...
Self-hosted engineering memory for coding agents with CLI, MCP, proof gates and benchmark artifacts.
A benchmarking harness for coding agents.
An Agent Infra Benchmark Suite for Agentic Tool Serving (AgentCore, Lambda, GCP, AgentRun, E2b etc.)
An MCP server that lets LLM agents play Civilization VI.