Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic...
Models
Model releases, capabilities, pricing, and benchmarks.
293 posts
Open-source reverse engineering lab: 197-article knowledge base + MCP tools + CTF/APK/PE automation toolchain. Agent-native. Note:由于场景...
Run production apps without thinking about infrastructure. On your server or ours. Fully agentic.
The open testing standard for voice AI agents. Deterministic + semantic + RAG augmented evaluation. Local first. Zero telemetry.
Industrial benchmark for AI memory – measuring how accurately LLM and agent memory systems recall long-term context, with uniform, reproduci...
Quickly find bottlenecks in Rust - one profiler for CPU, time, memory, SQL and async code.
Opinionated agentic RAG powered by LanceDB, Pydantic AI, and Docling
MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.
AssetOpsBench - Industry 4.0: A unified benchmark and framework for building, orchestrating, and evaluating domain-specific AI agents for In...
AI-powered offensive security agent with 7,300+ actionable security skills. Autonomous pentesting powered by MITRE ATT&CK (2,000+ Atomic tes...
The best-benchmarked open-source AI memory system. And it's free.
Octopus Deploy Official MCP Server