Eco mode for Claude Code. /eco: -31% to -73% output tokens with critical findings intact; /eco-max: up to -75% with lowered effort. Measured...
Models
Model releases, capabilities, pricing, and benchmarks.
375 posts
Architecture-first skill lifecycle for AI agents. 5 modes: CREATE → EVAL → EDIT → REVIEW → PACKAGE. Integrates Anthropic's eval engine (grad...
Six Claude Code skills that harden Opus 4.8 toward frontier behavior — written by Fable 5, pressure-tested on the target model with transcri...
Project-local agent harness skill for traceable AI workflows
Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.
Fable 5-grade work discipline for any Claude model — a Claude Code skill + guard hooks (plan gate, model ceiling, per-task enforcement) that...
[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
Scopus MCP provides researchers with tools shaped by real workflows
Private, local-first AI assistant for Windows - use your own model, keep durable memory, and approve every sensitive tool.
Grant the AI octopus access to a portion of your desktop
Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, lay...