← All topics

Models

Model releases, capabilities, pricing, and benchmarks.

375 posts

Claude Skill 33

Eco mode for Claude Code. /eco: -31% to -73% output tokens with critical findings intact; /eco-max: up to -75% with lowered effort. Measured...

JavaScript MIT Updated 2mos ago
Claude Skill 167

Architecture-first skill lifecycle for AI agents. 5 modes: CREATE → EVAL → EDIT → REVIEW → PACKAGE. Integrates Anthropic's eval engine (grad...

Python MIT Updated 1mo ago
Claude Skill 51

Six Claude Code skills that harden Opus 4.8 toward frontier behavior — written by Fable 5, pressure-tested on the target model with transcri...

TypeScript MIT Updated 1mo ago
Claude Skill 13

Project-local agent harness skill for traceable AI workflows

Python MIT Updated 2mos ago
Claude Skill 37

Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

Python MIT Updated 4w ago
Claude Skill 111

Fable 5-grade work discipline for any Claude model — a Claude Code skill + guard hooks (plan gate, model ceiling, per-task enforcement) that...

Python MIT Updated 1mo ago
Claude Skill 83

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

Python MIT Updated 1mo ago
MCP Server 35

MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols

Python MIT Updated 6mos ago
MCP Server 8

Private, local-first AI assistant for Windows - use your own model, keep durable memory, and approve every sensitive tool.

C# Apache-2.0 Updated 1mo ago
MCP Server 6

Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, lay...

Python MIT Updated 4mos ago