← All topics

Models

Model releases, capabilities, pricing, and benchmarks.

292 posts

Claude Skill 1.7k

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / pr...

JavaScript MIT Updated 1w ago
Claude Skill 24

Eco mode for Claude Code. /eco: -31% to -73% output tokens with critical findings intact; /eco-max: up to -75% with lowered effort. Measured...

JavaScript MIT Updated 2w ago
Claude Skill 110

Architecture-first skill lifecycle for AI agents. 5 modes: CREATE → EVAL → EDIT → REVIEW → PACKAGE. Integrates Anthropic's eval engine (grad...

Python MIT Updated 3w ago
Claude Skill 48

Six Claude Code skills that harden Opus 4.8 toward frontier behavior — written by Fable 5, pressure-tested on the target model with transcri...

TypeScript MIT Updated 1w ago
Claude Skill 10

Project-local agent harness skill for traceable AI workflows

Python MIT Updated 3w ago
Claude Skill 111

Fable 5-grade work discipline for any Claude model — a Claude Code skill + guard hooks (plan gate, model ceiling, per-task enforcement) that...

Python MIT Updated 1w ago
Claude Skill 33

Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.

Python MIT Updated 1w ago
Claude Skill 71

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

Python MIT Updated 1w ago
MCP Server 35

MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols

Python MIT Updated 4mos ago
MCP Server 8

Private, local-first AI assistant for Windows - use your own model, keep durable memory, and approve every sensitive tool.

C# Apache-2.0 Updated 1w ago