We pointed it at the providers. Anthropic's default context-editing policy (keep=3) changed the
agent's next action in 95–100% of cases, against a 2.5% A/A noise floor. Keeping the 3 most recent
tool uses didn't lower the change rate at all — it turned stalling into acting on missing facts.
OpenAI's compaction changed 12.5–20%. Pre-registered, replicated, n=40 per run.
Read the study → · rerun it on your own config
What it does
- Wrap your agent —
distil wrap -- claude·codex·gemini·aider·opencode·qwen·goose·grok·openhands. Zero config, no code change. - Run a proxy — point any
base_urlclient at it. Python, TypeScript, any language, any framework. - Call it as a library —
from distil import compress_messagesin your own agent loop. - Give your agent a recall tool — MCP server: it compresses its own output and gets the exact bytes back on demand.
- Framework hooks — LangChain · LangGraph · LiteLLM · Agno · Strands, in-process, no network hop.
- On a subscription —
distil hook --install: Claude Code compresses its own tool output through the documentedPostToolUseextension point. No proxy, no credentials touched.distil quotashows the rate-limit window it buys back. Details → - See what it did — live status line, session dissect, per-request headers, OTel spans, Prometheus metrics.
pipx install distil-llm && distil onboard # detects your agent + billing, wires everything
Not sure which of those you want? Two questions pick your mode → — plain language, honest savings ranges, no jargon.
Will it save you money? On metered billing (an API key), yes — directly, off the bill. On a flat-rate Pro/Max subscription there is no per-token bill to cut, but there is a rate-limit window, and spending fewer tokens per turn leaves more of it for the next task.
distil quotashows that window live. Savings come from large, repetitive tool output: verbose JSON and duplicated log runs compress 25–99%, while prose and unique-line output compress ~0% — a short session that never reads a big file showing near 0% is the tool working correctly, not failing. Why →
🧩 Use it as a library
Building the agent yourself? Compress the message list where it lives — no proxy, no network hop:
from distil import compress_messages, expand_handle
result = compress_messages(messages) # OpenAI/Anthropic-style dicts
print(f"{result.saved_pct:.1f}% smaller")
response = client.messages.create(model=..., messages=result.messages)
original = expand_handle(result.handles[0]) # byte-exact, any time, any process
Tool results get the reversible digest; user and system text get lossless transforms only; the model's own turns are never rewritten. Handles resolve across processes and restarts, so a digest made by the proxy expands here and vice versa. verbatim=True disables digests entirely.
Named compress_messages/expand_handle rather than compress/expand because distil.compress and distil.expand are modules — a top-level export sharing those names would resolve to the function or the module depending on unrelated import order.
TypeScript too — compress(messages) from the npm package, byte-identical to the Python engine. Full reference: Library API → · runnable examples: python_library.py · js_library.ts.
Maintain a framework? docs/INTEGRATING.md is the ~20 lines and the four rules — we would rather the integration live in your repo than ours.
🚀 Use it now
One command sets you up and tells you what to do next:
pipx install distil-llm
distil onboard # detects your agent + billing, wires the status line, prints a guided tour
It detects your environment (Claude Code · Codex · Gemini CLI; metered vs subscription) and hands you the exact commands. Or wrap your agent directly — no config, no code change:
# Claude Code on a metered API key — saves real $$:
distil wrap --expand -- claude
# Claude Code on a Pro/Max subscription — flat-rate, ToS-safe (trims context, not $):
distil wrap --lossless-only -- claude
# Codex, Gemini CLI, aider — same pattern; env var auto-selected per agent:
distil wrap --expand -- codex # → OPENAI_BASE_URL
distil wrap --expand -- gemini # → GOOGLE_GEMINI_BASE_URL
distil wrap --expand -- aider # → OPENAI_BASE_URL
# Headless too — print mode, CI, and Agent SDK scripts route the same way:
distil wrap -- claude -p "summarise this diff"
distil wrap -- python my_agent_sdk_script.py
Using Cursor, Cline, Continue, or Windsurf? They are IDE extensions — no argv to wrap and no documented env var, so
distil wrapcannot reach them. Run a proxy and point the editor's base-URL setting at it: docs/IDE-AGENTS.md. (GitHub Copilot is not redirectable at all, and that page says so rather than wasting your afternoon.)
Each recognized agent (claude / codex / gemini / aider / opencode / qwen / goose) auto-selects the right env var and upstream — no --env-var or --upstream flag needed. Prints preset: <agent> detected → <VAR> on start. Explicit flags always win.
Tired of typing distil wrap every time? Make it the default — once:
distil default # adds a managed shell alias so `claude` always routes through distil
distil default --undo # remove it anytime (backed up before any change)
It detects your shell (zsh / bash / fish / PowerShell) and billing mode, writes the
right line to the rc file your shell actually reads, and tells you what it detected.
Want every SDK covered (not just the agent you type)? distil default --always-on
runs a persistent proxy service — powerful, but it pins ANTHROPIC_BASE_URL, so
every client on the machine goes through one local process.
That pin used to be a single point of failure: a proxy that was down for one
second meant sessions failing with ConnectionRefused, an error that names the
provider rather than distil. It no longer is. The service supervisor
(launchd/systemd) owns the listening socket, so a crash or a restart leaves
connections queued in the kernel backlog instead of refused — the client waits
about a second rather than dying. distil default --always-on also verifies the
service is genuinely registered and serving before it wires anything, and refuses
to wire at all if it isn't.
If you ever need out and distil is already uninstalled, sh ~/.distil/uninstall.sh
removes the pin, the service, and the shell block using nothing but sh.
Then watch genuine savings from your traffic — measured, not estimated:
distil leaderboard # cumulative tokens + $ saved, from the local ledger
distil dashboard # live terminal TUI — token-trim + decision-equiv bars, Ctrl-C to exit
distil dissect # per-session deep-dive: savings, digest inventory, anomalies (--html/--serve)
Validate it on your traffic. --shadow runs a fraction of requests twice (compressed and full) and compares the agent's chosen next action:
distil wrap --shadow 0.1 -- claude # wrap + shadow 10% of requests
distil shadow-stats # live decision-equivalence rate
Honest scope: that's next-action equivalence — a proxy, not task success (E7 shows it doesn't fully transfer under aggressive lossy compression). Distil fails safe to full context.
Will it save money? On metered billing (API key) — fewer tokens, fewer dollars, directly. On a flat-rate subscription there is no per-token bill, so the saving is rate-limit headroom: fewer tokens per turn means more turns before you hit the window (
distil quotashows it live). Coding agents: short sessions ~7%, big wins on long, many-turn sessions the model never re-reads.
💡 Why Distil is different
You don't need byte-equivalence — you need decision-equivalence: your agent taking the same actions with compressed context. That's measurable and certifiable.
- Certified, not estimated — a strategy ships only if a non-inferiority test passes; can't certify → full context.
- Certified end-to-end, too —
distil certify-trajectoriesbounds how many solvable tasks compression can cost (no other compressor certifies either level). - Reversible, not lossy — digests behind a handle, keeps the original, hands the agent a
distil_expandtool. Compress fearlessly. - Keeps the answer, folds the noise — a per-content-type keep policy pins each kind's load-bearing lines (a log's pass/fail verdict, a traceback's frames, a diff's hunk headers); repeated near-identical error spam is deduped, and on a green run dedup tightens further since that noise didn't fail anything.
- Query-aware — keeps the line you're actually asking about — distil is a proxy, so it sees the agent's intent (its tool_use args + latest ask) in the same request as the output. The line matching what you searched for (a grep hit, a config value, a SHA) is pinned even in arbitrary output — additively, so reversibility and the certificate are untouched. No post-hoc compressor has that query/output pairing. It also goes semantic, and always-on: a zero-dependency bridge — morphology, a curated technical synonym map, and char-trigram fuzz — pins lines that answer the query without sharing a word with it. Ask "the retry limit?" and it keeps
max_attempts = 5; ask "the connection timeout?" and it keepsdeadline_ms. Two more layers grow from your own traffic, never from a shipped blob: associations distil learns from its content-free expand flywheel (hashed pairs,--expandsessions), and a learned relevance model that is promoted only after its held-out recall beats the lexical baseline on your labels — until promotion, the lexical + bridge layers are exactly what runs. An optional distributional-vector table can be supplied too (pure-Python cosine; none ships). Every layer is additive — it can only widen keeps, so reversibility and the certificate are untouched — and it needs no embeddings or model to work. - Lossless even on a flat-rate plan — subscription/lossless mode isn't just verbatim: it minifies JSON, collapses duplicate runs, and folds tabular tool output into a compact self-describing table (~70–79% smaller, ToS-safe, no lossy digest). Recent tool outputs stay byte-exact.
- See exactly what happened —
distil dissectturns a wrap session into a report: savings by model/mechanism, the digest inventory, billed-usage calibration, latency by path, and a worth-your-attention anomaly list that catches silent failures automatically. - Compounds on outcomes — expansions and matched failures teach the policy what to protect (signatures only, never content) — always more conservative.
- Streams like it isn't there — SSE relays chunk-by-chunk; TTFT preserved — including recoverable digest, which speculatively streams and only intercepts an actual
distil_expandcall mid-stream, splicing the recovery in without buffering the turn (no TTFT tax on the reversible tier).
Fidelity tiers: lossless (
--verbatim) · reversible (byte-recoverable on demand — default) · lossy (every other tool). Only Distil certifies the reversible tier (Headroom ships an uncertified retrieve; Distil's recovery is agent-facing — the model expands mid-task — and gated by the decision-equivalence certificate).
⚡ Prove the numbers yourself — no API key
Don't take the table above on faith. distil bench re-certifies savings and decision-equivalence on a bundled 8-domain corpus, offline, in seconds — the same gate that runs in CI. How we evaluate — and why a compression ratio without a task-success delta is meaningless — is written up in docs/EVALUATION.md, including our own negative result:
uvx --from distil-llm distil bench # certify savings + quality across 9 domains, in seconds
distil verify # byte-fidelity: every compression is exactly reversible
distil validate # adversarial real-path gate: invariants on hostile inputs
distil retention # fact recall: what stays visible vs expand-recoverable
distil retention --dataset hotpotqa # graded against a PUBLIC benchmark's ground truth
distil fidelity # state probes: artifact state, overclaim, continuation
Five gates, all in CI: bench (non-inferiority on the corpus), verify (byte-fidelity), retention (fact-level recall), fidelity (state probes, below), and validate — which drives the compressor against adversarial inputs (huge/unicode/nested/malformed/marker-injection/secret-looking) and asserts reversibility, reject-if-bigger, recency-exactness, fail-open, and content-free telemetry hold on every one. That last gate exists because a green unit suite kept coexisting with real-traffic bugs; validate is the adversarial layer that catches them.
Recall is not enough, and here's the case that proves it. A trajectory creates net/scratch_bench.py at turn 2 and deletes it at turn 4. Compress away turn 4 and every path token is still present — string recall reads 100% — while the agent now believes a file exists that doesn't, and will plan around it. distil fidelity folds tool calls into a file-state ledger and grades the final state, separating lost (path gone — the agent can see the gap) from stale (path present, state wrong — the agent acts confidently on a falsehood). On that case: string recall 100%, state fidelity 0%.
It reports three more things recall can't see: overclaim ("approximately 4200 ms" → "4200 ms" — the value survives, its uncertainty doesn't), continuation (does the agent still know what's left to do?), and error propagation (does a loss at turn k show up as a behaviour change at turn k+n?). The gate is on silent failures only — CI runs --max-silent 15 — because loud loss is already retention --max-lost's job, and gating one regression twice hides which property broke. The bound is the measured one, not zero: Tier-1 digests hedged spans behind restore handles and drops the qualifier on 9 of 171 claims, so gating at zero would assert a property the compressor does not have. On top of that, distil suite grades twelve public benchmarks whose answer keys were written by someone else — including BFCL, which compresses the tool schema and checks that every name the gold call needs — the function and each argument — survives. At matched savings (90.1% vs 89.3%) truncation keeps 0 of 70 names; distil keeps all 70 — though none of them visibly: the schema sits behind a restore handle, one distil_expand away. The suite prints that gap (visible → true support: bfcl 0%→100%) rather than the flattering number alone, because a reader who assumes the model can see a schema it must actually expand first has been misled by figures that are individually correct. Names are matched as identifiers — a quoted JSON token, escaping tolerated — not as prose: the generic matcher was crediting 11 of 85 golds by accident ('a' matching inside "tool-schemas"). Fifteen golds BFCL genuinely names a, b, c are excluded and counted, since a one-letter token can be neither credited nor failed honestly. Every row is labelled rich or thin payload, because a benchmark with nothing to compress is a control, not evidence — and a run that grades only controls exits 1. It needs no API key and no spend, so it is wired into make gate and the CI gate job rather than run before a launch. Full methodology, including what these probes found wrong with our own corpus, in docs/EVALUATION.md §6; how to run everything, in docs/RUNNING-EVALS.md.
Recall, and a number you can check yourself. The three gates above are graded on our corpus against our oracle — rigorous, but not checkable by you. distil retention --dataset hotpotqa grades against ground truth written by someone else (HotpotQA's gold supporting sentences, amid 8 distractor paragraphs), next to a truncation baseline tuned to distil's own savings on the same case:
| HotpotQA, n=100 | savings | answer recall | gold-sentence recall |
|---|---|---|---|
| distil (reversible) | 14.3% | 100.0% | 100.0% |
| truncation @ matched savings | 14.1% | 91.6% | 82.7% |
distil retention also splits recall into visible (in front of the model) and recoverable (one distil_expand away, verified against the handle's restore bytes). On the corpus that's 100% true recall with 0 lost, and being reversible instead of lossy is worth 21.4% recall — the mean across all 9 domains, each counted once. That's deliberately the macro average: the fact-weighted one reads 62.6%, but it's set by whichever domain carries the most probes, and one HTML fixture moved it from 9.8% to 62.6% without the compressor changing at all — the moat, as a measurement rather than an argument. distil retention --live reports the same on your own traffic; the meter stores counts only, never content.
And it found a real hole. The first thing the recall harness caught was not a regression but a missing capability: distil was compressing 0.0% of HTML tool results — minified markup is one long line, so line-folding had nothing to fold. Agents with a fetch or browser tool were paying full price for <script>, <style>, and nav chrome. Now:
| real page | before | after | saved | facts lost |
|---|---|---|---|---|
| Wikipedia article | 281,093 tok | 14,260 tok | 94.9% | 0 |
| Python docs page | 32,322 tok | 4,229 tok | 86.9% | 0 |
Reversible, which is the part a lossy extractor can't offer: the exact original stays behind the handle, so a bad heuristic call costs one distil_expand instead of the content.
To be precise about what each layer proves: the per-commit gates grade decision-equivalence with an offline deterministic oracle over the committed corpus (fast, free, runs on every push — but synthetic). A nightly live-cert job re-certifies the same trajectories against a real model (distil certify --runner anthropic), budget-capped with a hard --max-live-calls ceiling so an unattended run can never spend silently. The empirical results above (SWE-bench n=500, live head-to-head n=200) were graded by real models; the per-commit badge alone doesn't claim that.
domain trajectory $ saved distil aggr pruned
---------------------------------------------------------------------------
ops/sre sre-disk-incident 32.8% PASS FAIL 615
coding coding-bugfix 25.5% PASS FAIL 736
support support-refund 32.6% PASS FAIL 765
research research-synthesis 25.7% PASS FAIL 809
data-analysis data-analysis-sql 18.1% PASS FAIL 965
devops devops-rollback 22.8% PASS FAIL 857
finance finance-reconcile 24.9% PASS FAIL 1014
web-research web-research 89.8% PASS FAIL 428
agent-worklog agent-worklog 35.3% PASS FAIL 891
---------------------------------------------------------------------------
aggregate: distil cuts $0.24052 -> $0.12400 (48.4% cheaper) reversibly; 7080 tokens causally prunable.
GATE: PASS — every trajectory certified non-inferior; aggressive rejected on all.
Why trust the number? Token-savings numbers are easy to fake — measure quality at low compression, advertise savings at high compression. Distil refuses that: accuracy and compression are measured on the same trajectories, and a strategy that can't pass non-inferiority doesn't ship.
distil certify --strategy distil # VERDICT: PASS (100% decision-equivalence) distil certify --strategy aggressive # VERDICT: FAIL (mean diff −1.0, blocked)
distil eval plots the certified compression frontier — a savings-vs-quality curve where every point carries its certification verdict, locating the cliff past which lossy compression drops decisions. The artifact no competitor publishes: benchmark.html.
📊 The proof
Three results, all reproducible, all published with caveats:
- Live head-to-head vs real
llmlingua/headroom-ai(graded byclaude-opus-4-8): 83.2% savings at 0% decision-change, ~1,000× faster (no ML model loaded vs. competitors' local transformer inference). The live proxy behavior is pinned to the certified strategy bytests/test_live_certified_equivalence.py; the one reviewed delta is a recency carve-out that keeps the freshest tool-result turns verbatim (an agent needs its freshest output byte-exact). Since 1.45 that carve-out applies only to content the provider has not cached — anchored to the client'scache_controlbreakpoint, and dropped entirely for providers that cache implicitly. A carve-out counted back from the end of the conversation slid forward as it grew, rewriting already-cached content one turn later and costing more in re-billed prefix than the digest saved. → benchmark - E7 (SWE-bench Verified): aggressive lossy compression craters task success (52% → 16%) — a per-step certificate doesn't transfer to multi-turn. The reversible tier survives (56% vs 52%). We publish it because it's true. → E7
- E8–E14 (500-instance agent): the reversible tier is the only compressor non-inferior to full context, generalizes across 5 models / 3 vendors, and the newest digest matches full within noise (42.0% vs 39.2%). → E8–E14
Full methodology, McNemar tests, per-instance data: docs/PAPER.md · PDF.
📡 See it working
Measured on your traffic, never estimated, nothing leaves your machine:
- Per request:
x-distil-*response headers (tokens-saved,mode,compressible-tokens,expanded). - Per machine:
distil leaderboard(--htmlfor a page). - Shadow mode:
distil proxy --shadow 0.05reports the live decision-change rate — streaming-aware. - Org-wide:
distil proxysidecar + setANTHROPIC_BASE_URLonce; every client routes through it. - Community: an opt-in census (
distil census on) shares your numbers-only totals — preview the exact payload withdistil census showbefore consenting;TELEMETRY.mdhas the frozen schema. Default remains: nothing is sent.
Dashboard, status-line plugin, federated leaderboard: Deploy & observability.
🔌 Works with every SDK
One proxy. Point any base_url-honoring client at it — Python, TypeScript, any language — and get cache-aware reversible compression with no code change.
distil proxy --upstream https://api.anthropic.com # localhost:8788
// JS/TS: npm i distil-llm → helper so you don't hardcode the URL
import Anthropic from "@anthropic-ai/sdk";
import { distilBaseURL } from "distil-llm";
const client = new Anthropic({ baseURL: distilBaseURL() });
| SDK / framework | Change | Example |
|---|---|---|
| Anthropic SDK (Py/TS) | base_url="http://127.0.0.1:8788" |
examples/python_anthropic.py · examples/js_anthropic.ts |
Claude Agent SDK / claude -p (headless) |
distil wrap -- <cmd> or ANTHROPIC_BASE_URL |
examples/python_claude_agent_sdk.py |
| OpenAI SDK (Chat + Responses) | base_url="http://127.0.0.1:8788/v1" |
examples/python_openai.py |
| Vercel AI SDK | createAnthropic({ baseURL: '…:8788' }) — or in-process: wrapLanguageModel({ model, middleware: distilMiddleware() }) |
examples/js_vercel_ai_sdk.ts |
| LangChain (py/js) · LangGraph | anthropicApiUrl / base URL · pre_model_hook |
examples/js_langchain.ts |
| LiteLLM | api_base="http://127.0.0.1:8788" |
examples/python_litellm.py |
| Google Gemini | --upstream https://generativelanguage.googleapis.com |
examples/python_gemini.py |
Codex · aider · Cursor-agent · any base_url client |
distil wrap -- <agent> or OPENAI_BASE_URL |
— |
Anything that speaks the Anthropic / OpenAI / Gemini wire format works — the proxy is framework-agnostic, so CrewAI, AutoGen, Agno, Strands, Bedrock, etc. route through it unchanged by pointing their client's base URL at distil.
Prefer in-process? Wrap the client directly — still no call-site change:
from distil.adapters.anthropic import wrap
client = wrap(anthropic.Anthropic()) # compresses the request, keeps the cache warm
(OpenAI — Chat Completions and Responses API — and Gemini route through the proxy: distil wrap -- codex, or point OPENAI_BASE_URL at it. An in-process client wrap exists for the Anthropic SDK only.)
Framework hooks (no proxy, no network hop) — for agent frameworks that own the message list, compress it where it lives:
| Framework | Hook | Example |
|---|---|---|
| LiteLLM | distil.integrations.litellm.compress(kwargs) |
examples/python_litellm.py |
| LangChain | distil.integrations.langchain.compress_messages(msgs) |
— |
| LangGraph | pre_model_hook=pre_model_hook() (compresses graph state before the model node) |
examples/python_langgraph.py |
| Agno | distil.integrations.agno.compressed_model(model) |
— |
| Strands | distil.integrations.strands.compressing_hook() |
— |
LangChain / LangGraph — langchain-distil
Listed in LangChain's own community middleware integrations. If you came from there, this is the package:
pip install langchain-distil
from langchain_distil import compress_messages, pre_model_hook, as_runnable
msgs = compress_messages(msgs) # compress a message list in place of the call
graph = create_react_agent(..., pre_model_hook=pre_model_hook()) # LangGraph: before the model node
chain = as_runnable() | llm # or drop it into a chain (lazy langchain-core import)
Tool and function messages get the reversible Tier-1 digest, human and system messages are Tier-0 lossless, and assistant messages are never rewritten — a model's own words are not distil's to edit. Every digest is byte-exact recoverable. Pass verbatim=True for Tier-0-only when no recovery tool is available.
It is a thin wrapper over the hooks in the table above, so it inherits the same certified compression path — nothing is re-implemented. distil-llm is a dependency; you do not install both by hand.
🎟️ Subscription — save the window, not the bill
On a flat-rate Pro/Max plan there is no per-token bill to cut, so distil's dollar figures are notional. The rate-limit window is not notional: tokens spent on a 40 KB test log are quota unavailable for the next task.
The proxy can't help much here. Anthropic's consumer terms (§3, item 7) restrict automated access on
subscription credentials, so distil deliberately runs --lossless-only there and measures 0.27%.
Your account isn't worth a few percent.
A PostToolUse hook is a different mechanism — a documented, first-party extension point. Claude
Code compresses its own tool output, in its own process, before the model reads it:
distil hook --install # writes ~/.claude/settings.json (idempotent, preserves your other hooks)
distil hook --selftest # verify the schema adapters — a live mismatch is SILENT
distil quota # the window it buys back
$ distil quota
Subscription quota (the currency a flat-rate plan actually spends):
five_hour [########............] 43.0% used resets 2026-08-16 15:49Z
seven_day [....................] 4.0% used resets 2026-08-23 07:59Z
Measured on a paired live A/B, both arms answering correctly: tool_result −38.6%,
cache_creation −67.4%, cost-weighted −68.3%, and decision-equivalence 5/5 across five
verifiable tasks. Critically cache_read did not collapse — a hook sees each result once and cannot
rewrite history, so compression is append-only by construction and the prompt cache survives.
Where it saves nothing. Tier-0 is JSON minification plus consecutive-run collapse, so savings are
shape-dependent: verbose JSON (npm/pip/kubectl/terraform) 28–33%, duplicated log runs up to
99%, and unique-line logs, prose, git log and git diff 0%. On distil's own eval corpus it
saves 0.00% — that corpus has no JSON and no consecutive duplicates. Published because quoting
only the favourable fixtures would be the overclaim we criticise in others.
Other agents: Gemini CLI's
AfterToolcan influence output indirectly (under evaluation); Codex CLI hooks are observe-only and reject output rewriting, so it's blocked upstream there.
Full page, with the method and the caveats →
🧠 MCP server — give your agent a recall tool
Distil ships a Model Context Protocol server so an agent can compress its own tool output and get the exact bytes back later. Zero dependencies (stdlib JSON-RPC over stdio, no SDK), fully local — content never leaves the machine.
Add it in one line:
claude mcp add distil -- distil mcp
{
"mcpServers": {
"distil": { "command": "distil", "args": ["mcp"] }
}
}
Haven't installed distil? Run it straight from PyPI — no install step:
{
"mcpServers": {
"distil": { "command": "uvx", "args": ["--from", "distil-llm", "distil", "mcp"] }
}
}
Config lives in ~/Library/Application Support/Claude/claude_desktop_config.json (Claude Desktop,
macOS), .cursor/mcp.json (Cursor), or .vscode/mcp.json (VS Code). Restart the client after editing.
Verify it's up — no c
No comments yet
Be the first to share your take.