Repository iconcurepo.dev
headroom preview

chopratejas / headroom

agentaianthropicclaude-code

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.

74.9k Stars
visibility217 Watchers
fork_right5.8k Forks
Python
historyUpdated recently

description README.md

Headroom — the context compression layer for AI agents. A 55,957 token agent prompt compresses to the 24,340 tokens actually sent to the model, and the FATAL line at item 67 survives byte for byte.

headroomlabs-ai/headroom | Trendshift — #1 Repository Of The Day

GitHub stars CI PyPI npm Model Docs License

Quickstart · Headroom for Teams · Install · Proof · Agents · Docs · Discord · llms.txt

AI agents / LLMs: read /llms.txt here, or fetch the live index · full docs blob.

Headroom compresses everything your AI agent reads — tool outputs, logs, RAG chunks, files, and conversation history — before it reaches the LLM. Same answers, fraction of the tokens. Compression runs on your machine; no prompt or file content is sent anywhere to be compressed.

Headroom Enterprise for teams: 3 months free with a 12-month plan. Claim the perk.

Headroom compressing a 10,144 token log dump to 1,260 tokens while preserving the FATAL line
10,144 → 1,260 tokens. The same FATAL found.

What it does

  • Library — compress(messages) in Python or TypeScript, inline in any app.
  • Proxy — headroom proxy --port 8787, zero code changes, any language.
  • Agent wrap — headroom wrap claude|codex|grok|copilot|cursor|aider|opencode|cline|continue|goose|openhands|openclaw|vibe|omp|zcode in one command; undo with headroom unwrap <tool>.
  • MCP server — headroom_compress, headroom_retrieve, headroom_stats for any MCP client.
  • Cross-agent memory — one shared store across Claude, Codex, Gemini and Grok, with automatic dedup.
  • headroom learn — mines failed sessions and writes corrections to CLAUDE.local.md (default, gitignored), CLAUDE.md, AGENTS.md, GEMINI.md or GROK.md.
  • Output token reduction — trims what the model writes back, not only what you send. See below.
  • Reversible (CCR) — originals are cached locally and retrieved on demand.

How it works

 Your agent / app
   (Claude Code, Cursor, Codex, LangChain, Agno, Strands, your own code…)
        │   prompts · tool outputs · logs · RAG results · files
        ▼
    ┌────────────────────────────────────────────────────┐
    │  Headroom   (runs locally — your data stays here)  │
    │  ────────────────────────────────────────────────  │
    │  CacheAligner  →  ContentRouter  →  CCR            │
    │                    ├─ SmartCrusher   (JSON)        │
    │                    ├─ CodeCompressor (AST)         │
    │                    └─ Kompress-v2-base (text, HF)  │
    │                                                    │
    │  Cross-agent memory  ·  headroom learn  ·  MCP     │
    └────────────────────────────────────────────────────┘
        │   compressed prompt  +  retrieval tool
        ▼
 LLM provider  (Anthropic · OpenAI · Bedrock · …)
  • ContentRouter detects the content type and selects a compressor for it.
  • SmartCrusher / CodeCompressor / Kompress-v2-base handle JSON, source code and prose respectively.
  • CacheAligner flags volatile content that would bust a provider KV-cache prefix. It never rewrites prompts.
  • CCR stores originals locally so the model can call headroom_retrieve when it needs the full text.

→ Architecture · CCR · Kompress-v2-base model card

Get started (60 seconds)

# 1 — Install
uv tool install --python 3.13 "headroom-ai[all]"  # CLI in a self-contained env
pip install "headroom-ai[all]"                    # Python — ships the `headroom` CLI
npm install headroom-ai                           # TypeScript SDK only — no CLI

# 2 — Pick a mode
headroom deploy                         # turnkey local deployment + agent config
headroom wrap claude                    # wrap a coding agent
headroom proxy --port 8787              # drop-in proxy, zero code changes
# or: from headroom import compress     # inline library

# 3 — Check it and watch the savings
headroom doctor                         # health check — confirms routing works
headroom perf
headroom dashboard                      # live savings (proxy must be running)

Inline, in Python:

from headroom import compress
from openai import OpenAI

messages = [{"role": "user", "content": "Analyze these results"}]
result = compress(messages, model="gpt-4o")

client = OpenAI()
response = client.chat.completions.create(model="gpt-4o", messages=result.messages)
print(f"Saved {result.tokens_saved} tokens ({result.compression_ratio:.0%})")

Launch a wrapped agent session each time, so the setup runs. headroom wrap starts a local proxy, installs Serena for semantic code navigation, and launches the agent configured to route through Headroom. For Claude Code, Serena is registered for the wrapped project only (as a local-scope MCP server in ~/.claude.json). Use --code-memory-scope user to make it available in every project, or --code-memory none to skip it. headroom unwrap removes either registration.

The headroom CLI ships only in the PyPI package. The npm headroom-ai package is the TypeScript SDK — a library you import (import { compress } from 'headroom-ai') — and provides no headroom command.

Proof

Four scenarios built from real MCP server output formats, measured with the provider tokenizer and the shipped compress(). Seeded and offline, so you get the same numbers we did:

uv run python benchmarks/index_proof_table.py --seed 20260902
ScenarioBeforeAfterSaved
Code search (100 results)17,19913,59721%
SRE incident debugging5