Spaces:
Build error
Build error
File size: 3,466 Bytes
b458389 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 | # Headroom
**The Context Optimization Layer for LLM Applications**
Tool outputs are 70-95% redundant. Headroom compresses that away—without losing information.
---
## Quick Install
```bash
pip install headroom-ai[all]
```
## Quick Start
### Option 1: Proxy (Zero Code Changes)
Start the proxy:
```bash
headroom proxy
```
Point your tools at it:
```bash
ANTHROPIC_BASE_URL=http://localhost:8787 claude
```
That's it. Your existing code works unchanged, with 40-90% fewer tokens.
### Option 2: Python SDK
```python
from headroom import Headroom
hr = Headroom()
# Compress tool output before sending to LLM
compressed = hr.compress(large_tool_output)
# If LLM needs the full data, retrieve it
original = hr.retrieve(compressed)
```
---
## Why Headroom?
| Problem | Solution |
|---------|----------|
| Tool outputs bloat context with repetitive JSON | Statistical compression removes redundancy |
| Dynamic content breaks provider caching | Cache alignment stabilizes prefixes |
| Long conversations exceed context limits | Intelligent scoring drops low-value messages |
| Compressed data might be needed later | CCR stores originals for on-demand retrieval |
---
## Results
**100 log entries. One critical error buried at position 67.**
| Metric | Baseline | Headroom |
|--------|----------|----------|
| Input tokens | 10,144 | 1,260 |
| Correct answers | 4/4 | 4/4 |
**87.6% fewer tokens. Same answer.**
The FATAL error was automatically preserved—no configuration needed.
---
## How It Works
```
Your App → Headroom → LLM Provider
↓
Compression
Caching
Retrieval
```
1. **Intercepts context** — Tool outputs, logs, search results
2. **Compresses intelligently** — Keeps errors, outliers, boundaries
3. **Stores originals** — Full data available if LLM requests it
4. **Aligns for caching** — Provider caches actually hit
---
## Integrations
=== "LangChain"
```python
from langchain_openai import ChatOpenAI
from headroom.integrations import HeadroomChatModel
llm = HeadroomChatModel(ChatOpenAI(model="gpt-4o"))
response = llm.invoke("Hello!")
```
=== "Agno"
```python
from agno.agent import Agent
from agno.models.openai import OpenAIChat
from headroom.integrations.agno import HeadroomAgnoModel
model = HeadroomAgnoModel(OpenAIChat(id="gpt-4o"))
agent = Agent(model=model)
```
=== "AWS Bedrock"
```bash
# Start proxy with Bedrock backend
headroom proxy --backend bedrock --region us-east-1
# Point Claude Code at it
ANTHROPIC_API_KEY="sk-ant-dummy" \
ANTHROPIC_BASE_URL=http://localhost:8787 \
claude
```
---
## Features
**Compression**
- Statistical JSON array compression (no hardcoded rules)
- ML-based text compression via LLMLingua
- AST-aware code compression
- Image optimization (40-90% reduction)
**Context Management**
- Intelligent message scoring and dropping
- Compress-Cache-Retrieve (CCR) for lossless compression
- Provider cache alignment for better hit rates
**Operations**
- Prometheus metrics endpoint
- Request logging and cost tracking
- Budget limits and rate limiting
---
## Next Steps
- [Quickstart Guide](quickstart.md) — Get running in 5 minutes
- [Proxy Documentation](proxy.md) — Configure the optimization proxy
- [Architecture](ARCHITECTURE.md) — Deep dive into how it works
---
## License
Apache 2.0 — Free for commercial use.
|