File size: 3,466 Bytes
b458389
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
# Headroom

**The Context Optimization Layer for LLM Applications**

Tool outputs are 70-95% redundant. Headroom compresses that away—without losing information.

---

## Quick Install

```bash
pip install headroom-ai[all]
```

## Quick Start

### Option 1: Proxy (Zero Code Changes)

Start the proxy:

```bash
headroom proxy
```

Point your tools at it:

```bash
ANTHROPIC_BASE_URL=http://localhost:8787 claude
```

That's it. Your existing code works unchanged, with 40-90% fewer tokens.

### Option 2: Python SDK

```python
from headroom import Headroom

hr = Headroom()

# Compress tool output before sending to LLM
compressed = hr.compress(large_tool_output)

# If LLM needs the full data, retrieve it
original = hr.retrieve(compressed)
```

---

## Why Headroom?

| Problem | Solution |
|---------|----------|
| Tool outputs bloat context with repetitive JSON | Statistical compression removes redundancy |
| Dynamic content breaks provider caching | Cache alignment stabilizes prefixes |
| Long conversations exceed context limits | Intelligent scoring drops low-value messages |
| Compressed data might be needed later | CCR stores originals for on-demand retrieval |

---

## Results

**100 log entries. One critical error buried at position 67.**

| Metric | Baseline | Headroom |
|--------|----------|----------|
| Input tokens | 10,144 | 1,260 |
| Correct answers | 4/4 | 4/4 |

**87.6% fewer tokens. Same answer.**

The FATAL error was automatically preserved—no configuration needed.

---

## How It Works

```
Your App → Headroom → LLM Provider
              ↓
         Compression
         Caching
         Retrieval
```

1. **Intercepts context** — Tool outputs, logs, search results
2. **Compresses intelligently** — Keeps errors, outliers, boundaries
3. **Stores originals** — Full data available if LLM requests it
4. **Aligns for caching** — Provider caches actually hit

---

## Integrations

=== "LangChain"

    ```python
    from langchain_openai import ChatOpenAI
    from headroom.integrations import HeadroomChatModel

    llm = HeadroomChatModel(ChatOpenAI(model="gpt-4o"))
    response = llm.invoke("Hello!")
    ```

=== "Agno"

    ```python
    from agno.agent import Agent
    from agno.models.openai import OpenAIChat
    from headroom.integrations.agno import HeadroomAgnoModel

    model = HeadroomAgnoModel(OpenAIChat(id="gpt-4o"))
    agent = Agent(model=model)
    ```

=== "AWS Bedrock"

    ```bash
    # Start proxy with Bedrock backend
    headroom proxy --backend bedrock --region us-east-1

    # Point Claude Code at it
    ANTHROPIC_API_KEY="sk-ant-dummy" \
    ANTHROPIC_BASE_URL=http://localhost:8787 \
    claude
    ```

---

## Features

**Compression**

- Statistical JSON array compression (no hardcoded rules)
- ML-based text compression via LLMLingua
- AST-aware code compression
- Image optimization (40-90% reduction)

**Context Management**

- Intelligent message scoring and dropping
- Compress-Cache-Retrieve (CCR) for lossless compression
- Provider cache alignment for better hit rates

**Operations**

- Prometheus metrics endpoint
- Request logging and cost tracking
- Budget limits and rate limiting

---

## Next Steps

- [Quickstart Guide](quickstart.md) — Get running in 5 minutes
- [Proxy Documentation](proxy.md) — Configure the optimization proxy
- [Architecture](ARCHITECTURE.md) — Deep dive into how it works

---

## License

Apache 2.0 — Free for commercial use.