Spaces:
Build error
Build error
File size: 8,030 Bytes
d724f14 485ea38 d724f14 485ea38 d724f14 485ea38 d724f14 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 | # CCR: Compress-Cache-Retrieve
Headroom's CCR architecture makes compression **reversible**. When content is compressed, the original data is cached. If the LLM needs more data, it can retrieve it instantly.
## The Problem with Traditional Compression
Traditional compression is lossy β if you guess wrong about what's important, data is lost forever. This creates a difficult tradeoff:
- **Aggressive compression**: Risk losing data the LLM needs
- **Conservative compression**: Miss out on token savings
CCR eliminates this tradeoff.
## CCR-Enabled Components
| Component | What it compresses | CCR integration |
|-----------|-------------------|-----------------|
| **SmartCrusher** | JSON arrays (tool outputs) | Stores original array, marker includes hash |
| **ContentRouter** | Code, logs, search results, text | Stores original content by strategy |
| **IntelligentContextManager** | Messages (conversation turns) | Stores dropped messages, marker includes hash |
## How CCR Works
```
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β TOOL OUTPUT (1000 items) β
β ββ SmartCrusher compresses to 20 items β
β ββ Original cached with hash=abc123 β
β ββ Retrieval tool injected into context β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β LLM PROCESSING β
β Option A: LLM solves task with 20 items β Done (90% savings) β
β Option B: LLM calls headroom_retrieve(hash=abc123) β
β β Response Handler executes retrieval automatically β
β β LLM receives full data, responds accurately β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
```
### Phase 1: Compression Store
When SmartCrusher compresses tool output:
1. Original content is stored in an LRU cache
2. A hash key is generated for retrieval
3. A marker is added to the compressed output: `[1000 items compressed to 20. Retrieve more: hash=abc123]`
### Phase 2: Tool Injection
Headroom injects a `headroom_retrieve` tool into the LLM's available tools:
```json
{
"name": "headroom_retrieve",
"description": "Retrieve original uncompressed data from Headroom cache",
"parameters": {
"hash": "The hash key from the compression marker",
"query": "Optional: search within the cached data"
}
}
```
### Phase 3: Response Handler
When the LLM calls `headroom_retrieve`:
1. Response Handler intercepts the tool call
2. Retrieves data from the local cache (~1ms)
3. Adds the result to the conversation
4. Continues the API call automatically
**The client never sees CCR tool calls** β they're handled transparently.
### Phase 4: Context Tracker
Across multiple turns, the Context Tracker:
1. Remembers what was compressed in earlier turns
2. Analyzes new queries for relevance to compressed content
3. Proactively expands relevant data before the LLM asks
**Example:**
```
Turn 1: User searches for files
β Tool returns 500 files
β SmartCrusher compresses to 15, caches original (hash=abc123)
β LLM sees 15 files, answers question
Turn 5: User asks "What about the auth middleware?"
β Context Tracker detects "auth" might be in abc123
β Proactively expands compressed content
β LLM sees full file list, finds auth_middleware.py
```
## Message-Level CCR (IntelligentContext)
IntelligentContextManager is a **message-level compressor**. When it drops low-importance messages to fit the context budget, those messages are stored in CCR:
```
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β LONG CONVERSATION (100 messages, 50K tokens) β
β ββ IntelligentContext scores messages by importance β
β ββ Drops 60 low-scoring messages β
β ββ Dropped messages cached with hash=def456 β
β ββ Marker inserted: "60 messages dropped, retrieve: def456" β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β LLM PROCESSING β
β Option A: LLM solves task with remaining messages β Done β
β Option B: LLM needs earlier context β
β β Calls headroom_retrieve(hash=def456) β
β β Full conversation restored β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
```
**The marker includes the CCR reference:**
```
[Earlier context compressed: 60 message(s) dropped by importance scoring.
Full content available via ccr_retrieve tool with reference 'def456'.]
```
**TOIN integration:** When users retrieve dropped messages, TOIN learns to score those message patterns higher next time, improving future drop decisions across all users.
## Features
| Feature | Description |
|---------|-------------|
| **Automatic Response Handling** | When LLM calls `headroom_retrieve`, the proxy handles it automatically |
| **Multi-Turn Context Tracking** | Tracks compressed content across turns, proactively expands when relevant |
| **BM25 Search** | LLM can search within compressed data: `headroom_retrieve(hash, query="errors")` |
| **Feedback Learning** | Learns from retrieval patterns to improve future compression |
## Configuration
```bash
# Proxy with CCR enabled (default)
headroom proxy --port 8787
# Disable CCR response handling
headroom proxy --no-ccr-responses
# Disable proactive expansion
headroom proxy --no-ccr-expansion
```
## Why This Matters
| Approach | Risk | Savings |
|----------|------|---------|
| No compression | None | 0% |
| Traditional compression | Data loss | 70-90% |
| CCR compression | None (reversible) | 70-90% |
CCR gives you the savings of aggressive compression with zero risk β the LLM can always retrieve the original data if needed.
## Demo
Run the CCR demonstration to see it in action:
```bash
python examples/ccr_demo.py
```
Output:
```
1. COMPRESSION STORE
Original: 100 items (7,059 chars)
Compressed: 8 items (633 chars)
Reduction: 91.0%
3. RESPONSE HANDLER
Detected CCR tool call: True
Retrieved 100 items automatically
4. CONTEXT TRACKER
Turn 5: User asks "show authentication middleware"
Tracker found 1 relevant context
β relevance=0.73
Proactively expanded: 100 items
```
## Architecture
For implementation details, see [ARCHITECTURE.md](ARCHITECTURE.md#ccr-compress-cache-retrieve).
|