File size: 4,291 Bytes
c0fb63f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
# Strands Integration

Headroom integrates with [Strands Agents](https://github.com/strands-agents/sdk-python) to provide automatic context optimization. Two integration patterns: wrap the model, or hook into tool calls.

---

## Installation

```bash
pip install headroom-ai strands-agents
```

---

## Quick Start

```python
from strands import Agent
from strands.models.bedrock import BedrockModel
from headroom.integrations.strands import HeadroomStrandsModel

# Wrap your model
model = BedrockModel(model_id="us.anthropic.claude-sonnet-4-20250514-v1:0")
optimized = HeadroomStrandsModel(wrapped_model=model)

# Create agent as usual
agent = Agent(model=optimized)
response = agent("Investigate the production incident")

# Check savings
print(f"Tokens saved: {optimized.total_tokens_saved}")
```

Every API call the agent makes — including tool result round-trips — gets compressed automatically.

---

## Integration Patterns

### 1. Model Wrapping

Wraps the Strands `Model` interface. Every call to `stream()` compresses the messages before they hit the provider.

```python
from strands.models.bedrock import BedrockModel
from headroom.integrations.strands import HeadroomStrandsModel

model = BedrockModel(model_id="us.anthropic.claude-sonnet-4-20250514-v1:0")
optimized = HeadroomStrandsModel(wrapped_model=model)

# Streaming works identically
agent = Agent(model=optimized)
response = agent("Analyze these logs")
```

With custom config:

```python
from headroom import HeadroomConfig

config = HeadroomConfig()
optimized = HeadroomStrandsModel(wrapped_model=model, config=config)
```

### 2. Hook Provider (Tool Output Compression)

Compresses tool call results via Strands' hook system. Uses SmartCrusher on JSON arrays returned by tools.

```python
from strands import Agent
from strands.models.bedrock import BedrockModel
from headroom.integrations.strands import HeadroomHookProvider

model = BedrockModel(model_id="us.anthropic.claude-sonnet-4-20250514-v1:0")
hooks = HeadroomHookProvider(
    compress_tool_outputs=True,
    min_tokens_to_compress=200,
    preserve_errors=True,
)

agent = Agent(model=model, hooks=[hooks])
response = agent("Search the database for recent failures")

# Check tool compression savings
print(f"Tokens saved by hooks: {hooks.total_tokens_saved}")
```

The hook preserves:

- Error items (error indicators, exceptions)
- Anomalous values (statistical outliers)
- Items matching the user's query context
- First/last items for boundary context

### 3. Both Together

Model wrapping compresses conversation history. Hooks compress individual tool results. Use both for maximum savings.

```python
from headroom.integrations.strands import HeadroomStrandsModel, HeadroomHookProvider

optimized = HeadroomStrandsModel(wrapped_model=model)
hooks = HeadroomHookProvider(compress_tool_outputs=True)

agent = Agent(model=optimized, hooks=[hooks])
```

---

## Structured Output

HeadroomStrandsModel supports Strands' structured output feature:

```python
from pydantic import BaseModel

class Analysis(BaseModel):
    severity: str
    root_cause: str
    recommendation: str

result = optimized.structured_output(Analysis, messages)
```

---

## Metrics

```python
# Per-request metrics
for m in optimized.metrics_history:
    print(f"  {m.tokens_before} → {m.tokens_after} ({m.tokens_saved} saved)")

# Running total
print(f"Total saved: {optimized.total_tokens_saved}")
```

---

## How It Works

```
Agent decides to call tool
    │
    â–¼
Tool executes, returns result
    │
    â–¼
HeadroomHookProvider (optional)
    compresses tool result JSON
    │
    â–¼
Agent builds next API request
    │
    â–¼
HeadroomStrandsModel.stream()
    compresses full message list
    │
    â–¼
Provider API (Bedrock, etc.)
```

The model wrapper uses Headroom's full pipeline (CacheAligner → ContentRouter → IntelligentContext). The hook provider uses SmartCrusher directly for fast JSON compression of individual tool results.

---

## Supported Providers

HeadroomStrandsModel auto-detects the provider from the wrapped model:

| Strands Model | Provider Detected |
|--------------|-------------------|
| `BedrockModel` | Anthropic (via Bedrock) |
| `OllamaModel` | OpenAI-compatible |
| Custom `Model` | Falls back to estimation |