File size: 10,665 Bytes
5e01d16
 
 
9748538
5e01d16
5f0ee81
5e01d16
 
9f1a442
 
 
5e01d16
 
 
 
c530129
a2a9fa1
9748538
5e01d16
9748538
 
9f1a442
 
 
 
 
 
 
9748538
5e01d16
9f1a442
 
 
 
 
 
 
 
 
 
5e01d16
9748538
5e01d16
 
 
 
 
 
 
9748538
 
5e01d16
9748538
5e01d16
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9f1a442
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5e01d16
9748538
 
 
5e01d16
9748538
5e01d16
9748538
 
 
 
 
 
 
 
5e01d16
 
 
9748538
 
5e01d16
9f1a442
b055325
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9f1a442
b055325
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9f1a442
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5e01d16
 
9748538
5e01d16
9748538
5e01d16
9748538
 
 
9f1a442
 
5e01d16
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
---
tags:
- blackhole
- p300x2
- tt-model-cache
- tt-model-catalog
- tt-model-container
- vllm-plugin
pipeline_tag: text-generation
base_model:
- meta-models/Muse-Glimmer-30B
---

# muse-glimmer-30b

> Derived from [meta-models/Muse-Glimmer-30B](https://huggingface.co/meta-models/Muse-Glimmer-30B). Weights revision [f84ecc3a](https://huggingface.co/meta-models/Muse-Glimmer-30B/tree/f84ecc3a0ea984a4c04542a84269e3d065350a6e). This represents the model implementation on Tenstorrent hardware. See the original model card for license, training, and evaluation details.

Runs on **p300x2** (mesh `P300x2`) — 131,072-token context, up to 32 concurrent sequences.

Packaged and published with [tt-model-manager](https://github.com/tenstorrent/tt-model-manager) 0.1.0 (manifest schema 5.1).

## At a glance

| | |
| --- | --- |
| Hardware | p300x2 |
| Context | 131,072 tokens |

## Quickstart

```bash
uv tool install tenstorrent   # once — the Tenstorrent CLI, `tt`
tt model pull tt-hous/muse-glimmer-30b
tt serve tt-hous/muse-glimmer-30b
```

`tt model pull` (or `tt-model pull --with-weights`) downloads the Docker image and the [`meta-models/Muse-Glimmer-30B`](https://huggingface.co/meta-models/Muse-Glimmer-30B) weights at `f84ecc3a0ea984a4c04542a84269e3d065350a6e` (into your HF cache; they are not in the image). `tt serve` (or `tt-model serve`) starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs `Application startup complete`.

Without tt-cli — tt-model alone does the whole job:

```bash
tt-model pull  tt-hous/muse-glimmer-30b --with-weights
tt-model serve tt-hous/muse-glimmer-30b
```

Muse-Glimmer-30B (~29.6 B dense, text-only) is served as an
OpenAI-compatible endpoint for **agentic coding**: long-context (131k)
tool-calling work driven by a coding agent.

On this model the first start takes about 4 minutes (weight loading + kernel
compilation). Verify it is running correctly with tool calling:
```bash
curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "meta-models/Muse-Glimmer-30B",
  "messages": [{"role": "user", "content": "What is the weather in Paris right now, in Celsius?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Get the current weather for a city.",
      "parameters": {
        "type": "object",
        "properties": {
          "city":   {"type": "string", "description": "City name"},
          "metric": {"type": "boolean", "description": "true for Celsius"}
        },
        "required": ["city"]
      }
    }
  }],
  "tool_choice": "auto",
  "max_tokens": 256,
  "temperature": 0
}' | python3 -c 'import sys, json; c = json.load(sys.stdin)["choices"][0]; print(c["finish_reason"], json.dumps(c["message"]["tool_calls"], indent=2))'
```
A correct serve prints `tool_calls` followed by a structured `get_weather` call
with JSON arguments (e.g. `{"city": "Paris", "metric": true}`). If the call comes
back as prose in `message.content` with `finish_reason` `stop`, the tool-call
parser is not active in the launch.

Then verify plain chat, streamed. The model always thinks first, so this is
the path that shows whether the reasoning parser is splitting the channels:
```bash
curl -sN localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "meta-models/Muse-Glimmer-30B",
  "messages": [{"role": "user", "content": "What is 17 * 23?"}],
  "max_tokens": 512, "temperature": 0, "stream": true
}' | grep -c '"reasoning"'
```
A correct serve prints a positive count: the analysis arrives as `reasoning`
deltas and only the answer arrives as `content`. Zero means the stream is raw
channel text (` to=self<|message|>...`), which is what a launch without the
two parser plugins produces. Give plain chat a real token budget: at the
template's default `Reasoning strength: high` a short factual question needs
roughly 500 to 700 completion tokens, and a turn cut off by `max_tokens`
inside the analysis returns that analysis as `reasoning` with empty `content`.

## Using it

The server speaks the OpenAI API at `http://127.0.0.1:20000/v1` (or whichever port your serve command reported; chat completions, completions, and `/v1/models`). Pass `"model": "meta-models/Muse-Glimmer-30B"` — the weights id, not this package's name.

Tool calling is enabled (`--tool-call-parser muse_glimmer`): a request that passes `tools` comes back with `tool_calls` and `finish_reason: tool_calls`.

Reasoning output is separated (`--reasoning_parser muse_glimmer`): the thinking text arrives in `reasoning_content`, apart from `content`.

## Expected performance

Release latency sweep on P300x2: one request at a time (batch 1), 512
output tokens, input length swept to the full context. Decode rate is per
user. Retried points show the median of three independent runs.

| input tokens | output tokens | TTFT | TPOT | end-to-end | tokens/s/user |
|---:|---:|---:|---:|---:|---:|
| 128 | 512 | 69.5 ms | 23.60 ms | 12.1 s | 42.38 |
| 1,024 | 512 | 144.6 ms | 24.99 ms | 12.9 s | 40.02 |
| 4,096 | 512 | 454.5 ms | 26.64 ms | 14.1 s | 37.54 |
| 8,192 | 512 | 912.8 ms | 27.86 ms | 15.1 s | 35.90 |
| 16,384 | 512 | 2.08 s | 30.26 ms | 17.5 s | 33.05 |
| 32,768 | 512 | 4.48 s | 35.29 ms | 22.5 s | 28.34 |
| 65,536 | 512 | 10.17 s | 45.10 ms | 33.2 s | 22.17 |
| 130,560 | 512 | 25.12 s | 64.76 ms | 58.2 s | 15.44 |

The last row saturates the advertised context (130,560 + 512 = 131,072).
Serve one request at a time: concurrency at long context is admission-limited
by the KV cache. The release passed the bounded latency gate: 2% per metric,
plus a 5 ms absolute TTFT allowance for short-input measurement variance.

### Prefix caching

Measured with vLLM's own `prefix_repetition` benchmark, the standard dataset
for this feature. Eight distinct 4,096-token prefixes, each reused across
eight requests, 64 requests at concurrency 1 -- the same package served twice,
one flag apart.

```
vllm bench serve --model meta-models/Muse-Glimmer-30B \
  --dataset-name prefix_repetition \
  --prefix-repetition-prefix-len 4096 --prefix-repetition-suffix-len 128 \
  --prefix-repetition-num-prefixes 8 --prefix-repetition-output-len 32 \
  --num-prompts 64 --max-concurrency 1 --ignore-eos --seed 1234
```

| metric | caching off | caching on |
|---|---:|---:|
| mean TTFT | 497.00 ms | 168.04 ms |
| median TTFT | 495.82 ms | 98.18 ms |
| p99 TTFT | 525.84 ms | 952.04 ms |
| benchmark duration | 80.14 s | 59.14 s |
| output throughput | 25.56 tok/s | 34.63 tok/s |
| prefix cache hit rate | 0 | 54.7% |

The distribution is the evidence, not the mean. With caching off every request
pays the full 4,096-token prefill and TTFT is flat at 496/497/526. With it on
the distribution splits: 98 ms median for the 56 requests that hit, 952 ms at
p99 for the 8 cold prefixes.

Note the p99 moves the wrong way, 526 ms to 952 ms. A cold prefix now compiles
its own SDPA program for its resume offset, so the first request at any
previously unseen offset is slower than it was. Median improves 5x; the tail
regresses 1.8x. Workloads that reuse a small set of prefixes gain; workloads
whose offsets keep changing may not.

### Evaluations

Run through tt-inference-server's eval workflow -- its lm-eval command, venv
and scoring -- against this package at 131,072 context. Sampling is this
card's recipe: temperature 1.0, top_p 0.95, top_k 64.

| task | samples | metric | score | reference |
|---|---:|---|---:|---:|
| `gpqa_diamond_cot_zeroshot` | 198 | exact_match, flexible-extract | 77.27 | 72.8 |
| `ifeval` | 541 | prompt_level_strict_acc | 88.72 | 77.0 |
| `aime25` | 30 | exact_match | not valid, see below | 94.7 |

The GPQA reference is the GPU reference score for `openai/gpt-oss-20b`, the
closest configured analogue. This model had no GPQA row before this run, while
every other reasoning model of its size in the catalogue has one. The `ifeval`
reference is an IFBench floor, not an equivalence target.

`aime25` is not reported as a score. 18 of its 30 responses contained no
extractable answer and one ran to 294,912 characters, while the same problems
put to the server directly return correct boxed answers -- so the figure
measures the eval path, not the model. 30 of the 198 GPQA responses show the
same pathology, which makes 77.27 a lower bound rather than a point estimate.

## Limitations

- Text-only. The checkpoint carries a perception encoder; this port serves text
  and does not accept images.
- Serve one request at a time at long context: admission is limited by the KV
  cache (see the latency sweep under Expected performance).
- The chat template's default `Reasoning strength: high` makes the model think
  before every reply. A short factual question needs roughly 500 to 700
  completion tokens; a `max_tokens` budget that ends inside the analysis returns
  the analysis as `reasoning` and an empty string as `content`. Put
  `Reasoning strength: low` in the system prompt, or send
  `chat_template_kwargs: {"reasoning_strength": "low"}`, when a short budget is
  required.
- A request that sends `tools` with `tool_choice: "none"` suppresses tool calls
  as the API requires, but if the model still writes a call, a non-streaming
  response carries that call's markup in `content` as text.
- Prefix caching improves median TTFT 5x but regresses p99 TTFT 1.8x on cold
  prefixes, as measured under Expected performance.
- `aime25` through the eval harness is not a valid score: 18 of 30 responses had
  no extractable answer and one ran to 294,912 characters, while the same
  problems put to the server directly return correct boxed answers.

## Feedback

Questions or problems with this package: open a discussion at https://huggingface.co/tt-hous/muse-glimmer-30b/discussions — that is what reaches its author. A problem with the `tt` tooling itself: `tt report issue` (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.

## Provenance

The exact sources the image was built from — `code/` in this repo is byte-identical to the model code inside the image:

| component | built from |
| --- | --- |
| tt-metal | a local checkout — commit not published |
| vLLM | [`v0.24.0`](https://github.com/vllm-project/vllm/releases/tag/v0.24.0) |
| vllm-tt-plugin | a local checkout — commit not published |
| `code/` digest | `7931f069dd3c0235` (sha256, first 16 hex digits) |
| built | 2026-10-05T16:40:10+00:00 by tt-model 0.1.0 |