muse-glimmer-30b

Runs on p300x2 (mesh P300x2) โ€” 131,072-token context, up to 32 concurrent sequences.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

tt-model pull  tt-hous/muse-glimmer-30b --with-weights
tt-model serve tt-hous/muse-glimmer-30b

pull --with-weights downloads the Docker image and the meta-models/Muse-Glimmer-30B weights at f84ecc3a0ea984a4c04542a84269e3d065350a6e (into your HF cache; they are not in the image). serve starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Muse-Glimmer-30B (~29.6 B dense, text-only) is served as an OpenAI-compatible endpoint for agentic coding: long-context (131k) tool-calling work driven by a coding agent.

On this model the first start takes about 4 minutes (weight loading + kernel compilation). Verify it is running correctly with tool calling:

curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "meta-models/Muse-Glimmer-30B",
  "messages": [{"role": "user", "content": "What is the weather in Paris right now, in Celsius?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Get the current weather for a city.",
      "parameters": {
        "type": "object",
        "properties": {
          "city":   {"type": "string", "description": "City name"},
          "metric": {"type": "boolean", "description": "true for Celsius"}
        },
        "required": ["city"]
      }
    }
  }],
  "tool_choice": "auto",
  "max_tokens": 256,
  "temperature": 0
}' | python3 -c 'import sys, json; c = json.load(sys.stdin)["choices"][0]; print(c["finish_reason"], json.dumps(c["message"]["tool_calls"], indent=2))'

A correct serve prints tool_calls followed by a structured get_weather call with JSON arguments (e.g. {"city": "Paris", "metric": true}). If the call comes back as prose in message.content with finish_reason stop, the tool-call parser is not active in the launch.

Performance

Release latency sweep on P300x2: one request at a time (batch 1), 512 output tokens, input length swept to the full context. Decode rate is per user. Retried points show the median of three independent runs.

input tokens output tokens TTFT TPOT end-to-end tokens/s/user
128 512 69.5 ms 23.60 ms 12.1 s 42.38
1,024 512 144.6 ms 24.99 ms 12.9 s 40.02
4,096 512 454.5 ms 26.64 ms 14.1 s 37.54
8,192 512 912.8 ms 27.86 ms 15.1 s 35.90
16,384 512 2.08 s 30.26 ms 17.5 s 33.05
32,768 512 4.48 s 35.29 ms 22.5 s 28.34
65,536 512 10.17 s 45.10 ms 33.2 s 22.17
130,560 512 25.12 s 64.76 ms 58.2 s 15.44

The last row saturates the advertised context (130,560 + 512 = 131,072). Serve one request at a time: concurrency at long context is admission-limited by the KV cache. The release passed the bounded latency gate: 2% per metric, plus a 5 ms absolute TTFT allowance for short-input measurement variance.

Prefix caching

Measured with vLLM's own prefix_repetition benchmark, the standard dataset for this feature. Eight distinct 4,096-token prefixes, each reused across eight requests, 64 requests at concurrency 1 -- the same package served twice, one flag apart.

vllm bench serve --model meta-models/Muse-Glimmer-30B \
  --dataset-name prefix_repetition \
  --prefix-repetition-prefix-len 4096 --prefix-repetition-suffix-len 128 \
  --prefix-repetition-num-prefixes 8 --prefix-repetition-output-len 32 \
  --num-prompts 64 --max-concurrency 1 --ignore-eos --seed 1234
metric caching off caching on
mean TTFT 497.00 ms 168.04 ms
median TTFT 495.82 ms 98.18 ms
p99 TTFT 525.84 ms 952.04 ms
benchmark duration 80.14 s 59.14 s
output throughput 25.56 tok/s 34.63 tok/s
prefix cache hit rate 0 54.7%

The distribution is the evidence, not the mean. With caching off every request pays the full 4,096-token prefill and TTFT is flat at 496/497/526. With it on the distribution splits: 98 ms median for the 56 requests that hit, 952 ms at p99 for the 8 cold prefixes.

Note the p99 moves the wrong way, 526 ms to 952 ms. A cold prefix now compiles its own SDPA program for its resume offset, so the first request at any previously unseen offset is slower than it was. Median improves 5x; the tail regresses 1.8x. Workloads that reuse a small set of prefixes gain; workloads whose offsets keep changing may not.

Evaluations

Run through tt-inference-server's eval workflow -- its lm-eval command, venv and scoring -- against this package at 131,072 context. Sampling is this card's recipe: temperature 1.0, top_p 0.95, top_k 64.

task samples metric score reference
gpqa_diamond_cot_zeroshot 198 exact_match, flexible-extract 77.27 72.8
ifeval 541 prompt_level_strict_acc 88.72 77.0
aime25 30 exact_match not valid, see below 94.7

The GPQA reference is the GPU reference score for openai/gpt-oss-20b, the closest configured analogue. This model had no GPQA row before this run, while every other reasoning model of its size in the catalogue has one. The ifeval reference is an IFBench floor, not an equivalence target.

aime25 is not reported as a score. 18 of its 30 responses contained no extractable answer and one ran to 294,912 characters, while the same problems put to the server directly return correct boxed answers -- so the figure measures the eval path, not the model. 30 of the 198 GPQA responses show the same pathology, which makes 77.27 a lower bound rather than a point estimate.

Provenance

The exact sources the image was built from โ€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal a local checkout โ€” commit not published
vLLM v0.24.0
vllm-tt-plugin a local checkout โ€” commit not published
code/ digest 1b2d3eba0ea623e8 (sha256, first 16 hex digits)
built 2026-09-23T12:35:51+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support