qwen3-5-122b-a10b-blackholex4
Qwen3.5-122B-A10B (256-expert MoE, 48 layers of Gated DeltaNet + gated attention, vision tower) served on four Blackhole dies through vLLM and the Tenstorrent plugin: text and image chat requests (up to 4 images per request), tensor-parallel mixers, experts sharded on the intermediate dimension, on-device sampling, single-user multi-token prediction (MTP) decode, per-width decode traces, 262144-token context, up to 32 concurrent sequences.
Runs on p300x2 or p300x2 or p300x2 or p300x2 or p150x4 or p150x4 or p150x4 or p150x4 โ see the serve profiles below.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart
tt-model pull raahemnabeel/qwen3-5-122b-a10b-blackholex4 --with-weights
tt-model serve raahemnabeel/qwen3-5-122b-a10b-blackholex4
pull --with-weights downloads the Docker image and the Qwen/Qwen3.5-122B-A10B weights at dc4d348443bc740c68e2d77492492c11606384d5 (into your HF cache; they are not in the image). serve starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.
Use with your client
Point an OpenAI-compatible client at http://127.0.0.1:20000 (or the port tt-model serve printed) with model
Qwen/Qwen3.5-122B-A10B. Thinking is ON by default: the reasoning arrives in the message's reasoning field (vLLM
0.26; older clients know it as reasoning_content) and the answer in content. Give thinking requests a large
max_tokens (thousands): when the budget ends inside the think block, content is empty. Turn thinking off per
request with "chat_template_kwargs": {"enable_thinking": false} for short, direct answers (the configuration the
release gates and the published latency numbers were measured in). Without sampling parameters a request uses the
checkpoint's defaults (temperature 0.6, top_p 0.95, top_k 20); for thinking on constraint-heavy prompts the model
card's general preset (temperature 1.0, top_p 0.95, top_k 20, presence_penalty 1.5) avoids repetition.
Tool calls use the qwen3_xml parser; images (up to 4 per request, resized to at most 4 MP) as image_url parts.
Profiles - which one for which use case
tt-model serve raahemnabeel/qwen3-5-122b-a10b-blackholex4 --profile <name> (p300x2 names below; the same four exist
as p150x4, p150x4-burst, p150x4-longctx, p150x4-text for four p150 boards):
p300x2(default) - interactive assistants and agents: one or a few users at a time who want the fastest answer. Speculative decoding gives a single user about 50 tok/s; images and tool calls work; thinking on.p300x2-burst- many short prompts arriving together (chat front-ends with dozens of users, batch classification, extraction jobs): up to four prompts are prefilled per pass, so a 32-request burst gets its first token in 3.8 s instead of 6.3 s with the same decode speed. Choose it when throughput matters more than bit-exact reproducibility of a request's first token across different batch compositions (seeded runs may differ by what arrived together).p300x2-longctx- document and codebase work where very long prompts (tens of thousands of tokens) mix with short ones: prefill runs in 2048-token chunks shared between prompts, so a short question is not stuck behind a 100K-token upload. Speculative decoding is off in this profile (about 30 tok/s for a lone user), and running users still pause while prefill chunks run.p300x2-text- text-only deployments: the vision tower is not loaded, giving a 25 % larger KV pool (1.8 M pooled tokens for 32 users) for longer conversations at high occupancy; image requests are rejected at the front end.
Serve profiles
One image serves every profile below; pick one with --profile.
| profile | hardware | mesh | max_num_seqs | max_model_len |
|---|---|---|---|---|
p300x2 (default) |
p300x2 | P300x2 | 32 | 262144 |
p300x2-burst |
p300x2 | P300x2 | 32 | 262144 |
p300x2-longctx |
p300x2 | P300x2 | 32 | 262144 |
p300x2-text |
p300x2 | P300x2 | 32 | 262144 |
p150x4 |
p150x4 | P150x4 | 32 | 262144 |
p150x4-burst |
p150x4 | P150x4 | 32 | 262144 |
p150x4-longctx |
p150x4 | P150x4 | 32 | 262144 |
p150x4-text |
p150x4 | P150x4 | 32 | 262144 |
Provenance
The exact sources the image was built from โ code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | a local checkout โ commit not published (dirty tree โ the image includes uncommitted changes) |
| vLLM | v0.26.0 |
| vllm-tt-plugin | a local checkout โ commit not published |
code/ digest |
2263b531006b1286 (sha256, first 16 hex digits) |
| built | 2026-10-02T02:16:54+00:00 by tt-model 0.1.0 |