qwen3-5-122b-a10b-blackholex4

Qwen3.5-122B-A10B (256-expert MoE, 48 layers of Gated DeltaNet + gated attention, vision tower) served on four Blackhole dies through vLLM and the Tenstorrent plugin: text and image chat requests (up to 4 images per request), tensor-parallel mixers, experts sharded on the intermediate dimension, on-device sampling, single-user multi-token prediction (MTP) decode, per-width decode traces, 262144-token context, up to 32 concurrent sequences.

Runs on p300x2 or p300x2 or p300x2 or p300x2 or p150x4 or p150x4 or p150x4 or p150x4 โ€” see the serve profiles below.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

tt-model pull  raahemnabeel/qwen3-5-122b-a10b-blackholex4 --with-weights
tt-model serve raahemnabeel/qwen3-5-122b-a10b-blackholex4

pull --with-weights downloads the Docker image and the Qwen/Qwen3.5-122B-A10B weights at dc4d348443bc740c68e2d77492492c11606384d5 (into your HF cache; they are not in the image). serve starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Use with your client

Point an OpenAI-compatible client at http://127.0.0.1:20000 (or the port tt-model serve printed) with model Qwen/Qwen3.5-122B-A10B. Thinking is ON by default: the reasoning arrives in the message's reasoning field (vLLM 0.26; older clients know it as reasoning_content) and the answer in content. Give thinking requests a large max_tokens (thousands): when the budget ends inside the think block, content is empty. Turn thinking off per request with "chat_template_kwargs": {"enable_thinking": false} for short, direct answers (the configuration the release gates and the published latency numbers were measured in). Without sampling parameters a request uses the checkpoint's defaults (temperature 0.6, top_p 0.95, top_k 20); for thinking on constraint-heavy prompts the model card's general preset (temperature 1.0, top_p 0.95, top_k 20, presence_penalty 1.5) avoids repetition. Tool calls use the qwen3_xml parser; images (up to 4 per request, resized to at most 4 MP) as image_url parts.

Profiles - which one for which use case

tt-model serve raahemnabeel/qwen3-5-122b-a10b-blackholex4 --profile <name> (p300x2 names below; the same four exist as p150x4, p150x4-burst, p150x4-longctx, p150x4-text for four p150 boards):

  • p300x2 (default) - interactive assistants and agents: one or a few users at a time who want the fastest answer. Speculative decoding gives a single user about 50 tok/s; images and tool calls work; thinking on.
  • p300x2-burst - many short prompts arriving together (chat front-ends with dozens of users, batch classification, extraction jobs): up to four prompts are prefilled per pass, so a 32-request burst gets its first token in 3.8 s instead of 6.3 s with the same decode speed. Choose it when throughput matters more than bit-exact reproducibility of a request's first token across different batch compositions (seeded runs may differ by what arrived together).
  • p300x2-longctx - document and codebase work where very long prompts (tens of thousands of tokens) mix with short ones: prefill runs in 2048-token chunks shared between prompts, so a short question is not stuck behind a 100K-token upload. Speculative decoding is off in this profile (about 30 tok/s for a lone user), and running users still pause while prefill chunks run.
  • p300x2-text - text-only deployments: the vision tower is not loaded, giving a 25 % larger KV pool (1.8 M pooled tokens for 32 users) for longer conversations at high occupancy; image requests are rejected at the front end.

Serve profiles

One image serves every profile below; pick one with --profile.

profile hardware mesh max_num_seqs max_model_len
p300x2 (default) p300x2 P300x2 32 262144
p300x2-burst p300x2 P300x2 32 262144
p300x2-longctx p300x2 P300x2 32 262144
p300x2-text p300x2 P300x2 32 262144
p150x4 p150x4 P150x4 32 262144
p150x4-burst p150x4 P150x4 32 262144
p150x4-longctx p150x4 P150x4 32 262144
p150x4-text p150x4 P150x4 32 262144

Provenance

The exact sources the image was built from โ€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal a local checkout โ€” commit not published (dirty tree โ€” the image includes uncommitted changes)
vLLM v0.26.0
vllm-tt-plugin a local checkout โ€” commit not published
code/ digest 2263b531006b1286 (sha256, first 16 hex digits)
built 2026-10-02T02:16:54+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support