qwen38-flash-next
A dense decoder-only LLM (Qwen3.8-Flash-Next, custom qwen4_exp architecture with a multi-token-prediction head) served on 4 Blackhole chips by its own OpenAI-compatible HTTP server: POST /v1/chat/completions (streaming SSE or JSON), GET /v1/models, GET /health. The server implements speculative multi-token decoding (MTP depth 4), 4-lane aggregate serving, chunked prefill with slab packing, and fail-loud boot admission (it refuses to serve unless the runtime contract and the acceptance corpus line up).
Runs on p300x2 (mesh (1, 4)) β 131,072-token context.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
At a glance
| Architecture | dense decoder-only, MTP head, custom qwen4_exp arch |
| Hardware | p300x2 |
| Context | 131,072 tokens |
| Status | Experimental community bring-up |
Intended use
Direct use: Single-user or small-batch agent/chat serving on a 4-chip Blackhole box, speaking the OpenAI chat-completions API (streaming supported).
Out-of-scope use: Not validated on other meshes, other boards, or other context lengths; not a vLLM server (no /v1/completions, no tool-parser plumbing beyond what the model's chat template carries).
Quickstart
uv tool install tenstorrent # once β the Tenstorrent CLI, `tt`
tt model pull DeAIafdeling/qwen38-flash-next
tt serve DeAIafdeling/qwen38-flash-next
tt model pull (or tt-model pull --with-weights) downloads the Docker image and the Qwen/Qwen3.8-Flash-Next weights at de4b8e4d43b917e7706784d8bb445c9af86a3540 (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs "event": "ready".
Without tt-cli β tt-model alone does the whole job:
tt-model pull DeAIafdeling/qwen38-flash-next --with-weights
tt-model serve DeAIafdeling/qwen38-flash-next
Using it
This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API β its request and response shapes are the model's own. See the author's notes above for the payload it expects.
POST /v1/chat/completions with {"messages": [...], "stream": true} β streaming is SSE. GET /health reports load and the in-flight request. The server speaks the OpenAI chat-completions shape; point any OpenAI client at the endpoint.
Expected performance
Published decode 74.4 tok/s single-stream (MTP depth 5, 64k-context stack, mean 74.2 +/-0.5% over 4 runs, 30 Sep 2026); the 131072-context lanes stack in this package measures 70.7 tok/s aggregate at lanes=4 and ~21 tok/s per stream on a 4946-token agent session; prefill 1995 tok/s real-work (3.3k-token class), 604.5 tok/s canonical cold, TTFT ~41 ms. All measured on tt-quietbox-2 (2x p300c, 4 chips).
Limitations
Validated ONLY on a 4-chip p300c QB2 (1x4 line mesh) at 131072 allocated context. Lanes mode refuses sampled requests, MTP GDN anchors and per-request draft counts by design (fail-closed). First boot JIT-compiles kernels (tens of minutes); the cache bind-mount makes later boots fast. The k=6 draft depth is a measured NO-GO (stay at 4-5). The 74.4 published record is the 64k single-stream stack, not this package's default.
Risks and safety considerations
The server is fail-loud by design: it refuses to boot if the profiler envs leak in, if the worktree is dirty, or if the acceptance corpus is absent β treat a refused boot as the contract working. A SIGKILLed container can leave the mesh dirty; use tt-model stop (SIGTERM), and reset the chips if a later boot wedges at the first large host->device DMA.
Licensing
Weights: Qwen/Qwen3.8-Flash-Next under its upstream licence. The TTNN port and serving code in code/ derive from tt-metal (Apache-2.0).
Feedback
Questions or problems with this package: open a discussion at https://huggingface.co/DeAIafdeling/qwen38-flash-next/discussions β that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.
Provenance
The exact sources the image was built from β code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | a local checkout β commit not published |
code/ digest |
b0d335a0158a6f46 (sha256, first 16 hex digits) |
| built | 2026-10-01T19:50:10+00:00 by tt-model 0.1.0 |
Model tree for DeAIafdeling/qwen38-flash-next
Base model
Qwen/Qwen3.8-Flash-Next