qwen38-flash-next

A dense decoder-only LLM (Qwen3.8-Flash-Next, custom qwen4_exp architecture with a multi-token-prediction head) served on 4 Blackhole chips by its own OpenAI-compatible HTTP server: POST /v1/chat/completions (streaming SSE or JSON), GET /v1/models, GET /health. The server implements speculative multi-token decoding (MTP depth 4), 4-lane aggregate serving, chunked prefill with slab packing, and fail-loud boot admission (it refuses to serve unless the runtime contract and the acceptance corpus line up).

Runs on p300x2 (mesh (1, 4)) β€” 131,072-token context.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

At a glance

Architecture dense decoder-only, MTP head, custom qwen4_exp arch
Hardware p300x2
Context 131,072 tokens
Status Experimental community bring-up

Intended use

Direct use: Single-user or small-batch agent/chat serving on a 4-chip Blackhole box, speaking the OpenAI chat-completions API (streaming supported).

Out-of-scope use: Not validated on other meshes, other boards, or other context lengths; not a vLLM server (no /v1/completions, no tool-parser plumbing beyond what the model's chat template carries).

Quickstart

uv tool install tenstorrent   # once β€” the Tenstorrent CLI, `tt`
tt model pull DeAIafdeling/qwen38-flash-next
tt serve DeAIafdeling/qwen38-flash-next

tt model pull (or tt-model pull --with-weights) downloads the Docker image and the Qwen/Qwen3.8-Flash-Next weights at de4b8e4d43b917e7706784d8bb445c9af86a3540 (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs "event": "ready".

Without tt-cli β€” tt-model alone does the whole job:

tt-model pull  DeAIafdeling/qwen38-flash-next --with-weights
tt-model serve DeAIafdeling/qwen38-flash-next

Using it

This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API β€” its request and response shapes are the model's own. See the author's notes above for the payload it expects.

POST /v1/chat/completions with {"messages": [...], "stream": true} β€” streaming is SSE. GET /health reports load and the in-flight request. The server speaks the OpenAI chat-completions shape; point any OpenAI client at the endpoint.

Expected performance

Published decode 74.4 tok/s single-stream (MTP depth 5, 64k-context stack, mean 74.2 +/-0.5% over 4 runs, 30 Sep 2026); the 131072-context lanes stack in this package measures 70.7 tok/s aggregate at lanes=4 and ~21 tok/s per stream on a 4946-token agent session; prefill 1995 tok/s real-work (3.3k-token class), 604.5 tok/s canonical cold, TTFT ~41 ms. All measured on tt-quietbox-2 (2x p300c, 4 chips).

Limitations

Validated ONLY on a 4-chip p300c QB2 (1x4 line mesh) at 131072 allocated context. Lanes mode refuses sampled requests, MTP GDN anchors and per-request draft counts by design (fail-closed). First boot JIT-compiles kernels (tens of minutes); the cache bind-mount makes later boots fast. The k=6 draft depth is a measured NO-GO (stay at 4-5). The 74.4 published record is the 64k single-stream stack, not this package's default.

Risks and safety considerations

The server is fail-loud by design: it refuses to boot if the profiler envs leak in, if the worktree is dirty, or if the acceptance corpus is absent β€” treat a refused boot as the contract working. A SIGKILLed container can leave the mesh dirty; use tt-model stop (SIGTERM), and reset the chips if a later boot wedges at the first large host->device DMA.

Licensing

Weights: Qwen/Qwen3.8-Flash-Next under its upstream licence. The TTNN port and serving code in code/ derive from tt-metal (Apache-2.0).

Feedback

Questions or problems with this package: open a discussion at https://huggingface.co/DeAIafdeling/qwen38-flash-next/discussions β€” that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.

Provenance

The exact sources the image was built from β€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal a local checkout β€” commit not published
code/ digest b0d335a0158a6f46 (sha256, first 16 hex digits)
built 2026-10-01T19:50:10+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for DeAIafdeling/qwen38-flash-next

Finetuned
(69)
this model