Qwen3.5-9B — TT-native mixed BFP4/BFP8, single P150, MTP-1

An experimental hardware-specific checkpoint and custom runtime derived from Qwen/Qwen3.5-9B for one Tenstorrent Blackhole P150. Release prefill-mtp-20260907 updates the runtime, native package, and download/launch workflow. It does not change the weights or their quantization format. This is not the earlier BFP4-only release.

MTP-1 speculative decoding is enabled by default. The native checkpoint includes the target model and all 15 original MTP tensors, and occupies approximately 9.63 GB. Serving downloads only the native package and its pinned Docker runtime: no original Hugging Face weights, HF Python packages, or source bind mounts are required.

This is not GGUF, AWQ, GPTQ, bitsandbytes, a standard Transformers checkpoint, or an ordinary CUDA/vLLM model. Transformers from_pretrained() and hosted Inference Providers are not supported launch paths. Use the launcher below.

Download and serve

Prerequisites

  • Linux with a supported P150 driver and hugepage mounts at /dev/hugepages and /dev/hugepages-1G.
  • Exclusive ownership of one P150. Stop your other accelerator workloads before confirming ownership. The launcher does not stop containers, reset the device, or acquire the device for you.
  • Python 3.11 or newer, curl, and Docker usable by your current account. The host launcher uses only Python's standard library; do not install huggingface_hub, Transformers, or other HF packages to launch it.
  • A large local Linux filesystem for the download and writable device/tensor caches. Budget approximately 9.6 GB for the checkpoint alone, plus substantial additional space for the large Docker runtime image, extracted Docker layers, and generated device/tensor caches. Free space only slightly above 9.6 GB is insufficient. Docker's data directory may be on a different filesystem from --cache-root; allow capacity on both. Do not place the working cache on a Windows-mounted filesystem.

Download the single launcher, then run it:

curl --fail --location --output launch.py \
  https://huggingface.co/Lottolabs/Qwen3.5-9B-TT-Mixed-BFP4-BFP8-P150/resolve/main/launch.py

python3 launch.py \
  --cache-root "$HOME/.cache/qwen9b-tt-native" \
  --device-ownership-confirmed

Choose a different --cache-root on your large local disk if needed. No Hugging Face login or token is needed for this public native download. The launcher uses anonymous HTTPS requests, resolves the requested HF revision to an immutable commit before fetching package files, and checks each file's byte size and SHA-256 against runtime-release.json. It checks the native manifest against that inventory, reuses verified cached files, and invokes serve_native.py with the release's digest-pinned image. It does not download the original BF16 model or the full repository's calibration/source/evaluation assets.

Defaults are MTP on, an 8,192-token context limit, one active sequence, and the API at http://127.0.0.1:8000. The API binds to loopback only. This serving path is text-only; image and video inputs are disabled. Use --port to select another loopback port.

Example request from another terminal:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"Qwen/Qwen3.5-9B","messages":[{"role":"user","content":"Explain why batch-one decoding is memory-bandwidth limited."}],"temperature":0,"max_tokens":128,"chat_template_kwargs":{"enable_thinking":false}}'

Pin a revision or disable MTP

The launcher prints Revision: with the resolved 40-character HF commit. Save that value. --revision accepts an HF branch, tag, or commit; use the printed immutable commit for repeatable downloads, rather than main. This interactive command uses the commit you saved and selects the non-speculative baseline:

read -r -p 'Paste the saved 40-character Revision: ' QWEN9B_REVISION
python3 launch.py \
  --revision "$QWEN9B_REVISION" \
  --cache-root "$HOME/.cache/qwen9b-tt-native" \
  --mtp off \
  --device-ownership-confirmed

For a fully pinned replay, retain the original launch.py as well, or download launch.py from that same commit's HF resolve URL. The older release predates this launcher contract; use the separate rollback instructions below for it.

--download-only downloads and verifies the native package without requiring Docker or device ownership; it does not pull the image. --print-command downloads/verifies the package and prints the serving command without starting Docker. --tensor-cache can place the writable runtime cache elsewhere; it must remain separate from the immutable checkpoint snapshot. Run python3 launch.py --help for the supported options.

Runtime identity and source build

Release tag:

ghcr.io/lottolottolotto/qwen35-tt-p150:prefill-mtp-20260907

Published immutable image reference:

ghcr.io/lottolottolotto/qwen35-tt-p150@sha256:e67da563792a45c5e3079d54a048d1cd1de0c9ca84325b284680148430b91bc9

Runtime source commit: f060e86fdfa999eff009884b27a4aa30865530d5. The launcher takes the immutable image reference from runtime-release.json, not from a mutable Docker tag. This image was pushed to GHCR and pulled successfully with an empty Docker credential configuration. That anonymous pull reused existing local image layers; it was not an empty-layer-cache or clean-machine download.

Ordinary users do not need to build the image. For source reproduction, obtain the complete runtime/ directory at the same HF commit, then run from the directory containing it:

docker build \
  -f runtime/Dockerfile \
  -t qwen35-tt-native-quant:prefill-mtp \
  runtime

The build context is runtime/, not the repository root. This Dockerfile layers the bundled runtime/backend source and native loader onto an immutable parent runtime containing its compiled dependencies. It is not a from-scratch build of every upstream dependency. A rebuild can produce a different image identity and does not inherit validation of the published image. The normal launch.py path intentionally uses the release manifest's image rather than this locally named rebuild.

The current reproduction.json records this release's reproduction commands. It is distinct from the old revision's reproduction metadata linked in the historical section below.

Checkpoint and proof scope

  • Native TT BFP4/BFP8 matrix storage with selective precision retention; BF16 embeddings and original FP32 nonlinear parameters remain lossless.
  • The checkpoint retains 475 target tensors and 15 MTP tensors. The MTP tensors are stored losslessly at their original dtype, totaling 486,582,864 bytes; the runtime performs its own draft-layer conversions.
  • Importance statistics rank precision promotions. They do not optimize rounding or implement an imatrix-aware quantizer.
  • Precision summary, selected overrides, and native manifest describe the checkpoint. MTP-off and MTP-on proof records are mode-bound: inspect their runtime identity and execution contract when interpreting them.

Native-reload equality compares original-source and restored-native representations under the recorded mixed-precision execution contract. It is not equality to the original BF16 model, cross-mode output equality, a quality benchmark, or proof of distribution-preserving speculative sampling. Unchanged tensor bytes do not establish unchanged behavior under an upgraded runtime.

Current release verification

On the published prefill-mtp-20260907 image, original-source and native-restored weights produced bit-exact full-vocabulary logits across 1,415 token positions in each mode, MTP off and MTP on. Each comparison used three complete validation records and independent original/native tensor caches. The native manifest and recorded model-source hashes match this release. MTP-enabled teacher forcing loads the MTP head but does not itself exercise speculative generation.

Evidence: mode comparison summary, 136-file image payload verification, and anonymous registry access. Input selection, runtime environments, commands, and original/native metadata are retained in the same evidence directory. Raw logit arrays were compared locally; their hashes are recorded in the mode-bound checkpoint proofs.

The public launcher was downloaded anonymously and used with package revision 22819c55e9c4dccdca8660c16a3a5489a115c7c9. All 513 required files (9,629,982,308 bytes) were downloaded into a new checkpoint cache and verified. Both serving modes then started from independently empty writable tensor caches. Container inspection confirmed the pinned image, read-only public native checkpoint, loopback binding, and no original-weight or runtime-source mounts. Existing Docker image layers were reused; this was not a clean-machine image transfer. The default MTP-on path was exercised without passing --mtp.

Both modes passed 11/11 behavior cases, near-8K retrieval, short/long/short request isolation, cancellation recovery, requested logprobs, and non-greedy sampling. Logit bias was honored with MTP off; MTP on retained the existing speculative API's explicit HTTP 400 rejection. These are local serving diagnostics, not a rerun of the historical 901-request quality benchmark.

Public serving measurement MTP off MTP on (default)
Decode, 128 output tokens, median of 3 runs 36.58 tok/s 50.58 tok/s
2,048-token prompt, one output token, median latency 0.759 s 0.761 s
7,936-token prompt, one output token, median latency 3.104 s 3.109 s
7,936-token prompt / elapsed time 2,557 tok/s 2,552 tok/s

Decode timing excludes the first output token. Prompt timing includes HTTP, scheduling, prefill, and the first output token; it is not isolated kernel timing. Prompt measurements used one warmup and three measured requests at each length. First startup compiles and initializes caches and can take several minutes.

See the public installation proof, MTP-off API report, MTP-on API report, and verification tools. Later evidence/documentation commits retain the tested launcher's serving inventory; the immutable tested package revision above remains directly reproducible.

Historical quality and performance — previous runtime

Every measurement in this section belongs to the previous release, not a rerun of prefill-mtp-20260907. The unchanged weights preserve the relevance of the checkpoint's history, but these results do not validate the new runtime, its prefill path, or its published image. The authoritative historical card and evidence remain pinned at HF revision ef7b059d14db3a532452baa0cef8d1f4431b5657.

Task quality with MTP enabled

Historical LocalMaxxing v0.1.39, shard 1, concurrency 1, temperature 0, thinking disabled. All 901 requests completed without errors, and prompt hashes matched the corresponding earlier no-MTP shard runs.

Dataset Correct / scored Accuracy Output cap
GSM8K 97 / 101 96.04% 2,048 tokens
ARC-Challenge 754 / 800 94.25% 512 tokens

These are complete declared shard evaluations, not full-dataset official benchmark scores. ARC uses generative exact-match scoring, not log-likelihood multiple-choice scoring. This is not an Unsloth Divergence-300 result. The request adapter only added chat_template_kwargs.enable_thinking=false; task prompts and scoring were the CLI defaults. The older MTP runtime also changed warmup-buffer lifetime and KV block alignment, so comparison with its preceding no-MTP runtime does not isolate MTP's effect on accuracy.

Pinned historical evidence: GSM8K, ARC-Challenge, and protocol, commands, and request/response traces. The CLI's eval_shard_dry_run label denotes those locally executed, unsubmitted evaluations; inference and scoring were performed.

Decode speed and earlier validation

Historical single-P150, concurrency-one greedy decoding with thinking disabled, 128 generated tokens per measured request after warmup:

Measurement MTP off MTP on Speedup
Fixed prompt, median of 5 runs 36.57 tok/s 50.56 tok/s 38.26%
LocalMaxxing v0.1.39, median of 3 runs 36.5 tok/s 55.1 tok/s 50.96%

The fixed prompt had 55 prompt tokens and accepted 47 of 80 drafts per MTP run (58.75%). LocalMaxxing used --prompt-tokens 512; actual token counts and timing details are in the pinned raw records. These are decode/output-rate measurements, not prefill-inclusive end-to-end throughput. MTP-on and MTP-off completions differed; repeatability did not establish cross-mode equivalence.

The previous release reported exact full-vocabulary native-reload logit equality over 1,415 token positions in each mode, 11/11 API behavior checks, near-8K retrieval, and short/long/short request isolation. Its post-upload container test verified an anonymous digest-pinned pull with an empty Docker credential configuration, reusing existing local layers. A container from that old registry reference passed the API/retrieval/isolation workload with median decode 50.34 tok/s and 230 accepted tokens from 370 proposed drafts. That was neither a clean-machine pull nor a test of the new image.

Historical quality-summary.json and reproduction.json retain that session's runtime and commands. Their absolute paths, old build context, Nothing submitted publicly, and serving-restoration fields are historical records, not current deployment instructions. Likewise, development-tree integration measurements apply only to the runtime recorded in those artifacts, not automatically to this release image.

Roll back to the previous immutable release

The previous package and image remain independently addressable:

  • HF revision: ef7b059d14db3a532452baa0cef8d1f4431b5657 (files and original instructions).
  • Old image tag: ghcr.io/lottolottolotto/qwen35-tt-p150:mixed-mtp-20260906.
  • Immutable old image: ghcr.io/lottolottolotto/qwen35-tt-p150@sha256:fe06066653f6c1dc3d68cf7c2ee6e31a7724aebec82f88ff7bf792b78a0b1b21.

Download that exact old HF revision into a separate directory using the HF file browser or your existing repository-download tooling. Stop the current device workload yourself. From the old release directory, use its own serve_native.py and checkpoint, with a separate writable cache:

docker pull ghcr.io/lottolottolotto/qwen35-tt-p150@sha256:fe06066653f6c1dc3d68cf7c2ee6e31a7724aebec82f88ff7bf792b78a0b1b21
python3 serve_native.py \
  --checkpoint checkpoint \
  --cache-root rollback-serve-cache \
  --port 8001 \
  --image ghcr.io/lottolottolotto/qwen35-tt-p150@sha256:fe06066653f6c1dc3d68cf7c2ee6e31a7724aebec82f88ff7bf792b78a0b1b21 \
  --mtp on \
  --device-ownership-confirmed

Do not combine the old image with the new package's execution proofs, or pass this old revision to the new single-file launcher: it predates runtime-release.json.

Contents and integrity

  • launch.py: standard-library native downloader and launcher.
  • runtime-release.json: release identity, digest-pinned image, and SHA-256/byte-size inventory consumed by the launcher.
  • checkpoint/: TT-native target/MTP weights, tokenizer/config assets, and mode-bound native-reload proofs.
  • runtime/: standalone Docker build context, runtime/backend source, native loader, provenance, and upstream licenses; not required on the host for normal serving.
  • evidence/mtp/ and evidence/previous-no-mtp-quality/: retained historical measurements; their recorded runtime identities define their scope.
  • Calibration, evaluation, export, and precision-selection tools and inputs: available for research reproduction, not downloaded by the single-launch serving path.
  • release-manifest.json and SHA256SUMS: full publication inventory and integrity metadata, distinct from the launcher's smaller native-serving inventory.

License and attribution

Derived from Qwen/Qwen3.5-9B, revision c202236235762e1c871ad0ccb60c8ee5ba337b9a. Apache-2.0; see LICENSE. Runtime sources retain their upstream copyright notices and bundled licenses. This is a community experimental release by Lottolabs, not an official Qwen or Tenstorrent release.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lottolabs/Qwen3.5-9B-TT-Mixed-BFP4-BFP8-P150

Finetuned
Qwen/Qwen3.5-9B
Quantized
(540)
this model