Qwen3.8-27B — TT-native mixed BFP4/BFP8 (GPTQ), single P150, native 262K context

A hardware-specific, text-only checkpoint and custom runtime derived from Qwen/Qwen3.8-27B, revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0, for one Tenstorrent Blackhole P150.

This is a community experimental release by Lottolabs, not an official Qwen or Tenstorrent release. It is not GGUF, AWQ, an AutoGPTQ/GPTQModel int4 checkpoint, bitsandbytes, NVFP4, or an ordinary Transformers checkpoint. transformers.from_pretrained() and hosted Inference Providers are not supported launch paths.

What changed in this release (gptq-ctx2-20261002)

  • GPTQ-quantized BFP4 weights. The 192 matrices stored as BFP4 (MLP gate/up/down of all 64 layers) now hold weights computed by a GPTQ variant written for Tenstorrent block floating point (sequential, layer by layer, with error compensation inside the 16-value BFP4 blocks, on a balanced chat/code/instruction calibration set). The previous release rounded the same matrices to nearest. Every value is exactly representable in BFP4, so the runtime's own conversion reproduces it bit for bit. The precision map, the tensor sizes and the bytes read per token are unchanged; everything that is not a BFP4 matrix is byte-identical to the previous checkpoint.
  • Runtime ctx2:
    • native 262,144-token context (paged KV cache in BFP8 above 64K tokens);
    • DFlash2 block-8 speculative decoding up to 131,072 tokens, with the bundled z-lab/Qwen3.8-27B-DFlash2 drafter;
    • MTP-3 above 131,072 tokens.
  • Launcher. --max-model-len now goes up to 262,144, and --spec auto|dflash|mtp|off selects the decoding mode. The default is 8,192 tokens with DFlash2.

Quality

These are teacher-forced device runs on the P150. The eval set has 34 records: 14 multi-turn chats, 12 Python files and 8 instructions, with 42,129 scored positions. Chat and instruction records are scored on the assistant tokens only. The native checkpoint was evaluated as shipped. Its logits are bit-identical to those of the evaluation that applied the GPTQ weights as an override (34 of 34 records).

Against the TT all-BFP8 build (same runtime, every matrix in BFP8): mean KL divergence, lower is better, and top-1 agreement, higher is better.

Corpus Previous release (round-to-nearest BFP4) KL / top-1 This release (GPTQ BFP4) KL / top-1
Chat 0.180 / 0.838 0.070 / 0.897
Code 0.053 / 0.940 0.047 / 0.943
Instruct 0.013 / 0.960 0.011 / 0.967
All 0.103 / 0.899 0.055 / 0.925

Against a CPU BF16 reference (Δ perplexity, top-1 agreement, KL):

Corpus TT all-BFP8 Previous release (round-to-nearest) This release (GPTQ)
Code +0.59% / 0.952 / 0.037 +4.94% / 0.925 / 0.091 +4.25% / 0.931 / 0.086
Instruct −0.45% / 0.968 / 0.010 +0.24% / 0.957 / 0.019 −0.09% / 0.960 / 0.019

Chat is compared with the TT all-BFP8 build rather than the CPU reference: on chat, even the all-BFP8 build sits at KL 0.28 from the CPU reference, so that reference does not separate the quantizations.

Speed

Single P150, one sequence, greedy, 512 streamed output tokens. Decode tok/s is the median over requests and counts client-observed tokens after the first one. "Tokens/cycle" is the number of tokens emitted per verification (1 + accepted drafts). For comparison, the previous checkpoint (round-to-nearest) was measured on the same runtime on 2026-09-28.

Profile Prompt tokens Decode tok/s TTFT Tokens/cycle Previous checkpoint, same runtime: tok/s (tokens/cycle)
8K, DFlash2 (default) 128 49.1 0.21 s 3.15 ¹ 50.9 (3.31 ¹)
8K, DFlash2 (default) 2,048 48.1 1.69 s 3.15 ¹ 51.3 (3.31 ¹)
128K, DFlash2 128 45.3 0.24 s 3.09 ¹ —
128K, DFlash2 2,048 49.2 1.67 s 3.09 ¹ 49.0 (3.21)
128K, DFlash2 65,024 44.6 70.5 s 3.36 43.2 (3.28)
128K, DFlash2 130,560 37.9 179.0 s 3.21 38.8 (3.31)
262K, MTP-3 128 38.6 0.22 s 2.23 ¹ —
262K, MTP-3 2,048 37.9 1.65 s 2.23 ¹ —
262K, MTP-3 65,024 34.9 70.5 s 2.34 32.6 (2.19)
262K, MTP-3 261,632 24.5 510.0 s 2.34 24.5 (2.33)
Non-speculative (--spec off) 128 / 2,048 20.2 / 20.0 0.21 / 1.66 s 1 —

¹ Combined over the 128- and 2,048-token runs of that profile.

The prompt is a technical essay request. The new weights change the greedy text, so the drafts are accepted at a slightly different rate. Speed follows that rate: decode is between 6% slower (8K, 2,048 tokens) and 7% faster (262K profile, 65K tokens) than with the previous checkpoint. The cost of each verification cycle is unchanged, because the bytes read per token are identical.

In this release's runs, the 128K DFlash2 profile retrieved all passkeys (beginning, middle and end) at 65,400 and 130,944 prompt tokens. The 262K MTP-3 profile retrieved them at 65,400, 130,944 and 262,016 tokens. It accepted the 262,016 + 128 = 262,144-token boundary request and rejected a 262,145-token request with HTTP 400 before generation.

Download and serve

Prerequisites:

  • Linux with a supported P150 driver and hugepage mounts at /dev/hugepages and /dev/hugepages-1G.
  • Docker and Python 3.11 or newer.
  • Exclusive ownership of one P150.
  • Disk space for:
    • the ~22.1 GB checkpoint;
    • the 3.8 GB drafter;
    • the 5.6 GB compressed runtime archive and the loaded image;
    • writable tensor/device caches (~23 GB per decoding mode).
curl --fail --location --output launch.py \
  https://huggingface.co/Lottolabs/Qwen3.8-27B-TT-Mixed-BFP4-BFP8-P150/resolve/main/launch.py

# Default: 8,192-token context, DFlash2 speculative decoding
python3 launch.py --cache-root /large-disk/qwen27b --device-ownership-confirmed

# 128K context, DFlash2
python3 launch.py --cache-root /large-disk/qwen27b --max-model-len 131072 --device-ownership-confirmed

# Native 262K context, MTP-3
python3 launch.py --cache-root /large-disk/qwen27b --max-model-len 262144 --device-ownership-confirmed

# Verified non-speculative target path
python3 launch.py --cache-root /large-disk/qwen27b --spec off --device-ownership-confirmed

# Download and verify only (no Docker, no device) / print the serving command
python3 launch.py --cache-root /large-disk/qwen27b --download-only
python3 launch.py --cache-root /large-disk/qwen27b --print-command

The launcher is standard-library Python. Before it serves, it:

  1. resolves the requested revision to an immutable commit;
  2. downloads the native package anonymously;
  3. verifies every byte against runtime-release.json;
  4. loads the checksum-pinned runtime-image.tar.gz and checks the Docker image ID (sha256:53a2ae1582b3…).

The OpenAI-compatible API (model Qwen/Qwen3.8-27B) binds to 127.0.0.1:8000 and serves one sequence at a time. The first start of each decoding mode converts weights into the tensor cache and takes several minutes. Keep the printed revision and pass it with --revision for exact replay.

Verification

  • Native reload proofs. Two fresh proofs are bound to the manifest: checkpoint/equivalence.json (MTP off) and checkpoint/equivalence-mtp.json (MTP on). Each compares full-vocabulary logits on 95 teacher-forced positions. The baseline is the original BF16 checkpoint with the same GPTQ weights substituted, quantized at load; the comparison run reloads the native checkpoint. The logits are exactly equal in both modes. The manifest's ttquant section records the hashes of the GPTQ weight files and the regex selecting them, and the proof tool refuses a baseline that used anything else.

  • HTTP behaviour, every profile.

    • /v1/completions with the chat prompt's token IDs reproduces the chat stream.
    • Requests are isolated: A, B, A gives the same A, and a greedy request after a sampled one is unchanged.
    • Cancellation mid-decode and mid-prefill leaves the next request exact.
    • Image inputs are rejected.
    • Evidence: evidence/verification/.
  • Speculative vs non-speculative output. Greedy output is deterministic and identical across the speculative profiles, but it is not token-identical to non-speculative decoding. Multi-row verification rounds differently from single-row (M=1) decode. This is a property of the Qwen3.8-27B runtime since the DFlash rounds, not of the GPTQ checkpoint: the previous round-to-nearest checkpoint on the same runtime shows the same difference. Measured on this release:

    • DFlash2 at 8K and MTP-1 at 8K produced identical texts (8/8 requests, BF16 KV).
    • DFlash2 at 128K and MTP-3 at 262K produced identical texts (6/6 requests, BFP8 KV).
    • Against the non-speculative stream of the same profile (--spec off), the 6 chat prompts first diverge after 3–117 tokens.
    • The previous checkpoint on the same runtime: its non-speculative 8K texts also differ from its DFlash2 8K texts, while its DFlash2 texts equal its earlier MTP-1 reference.

    The quality tables are measured on the teacher-forced path, so they are unaffected.

  • Public package. An anonymous fresh download of revision 8659ecb6 with its own launch.py re-hashed all 991 serving files (31.6 GB) and loaded the archive to the expected image ID (evidence/public-download-verification.json). The downloaded package then served its default profile: a chat request returned PUBLIC PACKAGE OK, and two 2,048-token probes ran at 48.0 tok/s with texts identical to the pre-publication DFlash2 run (evidence/public-serving-smoke.json).

These checks establish native/runtime equivalence and serving behaviour. They do not certify quality against the upstream BF16 model; the quality tables above are measured, not certified. Vision inputs are not supported.

Previous release

The previous release was dual-noc-mtp1-20260920: round-to-nearest BFP4, MTP-1 runtime lottolabs/qwen38-27b-tt-p150:dual-noc-20260920, 8,192-token context. It stays available at revision 50becf020f24ac03d6b2b69434a3d4e2655e0865. To replay it, download that revision's launch.py and pass --revision 50becf020f24ac03d6b2b69434a3d4e2655e0865. Its evidence is kept under evidence/previous-dual-noc-mtp1-20260920/. The LocalMaxxing submission (23.8 tok/s, MTP-1) was measured on that release, not on this one.

Integrity and provenance

  • runtime-release.json: the downloadable serving inventory, the runtime-archive checksum, the expected Docker image ID and the upstream drafter identity.
  • release-manifest.json and SHA256SUMS: the complete repository inventory and checksums.
  • checkpoint/native_manifest.json: the per-tensor format, precision, hashes and the ttquant section.
  • checkpoint/provenance/: the builder, the evaluator and the BFP unpacker used for the GPTQ weights.
  • precision-summary.json: the precision map (128 BFP4 and 113 BFP8 layer-family assignments, LM head BFP8).
  • runtime/: the build context of the ctx2 image. It contains Python/JIT-kernel overlays on an immutable development parent image. It is provenance: rebuilding needs that parent image. Serving uses the checksum-pinned archive.

License and attribution

  • The model is derived from Qwen/Qwen3.8-27B, Apache-2.0; see LICENSE.
  • The DFlash2 drafter in dflash-draft/ is redistributed byte-identical from z-lab/Qwen3.8-27B-DFlash2, revision 50307d4c4cde6860d4eee73e2547cd786fe8e8a4, Apache-2.0. Its model card is kept as dflash-draft/README.md.
  • Runtime sources keep their upstream notices and bundled licenses.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lottolabs/Qwen3.8-27B-TT-Mixed-BFP4-BFP8-P150

Base model

Qwen/Qwen3.8-27B
Quantized
(1300)
this model