Add model card, measured results, publication provenance and checksums
Browse files- README.md +128 -0
- SHA256SUMS +2 -0
- publication.json +8 -0
README.md
ADDED
|
@@ -0,0 +1,128 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen3.5-9B
|
| 4 |
+
base_model_relation: quantized
|
| 5 |
+
pipeline_tag: text-generation
|
| 6 |
+
tags:
|
| 7 |
+
- tenstorrent
|
| 8 |
+
- blackhole
|
| 9 |
+
- p150
|
| 10 |
+
- ttnn
|
| 11 |
+
- mixed-precision
|
| 12 |
+
- bfp4
|
| 13 |
+
- bfp8
|
| 14 |
+
- speculative-decoding
|
| 15 |
+
- custom-runtime
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
# Qwen3.5-9B — TT-native mixed BFP4/BFP8, single P150, MTP-1
|
| 19 |
+
|
| 20 |
+
An experimental **hardware-specific checkpoint and custom runtime** derived from [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), verified on **one Tenstorrent Blackhole P150**. This is the new mixed-BFP4/BFP8 quant, not the [earlier BFP4 release](https://huggingface.co/Lottolabs/Qwen3.5-9B-TT-BFP4-P150).
|
| 21 |
+
|
| 22 |
+
**MTP-1 speculative decoding is included and enabled by default.** The package contains the target checkpoint, all 15 original MTP tensors, runtime source and build recipe, launcher, calibration inputs and precision selections, and evaluation evidence. The checkpoint is approximately **9.63 GB**; the complete sealed release payload is approximately **9.65 GB**.
|
| 23 |
+
|
| 24 |
+
This is **not GGUF, AWQ, GPTQ, bitsandbytes, a standard Transformers checkpoint, or an ordinary CUDA/vLLM model**. Use the bundled TT runtime. Hugging Face Transformers `from_pretrained()` and hosted Inference Providers are not supported launch paths. Serving requires no original Hugging Face weight download after this repository has been downloaded.
|
| 25 |
+
|
| 26 |
+
## Measured decode speed
|
| 27 |
+
|
| 28 |
+
Same new checkpoint and corrected runtime image; single P150, concurrency 1, greedy decoding, thinking disabled, 128 generated tokens per measured request, after warmup.
|
| 29 |
+
|
| 30 |
+
| Measurement | MTP off | MTP on | Speedup |
|
| 31 |
+
|---|---:|---:|---:|
|
| 32 |
+
| Fixed prompt, median of 5 runs | 36.57 tok/s | **50.56 tok/s** | **38.26%** |
|
| 33 |
+
| LocalMaxxing v0.1.39, median of 3 runs | 36.5 tok/s | **55.1 tok/s** | **50.96%** |
|
| 34 |
+
|
| 35 |
+
The fixed prompt has 55 prompt tokens. LocalMaxxing used `--prompt-tokens 512`; actual prompt token counts and client timing details are in the raw records. These are decode/output-rate measurements, not prefill-inclusive end-to-end throughput. The fixed-prompt MTP runs accepted **47 of 80 drafts (58.75%)** per run.
|
| 36 |
+
|
| 37 |
+
Evidence: [fixed prompt on](evidence/mtp/native-mtp-on-speed.json), [off](evidence/mtp/native-mtp-off-speed.json), [LocalMaxxing on](evidence/mtp/lmx-native-mtp-on.json), [off](evidence/mtp/lmx-native-mtp-off.json), and [runtime provenance](evidence/mtp/fixed-runtime-provenance.json).
|
| 38 |
+
|
| 39 |
+
**MTP-on and MTP-off completions differ.** Repeatability and serialization equivalence do not establish cross-mode output equivalence or distribution-preserving speculative sampling. The reported measurements exercise greedy decoding only.
|
| 40 |
+
|
| 41 |
+
## Task quality with MTP enabled
|
| 42 |
+
|
| 43 |
+
LocalMaxxing v0.1.39, shard 1, concurrency 1, temperature 0, thinking disabled. All 901 requests completed without errors, and prompt hashes match the corresponding earlier no-MTP shard runs.
|
| 44 |
+
|
| 45 |
+
| Dataset | Correct / scored | Accuracy | Output cap |
|
| 46 |
+
|---|---:|---:|---:|
|
| 47 |
+
| GSM8K | **97 / 101** | **96.04%** | 2,048 tokens |
|
| 48 |
+
| ARC-Challenge | **754 / 800** | **94.25%** | 512 tokens |
|
| 49 |
+
|
| 50 |
+
These are **the complete declared shard evaluations, not full-dataset official benchmark scores**. ARC is generative exact-match scoring, not log-likelihood multiple-choice scoring. This is not an Unsloth Divergence-300 result. The request adapter only adds `chat_template_kwargs.enable_thinking=false`; task prompts and scoring remain the CLI defaults.
|
| 51 |
+
|
| 52 |
+
The corrected runtime also changes warmup-buffer lifetime and KV block alignment. Comparing these scores with the previous no-MTP runtime **does not isolate MTP's effect on accuracy**.
|
| 53 |
+
|
| 54 |
+
Evidence: [GSM8K results](evidence/mtp/gsm8k-shard1.json), [ARC results](evidence/mtp/arc-challenge-shard1.json), [protocol](evidence/mtp/quality-protocol.json), [commands](evidence/mtp/quality-commands.json), and [all request/response traces](evidence/mtp/quality-request-traces.jsonl). The benchmark CLI's `eval_shard_dry_run` event is its label for these locally executed, unsubmitted evaluations; actual inference and scoring were performed.
|
| 55 |
+
|
| 56 |
+
## Quantization and verification
|
| 57 |
+
|
| 58 |
+
- Native TT BFP4/BFP8 matrix storage with selectively retained precision; BF16 embeddings and original FP32 nonlinear parameters remain lossless.
|
| 59 |
+
- All **475 target tensors are byte-identical** to the preceding mixed-precision export.
|
| 60 |
+
- All **15 MTP tensors are stored losslessly at their original dtype**, adding 486,582,864 bytes. The runtime performs its own draft-layer conversions.
|
| 61 |
+
- Importance statistics rank precision promotions. They do **not** optimize rounding or implement an imatrix-aware quantizer.
|
| 62 |
+
- Exact full-vocabulary native-reload logit equality across **1,415 token positions in each mode**, MTP off and on, using three validation records per original/native pair.
|
| 63 |
+
- **11/11 API behavior checks**, near-8K retrieval, and the short/long/short request-isolation regression passed.
|
| 64 |
+
- The runtime fixes persistent MTP buffer reallocation across warmup phases and retains **64-token attention cache blocks**; GDN state is owned by the TT model rather than paged by vLLM.
|
| 65 |
+
|
| 66 |
+
Reload equality proves parity between the original-source and restored-native representations under the recorded mixed-precision execution contract. It is **not equivalence to the original BF16 model**, nor a proof of all speculative-decoding behavior. The API and task evaluations provide separate evidence for the exercised serving paths.
|
| 67 |
+
|
| 68 |
+
See [precision summary](precision-summary.json), [selected overrides](selected-overrides.json), [checkpoint manifest](checkpoint/native_manifest.json), [MTP-off proof](checkpoint/equivalence.json), [MTP-on proof](checkpoint/equivalence-mtp.json), and [isolation regression](evidence/mtp/request-isolation-fixed.json).
|
| 69 |
+
|
| 70 |
+
## Download and serve
|
| 71 |
+
|
| 72 |
+
Prerequisites: supported P150 driver, exclusive ownership of the accelerator, Docker, host Python 3, and hugepage mounts at `/dev/hugepages` and `/dev/hugepages-1G`. Allow additional disk space for the Docker build and writable runtime cache. The verified serving context limit is 8,192 tokens, with one active sequence. Text-only: image and video inputs are disabled.
|
| 73 |
+
|
| 74 |
+
```bash
|
| 75 |
+
python3 -m pip install -U huggingface_hub
|
| 76 |
+
hf download Lottolabs/Qwen3.5-9B-TT-Mixed-BFP4-BFP8-P150 --local-dir qwen35-tt-mixed-mtp
|
| 77 |
+
cd qwen35-tt-mixed-mtp
|
| 78 |
+
sha256sum --check --quiet SHA256SUMS
|
| 79 |
+
|
| 80 |
+
docker build --build-arg BUILD_JOBS=8 \
|
| 81 |
+
-f runtime/Dockerfile -t qwen35-tt-native-quant:mtp .
|
| 82 |
+
|
| 83 |
+
# Stop other P150 workloads before launching.
|
| 84 |
+
python3 serve_native.py \
|
| 85 |
+
--checkpoint checkpoint \
|
| 86 |
+
--cache-root serve-cache \
|
| 87 |
+
--port 8001 \
|
| 88 |
+
--mtp on \
|
| 89 |
+
--device-ownership-confirmed
|
| 90 |
+
```
|
| 91 |
+
|
| 92 |
+
The runtime is rebuilt from the pinned public base and included source overlay. The image SHA in the evidence identifies the tested local image; it is **not a separately published Docker image pull address**. A fresh build can have a different image ID. The launcher validates the checkpoint's mode-bound proof and runtime execution contract.
|
| 93 |
+
|
| 94 |
+
Use `--mtp off` for the explicit non-speculative baseline. The default is `--mtp on`. Keep `serve-cache` writable and outside `checkpoint/`. The launcher does not stop other accelerator workloads for you.
|
| 95 |
+
|
| 96 |
+
Example request:
|
| 97 |
+
|
| 98 |
+
```bash
|
| 99 |
+
curl http://127.0.0.1:8001/v1/chat/completions \
|
| 100 |
+
-H 'Content-Type: application/json' \
|
| 101 |
+
-d '{"model":"Qwen/Qwen3.5-9B","messages":[{"role":"user","content":"Explain why batch-one decoding is memory-bandwidth limited."}],"temperature":0,"max_tokens":128,"chat_template_kwargs":{"enable_thinking":false}}'
|
| 102 |
+
```
|
| 103 |
+
|
| 104 |
+
Run the behavioral and request-isolation checks:
|
| 105 |
+
|
| 106 |
+
```bash
|
| 107 |
+
python3 api_validate.py \
|
| 108 |
+
--base-url http://127.0.0.1:8001 \
|
| 109 |
+
--cases evaluation/behavior.jsonl \
|
| 110 |
+
--long-case evaluation/long-context.jsonl \
|
| 111 |
+
--output api-validation.json
|
| 112 |
+
```
|
| 113 |
+
|
| 114 |
+
For other reproduction steps, see [reproduction.json](reproduction.json). Raw evidence preserves the original machine's absolute paths and container names; substitute your local paths when replaying the recorded commands. Historical `Nothing submitted publicly` and serving-restoration fields describe the measurement session before this Hugging Face publication; they are not current deployment instructions.
|
| 115 |
+
|
| 116 |
+
## Contents and integrity
|
| 117 |
+
|
| 118 |
+
- `checkpoint/`: standalone TT-native target and MTP weights, tokenizer/config assets, and mode-bound reload proofs.
|
| 119 |
+
- `runtime/`: pinned Docker build recipe, custom runtime/backend sources, and upstream licenses.
|
| 120 |
+
- `evidence/mtp/`: final MTP quality, speed, isolation, and runtime provenance data.
|
| 121 |
+
- `evidence/previous-no-mtp-quality/`: explicitly historical evidence; not relabeled as MTP results.
|
| 122 |
+
- Calibration, evaluation, export, and precision-selection tools and inputs.
|
| 123 |
+
- [quality-summary.json](quality-summary.json): machine-readable final results.
|
| 124 |
+
- [release-manifest.json](release-manifest.json) and [SHA256SUMS](SHA256SUMS): sealed release payload integrity. Publication adds this model card and publication metadata to the checksum list without changing the measured release files.
|
| 125 |
+
|
| 126 |
+
## License and attribution
|
| 127 |
+
|
| 128 |
+
Derived from [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), revision `c202236235762e1c871ad0ccb60c8ee5ba337b9a`. Apache-2.0; see [LICENSE](LICENSE). Runtime sources retain their upstream copyright notices and bundled licenses. This is a community experimental release by **Lottolabs**, not an official Qwen or Tenstorrent release.
|
SHA256SUMS
CHANGED
|
@@ -693,3 +693,5 @@ a87126812d5e9c761a7670589175e3be6b855f60df76a641bf2691789f7899a1 score_behavior
|
|
| 693 |
59d35218b4780c71e90f2a7cf148cc1f8a0db44a4a5c5dbd90cda624d6b54b1a selected-overrides.json
|
| 694 |
0118c611f52574043f9fe3b9a32657e59a663b5e92e051f3159e91fa0e1b1e01 serve_native.py
|
| 695 |
2939042f81d63dd6d90b1b13574c2a505623a07c1684bb9e3d048182c25eed59 tt_eval.py
|
|
|
|
|
|
|
|
|
| 693 |
59d35218b4780c71e90f2a7cf148cc1f8a0db44a4a5c5dbd90cda624d6b54b1a selected-overrides.json
|
| 694 |
0118c611f52574043f9fe3b9a32657e59a663b5e92e051f3159e91fa0e1b1e01 serve_native.py
|
| 695 |
2939042f81d63dd6d90b1b13574c2a505623a07c1684bb9e3d048182c25eed59 tt_eval.py
|
| 696 |
+
12bd27ea062de667042e24919143b6c99002cd05b81ae77c054c76a47399e594 README.md
|
| 697 |
+
f3ff93cffc27be23dfc81acf7417ef3ab6d26164a3b8e64aa32f3aab47db44d5 publication.json
|
publication.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schema_version": 1,
|
| 3 |
+
"repository": "Lottolabs/Qwen3.5-9B-TT-Mixed-BFP4-BFP8-P150",
|
| 4 |
+
"release_manifest_sha256": "0d7ea076699733ed08d20519dc75d62e342b4dd43d53c9dacef2d1d08dd6b50b",
|
| 5 |
+
"native_manifest_sha256": "34f774cf110e3ff94726171b52f6d462cb9b36651295b941efa85098ddb96ca2",
|
| 6 |
+
"scope": "Public upload of the sealed MTP-enabled mixed-precision release and its measured evidence; no checkpoint or measured release file modified",
|
| 7 |
+
"historical_metadata_note": "Measurement-time publication and serving-status fields are retained verbatim. This metadata records the Hugging Face publication separately."
|
| 8 |
+
}
|