Lottolabs commited on
Commit
17b70b7
·
verified ·
1 Parent(s): 12f320c

Add model card, measured results, publication provenance and checksums

Browse files
Files changed (3) hide show
  1. README.md +128 -0
  2. SHA256SUMS +2 -0
  3. publication.json +8 -0
README.md ADDED
@@ -0,0 +1,128 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen3.5-9B
4
+ base_model_relation: quantized
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - tenstorrent
8
+ - blackhole
9
+ - p150
10
+ - ttnn
11
+ - mixed-precision
12
+ - bfp4
13
+ - bfp8
14
+ - speculative-decoding
15
+ - custom-runtime
16
+ ---
17
+
18
+ # Qwen3.5-9B — TT-native mixed BFP4/BFP8, single P150, MTP-1
19
+
20
+ An experimental **hardware-specific checkpoint and custom runtime** derived from [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), verified on **one Tenstorrent Blackhole P150**. This is the new mixed-BFP4/BFP8 quant, not the [earlier BFP4 release](https://huggingface.co/Lottolabs/Qwen3.5-9B-TT-BFP4-P150).
21
+
22
+ **MTP-1 speculative decoding is included and enabled by default.** The package contains the target checkpoint, all 15 original MTP tensors, runtime source and build recipe, launcher, calibration inputs and precision selections, and evaluation evidence. The checkpoint is approximately **9.63 GB**; the complete sealed release payload is approximately **9.65 GB**.
23
+
24
+ This is **not GGUF, AWQ, GPTQ, bitsandbytes, a standard Transformers checkpoint, or an ordinary CUDA/vLLM model**. Use the bundled TT runtime. Hugging Face Transformers `from_pretrained()` and hosted Inference Providers are not supported launch paths. Serving requires no original Hugging Face weight download after this repository has been downloaded.
25
+
26
+ ## Measured decode speed
27
+
28
+ Same new checkpoint and corrected runtime image; single P150, concurrency 1, greedy decoding, thinking disabled, 128 generated tokens per measured request, after warmup.
29
+
30
+ | Measurement | MTP off | MTP on | Speedup |
31
+ |---|---:|---:|---:|
32
+ | Fixed prompt, median of 5 runs | 36.57 tok/s | **50.56 tok/s** | **38.26%** |
33
+ | LocalMaxxing v0.1.39, median of 3 runs | 36.5 tok/s | **55.1 tok/s** | **50.96%** |
34
+
35
+ The fixed prompt has 55 prompt tokens. LocalMaxxing used `--prompt-tokens 512`; actual prompt token counts and client timing details are in the raw records. These are decode/output-rate measurements, not prefill-inclusive end-to-end throughput. The fixed-prompt MTP runs accepted **47 of 80 drafts (58.75%)** per run.
36
+
37
+ Evidence: [fixed prompt on](evidence/mtp/native-mtp-on-speed.json), [off](evidence/mtp/native-mtp-off-speed.json), [LocalMaxxing on](evidence/mtp/lmx-native-mtp-on.json), [off](evidence/mtp/lmx-native-mtp-off.json), and [runtime provenance](evidence/mtp/fixed-runtime-provenance.json).
38
+
39
+ **MTP-on and MTP-off completions differ.** Repeatability and serialization equivalence do not establish cross-mode output equivalence or distribution-preserving speculative sampling. The reported measurements exercise greedy decoding only.
40
+
41
+ ## Task quality with MTP enabled
42
+
43
+ LocalMaxxing v0.1.39, shard 1, concurrency 1, temperature 0, thinking disabled. All 901 requests completed without errors, and prompt hashes match the corresponding earlier no-MTP shard runs.
44
+
45
+ | Dataset | Correct / scored | Accuracy | Output cap |
46
+ |---|---:|---:|---:|
47
+ | GSM8K | **97 / 101** | **96.04%** | 2,048 tokens |
48
+ | ARC-Challenge | **754 / 800** | **94.25%** | 512 tokens |
49
+
50
+ These are **the complete declared shard evaluations, not full-dataset official benchmark scores**. ARC is generative exact-match scoring, not log-likelihood multiple-choice scoring. This is not an Unsloth Divergence-300 result. The request adapter only adds `chat_template_kwargs.enable_thinking=false`; task prompts and scoring remain the CLI defaults.
51
+
52
+ The corrected runtime also changes warmup-buffer lifetime and KV block alignment. Comparing these scores with the previous no-MTP runtime **does not isolate MTP's effect on accuracy**.
53
+
54
+ Evidence: [GSM8K results](evidence/mtp/gsm8k-shard1.json), [ARC results](evidence/mtp/arc-challenge-shard1.json), [protocol](evidence/mtp/quality-protocol.json), [commands](evidence/mtp/quality-commands.json), and [all request/response traces](evidence/mtp/quality-request-traces.jsonl). The benchmark CLI's `eval_shard_dry_run` event is its label for these locally executed, unsubmitted evaluations; actual inference and scoring were performed.
55
+
56
+ ## Quantization and verification
57
+
58
+ - Native TT BFP4/BFP8 matrix storage with selectively retained precision; BF16 embeddings and original FP32 nonlinear parameters remain lossless.
59
+ - All **475 target tensors are byte-identical** to the preceding mixed-precision export.
60
+ - All **15 MTP tensors are stored losslessly at their original dtype**, adding 486,582,864 bytes. The runtime performs its own draft-layer conversions.
61
+ - Importance statistics rank precision promotions. They do **not** optimize rounding or implement an imatrix-aware quantizer.
62
+ - Exact full-vocabulary native-reload logit equality across **1,415 token positions in each mode**, MTP off and on, using three validation records per original/native pair.
63
+ - **11/11 API behavior checks**, near-8K retrieval, and the short/long/short request-isolation regression passed.
64
+ - The runtime fixes persistent MTP buffer reallocation across warmup phases and retains **64-token attention cache blocks**; GDN state is owned by the TT model rather than paged by vLLM.
65
+
66
+ Reload equality proves parity between the original-source and restored-native representations under the recorded mixed-precision execution contract. It is **not equivalence to the original BF16 model**, nor a proof of all speculative-decoding behavior. The API and task evaluations provide separate evidence for the exercised serving paths.
67
+
68
+ See [precision summary](precision-summary.json), [selected overrides](selected-overrides.json), [checkpoint manifest](checkpoint/native_manifest.json), [MTP-off proof](checkpoint/equivalence.json), [MTP-on proof](checkpoint/equivalence-mtp.json), and [isolation regression](evidence/mtp/request-isolation-fixed.json).
69
+
70
+ ## Download and serve
71
+
72
+ Prerequisites: supported P150 driver, exclusive ownership of the accelerator, Docker, host Python 3, and hugepage mounts at `/dev/hugepages` and `/dev/hugepages-1G`. Allow additional disk space for the Docker build and writable runtime cache. The verified serving context limit is 8,192 tokens, with one active sequence. Text-only: image and video inputs are disabled.
73
+
74
+ ```bash
75
+ python3 -m pip install -U huggingface_hub
76
+ hf download Lottolabs/Qwen3.5-9B-TT-Mixed-BFP4-BFP8-P150 --local-dir qwen35-tt-mixed-mtp
77
+ cd qwen35-tt-mixed-mtp
78
+ sha256sum --check --quiet SHA256SUMS
79
+
80
+ docker build --build-arg BUILD_JOBS=8 \
81
+ -f runtime/Dockerfile -t qwen35-tt-native-quant:mtp .
82
+
83
+ # Stop other P150 workloads before launching.
84
+ python3 serve_native.py \
85
+ --checkpoint checkpoint \
86
+ --cache-root serve-cache \
87
+ --port 8001 \
88
+ --mtp on \
89
+ --device-ownership-confirmed
90
+ ```
91
+
92
+ The runtime is rebuilt from the pinned public base and included source overlay. The image SHA in the evidence identifies the tested local image; it is **not a separately published Docker image pull address**. A fresh build can have a different image ID. The launcher validates the checkpoint's mode-bound proof and runtime execution contract.
93
+
94
+ Use `--mtp off` for the explicit non-speculative baseline. The default is `--mtp on`. Keep `serve-cache` writable and outside `checkpoint/`. The launcher does not stop other accelerator workloads for you.
95
+
96
+ Example request:
97
+
98
+ ```bash
99
+ curl http://127.0.0.1:8001/v1/chat/completions \
100
+ -H 'Content-Type: application/json' \
101
+ -d '{"model":"Qwen/Qwen3.5-9B","messages":[{"role":"user","content":"Explain why batch-one decoding is memory-bandwidth limited."}],"temperature":0,"max_tokens":128,"chat_template_kwargs":{"enable_thinking":false}}'
102
+ ```
103
+
104
+ Run the behavioral and request-isolation checks:
105
+
106
+ ```bash
107
+ python3 api_validate.py \
108
+ --base-url http://127.0.0.1:8001 \
109
+ --cases evaluation/behavior.jsonl \
110
+ --long-case evaluation/long-context.jsonl \
111
+ --output api-validation.json
112
+ ```
113
+
114
+ For other reproduction steps, see [reproduction.json](reproduction.json). Raw evidence preserves the original machine's absolute paths and container names; substitute your local paths when replaying the recorded commands. Historical `Nothing submitted publicly` and serving-restoration fields describe the measurement session before this Hugging Face publication; they are not current deployment instructions.
115
+
116
+ ## Contents and integrity
117
+
118
+ - `checkpoint/`: standalone TT-native target and MTP weights, tokenizer/config assets, and mode-bound reload proofs.
119
+ - `runtime/`: pinned Docker build recipe, custom runtime/backend sources, and upstream licenses.
120
+ - `evidence/mtp/`: final MTP quality, speed, isolation, and runtime provenance data.
121
+ - `evidence/previous-no-mtp-quality/`: explicitly historical evidence; not relabeled as MTP results.
122
+ - Calibration, evaluation, export, and precision-selection tools and inputs.
123
+ - [quality-summary.json](quality-summary.json): machine-readable final results.
124
+ - [release-manifest.json](release-manifest.json) and [SHA256SUMS](SHA256SUMS): sealed release payload integrity. Publication adds this model card and publication metadata to the checksum list without changing the measured release files.
125
+
126
+ ## License and attribution
127
+
128
+ Derived from [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), revision `c202236235762e1c871ad0ccb60c8ee5ba337b9a`. Apache-2.0; see [LICENSE](LICENSE). Runtime sources retain their upstream copyright notices and bundled licenses. This is a community experimental release by **Lottolabs**, not an official Qwen or Tenstorrent release.
SHA256SUMS CHANGED
@@ -693,3 +693,5 @@ a87126812d5e9c761a7670589175e3be6b855f60df76a641bf2691789f7899a1 score_behavior
693
  59d35218b4780c71e90f2a7cf148cc1f8a0db44a4a5c5dbd90cda624d6b54b1a selected-overrides.json
694
  0118c611f52574043f9fe3b9a32657e59a663b5e92e051f3159e91fa0e1b1e01 serve_native.py
695
  2939042f81d63dd6d90b1b13574c2a505623a07c1684bb9e3d048182c25eed59 tt_eval.py
 
 
 
693
  59d35218b4780c71e90f2a7cf148cc1f8a0db44a4a5c5dbd90cda624d6b54b1a selected-overrides.json
694
  0118c611f52574043f9fe3b9a32657e59a663b5e92e051f3159e91fa0e1b1e01 serve_native.py
695
  2939042f81d63dd6d90b1b13574c2a505623a07c1684bb9e3d048182c25eed59 tt_eval.py
696
+ 12bd27ea062de667042e24919143b6c99002cd05b81ae77c054c76a47399e594 README.md
697
+ f3ff93cffc27be23dfc81acf7417ef3ab6d26164a3b8e64aa32f3aab47db44d5 publication.json
publication.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 1,
3
+ "repository": "Lottolabs/Qwen3.5-9B-TT-Mixed-BFP4-BFP8-P150",
4
+ "release_manifest_sha256": "0d7ea076699733ed08d20519dc75d62e342b4dd43d53c9dacef2d1d08dd6b50b",
5
+ "native_manifest_sha256": "34f774cf110e3ff94726171b52f6d462cb9b36651295b941efa85098ddb96ca2",
6
+ "scope": "Public upload of the sealed MTP-enabled mixed-precision release and its measured evidence; no checkpoint or measured release file modified",
7
+ "historical_metadata_note": "Measurement-time publication and serving-status fields are retained verbatim. This metadata records the Hugging Face publication separately."
8
+ }