Lottolabs commited on
Commit
8e6b0e0
·
verified ·
1 Parent(s): e469d86

Add files using upload-large-folder tool

Browse files
Files changed (50) hide show
  1. LICENSE +202 -0
  2. README.md +139 -0
  3. SHA256SUMS +0 -0
  4. evidence/http-p3c/exact.json +185 -0
  5. evidence/http-p3c/long.json +231 -0
  6. evidence/http-p3c/tput-128/requests-128-c1.json +34 -0
  7. evidence/http-p3c/tput-128/summary.json +24 -0
  8. evidence/http-p3c/tput-128/warmup-128.json +16 -0
  9. evidence/http-p3c/tput-131072/requests-131072-c1.json +34 -0
  10. evidence/http-p3c/tput-131072/summary.json +24 -0
  11. evidence/http-p3c/tput-131072/warmup-131072.json +16 -0
  12. evidence/http-p3c/tput-2048/requests-2048-c1.json +34 -0
  13. evidence/http-p3c/tput-2048/summary.json +24 -0
  14. evidence/http-p3c/tput-2048/warmup-2048.json +16 -0
  15. evidence/http-p3c/tput-261632/requests-261632-c1.json +34 -0
  16. evidence/http-p3c/tput-261632/summary.json +24 -0
  17. evidence/http-p3c/tput-261632/warmup-261632.json +16 -0
  18. evidence/http-p3c/tput-32768/requests-32768-c1.json +34 -0
  19. evidence/http-p3c/tput-32768/summary.json +24 -0
  20. evidence/http-p3c/tput-32768/warmup-32768.json +16 -0
  21. evidence/http-p3c/tput-8192/summary.json +24 -0
  22. evidence/http-p3c/tput-8192/warmup-8192.json +16 -0
  23. evidence/http-p3c/tput_vs_direct.json +22 -0
  24. evidence/localmaxxing/speed-test.json +175 -0
  25. evidence/quality/quant_table.json +533 -0
  26. evidence/quality/runtime-gate-dtf-score.json +41 -0
  27. gemma-4-12B-it-assistant/config.json +88 -0
  28. gemma-4-12B-it-assistant/drafter_manifest.json +285 -0
  29. gemma-4-12B-it-assistant/generation_config.json +14 -0
  30. gemma-4-12B-it-assistant/spec_equivalence.json +86 -0
  31. gemma-4-12B-it/chat_template.jinja +390 -0
  32. gemma-4-12B-it/config.json +172 -0
  33. gemma-4-12B-it/equivalence.json +106 -0
  34. gemma-4-12B-it/generation_config.json +18 -0
  35. gemma-4-12B-it/native_manifest.json +0 -0
  36. gemma-4-12B-it/precision_plan.json +8 -0
  37. gemma-4-12B-it/tokenizer_config.json +120 -0
  38. launch.py +300 -0
  39. native_checkpoint.py +476 -0
  40. provenance/build_native_checkpoint.py +54 -0
  41. release-manifest.json +0 -0
  42. reproduction.json +65 -0
  43. runtime-release.json +0 -0
  44. runtime/Dockerfile.fast1 +7 -0
  45. runtime/Dockerfile.p2b +14 -0
  46. runtime/Dockerfile.p2c +6 -0
  47. runtime/Dockerfile.p3c +7 -0
  48. runtime/build.sh +46 -0
  49. runtime/source.json +59 -0
  50. serve_native.py +128 -0
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright [yyyy] [name of copyright owner]
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
README.md ADDED
@@ -0,0 +1,139 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: google/gemma-4-12B-it
4
+ base_model_relation: quantized
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - tenstorrent
8
+ - blackhole
9
+ - p150
10
+ - ttnn
11
+ - tt-metal
12
+ - bfp8
13
+ - gemma4
14
+ - speculative-decoding
15
+ - long-context
16
+ - custom-runtime
17
+ ---
18
+
19
+ # Gemma 4 12B IT — TT-native all-BFP8, single P150, 262K context, speculative decoding
20
+
21
+ A hardware-specific, text-only checkpoint and custom runtime derived from [google/gemma-4-12B-it](https://huggingface.co/google/gemma-4-12B-it), revision `707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7`, for one Tenstorrent Blackhole P150. The drafter for speculative decoding is derived from [google/gemma-4-12B-it-assistant](https://huggingface.co/google/gemma-4-12B-it-assistant), revision `46d4c6f13f0ac0ad827b915669b8df9b81c64c51`.
22
+
23
+ This is a community experimental release by **Lottolabs**, not an official Google or Tenstorrent release. It is not GGUF, AWQ, GPTQ, bitsandbytes, or an ordinary Transformers checkpoint. `transformers.from_pretrained()`, llama.cpp and hosted Inference Providers cannot load it. It needs the bundled runtime.
24
+
25
+ ## Quantization
26
+
27
+ All-BFP8: every text matrix (attention, MLP, LM head) is stored as TTNN `bfloat8_b`. The embedding is BF16; 1-D tensors (norms, layer scalars) are lossless BF16 in `small.safetensors`. See [`gemma-4-12B-it/precision_plan.json`](gemma-4-12B-it/precision_plan.json) and [`gemma-4-12B-it/native_manifest.json`](gemma-4-12B-it/native_manifest.json).
28
+
29
+ `tensors/tensor_cache_bf16/*.tensorbin` holds the exact TTNN host tensors that the runtime uploads to device DRAM. The directory name comes from the runtime's tensor-cache writer. The files are not a regenerable cache: they are the checkpoint. Every file is hashed in the manifest, and the loader refuses to start if any file is missing or changed. The drafter (`gemma-4-12B-it-assistant/`) stores its tensors the same way. Its recorded dtype is `bfp8-lm4`: BFP8 weights and a BFP4 drafter LM head.
30
+
31
+ The loader requires two proofs, both bound to the manifests:
32
+
33
+ - [`gemma-4-12B-it/equivalence.json`](gemma-4-12B-it/equivalence.json): the full logits of the native reload are bit-exact to quantizing the original weights at load time, over 4,093 teacher-forced tokens (`exact_logits_equal: true`).
34
+ - [`gemma-4-12B-it-assistant/spec_equivalence.json`](gemma-4-12B-it-assistant/spec_equivalence.json): greedy speculative decoding with draft length 5 produced token streams identical to non-speculative greedy decoding on the same proven runtime (`streams_identical: true`, 3,130 tokens compared).
35
+
36
+ ## Quality versus the original BF16 model
37
+
38
+ Teacher-forced comparison against the original model run in BF16 on CPU ([`evidence/quality/quant_table.json`](evidence/quality/quant_table.json)). dPPL is the perplexity change relative to BF16, top-1 is argmax agreement with BF16, and KL is the mean KL divergence on the short-prompt set:
39
+
40
+ | Plan | Weight read | chat dPPL / top-1 | code dPPL / top-1 | 16K book dPPL / top-1 | KL | GSM8K-50 |
41
+ |---|---:|---:|---:|---:|---:|---:|
42
+ | **all-BFP8 (this release)** | 12.7 GB | **−0.3% / 97.1%** | **+1.0% / 98.3%** | **+1.6% / 97.2%** | **0.009** | **50/50** |
43
+ | BFP8 + BFP4 LM head | 12.2 GB | +0.6% / 95.5% | +2.0% / 97.0% | +2.6% / 96.1% | 0.019 | 49/50 |
44
+ | MLP in BFP4 | 8.4 GB | +5.2% / 87.1% | +3.9% / 90.8% | +9.9% / 87.7% | 0.175 | 49/50 |
45
+ | all-BFP4 | 7.2 GB | +7.2% / 82.7% | +8.3% / 87.4% | +19.2% / 80.1% | 0.296 | — |
46
+
47
+ **Why BFP8 and not BFP4:** with the MLP in BFP4, perplexity rises about 5% and top-1 agreement falls to about 87%. All-BFP4 is worse still. BFP8 stays within about 1–2% of BF16 perplexity.
48
+
49
+ The shipped runtime adds batched speculative verification and a fused GELU×up kernel. It passed a decode-path quality gate against the exact (unfused) build ([`evidence/quality/runtime-gate-dtf-score.json`](evidence/quality/runtime-gate-dtf-score.json)): argmax agreement with the exact build was 98.8% on chat, 98.8% on code and 98.4% on long-book text. The NLL change versus the exact build was +0.0017 ± 0.0012 on chat, −0.0004 ± 0.0030 on code and −0.0061 ± 0.0096 on long-book. The two proofs above were recorded with this shipped runtime configuration.
50
+
51
+ ## Performance
52
+
53
+ These numbers were measured over HTTP on one P150 with greedy decoding, 512-token outputs, chat prompts, one user and speculative decoding on ([`evidence/http-p3c/`](evidence/http-p3c/)):
54
+
55
+ | Prompt tokens | 128 | 2K | 8K | 32K | 131K | 262K (261,632) |
56
+ |---|---:|---:|---:|---:|---:|---:|
57
+ | Decode tok/s | 50.6 | 49.3 | 46.9 | 45.1 | 29.7 | 21.8 |
58
+ | TTFT (s) | 0.09 | 0.64 | 3.2 | 17.0 | 124 | 395 |
59
+
60
+ With the LocalMaxxing official prompt (a local run of the official protocol, not submitted), the server reached **49.9 output tok/s** with a **95 ms** TTFT (median of 5, 256 output tokens; [`evidence/localmaxxing/speed-test.json`](evidence/localmaxxing/speed-test.json)).
61
+
62
+ ## Serving correctness
63
+
64
+ - **HTTP matches the direct runtime token for token.** All 28 of 28 checks passed ([`evidence/http-p3c/exact.json`](evidence/http-p3c/exact.json)). They cover chat and completion outputs on 10 prompts, request isolation, sampled requests followed by greedy ones, cancellation mid-prefill and mid-decode, and rejection of image input. The 512-token outputs at 32K and 261,632 prompt tokens were also identical to the direct runtime ([`evidence/http-p3c/tput_vs_direct.json`](evidence/http-p3c/tput_vs_direct.json)).
65
+ - **Long context works to the native limit.** Passkeys at the beginning, middle and end of the prompt were retrieved at 32,768, 131,072 and 262,016 prompt tokens. A request totalling 262,145 tokens was rejected with HTTP 400 before generation started ([`evidence/http-p3c/long.json`](evidence/http-p3c/long.json)).
66
+ - **The public package was verified end to end.** It was downloaded fresh and served on a P150 ([`evidence/public-download-verification.json`](evidence/public-download-verification.json), [`evidence/public-serving-smoke.json`](evidence/public-serving-smoke.json)).
67
+
68
+ ## Download and serve
69
+
70
+ Prerequisites:
71
+
72
+ - Linux with a supported P150 driver and hugepage mounts at `/dev/hugepages` and `/dev/hugepages-1G`.
73
+ - Docker and Python 3.11 or newer.
74
+ - Exclusive ownership of one P150.
75
+ - Disk space for about 21.3 GB of downloads: a 14.7 GB checkpoint, a 1.2 GB drafter and a 5.4 GB compressed runtime archive. You also need room for the loaded image and a writable kernel cache.
76
+
77
+ Download the standard-library launcher:
78
+
79
+ ```bash
80
+ curl --fail --location --output launch.py \
81
+ https://huggingface.co/Lottolabs/gemma-4-12B-it-TT-BFP8-P150/resolve/main/launch.py
82
+
83
+ python3 launch.py \
84
+ --cache-root "$HOME/.cache/gemma4-12b-tt-native" \
85
+ --device-ownership-confirmed
86
+ ```
87
+
88
+ The launcher does the following before it serves anything:
89
+
90
+ 1. Resolves the requested Hugging Face revision to an immutable commit.
91
+ 2. Downloads only the serving package (the checkpoint, drafter, server script and runtime archive) and verifies every byte against `runtime-release.json`.
92
+ 3. Checks that the published file sets match both native manifests.
93
+ 4. Loads the checksum-pinned runtime archive into Docker and checks the resulting image ID.
94
+
95
+ `serve_native.py` then verifies the manifests, file hashes and both proofs again, checks the image ID once more, and starts vLLM. Inside the container, the loader repeats the checks against the image's runtime identity.
96
+
97
+ The original BF16 weights are not needed. The API is OpenAI-compatible, binds to `127.0.0.1:8000`, and serves the model name `google/gemma-4-12B-it`. The first launch compiles kernels into the kernel cache, so it takes longer than later launches.
98
+
99
+ Useful modes:
100
+
101
+ ```bash
102
+ # Download and verify without Docker or a device
103
+ python3 launch.py --cache-root /large-disk/gemma4-12b --download-only
104
+
105
+ # Print the exact serving command without starting it
106
+ python3 launch.py --cache-root /large-disk/gemma4-12b --print-command
107
+
108
+ # Non-speculative serving (no drafter)
109
+ python3 launch.py --cache-root /large-disk/gemma4-12b --speculation off --device-ownership-confirmed
110
+ ```
111
+
112
+ Save the immutable revision that the launcher prints, and pass it with `--revision` to reproduce the same setup exactly.
113
+
114
+ ```bash
115
+ curl http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
116
+ "model": "google/gemma-4-12B-it", "temperature": 0, "max_tokens": 256,
117
+ "messages": [{"role": "user", "content": "Explain database transactions in three sentences."}]}'
118
+ ```
119
+
120
+ ## Limits
121
+
122
+ - One user at a time (`max_num_seqs 1`), one P150 (MeshShape 1x1), no multi-device support.
123
+ - Text only. Requests with images are rejected.
124
+ - Context up to 262,144 tokens in total. Sliding-window layers keep their KV cache as 18,432-token rings. Prefill runs in 16,384-token chunks, and prefix caching is off.
125
+ - Only greedy requests (`temperature: 0`) decode speculatively (drafter K=5, exact per-row verification). Requests with other sampling settings are served but decode one token per step with vLLM's host sampler, so they are slower.
126
+ - Long prompts have a long time to first token: about 17 s at 32K and about 6.6 min at 262K.
127
+
128
+ ## Integrity and provenance
129
+
130
+ - [`runtime-release.json`](runtime-release.json): the serving inventory the launcher downloads, the runtime archive checksum, the expected Docker image ID and the upstream revisions.
131
+ - [`release-manifest.json`](release-manifest.json): the complete repository inventory with size and sha256 for every file.
132
+ - [`SHA256SUMS`](SHA256SUMS): file checksums.
133
+ - [`reproduction.json`](reproduction.json): source, runtime, serving and evidence contract.
134
+ - [`runtime/`](runtime/): source of the runtime image layers, recorded in [`runtime/source.json`](runtime/source.json). It contains the Gemma 4 TT model tree (`runtime/gemma4/`, git commit `ce67960382fffed8550b38717cedf574e821da06`, byte-identical to the tree in the image), the vLLM/TT-plugin serving overlay, the TTNN matmul overlay, the Dockerfile chain and the build script. This is provenance: serving uses the checksum-pinned `runtime-image.tar.gz`.
135
+ - [`native_checkpoint.py`](native_checkpoint.py) and [`provenance/build_native_checkpoint.py`](provenance/build_native_checkpoint.py): the manifest, proof and verification tool, and the checkpoint builder.
136
+
137
+ ## License and attribution
138
+
139
+ Derived from [google/gemma-4-12B-it](https://huggingface.co/google/gemma-4-12B-it) and [google/gemma-4-12B-it-assistant](https://huggingface.co/google/gemma-4-12B-it-assistant). Apache-2.0; see [`LICENSE`](LICENSE). The runtime includes tt-metal and vLLM, which keep their upstream notices; see [`runtime/licenses/`](runtime/licenses/).
SHA256SUMS ADDED
The diff for this file is too large to render. See raw diff
 
evidence/http-p3c/exact.json ADDED
@@ -0,0 +1,185 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "checks": [
3
+ {
4
+ "name": "health",
5
+ "ok": true
6
+ },
7
+ {
8
+ "name": "chat_exact_story",
9
+ "ok": true,
10
+ "n": 250,
11
+ "ref_n": 250,
12
+ "finish": "stop",
13
+ "events": 104,
14
+ "first_divergence": null
15
+ },
16
+ {
17
+ "name": "chat_exact_explain",
18
+ "ok": true,
19
+ "n": 320,
20
+ "ref_n": 320,
21
+ "finish": "length",
22
+ "events": 109,
23
+ "first_divergence": null
24
+ },
25
+ {
26
+ "name": "chat_exact_code",
27
+ "ok": true,
28
+ "n": 320,
29
+ "ref_n": 320,
30
+ "finish": "length",
31
+ "events": 102,
32
+ "first_divergence": null
33
+ },
34
+ {
35
+ "name": "chat_exact_math",
36
+ "ok": true,
37
+ "n": 320,
38
+ "ref_n": 320,
39
+ "finish": "length",
40
+ "events": 66,
41
+ "first_divergence": null
42
+ },
43
+ {
44
+ "name": "chat_exact_logic",
45
+ "ok": true,
46
+ "n": 320,
47
+ "ref_n": 320,
48
+ "finish": "length",
49
+ "events": 96,
50
+ "first_divergence": null
51
+ },
52
+ {
53
+ "name": "chat_exact_code-lru",
54
+ "ok": true,
55
+ "n": 320,
56
+ "ref_n": 320,
57
+ "finish": "length",
58
+ "events": 86,
59
+ "first_divergence": null
60
+ },
61
+ {
62
+ "name": "chat_exact_code-c",
63
+ "ok": true,
64
+ "n": 320,
65
+ "ref_n": 320,
66
+ "finish": "length",
67
+ "events": 74,
68
+ "first_divergence": null
69
+ },
70
+ {
71
+ "name": "chat_exact_code-sql",
72
+ "ok": true,
73
+ "n": 320,
74
+ "ref_n": 320,
75
+ "finish": "length",
76
+ "events": 76,
77
+ "first_divergence": null
78
+ },
79
+ {
80
+ "name": "chat_exact_long-summary",
81
+ "ok": true,
82
+ "n": 320,
83
+ "ref_n": 320,
84
+ "finish": "length",
85
+ "events": 108,
86
+ "first_divergence": null
87
+ },
88
+ {
89
+ "name": "chat_exact_long-code",
90
+ "ok": true,
91
+ "n": 320,
92
+ "ref_n": 320,
93
+ "finish": "length",
94
+ "events": 108,
95
+ "first_divergence": null
96
+ },
97
+ {
98
+ "name": "completion_exact_story",
99
+ "ok": true,
100
+ "n": 250
101
+ },
102
+ {
103
+ "name": "completion_exact_explain",
104
+ "ok": true,
105
+ "n": 320
106
+ },
107
+ {
108
+ "name": "completion_exact_code",
109
+ "ok": true,
110
+ "n": 320
111
+ },
112
+ {
113
+ "name": "completion_exact_math",
114
+ "ok": true,
115
+ "n": 320
116
+ },
117
+ {
118
+ "name": "completion_exact_logic",
119
+ "ok": true,
120
+ "n": 320
121
+ },
122
+ {
123
+ "name": "completion_exact_code-lru",
124
+ "ok": true,
125
+ "n": 320
126
+ },
127
+ {
128
+ "name": "completion_exact_code-c",
129
+ "ok": true,
130
+ "n": 320
131
+ },
132
+ {
133
+ "name": "completion_exact_code-sql",
134
+ "ok": true,
135
+ "n": 320
136
+ },
137
+ {
138
+ "name": "completion_exact_long-summary",
139
+ "ok": true,
140
+ "n": 320
141
+ },
142
+ {
143
+ "name": "completion_exact_long-code",
144
+ "ok": true,
145
+ "n": 320
146
+ },
147
+ {
148
+ "name": "isolation_ABA",
149
+ "ok": true,
150
+ "n": 128
151
+ },
152
+ {
153
+ "name": "sampled_request",
154
+ "ok": true,
155
+ "n": 64
156
+ },
157
+ {
158
+ "name": "greedy_after_sampled",
159
+ "ok": true
160
+ },
161
+ {
162
+ "name": "cancel_mid_decode",
163
+ "ok": true,
164
+ "cancelled_after": 42,
165
+ "next_request_s": 3.266689165000571
166
+ },
167
+ {
168
+ "name": "cancel_mid_prefill",
169
+ "ok": true,
170
+ "cancelled_after_s": 8.170486249990063,
171
+ "next_request_s": 11.995801645010943
172
+ },
173
+ {
174
+ "name": "image_rejected",
175
+ "ok": true,
176
+ "status": 400,
177
+ "body": "{\"error\":{\"message\":\"/model is not a multimodal model\",\"type\":\"BadRequestError\",\"param\":null,\"code\":400}}"
178
+ },
179
+ {
180
+ "name": "health_end",
181
+ "ok": true
182
+ }
183
+ ],
184
+ "exact_seconds": 110.89393173001008
185
+ }
evidence/http-p3c/long.json ADDED
@@ -0,0 +1,231 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "name": "configuration",
4
+ "url": "http://127.0.0.1:8000",
5
+ "model": "google/gemma-4-12B-it",
6
+ "contexts": [
7
+ 32768,
8
+ 131072,
9
+ 262016
10
+ ],
11
+ "tokens": 128,
12
+ "native_max_len": 262144,
13
+ "request_timeout_seconds": 3600,
14
+ "cases": [
15
+ {
16
+ "input_tokens": 32768,
17
+ "output_tokens": 128
18
+ },
19
+ {
20
+ "input_tokens": 131072,
21
+ "output_tokens": 128
22
+ },
23
+ {
24
+ "input_tokens": 262016,
25
+ "output_tokens": 128
26
+ }
27
+ ],
28
+ "begin_key_token_region": [
29
+ 0,
30
+ 43
31
+ ],
32
+ "end_key_suffix_tokens": 65,
33
+ "timing_basis": "Client first/last non-empty SSE text arrival; MTP may emit multiple tokens per event"
34
+ },
35
+ {
36
+ "name": "short_before",
37
+ "input_tokens": 25,
38
+ "requested_output_tokens": 32,
39
+ "requested_total_tokens": 57,
40
+ "ttft_seconds": 0.09288164801546372,
41
+ "decode_wall_seconds": 0.054585256992140785,
42
+ "seconds": 0.20159582499763928,
43
+ "decode_tokens_per_second": 128.2397553062334,
44
+ "finish_reason": "stop",
45
+ "sse_done": true,
46
+ "response": {
47
+ "text": "The capital of France is Paris.",
48
+ "usage": {
49
+ "prompt_tokens": 25,
50
+ "total_tokens": 33,
51
+ "completion_tokens": 8
52
+ }
53
+ },
54
+ "expected_keys": {
55
+ "known_answer": "Paris"
56
+ },
57
+ "retrieved_keys": {
58
+ "known_answer": true
59
+ },
60
+ "accuracy": true
61
+ },
62
+ {
63
+ "name": "layout_32768",
64
+ "input_tokens": 32768,
65
+ "begin_key_offset": 35,
66
+ "middle_key_offset": 16021,
67
+ "end_key_offset": 32707
68
+ },
69
+ {
70
+ "name": "stream_32768_128",
71
+ "input_tokens": 32768,
72
+ "requested_output_tokens": 128,
73
+ "requested_total_tokens": 32896,
74
+ "ttft_seconds": 16.92090609599836,
75
+ "decode_wall_seconds": 0.47541975299827754,
76
+ "seconds": 17.39640123900608,
77
+ "decode_tokens_per_second": 71.51574957829567,
78
+ "finish_reason": "stop",
79
+ "sse_done": true,
80
+ "response": {
81
+ "text": "CEDAR-7419-QUARTZ\nHARBOR-5096-VIOLET\nORBIT-2863-MARBLE",
82
+ "usage": {
83
+ "prompt_tokens": 32768,
84
+ "total_tokens": 32803,
85
+ "completion_tokens": 35
86
+ }
87
+ },
88
+ "expected_keys": {
89
+ "begin": "CEDAR-7419-QUARTZ",
90
+ "middle": "HARBOR-5096-VIOLET",
91
+ "end": "ORBIT-2863-MARBLE"
92
+ },
93
+ "retrieved_keys": {
94
+ "begin": true,
95
+ "middle": true,
96
+ "end": true
97
+ },
98
+ "accuracy": true
99
+ },
100
+ {
101
+ "name": "layout_131072",
102
+ "input_tokens": 131072,
103
+ "begin_key_offset": 35,
104
+ "middle_key_offset": 65832,
105
+ "end_key_offset": 131011
106
+ },
107
+ {
108
+ "name": "stream_131072_128",
109
+ "input_tokens": 131072,
110
+ "requested_output_tokens": 128,
111
+ "requested_total_tokens": 131200,
112
+ "ttft_seconds": 123.78426550599397,
113
+ "decode_wall_seconds": 0.6346804780187085,
114
+ "seconds": 124.41902046301402,
115
+ "decode_tokens_per_second": 53.57026279764946,
116
+ "finish_reason": "stop",
117
+ "sse_done": true,
118
+ "response": {
119
+ "text": "CEDAR-7419-QUARTZ\nHARBOR-5096-VIOLET\nORBIT-2863-MARBLE",
120
+ "usage": {
121
+ "prompt_tokens": 131072,
122
+ "total_tokens": 131107,
123
+ "completion_tokens": 35
124
+ }
125
+ },
126
+ "expected_keys": {
127
+ "begin": "CEDAR-7419-QUARTZ",
128
+ "middle": "HARBOR-5096-VIOLET",
129
+ "end": "ORBIT-2863-MARBLE"
130
+ },
131
+ "retrieved_keys": {
132
+ "begin": true,
133
+ "middle": true,
134
+ "end": true
135
+ },
136
+ "accuracy": true
137
+ },
138
+ {
139
+ "name": "layout_262016",
140
+ "input_tokens": 262016,
141
+ "begin_key_offset": 35,
142
+ "middle_key_offset": 130632,
143
+ "end_key_offset": 261955
144
+ },
145
+ {
146
+ "name": "stream_262016_128",
147
+ "input_tokens": 262016,
148
+ "requested_output_tokens": 128,
149
+ "requested_total_tokens": 262144,
150
+ "ttft_seconds": 394.87144183402415,
151
+ "decode_wall_seconds": 0.9865862159931567,
152
+ "seconds": 395.98883561600815,
153
+ "decode_tokens_per_second": 34.462269438635495,
154
+ "finish_reason": "stop",
155
+ "sse_done": true,
156
+ "response": {
157
+ "text": "CEDAR-7419-QUARTZ\nHARBOR-5096-VIOLET\nORBIT-2863-MARBLE",
158
+ "usage": {
159
+ "prompt_tokens": 262016,
160
+ "total_tokens": 262051,
161
+ "completion_tokens": 35
162
+ }
163
+ },
164
+ "expected_keys": {
165
+ "begin": "CEDAR-7419-QUARTZ",
166
+ "middle": "HARBOR-5096-VIOLET",
167
+ "end": "ORBIT-2863-MARBLE"
168
+ },
169
+ "retrieved_keys": {
170
+ "begin": true,
171
+ "middle": true,
172
+ "end": true
173
+ },
174
+ "accuracy": true
175
+ },
176
+ {
177
+ "name": "native_overflow_rejection",
178
+ "input_tokens": 262016,
179
+ "requested_output_tokens": 129,
180
+ "requested_total_tokens": 262145,
181
+ "native_max_len": 262144,
182
+ "status": 400,
183
+ "seconds": 0.014193854993209243,
184
+ "rejected_before_generation": true,
185
+ "response": {
186
+ "error": {
187
+ "message": "You passed 262016 input tokens and requested 129 output tokens. However, the model's context length is only 262144 tokens, resulting in a maximum input length of 262015 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=262016)",
188
+ "type": "BadRequestError",
189
+ "param": "input_tokens",
190
+ "code": 400
191
+ }
192
+ }
193
+ },
194
+ {
195
+ "name": "mtp_acceptance",
196
+ "accepted_tokens_before": 10883.0,
197
+ "accepted_tokens_after": 10964.0,
198
+ "accepted_tokens_delta": 81.0
199
+ },
200
+ {
201
+ "name": "short_after",
202
+ "input_tokens": 25,
203
+ "requested_output_tokens": 32,
204
+ "requested_total_tokens": 57,
205
+ "ttft_seconds": 0.0965569649997633,
206
+ "decode_wall_seconds": 0.05454791599186137,
207
+ "seconds": 0.20529347998672165,
208
+ "decode_tokens_per_second": 128.32754235825269,
209
+ "finish_reason": "stop",
210
+ "sse_done": true,
211
+ "response": {
212
+ "text": "The capital of France is Paris.",
213
+ "usage": {
214
+ "prompt_tokens": 25,
215
+ "total_tokens": 33,
216
+ "completion_tokens": 8
217
+ }
218
+ },
219
+ "expected_keys": {
220
+ "known_answer": "Paris"
221
+ },
222
+ "retrieved_keys": {
223
+ "known_answer": true
224
+ },
225
+ "accuracy": true
226
+ },
227
+ {
228
+ "name": "state_isolation",
229
+ "exact_text_match": true
230
+ }
231
+ ]
evidence/http-p3c/tput-128/requests-128-c1.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "index": 0,
4
+ "start_seconds": 1.4730001566931605e-05,
5
+ "end_seconds": 10.182251366000855,
6
+ "ttft_seconds": 0.09181618399452418,
7
+ "latency_seconds": 10.182236635999288,
8
+ "decode_tokens_per_second": 50.64247901615711,
9
+ "text_sha256": "63c80dca9b9d4cea4f271acb863cf1aa059fb2b80e4fdb79796fb57ba73337e3",
10
+ "usage": {
11
+ "prompt_tokens": 128,
12
+ "total_tokens": 640,
13
+ "completion_tokens": 512
14
+ },
15
+ "error": null,
16
+ "text": "# Understanding Database Transactions: The Foundation of Data Integrity\n\nIn the world of software engineering and data management, the reliability of information is paramount. Whether you are transferring money between bank accounts, booking a flight, or updating an inventory system, the underlying database must ensure that data remains accurate, consistent, and reliable even in the face of system crashes, power failures, or concurrent access by thousands of users. This reliability is achieved through a mechanism known as a **Database Transaction**.\n\nA transaction is a logical unit of work that consists of one or more database operations (such as Insert, Update, Delete, or Select). To be considered valid, a transaction must be treated as a single, indivisible entity: either every operation within the transaction succeeds, or none of them do.\n\nTo define the behavior of these transactions, the database community relies on a set of four fundamental properties known as **ACID**.\n\n---\n\n### The ACID Properties\n\nTo understand transactions in detail, we must break down the ACID acronym, which serves as the gold standard for database reliability.\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as a single \"atom.\" If a transaction contains five different SQL statements, and the fourth one fails (due to a constraint violation, a lost connection, or a crash), the database must \"roll back\" the first three statements. The database is returned to the exact state it was in before the transaction started.\n\n* **Example:** Imagine a bank transfer where User A sends $100 to User B. This involves two steps:\n 1. Subtract $100 from User A\u2019s balance.\n 2. Add $100 to User B\u2019s balance.\n If the system crashes after step 1 but before step 2, User A loses money, but User B never receives it. Atomicity prevents this by ensuring that if step 2 fails, step 1 is undone automatically.\n\n#### 2. Consistency (Preserving Rules)\nConsistency ensures that a transaction brings the database from one valid state to another valid state, maintaining all predefined rules, including constraints, foreign keys, and triggers. A transaction cannot leave the database in a \"broken\" state where data violates the business logic.\n\n* **Example:** Suppose a database has a rule that a \"Product Price\" cannot be a negative number. If a transaction attempts to update a price to -$50, the database will detect this violation of the"
17
+ },
18
+ {
19
+ "index": 1,
20
+ "start_seconds": 10.182313314988278,
21
+ "end_seconds": 20.363098062982317,
22
+ "ttft_seconds": 0.09173648501746356,
23
+ "latency_seconds": 10.18078474799404,
24
+ "decode_tokens_per_second": 50.64928144502961,
25
+ "text_sha256": "63c80dca9b9d4cea4f271acb863cf1aa059fb2b80e4fdb79796fb57ba73337e3",
26
+ "usage": {
27
+ "prompt_tokens": 128,
28
+ "total_tokens": 640,
29
+ "completion_tokens": 512
30
+ },
31
+ "error": null,
32
+ "text": "# Understanding Database Transactions: The Foundation of Data Integrity\n\nIn the world of software engineering and data management, the reliability of information is paramount. Whether you are transferring money between bank accounts, booking a flight, or updating an inventory system, the underlying database must ensure that data remains accurate, consistent, and reliable even in the face of system crashes, power failures, or concurrent access by thousands of users. This reliability is achieved through a mechanism known as a **Database Transaction**.\n\nA transaction is a logical unit of work that consists of one or more database operations (such as Insert, Update, Delete, or Select). To be considered valid, a transaction must be treated as a single, indivisible entity: either every operation within the transaction succeeds, or none of them do.\n\nTo define the behavior of these transactions, the database community relies on a set of four fundamental properties known as **ACID**.\n\n---\n\n### The ACID Properties\n\nTo understand transactions in detail, we must break down the ACID acronym, which serves as the gold standard for database reliability.\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as a single \"atom.\" If a transaction contains five different SQL statements, and the fourth one fails (due to a constraint violation, a lost connection, or a crash), the database must \"roll back\" the first three statements. The database is returned to the exact state it was in before the transaction started.\n\n* **Example:** Imagine a bank transfer where User A sends $100 to User B. This involves two steps:\n 1. Subtract $100 from User A\u2019s balance.\n 2. Add $100 to User B\u2019s balance.\n If the system crashes after step 1 but before step 2, User A loses money, but User B never receives it. Atomicity prevents this by ensuring that if step 2 fails, step 1 is undone automatically.\n\n#### 2. Consistency (Preserving Rules)\nConsistency ensures that a transaction brings the database from one valid state to another valid state, maintaining all predefined rules, including constraints, foreign keys, and triggers. A transaction cannot leave the database in a \"broken\" state where data violates the business logic.\n\n* **Example:** Suppose a database has a rule that a \"Product Price\" cannot be a negative number. If a transaction attempts to update a price to -$50, the database will detect this violation of the"
33
+ }
34
+ ]
evidence/http-p3c/tput-128/summary.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "google/gemma-4-12B-it",
3
+ "output_tokens": 512,
4
+ "method": "Closed loop; synchronized initial clients; max(min_requests,2*concurrency) requests; nearest-rank percentiles; aggregate includes prefill and queue drain",
5
+ "cases": [
6
+ {
7
+ "input_tokens": 128,
8
+ "concurrency": 1,
9
+ "requests": 2,
10
+ "successes": 2,
11
+ "errors": 0,
12
+ "wall_seconds": 20.36308333298075,
13
+ "matching_single_request_outputs": 2,
14
+ "aggregate_output_tokens_per_second": 50.28707996993237,
15
+ "requests_per_second": 0.09821695306627416,
16
+ "ttft_seconds_p50": 0.09173648501746356,
17
+ "ttft_seconds_p95": 0.09181618399452418,
18
+ "latency_seconds_p50": 10.18078474799404,
19
+ "latency_seconds_p95": 10.182236635999288,
20
+ "decode_tokens_per_second_p50": 50.64247901615711,
21
+ "decode_tokens_per_second_p95": 50.64928144502961
22
+ }
23
+ ]
24
+ }
evidence/http-p3c/tput-128/warmup-128.json ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "index": -1,
3
+ "start_seconds": 7.700000423938036e-07,
4
+ "end_seconds": 10.291406103991903,
5
+ "ttft_seconds": 0.09075412299716845,
6
+ "latency_seconds": 10.291405333991861,
7
+ "decode_tokens_per_second": 50.09516148932038,
8
+ "text_sha256": "63c80dca9b9d4cea4f271acb863cf1aa059fb2b80e4fdb79796fb57ba73337e3",
9
+ "usage": {
10
+ "prompt_tokens": 128,
11
+ "total_tokens": 640,
12
+ "completion_tokens": 512
13
+ },
14
+ "error": null,
15
+ "text": "# Understanding Database Transactions: The Foundation of Data Integrity\n\nIn the world of software engineering and data management, the reliability of information is paramount. Whether you are transferring money between bank accounts, booking a flight, or updating an inventory system, the underlying database must ensure that data remains accurate, consistent, and reliable even in the face of system crashes, power failures, or concurrent access by thousands of users. This reliability is achieved through a mechanism known as a **Database Transaction**.\n\nA transaction is a logical unit of work that consists of one or more database operations (such as Insert, Update, Delete, or Select). To be considered valid, a transaction must be treated as a single, indivisible entity: either every operation within the transaction succeeds, or none of them do.\n\nTo define the behavior of these transactions, the database community relies on a set of four fundamental properties known as **ACID**.\n\n---\n\n### The ACID Properties\n\nTo understand transactions in detail, we must break down the ACID acronym, which serves as the gold standard for database reliability.\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as a single \"atom.\" If a transaction contains five different SQL statements, and the fourth one fails (due to a constraint violation, a lost connection, or a crash), the database must \"roll back\" the first three statements. The database is returned to the exact state it was in before the transaction started.\n\n* **Example:** Imagine a bank transfer where User A sends $100 to User B. This involves two steps:\n 1. Subtract $100 from User A\u2019s balance.\n 2. Add $100 to User B\u2019s balance.\n If the system crashes after step 1 but before step 2, User A loses money, but User B never receives it. Atomicity prevents this by ensuring that if step 2 fails, step 1 is undone automatically.\n\n#### 2. Consistency (Preserving Rules)\nConsistency ensures that a transaction brings the database from one valid state to another valid state, maintaining all predefined rules, including constraints, foreign keys, and triggers. A transaction cannot leave the database in a \"broken\" state where data violates the business logic.\n\n* **Example:** Suppose a database has a rule that a \"Product Price\" cannot be a negative number. If a transaction attempts to update a price to -$50, the database will detect this violation of the"
16
+ }
evidence/http-p3c/tput-131072/requests-131072-c1.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "index": 0,
4
+ "start_seconds": 3.7209014408290386e-05,
5
+ "end_seconds": 140.95895639900118,
6
+ "ttft_seconds": 123.78377083898522,
7
+ "latency_seconds": 140.95891918998677,
8
+ "decode_tokens_per_second": 29.752375978499117,
9
+ "text_sha256": "3c3c3d343fcdd17e18b186d6c873612d183159c8add64b188f1e1594b957a4f6",
10
+ "usage": {
11
+ "prompt_tokens": 131072,
12
+ "total_tokens": 131584,
13
+ "completion_tokens": 512
14
+ },
15
+ "error": null,
16
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is the fundamental unit of work. It is a sequence of one or more operations performed as a single logical unit. The core philosophy of a transaction is \"all or nothing\": either every operation within the transaction succeeds and the database is updated, or if even one operation fails, the entire transaction is undone, leaving the database in its original state.\n\nTo understand why this is critical, imagine an ATM withdrawal. You want to take out $100. The system must perform two distinct actions:\n1. Subtract $100 from your account balance.\n2. Dispense $100 in cash.\n\nIf the system subtracts the money but the machine fails to dispense the cash, you lose your money. If the machine dispenses the cash but the system fails to update your balance, the bank loses money. A transaction ensures that these two steps are bound together so that one cannot happen without the other.\n\n---\n\n### The ACID Properties: The Gold Standard of Transactions\n\nFor a database to be considered reliable, it must adhere to the **ACID** properties. These four principles ensure that data remains accurate and consistent, even in the event of system crashes, power failures, or concurrent access by thousands of users.\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity guarantees that a transaction is treated as a single, indivisible unit. If a transaction consists of five different SQL queries, and the fourth query fails due to a constraint violation or a lost connection, the first three queries are \"rolled back.\" The database reverts to the state it was in before the transaction started.\n\n* **Example:** In a hotel booking system, a transaction involves:\n * Checking room availability.\n * Creating a reservation record.\n * Charging the customer's credit card.\n * Sending a confirmation email.\n If the credit card charge fails, the reservation record must be deleted, and the room must be marked as available again. Atomicity prevents \"ghost\" reservations where a room is held but no payment was received.\n\n#### 2. Consistency (Rules of the Game)\nConsistency ensures that a transaction brings the database from one valid state to another. Every database has predefined rules (constraints), such as \"account balances cannot be negative\" or \"every order must have a valid Customer ID.\" A transaction must follow these rules. If"
17
+ },
18
+ {
19
+ "index": 1,
20
+ "start_seconds": 140.95899165899027,
21
+ "end_seconds": 281.9431601790129,
22
+ "ttft_seconds": 123.80123381401063,
23
+ "latency_seconds": 140.98416852002265,
24
+ "decode_tokens_per_second": 29.738932975990465,
25
+ "text_sha256": "3c3c3d343fcdd17e18b186d6c873612d183159c8add64b188f1e1594b957a4f6",
26
+ "usage": {
27
+ "prompt_tokens": 131072,
28
+ "total_tokens": 131584,
29
+ "completion_tokens": 512
30
+ },
31
+ "error": null,
32
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is the fundamental unit of work. It is a sequence of one or more operations performed as a single logical unit. The core philosophy of a transaction is \"all or nothing\": either every operation within the transaction succeeds and the database is updated, or if even one operation fails, the entire transaction is undone, leaving the database in its original state.\n\nTo understand why this is critical, imagine an ATM withdrawal. You want to take out $100. The system must perform two distinct actions:\n1. Subtract $100 from your account balance.\n2. Dispense $100 in cash.\n\nIf the system subtracts the money but the machine fails to dispense the cash, you lose your money. If the machine dispenses the cash but the system fails to update your balance, the bank loses money. A transaction ensures that these two steps are bound together so that one cannot happen without the other.\n\n---\n\n### The ACID Properties: The Gold Standard of Transactions\n\nFor a database to be considered reliable, it must adhere to the **ACID** properties. These four principles ensure that data remains accurate and consistent, even in the event of system crashes, power failures, or concurrent access by thousands of users.\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity guarantees that a transaction is treated as a single, indivisible unit. If a transaction consists of five different SQL queries, and the fourth query fails due to a constraint violation or a lost connection, the first three queries are \"rolled back.\" The database reverts to the state it was in before the transaction started.\n\n* **Example:** In a hotel booking system, a transaction involves:\n * Checking room availability.\n * Creating a reservation record.\n * Charging the customer's credit card.\n * Sending a confirmation email.\n If the credit card charge fails, the reservation record must be deleted, and the room must be marked as available again. Atomicity prevents \"ghost\" reservations where a room is held but no payment was received.\n\n#### 2. Consistency (Rules of the Game)\nConsistency ensures that a transaction brings the database from one valid state to another. Every database has predefined rules (constraints), such as \"account balances cannot be negative\" or \"every order must have a valid Customer ID.\" A transaction must follow these rules. If"
33
+ }
34
+ ]
evidence/http-p3c/tput-131072/summary.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "google/gemma-4-12B-it",
3
+ "output_tokens": 512,
4
+ "method": "Closed loop; synchronized initial clients; max(min_requests,2*concurrency) requests; nearest-rank percentiles; aggregate includes prefill and queue drain",
5
+ "cases": [
6
+ {
7
+ "input_tokens": 131072,
8
+ "concurrency": 1,
9
+ "requests": 2,
10
+ "successes": 2,
11
+ "errors": 0,
12
+ "wall_seconds": 281.9431229699985,
13
+ "matching_single_request_outputs": 2,
14
+ "aggregate_output_tokens_per_second": 3.6319382051711315,
15
+ "requests_per_second": 0.007093629306974866,
16
+ "ttft_seconds_p50": 123.78377083898522,
17
+ "ttft_seconds_p95": 123.80123381401063,
18
+ "latency_seconds_p50": 140.95891918998677,
19
+ "latency_seconds_p95": 140.98416852002265,
20
+ "decode_tokens_per_second_p50": 29.738932975990465,
21
+ "decode_tokens_per_second_p95": 29.752375978499117
22
+ }
23
+ ]
24
+ }
evidence/http-p3c/tput-131072/warmup-131072.json ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "index": -1,
3
+ "start_seconds": 1.1700030881911516e-06,
4
+ "end_seconds": 143.83761601298465,
5
+ "ttft_seconds": 126.47853563597891,
6
+ "latency_seconds": 143.83761484298157,
7
+ "decode_tokens_per_second": 29.43717637800883,
8
+ "text_sha256": "3c3c3d343fcdd17e18b186d6c873612d183159c8add64b188f1e1594b957a4f6",
9
+ "usage": {
10
+ "prompt_tokens": 131072,
11
+ "total_tokens": 131584,
12
+ "completion_tokens": 512
13
+ },
14
+ "error": null,
15
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is the fundamental unit of work. It is a sequence of one or more operations performed as a single logical unit. The core philosophy of a transaction is \"all or nothing\": either every operation within the transaction succeeds and the database is updated, or if even one operation fails, the entire transaction is undone, leaving the database in its original state.\n\nTo understand why this is critical, imagine an ATM withdrawal. You want to take out $100. The system must perform two distinct actions:\n1. Subtract $100 from your account balance.\n2. Dispense $100 in cash.\n\nIf the system subtracts the money but the machine fails to dispense the cash, you lose your money. If the machine dispenses the cash but the system fails to update your balance, the bank loses money. A transaction ensures that these two steps are bound together so that one cannot happen without the other.\n\n---\n\n### The ACID Properties: The Gold Standard of Transactions\n\nFor a database to be considered reliable, it must adhere to the **ACID** properties. These four principles ensure that data remains accurate and consistent, even in the event of system crashes, power failures, or concurrent access by thousands of users.\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity guarantees that a transaction is treated as a single, indivisible unit. If a transaction consists of five different SQL queries, and the fourth query fails due to a constraint violation or a lost connection, the first three queries are \"rolled back.\" The database reverts to the state it was in before the transaction started.\n\n* **Example:** In a hotel booking system, a transaction involves:\n * Checking room availability.\n * Creating a reservation record.\n * Charging the customer's credit card.\n * Sending a confirmation email.\n If the credit card charge fails, the reservation record must be deleted, and the room must be marked as available again. Atomicity prevents \"ghost\" reservations where a room is held but no payment was received.\n\n#### 2. Consistency (Rules of the Game)\nConsistency ensures that a transaction brings the database from one valid state to another. Every database has predefined rules (constraints), such as \"account balances cannot be negative\" or \"every order must have a valid Customer ID.\" A transaction must follow these rules. If"
16
+ }
evidence/http-p3c/tput-2048/requests-2048-c1.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "index": 0,
4
+ "start_seconds": 1.6479985788464546e-05,
5
+ "end_seconds": 11.016043118987,
6
+ "ttft_seconds": 0.6452015560062137,
7
+ "latency_seconds": 11.016026639001211,
8
+ "decode_tokens_per_second": 49.27316326165929,
9
+ "text_sha256": "271e3153ec0cf27b50e15284295499fdc4e38a58b0e7034a928379447e546564",
10
+ "usage": {
11
+ "prompt_tokens": 2048,
12
+ "total_tokens": 2560,
13
+ "completion_tokens": 512
14
+ },
15
+ "error": null,
16
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is a fundamental concept that ensures data integrity and reliability. At its simplest, a transaction is a sequence of one or more operations (such as inserting, updating, deleting, or querying data) that are treated as a single, indivisible unit of work. \n\nThe core philosophy behind transactions is \"all or nothing.\" If every operation within the transaction succeeds, the entire transaction is \"committed,\" and the changes become permanent. If even a single operation fails, the entire transaction is \"rolled back,\" and the database is returned to the state it was in before the transaction began.\n\nTo understand how transactions work, we must explore the **ACID properties**, the lifecycle of a transaction, concurrency control, and real-world applications.\n\n---\n\n### 1. The ACID Properties\nThe reliability of a database system is measured by its adherence to the ACID properties. These four principles ensure that even in the event of system crashes, power failures, or concurrent access by thousands of users, the data remains accurate.\n\n#### A - Atomicity\nAtomicity ensures that a transaction is treated as a single \"atom\"\u2014it cannot be broken into smaller pieces. If a transaction involves five different SQL statements, and the fourth one fails due to a constraint violation or a lost connection, the first three must be undone.\n* **Example:** Imagine a bank transfer where \\$100 is moved from Account A to Account B. This involves two steps: subtracting \\$100 from A and adding \\$100 to B. If the system crashes after subtracting from A but before adding to B, the money would effectively vanish. Atomicity prevents this by ensuring that if the addition to B fails, the subtraction from A is reversed.\n\n#### C - Consistency\nConsistency ensures that a transaction brings the database from one valid state to another, maintaining all predefined rules, including constraints, cascades, and triggers. A transaction must not violate the \"rules\" of the database.\n* **Example:** If a database has a rule that a bank balance cannot fall below zero, and a transaction attempts to withdraw \\$500 from an account containing only \\$200, the database will reject the transaction to maintain consistency.\n\n#### I - Isolation\nIn a multi-user environment, many transactions happen simultaneously. Isolation ensures that the concurrent execution of transactions results in a system state that is the same as if the transactions were executed sequentially ("
17
+ },
18
+ {
19
+ "index": 1,
20
+ "start_seconds": 11.016083038994111,
21
+ "end_seconds": 22.032156578003196,
22
+ "ttft_seconds": 0.6449967880034819,
23
+ "latency_seconds": 11.016073539009085,
24
+ "decode_tokens_per_second": 49.27197280097986,
25
+ "text_sha256": "271e3153ec0cf27b50e15284295499fdc4e38a58b0e7034a928379447e546564",
26
+ "usage": {
27
+ "prompt_tokens": 2048,
28
+ "total_tokens": 2560,
29
+ "completion_tokens": 512
30
+ },
31
+ "error": null,
32
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is a fundamental concept that ensures data integrity and reliability. At its simplest, a transaction is a sequence of one or more operations (such as inserting, updating, deleting, or querying data) that are treated as a single, indivisible unit of work. \n\nThe core philosophy behind transactions is \"all or nothing.\" If every operation within the transaction succeeds, the entire transaction is \"committed,\" and the changes become permanent. If even a single operation fails, the entire transaction is \"rolled back,\" and the database is returned to the state it was in before the transaction began.\n\nTo understand how transactions work, we must explore the **ACID properties**, the lifecycle of a transaction, concurrency control, and real-world applications.\n\n---\n\n### 1. The ACID Properties\nThe reliability of a database system is measured by its adherence to the ACID properties. These four principles ensure that even in the event of system crashes, power failures, or concurrent access by thousands of users, the data remains accurate.\n\n#### A - Atomicity\nAtomicity ensures that a transaction is treated as a single \"atom\"\u2014it cannot be broken into smaller pieces. If a transaction involves five different SQL statements, and the fourth one fails due to a constraint violation or a lost connection, the first three must be undone.\n* **Example:** Imagine a bank transfer where \\$100 is moved from Account A to Account B. This involves two steps: subtracting \\$100 from A and adding \\$100 to B. If the system crashes after subtracting from A but before adding to B, the money would effectively vanish. Atomicity prevents this by ensuring that if the addition to B fails, the subtraction from A is reversed.\n\n#### C - Consistency\nConsistency ensures that a transaction brings the database from one valid state to another, maintaining all predefined rules, including constraints, cascades, and triggers. A transaction must not violate the \"rules\" of the database.\n* **Example:** If a database has a rule that a bank balance cannot fall below zero, and a transaction attempts to withdraw \\$500 from an account containing only \\$200, the database will reject the transaction to maintain consistency.\n\n#### I - Isolation\nIn a multi-user environment, many transactions happen simultaneously. Isolation ensures that the concurrent execution of transactions results in a system state that is the same as if the transactions were executed sequentially ("
33
+ }
34
+ ]
evidence/http-p3c/tput-2048/summary.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "google/gemma-4-12B-it",
3
+ "output_tokens": 512,
4
+ "method": "Closed loop; synchronized initial clients; max(min_requests,2*concurrency) requests; nearest-rank percentiles; aggregate includes prefill and queue drain",
5
+ "cases": [
6
+ {
7
+ "input_tokens": 2048,
8
+ "concurrency": 1,
9
+ "requests": 2,
10
+ "successes": 2,
11
+ "errors": 0,
12
+ "wall_seconds": 22.032140098017408,
13
+ "matching_single_request_outputs": 2,
14
+ "aggregate_output_tokens_per_second": 46.477554855969075,
15
+ "requests_per_second": 0.0907764743280646,
16
+ "ttft_seconds_p50": 0.6449967880034819,
17
+ "ttft_seconds_p95": 0.6452015560062137,
18
+ "latency_seconds_p50": 11.016026639001211,
19
+ "latency_seconds_p95": 11.016073539009085,
20
+ "decode_tokens_per_second_p50": 49.27197280097986,
21
+ "decode_tokens_per_second_p95": 49.27316326165929
22
+ }
23
+ ]
24
+ }
evidence/http-p3c/tput-2048/warmup-2048.json ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "index": -1,
3
+ "start_seconds": 8.799834176898003e-07,
4
+ "end_seconds": 11.195358822995331,
5
+ "ttft_seconds": 0.6444937330088578,
6
+ "latency_seconds": 11.195357943011913,
7
+ "decode_tokens_per_second": 48.43237123768412,
8
+ "text_sha256": "271e3153ec0cf27b50e15284295499fdc4e38a58b0e7034a928379447e546564",
9
+ "usage": {
10
+ "prompt_tokens": 2048,
11
+ "total_tokens": 2560,
12
+ "completion_tokens": 512
13
+ },
14
+ "error": null,
15
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is a fundamental concept that ensures data integrity and reliability. At its simplest, a transaction is a sequence of one or more operations (such as inserting, updating, deleting, or querying data) that are treated as a single, indivisible unit of work. \n\nThe core philosophy behind transactions is \"all or nothing.\" If every operation within the transaction succeeds, the entire transaction is \"committed,\" and the changes become permanent. If even a single operation fails, the entire transaction is \"rolled back,\" and the database is returned to the state it was in before the transaction began.\n\nTo understand how transactions work, we must explore the **ACID properties**, the lifecycle of a transaction, concurrency control, and real-world applications.\n\n---\n\n### 1. The ACID Properties\nThe reliability of a database system is measured by its adherence to the ACID properties. These four principles ensure that even in the event of system crashes, power failures, or concurrent access by thousands of users, the data remains accurate.\n\n#### A - Atomicity\nAtomicity ensures that a transaction is treated as a single \"atom\"\u2014it cannot be broken into smaller pieces. If a transaction involves five different SQL statements, and the fourth one fails due to a constraint violation or a lost connection, the first three must be undone.\n* **Example:** Imagine a bank transfer where \\$100 is moved from Account A to Account B. This involves two steps: subtracting \\$100 from A and adding \\$100 to B. If the system crashes after subtracting from A but before adding to B, the money would effectively vanish. Atomicity prevents this by ensuring that if the addition to B fails, the subtraction from A is reversed.\n\n#### C - Consistency\nConsistency ensures that a transaction brings the database from one valid state to another, maintaining all predefined rules, including constraints, cascades, and triggers. A transaction must not violate the \"rules\" of the database.\n* **Example:** If a database has a rule that a bank balance cannot fall below zero, and a transaction attempts to withdraw \\$500 from an account containing only \\$200, the database will reject the transaction to maintain consistency.\n\n#### I - Isolation\nIn a multi-user environment, many transactions happen simultaneously. Isolation ensures that the concurrent execution of transactions results in a system state that is the same as if the transactions were executed sequentially ("
16
+ }
evidence/http-p3c/tput-261632/requests-261632-c1.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "index": 0,
4
+ "start_seconds": 4.0059007005766034e-05,
5
+ "end_seconds": 418.1412177470047,
6
+ "ttft_seconds": 394.7243818969873,
7
+ "latency_seconds": 418.1411776879977,
8
+ "decode_tokens_per_second": 21.822016190474002,
9
+ "text_sha256": "5f68fd62ed476b10bc5c9080ee9d9b8e764722ada2fe1e961fdfb8404c7444b4",
10
+ "usage": {
11
+ "prompt_tokens": 261632,
12
+ "total_tokens": 262144,
13
+ "completion_tokens": 512
14
+ },
15
+ "error": null,
16
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **transaction** is the fundamental unit of work. Whether you are transferring money between bank accounts, booking a flight, or updating an inventory system, these actions are rarely single, isolated events. They are usually composed of multiple steps that must all succeed together or fail together.\n\nTo ensure that a database remains reliable and accurate despite system crashes, concurrent access by multiple users, or network failures, database systems adhere to a set of principles known as **ACID properties**.\n\n---\n\n### 1. The Core Concept: What is a Transaction?\nA transaction is a sequence of operations performed as a single logical unit. The defining characteristic of a transaction is that it is **indivisible**. If any part of the sequence fails, the entire transaction must be undone, returning the database to the state it was in before the transaction started.\n\n**Example: The Bank Transfer**\nImagine you want to transfer \\$100 from Account A to Account B. This involves two distinct steps:\n1. **Debit:** Subtract \\$100 from Account A.\n2. **Credit:** Add \\$100 to Account B.\n\nIf the system crashes *after* step 1 but *before* step 2, Account A loses money, but Account B never receives it. The money effectively vanishes. A database transaction ensures that either both steps happen, or neither happens.\n\n---\n\n### 2. The ACID Properties\nTo guarantee reliability, every database management system (DBMS) must follow the ACID model:\n\n#### A - Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as a single \"atom.\" It cannot be broken down into smaller parts. If a transaction contains five SQL statements and the fourth one fails, the first three are \"rolled back\" (undone), and the fifth is never executed.\n* **Mechanism:** This is usually managed by a **Transaction Log**. The database records every change before it happens. If a failure occurs, it uses the log to reverse the partial changes.\n\n#### C - Consistency (\"Valid State to Valid State\")\nConsistency ensures that a transaction brings the database from one valid state to another, maintaining all predefined rules (constraints). These rules include data types, unique keys, and foreign key relationships.\n* **Example:** If a database rule states that an account balance cannot drop below zero, and a transaction tries to withdraw \\$500"
17
+ },
18
+ {
19
+ "index": 1,
20
+ "start_seconds": 418.1412568560045,
21
+ "end_seconds": 836.2879148600041,
22
+ "ttft_seconds": 394.70167387797846,
23
+ "latency_seconds": 418.1466580039996,
24
+ "decode_tokens_per_second": 21.795770483933808,
25
+ "text_sha256": "5f68fd62ed476b10bc5c9080ee9d9b8e764722ada2fe1e961fdfb8404c7444b4",
26
+ "usage": {
27
+ "prompt_tokens": 261632,
28
+ "total_tokens": 262144,
29
+ "completion_tokens": 512
30
+ },
31
+ "error": null,
32
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **transaction** is the fundamental unit of work. Whether you are transferring money between bank accounts, booking a flight, or updating an inventory system, these actions are rarely single, isolated events. They are usually composed of multiple steps that must all succeed together or fail together.\n\nTo ensure that a database remains reliable and accurate despite system crashes, concurrent access by multiple users, or network failures, database systems adhere to a set of principles known as **ACID properties**.\n\n---\n\n### 1. The Core Concept: What is a Transaction?\nA transaction is a sequence of operations performed as a single logical unit. The defining characteristic of a transaction is that it is **indivisible**. If any part of the sequence fails, the entire transaction must be undone, returning the database to the state it was in before the transaction started.\n\n**Example: The Bank Transfer**\nImagine you want to transfer \\$100 from Account A to Account B. This involves two distinct steps:\n1. **Debit:** Subtract \\$100 from Account A.\n2. **Credit:** Add \\$100 to Account B.\n\nIf the system crashes *after* step 1 but *before* step 2, Account A loses money, but Account B never receives it. The money effectively vanishes. A database transaction ensures that either both steps happen, or neither happens.\n\n---\n\n### 2. The ACID Properties\nTo guarantee reliability, every database management system (DBMS) must follow the ACID model:\n\n#### A - Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as a single \"atom.\" It cannot be broken down into smaller parts. If a transaction contains five SQL statements and the fourth one fails, the first three are \"rolled back\" (undone), and the fifth is never executed.\n* **Mechanism:** This is usually managed by a **Transaction Log**. The database records every change before it happens. If a failure occurs, it uses the log to reverse the partial changes.\n\n#### C - Consistency (\"Valid State to Valid State\")\nConsistency ensures that a transaction brings the database from one valid state to another, maintaining all predefined rules (constraints). These rules include data types, unique keys, and foreign key relationships.\n* **Example:** If a database rule states that an account balance cannot drop below zero, and a transaction tries to withdraw \\$500"
33
+ }
34
+ ]
evidence/http-p3c/tput-261632/summary.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "google/gemma-4-12B-it",
3
+ "output_tokens": 512,
4
+ "method": "Closed loop; synchronized initial clients; max(min_requests,2*concurrency) requests; nearest-rank percentiles; aggregate includes prefill and queue drain",
5
+ "cases": [
6
+ {
7
+ "input_tokens": 261632,
8
+ "concurrency": 1,
9
+ "requests": 2,
10
+ "successes": 2,
11
+ "errors": 0,
12
+ "wall_seconds": 836.2878748009971,
13
+ "matching_single_request_outputs": 2,
14
+ "aggregate_output_tokens_per_second": 1.2244587430418872,
15
+ "requests_per_second": 0.002391520982503686,
16
+ "ttft_seconds_p50": 394.70167387797846,
17
+ "ttft_seconds_p95": 394.7243818969873,
18
+ "latency_seconds_p50": 418.1411776879977,
19
+ "latency_seconds_p95": 418.1466580039996,
20
+ "decode_tokens_per_second_p50": 21.795770483933808,
21
+ "decode_tokens_per_second_p95": 21.822016190474002
22
+ }
23
+ ]
24
+ }
evidence/http-p3c/tput-261632/warmup-261632.json ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "index": -1,
3
+ "start_seconds": 8.700008038431406e-07,
4
+ "end_seconds": 419.97353838098934,
5
+ "ttft_seconds": 396.40250135198585,
6
+ "latency_seconds": 419.97353751098854,
7
+ "decode_tokens_per_second": 21.679220519427187,
8
+ "text_sha256": "5f68fd62ed476b10bc5c9080ee9d9b8e764722ada2fe1e961fdfb8404c7444b4",
9
+ "usage": {
10
+ "prompt_tokens": 261632,
11
+ "total_tokens": 262144,
12
+ "completion_tokens": 512
13
+ },
14
+ "error": null,
15
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **transaction** is the fundamental unit of work. Whether you are transferring money between bank accounts, booking a flight, or updating an inventory system, these actions are rarely single, isolated events. They are usually composed of multiple steps that must all succeed together or fail together.\n\nTo ensure that a database remains reliable and accurate despite system crashes, concurrent access by multiple users, or network failures, database systems adhere to a set of principles known as **ACID properties**.\n\n---\n\n### 1. The Core Concept: What is a Transaction?\nA transaction is a sequence of operations performed as a single logical unit. The defining characteristic of a transaction is that it is **indivisible**. If any part of the sequence fails, the entire transaction must be undone, returning the database to the state it was in before the transaction started.\n\n**Example: The Bank Transfer**\nImagine you want to transfer \\$100 from Account A to Account B. This involves two distinct steps:\n1. **Debit:** Subtract \\$100 from Account A.\n2. **Credit:** Add \\$100 to Account B.\n\nIf the system crashes *after* step 1 but *before* step 2, Account A loses money, but Account B never receives it. The money effectively vanishes. A database transaction ensures that either both steps happen, or neither happens.\n\n---\n\n### 2. The ACID Properties\nTo guarantee reliability, every database management system (DBMS) must follow the ACID model:\n\n#### A - Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as a single \"atom.\" It cannot be broken down into smaller parts. If a transaction contains five SQL statements and the fourth one fails, the first three are \"rolled back\" (undone), and the fifth is never executed.\n* **Mechanism:** This is usually managed by a **Transaction Log**. The database records every change before it happens. If a failure occurs, it uses the log to reverse the partial changes.\n\n#### C - Consistency (\"Valid State to Valid State\")\nConsistency ensures that a transaction brings the database from one valid state to another, maintaining all predefined rules (constraints). These rules include data types, unique keys, and foreign key relationships.\n* **Example:** If a database rule states that an account balance cannot drop below zero, and a transaction tries to withdraw \\$500"
16
+ }
evidence/http-p3c/tput-32768/requests-32768-c1.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "index": 0,
4
+ "start_seconds": 1.7960002878680825e-05,
5
+ "end_seconds": 28.333977759990375,
6
+ "ttft_seconds": 16.9979730900086,
7
+ "latency_seconds": 28.333959799987497,
8
+ "decode_tokens_per_second": 45.07796922419997,
9
+ "text_sha256": "776449f071272b7ffcaa722fc9113043c6475ff0c934ab95844e4884333a073f",
10
+ "usage": {
11
+ "prompt_tokens": 32768,
12
+ "total_tokens": 33280,
13
+ "completion_tokens": 512
14
+ },
15
+ "error": null,
16
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is a fundamental concept that ensures data integrity and reliability. At its simplest, a transaction is a sequence of one or more operations performed as a single logical unit of work. The core philosophy of a transaction is \"all or nothing\": either every operation within the transaction succeeds and is permanently saved to the database, or none of them are applied.\n\nTo understand why this is critical, consider a banking system. If you transfer \\$100 from Account A to Account B, two distinct operations occur:\n1. Subtract \\$100 from Account A.\n2. Add \\$100 to Account B.\n\nIf the system crashes after the first step but before the second, Account A loses money while Account B receives nothing. The money effectively vanishes into thin air. A database transaction prevents this by ensuring that both steps are treated as a single unit. If the second step fails, the first step is \"rolled back,\" and Account A\u2019s balance is restored.\n\n---\n\n### The ACID Properties: The Pillars of Transactions\n\nTo guarantee that transactions are processed reliably, database management systems (DBMS) must adhere to a set of four properties known as **ACID**.\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as an indivisible unit. If a transaction consists of five different SQL queries, and the fourth one fails due to a power outage or a constraint violation, the database must undo the first three queries. The database state must return to exactly how it was before the transaction started.\n* **Mechanism:** This is usually managed via a **Transaction Log** (or Write-Ahead Log). The system records every change before it happens. If a failure occurs, the system reads the log to \"undo\" (rollback) the partial changes.\n\n#### 2. Consistency (\"Follow the Rules\")\nConsistency ensures that a transaction brings the database from one valid state to another. Every database has predefined rules, such as \"account balances cannot be negative\" or \"every order must have a valid customer ID.\" A transaction must not violate these rules. If a transaction would result in an illegal state, the DBMS will abort the transaction.\n* **Example:** If you try to transfer \\$500 from an account that only has \\$200, and the database has a \"no negative balance\" rule, the transaction will fail the consistency check and be"
17
+ },
18
+ {
19
+ "index": 1,
20
+ "start_seconds": 28.334020109003177,
21
+ "end_seconds": 56.67031894001411,
22
+ "ttft_seconds": 17.010497626004508,
23
+ "latency_seconds": 28.336298831010936,
24
+ "decode_tokens_per_second": 45.11846273339815,
25
+ "text_sha256": "776449f071272b7ffcaa722fc9113043c6475ff0c934ab95844e4884333a073f",
26
+ "usage": {
27
+ "prompt_tokens": 32768,
28
+ "total_tokens": 33280,
29
+ "completion_tokens": 512
30
+ },
31
+ "error": null,
32
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is a fundamental concept that ensures data integrity and reliability. At its simplest, a transaction is a sequence of one or more operations performed as a single logical unit of work. The core philosophy of a transaction is \"all or nothing\": either every operation within the transaction succeeds and is permanently saved to the database, or none of them are applied.\n\nTo understand why this is critical, consider a banking system. If you transfer \\$100 from Account A to Account B, two distinct operations occur:\n1. Subtract \\$100 from Account A.\n2. Add \\$100 to Account B.\n\nIf the system crashes after the first step but before the second, Account A loses money while Account B receives nothing. The money effectively vanishes into thin air. A database transaction prevents this by ensuring that both steps are treated as a single unit. If the second step fails, the first step is \"rolled back,\" and Account A\u2019s balance is restored.\n\n---\n\n### The ACID Properties: The Pillars of Transactions\n\nTo guarantee that transactions are processed reliably, database management systems (DBMS) must adhere to a set of four properties known as **ACID**.\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as an indivisible unit. If a transaction consists of five different SQL queries, and the fourth one fails due to a power outage or a constraint violation, the database must undo the first three queries. The database state must return to exactly how it was before the transaction started.\n* **Mechanism:** This is usually managed via a **Transaction Log** (or Write-Ahead Log). The system records every change before it happens. If a failure occurs, the system reads the log to \"undo\" (rollback) the partial changes.\n\n#### 2. Consistency (\"Follow the Rules\")\nConsistency ensures that a transaction brings the database from one valid state to another. Every database has predefined rules, such as \"account balances cannot be negative\" or \"every order must have a valid customer ID.\" A transaction must not violate these rules. If a transaction would result in an illegal state, the DBMS will abort the transaction.\n* **Example:** If you try to transfer \\$500 from an account that only has \\$200, and the database has a \"no negative balance\" rule, the transaction will fail the consistency check and be"
33
+ }
34
+ ]
evidence/http-p3c/tput-32768/summary.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "google/gemma-4-12B-it",
3
+ "output_tokens": 512,
4
+ "method": "Closed loop; synchronized initial clients; max(min_requests,2*concurrency) requests; nearest-rank percentiles; aggregate includes prefill and queue drain",
5
+ "cases": [
6
+ {
7
+ "input_tokens": 32768,
8
+ "concurrency": 1,
9
+ "requests": 2,
10
+ "successes": 2,
11
+ "errors": 0,
12
+ "wall_seconds": 56.670300980011234,
13
+ "matching_single_request_outputs": 2,
14
+ "aggregate_output_tokens_per_second": 18.069429353501857,
15
+ "requests_per_second": 0.035291854206058314,
16
+ "ttft_seconds_p50": 16.9979730900086,
17
+ "ttft_seconds_p95": 17.010497626004508,
18
+ "latency_seconds_p50": 28.333959799987497,
19
+ "latency_seconds_p95": 28.336298831010936,
20
+ "decode_tokens_per_second_p50": 45.07796922419997,
21
+ "decode_tokens_per_second_p95": 45.11846273339815
22
+ }
23
+ ]
24
+ }
evidence/http-p3c/tput-32768/warmup-32768.json ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "index": -1,
3
+ "start_seconds": 6.599875632673502e-07,
4
+ "end_seconds": 28.28398171698791,
5
+ "ttft_seconds": 16.959950836986536,
6
+ "latency_seconds": 28.283981057000346,
7
+ "decode_tokens_per_second": 45.125468318013986,
8
+ "text_sha256": "776449f071272b7ffcaa722fc9113043c6475ff0c934ab95844e4884333a073f",
9
+ "usage": {
10
+ "prompt_tokens": 32768,
11
+ "total_tokens": 33280,
12
+ "completion_tokens": 512
13
+ },
14
+ "error": null,
15
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is a fundamental concept that ensures data integrity and reliability. At its simplest, a transaction is a sequence of one or more operations performed as a single logical unit of work. The core philosophy of a transaction is \"all or nothing\": either every operation within the transaction succeeds and is permanently saved to the database, or none of them are applied.\n\nTo understand why this is critical, consider a banking system. If you transfer \\$100 from Account A to Account B, two distinct operations occur:\n1. Subtract \\$100 from Account A.\n2. Add \\$100 to Account B.\n\nIf the system crashes after the first step but before the second, Account A loses money while Account B receives nothing. The money effectively vanishes into thin air. A database transaction prevents this by ensuring that both steps are treated as a single unit. If the second step fails, the first step is \"rolled back,\" and Account A\u2019s balance is restored.\n\n---\n\n### The ACID Properties: The Pillars of Transactions\n\nTo guarantee that transactions are processed reliably, database management systems (DBMS) must adhere to a set of four properties known as **ACID**.\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as an indivisible unit. If a transaction consists of five different SQL queries, and the fourth one fails due to a power outage or a constraint violation, the database must undo the first three queries. The database state must return to exactly how it was before the transaction started.\n* **Mechanism:** This is usually managed via a **Transaction Log** (or Write-Ahead Log). The system records every change before it happens. If a failure occurs, the system reads the log to \"undo\" (rollback) the partial changes.\n\n#### 2. Consistency (\"Follow the Rules\")\nConsistency ensures that a transaction brings the database from one valid state to another. Every database has predefined rules, such as \"account balances cannot be negative\" or \"every order must have a valid customer ID.\" A transaction must not violate these rules. If a transaction would result in an illegal state, the DBMS will abort the transaction.\n* **Example:** If you try to transfer \\$500 from an account that only has \\$200, and the database has a \"no negative balance\" rule, the transaction will fail the consistency check and be"
16
+ }
evidence/http-p3c/tput-8192/summary.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "google/gemma-4-12B-it",
3
+ "output_tokens": 512,
4
+ "method": "Closed loop; synchronized initial clients; max(min_requests,2*concurrency) requests; nearest-rank percentiles; aggregate includes prefill and queue drain",
5
+ "cases": [
6
+ {
7
+ "input_tokens": 8192,
8
+ "concurrency": 1,
9
+ "requests": 2,
10
+ "successes": 2,
11
+ "errors": 0,
12
+ "wall_seconds": 28.218357334000757,
13
+ "matching_single_request_outputs": 2,
14
+ "aggregate_output_tokens_per_second": 36.28843408139019,
15
+ "requests_per_second": 0.07087584781521522,
16
+ "ttft_seconds_p50": 3.2046391630137805,
17
+ "ttft_seconds_p95": 3.205977500998415,
18
+ "latency_seconds_p50": 14.10873377599637,
19
+ "latency_seconds_p95": 14.109585438010981,
20
+ "decode_tokens_per_second_p50": 46.86344028651859,
21
+ "decode_tokens_per_second_p95": 46.86549853892436
22
+ }
23
+ ]
24
+ }
evidence/http-p3c/tput-8192/warmup-8192.json ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "index": -1,
3
+ "start_seconds": 7.8001176007092e-07,
4
+ "end_seconds": 14.163157122005941,
5
+ "ttft_seconds": 3.2000005510053597,
6
+ "latency_seconds": 14.163156341994181,
7
+ "decode_tokens_per_second": 46.61093577559656,
8
+ "text_sha256": "252ca3663ced10787639340f33bf875901dae772ac0d46b4b1f10ae78b0fcea9",
9
+ "usage": {
10
+ "prompt_tokens": 8192,
11
+ "total_tokens": 8704,
12
+ "completion_tokens": 512
13
+ },
14
+ "error": null,
15
+ "text": "### Understanding Database Transactions: A Comprehensive Guide\n\nIn the world of data management, a **database transaction** is a fundamental concept that ensures data integrity and reliability. At its simplest, a transaction is a sequence of one or more operations performed as a single logical unit of work. The core philosophy of a transaction is \"all or nothing\": either every operation within the transaction succeeds and is permanently saved, or none of them are applied to the database.\n\nThis concept is vital because, in real-world applications, a single action (like buying a product online) often involves multiple underlying database updates. If one update succeeds but another fails, the data becomes inconsistent, leading to errors like \"ghost\" inventory or missing payments. To prevent this, databases adhere to a set of properties known as **ACID**.\n\n---\n\n### The ACID Properties\n\nTo be considered a valid transaction, a database operation must satisfy four key properties:\n\n#### 1. Atomicity (\"All or Nothing\")\nAtomicity ensures that a transaction is treated as a single, indivisible unit. If a transaction consists of five steps and the system crashes at step four, the database must \"roll back\" the first three steps so that the database remains in its original state.\n* **Example:** Imagine a bank transfer where \\$100 is moved from Account A to Account B. This involves two steps: (1) Subtract \\$100 from A, and (2) Add \\$100 to B. If the system fails after step one, Account A loses money while Account B gains nothing. Atomicity ensures that if step two fails, step one is undone.\n\n#### 2. Consistency (Valid State to Valid State)\nConsistency ensures that a transaction brings the database from one valid state to another, maintaining all predefined rules, including constraints, cascades, and triggers. A transaction cannot leave the database in a \"broken\" state (e.g., a negative balance in an account that doesn't allow it).\n* **Example:** If a database has a rule that \"Total Assets must equal Total Liabilities,\" any transaction that would cause this equation to be unbalanced will be rejected by the system.\n\n#### 3. Isolation (Independence of Transactions)\nIn a multi-user environment, hundreds of people might be accessing the same database simultaneously. Isolation ensures that concurrent transactions do not interfere with each other. The result of running multiple transactions at the same time should be the same as if they were run one after another (sequentially)."
16
+ }
evidence/http-p3c/tput_vs_direct.json ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "ctx": 32768,
4
+ "direct_tokens": 512,
5
+ "http_requests": 3,
6
+ "identical": [
7
+ true,
8
+ true,
9
+ true
10
+ ]
11
+ },
12
+ {
13
+ "ctx": 261632,
14
+ "direct_tokens": 512,
15
+ "http_requests": 3,
16
+ "identical": [
17
+ true,
18
+ true,
19
+ true
20
+ ]
21
+ }
22
+ ]
evidence/localmaxxing/speed-test.json ADDED
@@ -0,0 +1,175 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "agentFeedback": {
3
+ "benchmarkStatus": "completed",
4
+ "canApiValidate": true,
5
+ "canSubmit": true,
6
+ "engine": "vllm",
7
+ "message": "Speed-test payload is ready for API validation.",
8
+ "mode": "remote",
9
+ "nextCommand": "lmx speed-test dry-run /var/tmp/gemma4-12b/runs/lmx-official/speed-test.json",
10
+ "outputPath": "/var/tmp/gemma4-12b/runs/lmx-official/speed-test.json",
11
+ "outputPathAbsolute": "/var/tmp/gemma4-12b/runs/lmx-official/speed-test.json",
12
+ "requiresMetrics": false,
13
+ "runPersisted": true,
14
+ "savedRunPath": "/var/tmp/gemma4-12b/runs/lmx-official/runs/google-gemma-4-12B-it/20260930T163215Z.json",
15
+ "savedRunPathAbsolute": "/var/tmp/gemma4-12b/runs/lmx-official/runs/google-gemma-4-12B-it/20260930T163215Z.json",
16
+ "savedRunPathRelative": "/var/tmp/gemma4-12b/runs/lmx-official/runs/google-gemma-4-12B-it/20260930T163215Z.json",
17
+ "status": "ready_for_api_validation",
18
+ "submissionStatus": "ready",
19
+ "submitCommand": "lmx speed-test submit /var/tmp/gemma4-12b/runs/lmx-official/speed-test.json"
20
+ },
21
+ "backend": "tt-metal",
22
+ "batchSize": 1,
23
+ "benchmarkMode": "remote",
24
+ "contextLength": 262144,
25
+ "detectedEngines": [
26
+ {
27
+ "name": "llama.cpp",
28
+ "installed": true,
29
+ "binaries": {
30
+ "llama-bench": "/home/lotto/llama.cpp/build/bin/llama-bench",
31
+ "llama-cli": "/home/lotto/llama.cpp/build/bin/llama-cli",
32
+ "llama-server": "/home/lotto/llama.cpp/build/bin/llama-server"
33
+ }
34
+ }
35
+ ],
36
+ "engineFlags": {
37
+ "baseUrl": "http://127.0.0.1:8002",
38
+ "concurrency": 1,
39
+ "iterations": 5,
40
+ "maxTokens": 256,
41
+ "mode": "remote",
42
+ "prefixCacheBust": "leading_nonce_per_request",
43
+ "promptFile": "",
44
+ "promptSource": "default",
45
+ "servedModel": "google/gemma-4-12B-it",
46
+ "servedModelSource": "explicit",
47
+ "specDecoding": true,
48
+ "specDraftModel": "google/gemma-4-12B-it-assistant",
49
+ "specMethod": "assistant",
50
+ "specNumTokens": 5,
51
+ "stream": true,
52
+ "timeoutSeconds": 600,
53
+ "warmup": 2
54
+ },
55
+ "engineName": "vllm",
56
+ "hardware": {
57
+ "cpu": "AMD Ryzen 9 9950X",
58
+ "gpuCount": 1,
59
+ "gpuName": "Tenstorrent P150",
60
+ "hwClass": "DISCRETE_GPU",
61
+ "os": "Ubuntu 26.04 LTS",
62
+ "ramGb": 96,
63
+ "vramGb": 32
64
+ },
65
+ "hardwareSource": "file",
66
+ "hfId": "google/gemma-4-12B-it",
67
+ "metricSource": "remote_endpoint",
68
+ "modelRevision": "main",
69
+ "notes": "Single Tenstorrent P150, gemma4-12b:p3c, all-BFP8 native weights with exact proof, it-assistant drafter K=5, greedy. Local run, not submitted.",
70
+ "outputText": "To accurately evaluate the performance of a local Large Language Model (LLM) inference engine, reporting only a single \"tokens per second\" (TPS) metric is misleading. Because the computational demands of LLMs change drastically depending on whether the model is \"thinking\" (processing input) or \"speaking\" (generating output), you must break down performance into three distinct metrics: **Prompt Prefill Throughput**, **Decode Throughput**, and **Time to First Token (TTFT)**.\n\nHere is the technical breakdown of why each is essential:\n\n---\n\n### 1. Prompt Prefill Throughput (The \"Ingestion\" Phase)\n**What it is:** The speed at which the model processes the input prompt and converts it into the initial KV (Key-Value) cache.\n\n* **Why it matters:** This is a **compute-bound** operation. During prefill, the GPU processes all input tokens in parallel. If you are summarizing a 10,000-word document or analyzing a massive codebase, the prefill speed determines how long the user has to wait before the model starts \"typing.\"\n* **The Hardware Context:** High prefill throughput indicates efficient utilization of the GPU's CUDA cores and high memory bandwidth",
71
+ "outputTokens": 256,
72
+ "prefillTokens": 0,
73
+ "prompt": "Explain why local inference speed tests should report prompt prefill throughput, decode throughput, and time to first token.",
74
+ "promptTokens": 75,
75
+ "provenance": {
76
+ "benchmarkMode": "remote",
77
+ "cli": "localmaxxing-go",
78
+ "createdAt": "2026-09-30T16:32:15Z",
79
+ "metricSource": "remote_endpoint",
80
+ "timingSource": "client_observed_http",
81
+ "ttftSource": "stream_first_token"
82
+ },
83
+ "quantization": "BFP8",
84
+ "quantizationResolution": {
85
+ "cli": "BFP8",
86
+ "status": "matched",
87
+ "trusted": "BFP8",
88
+ "trustedSource": "cli"
89
+ },
90
+ "sampleStats": {
91
+ "tokSOut": {
92
+ "count": 5,
93
+ "max": 52.2,
94
+ "mean": 50.46,
95
+ "min": 49.3,
96
+ "p50": 49.9,
97
+ "stddev": 1.15
98
+ },
99
+ "tokSTotal": {
100
+ "count": 5,
101
+ "max": 66.2,
102
+ "mean": 64.24,
103
+ "min": 63.2,
104
+ "p50": 63.5,
105
+ "stddev": 1.35
106
+ },
107
+ "ttftMs": {
108
+ "count": 5,
109
+ "max": 95.37,
110
+ "mean": 94.96,
111
+ "min": 94.31,
112
+ "p50": 95.29,
113
+ "stddev": 0.51
114
+ }
115
+ },
116
+ "samples": [
117
+ {
118
+ "iteration": 1,
119
+ "outputTokens": 256,
120
+ "promptTokens": 75,
121
+ "request": 1,
122
+ "tokSOut": 49.9,
123
+ "tokSPrefill": 795.3,
124
+ "ttftMs": 94.31
125
+ },
126
+ {
127
+ "iteration": 2,
128
+ "outputTokens": 256,
129
+ "promptTokens": 77,
130
+ "request": 1,
131
+ "tokSOut": 49.3,
132
+ "tokSPrefill": 808.1,
133
+ "ttftMs": 95.29
134
+ },
135
+ {
136
+ "iteration": 3,
137
+ "outputTokens": 256,
138
+ "promptTokens": 76,
139
+ "request": 1,
140
+ "tokSOut": 51,
141
+ "tokSPrefill": 797.4,
142
+ "ttftMs": 95.31
143
+ },
144
+ {
145
+ "iteration": 4,
146
+ "outputTokens": 256,
147
+ "promptTokens": 73,
148
+ "request": 1,
149
+ "tokSOut": 49.9,
150
+ "tokSPrefill": 772.5,
151
+ "ttftMs": 94.5
152
+ },
153
+ {
154
+ "iteration": 5,
155
+ "outputTokens": 256,
156
+ "promptTokens": 74,
157
+ "request": 1,
158
+ "tokSOut": 52.2,
159
+ "tokSPrefill": 775.9,
160
+ "ttftMs": 95.37
161
+ }
162
+ ],
163
+ "timingSource": "client_observed_http",
164
+ "tokSOut": 49.9,
165
+ "tokSOutSource": "inter_token",
166
+ "tokSPrefill": 795.3,
167
+ "tokSPrefillSource": "estimated_from_ttft",
168
+ "tokSTotal": 63.5,
169
+ "tokenSources": {
170
+ "output": "endpoint_usage",
171
+ "prompt": "endpoint_usage"
172
+ },
173
+ "ttftMs": 95.29,
174
+ "ttftSource": "stream_first_token"
175
+ }
evidence/quality/quant_table.json ADDED
@@ -0,0 +1,533 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "ref_ppl": {
3
+ "chat": 8.5587,
4
+ "code": 5.6556,
5
+ "long-book": 12.1535,
6
+ "short-all": 7.1382
7
+ },
8
+ "gsm8k_bf16_cpu_acc": 0.98,
9
+ "rows": [
10
+ {
11
+ "config": "bf16",
12
+ "plan": "plans/bf16.json",
13
+ "weights": "IT",
14
+ "read_GB": 23.8446,
15
+ "floor_ms@128": 46.66,
16
+ "chat": {
17
+ "ppl_delta_pct": -0.77,
18
+ "dnll": -0.0077,
19
+ "se": 0.0014,
20
+ "top1": 0.975,
21
+ "kl": 0.0095,
22
+ "kl_p99": 0.103
23
+ },
24
+ "code": {
25
+ "ppl_delta_pct": -0.05,
26
+ "dnll": -0.0005,
27
+ "se": 0.0014,
28
+ "top1": 0.9843,
29
+ "kl": 0.0139,
30
+ "kl_p99": 0.061
31
+ },
32
+ "long-book": {
33
+ "ppl_delta_pct": -0.14,
34
+ "dnll": -0.0014,
35
+ "se": 0.0013,
36
+ "top1": 0.9764,
37
+ "kl": 0.0048,
38
+ "kl_p99": 0.053
39
+ },
40
+ "short-all": {
41
+ "ppl_delta_pct": -0.45,
42
+ "dnll": -0.0045,
43
+ "se": 0.001,
44
+ "top1": 0.9791,
45
+ "kl": 0.0114,
46
+ "kl_p99": 0.081
47
+ }
48
+ },
49
+ {
50
+ "config": "a-bfp8",
51
+ "plan": "plans/a-bfp8.json",
52
+ "weights": "IT",
53
+ "read_GB": 12.6675,
54
+ "floor_ms@128": 24.83,
55
+ "chat": {
56
+ "ppl_delta_pct": -0.29,
57
+ "dnll": -0.0029,
58
+ "se": 0.0015,
59
+ "top1": 0.9713,
60
+ "kl": 0.0122,
61
+ "kl_p99": 0.139
62
+ },
63
+ "code": {
64
+ "ppl_delta_pct": 1.01,
65
+ "dnll": 0.01,
66
+ "se": 0.0014,
67
+ "top1": 0.9827,
68
+ "kl": 0.0057,
69
+ "kl_p99": 0.08
70
+ },
71
+ "long-book": {
72
+ "ppl_delta_pct": 1.61,
73
+ "dnll": 0.016,
74
+ "se": 0.0015,
75
+ "top1": 0.9723,
76
+ "kl": 0.007,
77
+ "kl_p99": 0.068
78
+ },
79
+ "short-all": {
80
+ "ppl_delta_pct": 0.28,
81
+ "dnll": 0.0028,
82
+ "se": 0.001,
83
+ "top1": 0.9763,
84
+ "kl": 0.0093,
85
+ "kl_p99": 0.105
86
+ }
87
+ },
88
+ {
89
+ "config": "a-lm4",
90
+ "plan": "plans/a-lm4.json",
91
+ "weights": "IT",
92
+ "read_GB": 12.1641,
93
+ "floor_ms@128": 23.84,
94
+ "chat": {
95
+ "ppl_delta_pct": 0.61,
96
+ "dnll": 0.0061,
97
+ "se": 0.0018,
98
+ "top1": 0.955,
99
+ "kl": 0.0208,
100
+ "kl_p99": 0.244
101
+ },
102
+ "code": {
103
+ "ppl_delta_pct": 1.98,
104
+ "dnll": 0.0196,
105
+ "se": 0.0022,
106
+ "top1": 0.9703,
107
+ "kl": 0.0169,
108
+ "kl_p99": 0.289
109
+ },
110
+ "long-book": {
111
+ "ppl_delta_pct": 2.55,
112
+ "dnll": 0.0252,
113
+ "se": 0.0023,
114
+ "top1": 0.9605,
115
+ "kl": 0.0149,
116
+ "kl_p99": 0.128
117
+ },
118
+ "short-all": {
119
+ "ppl_delta_pct": 1.21,
120
+ "dnll": 0.012,
121
+ "se": 0.0014,
122
+ "top1": 0.9617,
123
+ "kl": 0.0191,
124
+ "kl_p99": 0.264
125
+ }
126
+ },
127
+ {
128
+ "config": "f-dn4",
129
+ "plan": "plans/f-dn4.json",
130
+ "weights": "IT",
131
+ "read_GB": 11.2519,
132
+ "floor_ms@128": 22.06,
133
+ "chat": {
134
+ "ppl_delta_pct": 2.61,
135
+ "dnll": 0.0257,
136
+ "se": 0.004,
137
+ "top1": 0.9102,
138
+ "kl": 0.0942,
139
+ "kl_p99": 1.456
140
+ },
141
+ "code": {
142
+ "ppl_delta_pct": 2.43,
143
+ "dnll": 0.024,
144
+ "se": 0.005,
145
+ "top1": 0.9333,
146
+ "kl": 0.101,
147
+ "kl_p99": 1.949
148
+ },
149
+ "long-book": {
150
+ "ppl_delta_pct": -0.64,
151
+ "dnll": -0.0065,
152
+ "se": 0.0049,
153
+ "top1": 0.912,
154
+ "kl": 0.0732,
155
+ "kl_p99": 0.819
156
+ },
157
+ "short-all": {
158
+ "ppl_delta_pct": 2.53,
159
+ "dnll": 0.025,
160
+ "se": 0.0031,
161
+ "top1": 0.9203,
162
+ "kl": 0.0972,
163
+ "kl_p99": 1.65
164
+ }
165
+ },
166
+ {
167
+ "config": "f-gu4",
168
+ "plan": "plans/f-gu4.json",
169
+ "weights": "IT",
170
+ "read_GB": 9.8363,
171
+ "floor_ms@128": 19.3,
172
+ "chat": {
173
+ "ppl_delta_pct": 1.56,
174
+ "dnll": 0.0155,
175
+ "se": 0.004,
176
+ "top1": 0.9021,
177
+ "kl": 0.0962,
178
+ "kl_p99": 1.461
179
+ },
180
+ "code": {
181
+ "ppl_delta_pct": 2.43,
182
+ "dnll": 0.024,
183
+ "se": 0.0051,
184
+ "top1": 0.9286,
185
+ "kl": 0.1111,
186
+ "kl_p99": 2.099
187
+ },
188
+ "long-book": {
189
+ "ppl_delta_pct": 9.73,
190
+ "dnll": 0.0929,
191
+ "se": 0.0053,
192
+ "top1": 0.9053,
193
+ "kl": 0.0848,
194
+ "kl_p99": 0.863
195
+ },
196
+ "short-all": {
197
+ "ppl_delta_pct": 1.94,
198
+ "dnll": 0.0192,
199
+ "se": 0.0032,
200
+ "top1": 0.9137,
201
+ "kl": 0.1028,
202
+ "kl_p99": 1.702
203
+ }
204
+ },
205
+ {
206
+ "config": "g-rank3072",
207
+ "plan": "plans/g-rank3072.json",
208
+ "weights": "IT",
209
+ "read_GB": 11.6058,
210
+ "floor_ms@128": 22.75,
211
+ "chat": {
212
+ "ppl_delta_pct": -0.04,
213
+ "dnll": -0.0004,
214
+ "se": 0.0031,
215
+ "top1": 0.9308,
216
+ "kl": 0.0548,
217
+ "kl_p99": 0.837
218
+ },
219
+ "code": {
220
+ "ppl_delta_pct": 2.05,
221
+ "dnll": 0.0203,
222
+ "se": 0.0041,
223
+ "top1": 0.9478,
224
+ "kl": 0.0688,
225
+ "kl_p99": 1.243
226
+ },
227
+ "long-book": {
228
+ "ppl_delta_pct": 7.3,
229
+ "dnll": 0.0704,
230
+ "se": 0.0037,
231
+ "top1": 0.9422,
232
+ "kl": 0.036,
233
+ "kl_p99": 0.353
234
+ },
235
+ "short-all": {
236
+ "ppl_delta_pct": 0.87,
237
+ "dnll": 0.0087,
238
+ "se": 0.0025,
239
+ "top1": 0.9382,
240
+ "kl": 0.061,
241
+ "kl_p99": 1.024
242
+ }
243
+ },
244
+ {
245
+ "config": "g-rank2048",
246
+ "plan": "plans/g-rank2048.json",
247
+ "weights": "IT",
248
+ "read_GB": 10.5441,
249
+ "floor_ms@128": 20.68,
250
+ "chat": {
251
+ "ppl_delta_pct": 1.25,
252
+ "dnll": 0.0125,
253
+ "se": 0.0037,
254
+ "top1": 0.9126,
255
+ "kl": 0.082,
256
+ "kl_p99": 1.28
257
+ },
258
+ "code": {
259
+ "ppl_delta_pct": 1.82,
260
+ "dnll": 0.018,
261
+ "se": 0.0048,
262
+ "top1": 0.9349,
263
+ "kl": 0.0971,
264
+ "kl_p99": 1.813
265
+ },
266
+ "long-book": {
267
+ "ppl_delta_pct": 8.86,
268
+ "dnll": 0.0849,
269
+ "se": 0.0048,
270
+ "top1": 0.9169,
271
+ "kl": 0.0671,
272
+ "kl_p99": 0.713
273
+ },
274
+ "short-all": {
275
+ "ppl_delta_pct": 1.5,
276
+ "dnll": 0.0149,
277
+ "se": 0.003,
278
+ "top1": 0.9224,
279
+ "kl": 0.0886,
280
+ "kl_p99": 1.542
281
+ }
282
+ },
283
+ {
284
+ "config": "b-mlp4",
285
+ "plan": "plans/b-mlp4.json",
286
+ "weights": "IT",
287
+ "read_GB": 8.4207,
288
+ "floor_ms@128": 16.53,
289
+ "chat": {
290
+ "ppl_delta_pct": 5.17,
291
+ "dnll": 0.0504,
292
+ "se": 0.0052,
293
+ "top1": 0.8706,
294
+ "kl": 0.1747,
295
+ "kl_p99": 2.743
296
+ },
297
+ "code": {
298
+ "ppl_delta_pct": 3.91,
299
+ "dnll": 0.0383,
300
+ "se": 0.0066,
301
+ "top1": 0.9083,
302
+ "kl": 0.1758,
303
+ "kl_p99": 3.069
304
+ },
305
+ "long-book": {
306
+ "ppl_delta_pct": 9.91,
307
+ "dnll": 0.0945,
308
+ "se": 0.007,
309
+ "top1": 0.8769,
310
+ "kl": 0.1502,
311
+ "kl_p99": 1.589
312
+ },
313
+ "short-all": {
314
+ "ppl_delta_pct": 4.61,
315
+ "dnll": 0.0451,
316
+ "se": 0.0041,
317
+ "top1": 0.8871,
318
+ "kl": 0.1752,
319
+ "kl_p99": 2.848
320
+ }
321
+ },
322
+ {
323
+ "config": "c-mlp4-lm4",
324
+ "plan": "plans/c-mlp4-lm4.json",
325
+ "weights": "IT",
326
+ "read_GB": 7.9174,
327
+ "floor_ms@128": 15.55,
328
+ "chat": {
329
+ "ppl_delta_pct": 6.21,
330
+ "dnll": 0.0603,
331
+ "se": 0.0054,
332
+ "top1": 0.8667,
333
+ "kl": 0.186,
334
+ "kl_p99": 2.85
335
+ },
336
+ "code": {
337
+ "ppl_delta_pct": 4.99,
338
+ "dnll": 0.0487,
339
+ "se": 0.0069,
340
+ "top1": 0.9057,
341
+ "kl": 0.1882,
342
+ "kl_p99": 3.352
343
+ },
344
+ "long-book": {
345
+ "ppl_delta_pct": 10.8,
346
+ "dnll": 0.1026,
347
+ "se": 0.0072,
348
+ "top1": 0.874,
349
+ "kl": 0.1582,
350
+ "kl_p99": 1.699
351
+ },
352
+ "short-all": {
353
+ "ppl_delta_pct": 5.67,
354
+ "dnll": 0.0552,
355
+ "se": 0.0043,
356
+ "top1": 0.8838,
357
+ "kl": 0.1869,
358
+ "kl_p99": 3.089
359
+ }
360
+ },
361
+ {
362
+ "config": "e-all4",
363
+ "plan": "plans/e-all4.json",
364
+ "weights": "IT",
365
+ "read_GB": 7.2096,
366
+ "floor_ms@128": 14.17,
367
+ "chat": {
368
+ "ppl_delta_pct": 7.24,
369
+ "dnll": 0.0699,
370
+ "se": 0.0068,
371
+ "top1": 0.8266,
372
+ "kl": 0.2946,
373
+ "kl_p99": 4.066
374
+ },
375
+ "code": {
376
+ "ppl_delta_pct": 8.32,
377
+ "dnll": 0.0799,
378
+ "se": 0.0085,
379
+ "top1": 0.8736,
380
+ "kl": 0.2975,
381
+ "kl_p99": 4.776
382
+ },
383
+ "long-book": {
384
+ "ppl_delta_pct": 19.16,
385
+ "dnll": 0.1753,
386
+ "se": 0.0107,
387
+ "top1": 0.8009,
388
+ "kl": 0.3806,
389
+ "kl_p99": 3.648
390
+ },
391
+ "short-all": {
392
+ "ppl_delta_pct": 7.71,
393
+ "dnll": 0.0743,
394
+ "se": 0.0053,
395
+ "top1": 0.8472,
396
+ "kl": 0.2959,
397
+ "kl_p99": 4.394
398
+ }
399
+ },
400
+ {
401
+ "config": "e-all4-lm4",
402
+ "plan": "plans/e-all4-lm4.json",
403
+ "weights": "IT",
404
+ "read_GB": 6.7063,
405
+ "floor_ms@128": 13.18,
406
+ "chat": {
407
+ "ppl_delta_pct": 8.48,
408
+ "dnll": 0.0814,
409
+ "se": 0.0069,
410
+ "top1": 0.824,
411
+ "kl": 0.3089,
412
+ "kl_p99": 4.273
413
+ },
414
+ "code": {
415
+ "ppl_delta_pct": 9.76,
416
+ "dnll": 0.0931,
417
+ "se": 0.0087,
418
+ "top1": 0.8709,
419
+ "kl": 0.3114,
420
+ "kl_p99": 5.03
421
+ },
422
+ "long-book": {
423
+ "ppl_delta_pct": 19.92,
424
+ "dnll": 0.1816,
425
+ "se": 0.0108,
426
+ "top1": 0.7996,
427
+ "kl": 0.3899,
428
+ "kl_p99": 3.703
429
+ },
430
+ "short-all": {
431
+ "ppl_delta_pct": 9.04,
432
+ "dnll": 0.0865,
433
+ "se": 0.0055,
434
+ "top1": 0.8446,
435
+ "kl": 0.31,
436
+ "kl_p99": 4.618
437
+ }
438
+ },
439
+ {
440
+ "config": "d-qat-a",
441
+ "plan": "plans/a-bfp8.json",
442
+ "weights": "QAT",
443
+ "read_GB": 12.6675,
444
+ "floor_ms@128": 24.83,
445
+ "chat": {
446
+ "ppl_delta_pct": -2.86,
447
+ "dnll": -0.029,
448
+ "se": 0.0053,
449
+ "top1": 0.9074,
450
+ "kl": 0.1289,
451
+ "kl_p99": 2.267
452
+ },
453
+ "code": {
454
+ "ppl_delta_pct": -8.9,
455
+ "dnll": -0.0932,
456
+ "se": 0.0059,
457
+ "top1": 0.928,
458
+ "kl": 0.1018,
459
+ "kl_p99": 1.497
460
+ },
461
+ "long-book": {
462
+ "ppl_delta_pct": -4.87,
463
+ "dnll": -0.0499,
464
+ "se": 0.0054,
465
+ "top1": 0.9055,
466
+ "kl": 0.0771,
467
+ "kl_p99": 0.782
468
+ },
469
+ "short-all": {
470
+ "ppl_delta_pct": -5.55,
471
+ "dnll": -0.0571,
472
+ "se": 0.0039,
473
+ "top1": 0.9164,
474
+ "kl": 0.117,
475
+ "kl_p99": 2.004
476
+ }
477
+ },
478
+ {
479
+ "config": "d-qat-b",
480
+ "plan": "plans/b-mlp4.json",
481
+ "weights": "QAT",
482
+ "read_GB": 8.4207,
483
+ "floor_ms@128": 16.53,
484
+ "chat": {
485
+ "ppl_delta_pct": 0.98,
486
+ "dnll": 0.0098,
487
+ "se": 0.0062,
488
+ "top1": 0.8658,
489
+ "kl": 0.2104,
490
+ "kl_p99": 3.608
491
+ },
492
+ "code": {
493
+ "ppl_delta_pct": -0.85,
494
+ "dnll": -0.0086,
495
+ "se": 0.0075,
496
+ "top1": 0.8987,
497
+ "kl": 0.1855,
498
+ "kl_p99": 3.064
499
+ },
500
+ "long-book": {
501
+ "ppl_delta_pct": 6.68,
502
+ "dnll": 0.0647,
503
+ "se": 0.0074,
504
+ "top1": 0.8719,
505
+ "kl": 0.1512,
506
+ "kl_p99": 1.517
507
+ },
508
+ "short-all": {
509
+ "ppl_delta_pct": 0.18,
510
+ "dnll": 0.0018,
511
+ "se": 0.0048,
512
+ "top1": 0.8802,
513
+ "kl": 0.1995,
514
+ "kl_p99": 3.406
515
+ }
516
+ }
517
+ ],
518
+ "gsm8k_50": {
519
+ "a-lm4": {
520
+ "acc": 0.98,
521
+ "agree": 0.96
522
+ },
523
+ "a-bfp8": {
524
+ "acc": 1.0,
525
+ "agree": 0.98
526
+ },
527
+ "b-mlp4": {
528
+ "acc": 0.98,
529
+ "agree": 0.96
530
+ }
531
+ },
532
+ "pick": "a-lm4"
533
+ }
evidence/quality/runtime-gate-dtf-score.json ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "chat": {
3
+ "n": 5888,
4
+ "dNLL": 0.010796350426971912,
5
+ "se": 0.002044405456103154,
6
+ "dPPL_pct": 1.0854721069335938,
7
+ "top1": 0.9799592391304348,
8
+ "kl_mean": 0.0037473044358193874,
9
+ "kl_p99": 0.04619763791561127,
10
+ "vs_base_dNLL": 0.001731055905111134,
11
+ "vs_base_se": 0.0011998895297415893,
12
+ "vs_base_top1_pts": 0.06793478260870289,
13
+ "argmax_agree_base": 0.98828125
14
+ },
15
+ "code": {
16
+ "n": 3328,
17
+ "dNLL": 0.004894797224551439,
18
+ "se": 0.00325290915414018,
19
+ "dPPL_pct": 0.4906773567199707,
20
+ "top1": 0.9846754807692307,
21
+ "kl_mean": 0.005180150270462036,
22
+ "kl_p99": 0.08787598460912704,
23
+ "vs_base_dNLL": -0.0004120908852200955,
24
+ "vs_base_se": 0.0030019859298719472,
25
+ "vs_base_top1_pts": -0.09014423076924016,
26
+ "argmax_agree_base": 0.9879807692307693
27
+ },
28
+ "long-book": {
29
+ "n": 256,
30
+ "dNLL": 0.0037011178210377693,
31
+ "se": 0.014648807235062122,
32
+ "dPPL_pct": 0.37078857421875,
33
+ "top1": 0.96484375,
34
+ "kl_mean": 0.008751568384468555,
35
+ "kl_p99": 0.09284984320402145,
36
+ "vs_base_dNLL": -0.006091505289077759,
37
+ "vs_base_se": 0.009644112549722195,
38
+ "vs_base_top1_pts": -0.390625,
39
+ "argmax_agree_base": 0.984375
40
+ }
41
+ }
gemma-4-12B-it-assistant/config.json ADDED
@@ -0,0 +1,88 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Gemma4UnifiedAssistantForCausalLM"
4
+ ],
5
+ "audio_token_id": 258881,
6
+ "backbone_hidden_size": 3840,
7
+ "boa_token_id": 256000,
8
+ "boi_token_id": 255999,
9
+ "centroid_intermediate_top_k": 32,
10
+ "dtype": "bfloat16",
11
+ "eoa_token_index": 258883,
12
+ "eoi_token_id": 258882,
13
+ "image_token_id": 258880,
14
+ "model_type": "gemma4_unified_assistant",
15
+ "num_centroids": 2048,
16
+ "text_config": {
17
+ "_name_or_path": "",
18
+ "architectures": null,
19
+ "attention_bias": false,
20
+ "attention_dropout": 0.0,
21
+ "attention_k_eq_v": true,
22
+ "bos_token_id": 2,
23
+ "chunk_size_feed_forward": 0,
24
+ "dtype": "bfloat16",
25
+ "enable_moe_block": false,
26
+ "eos_token_id": 1,
27
+ "final_logit_softcapping": null,
28
+ "global_head_dim": 512,
29
+ "head_dim": 256,
30
+ "hidden_activation": "gelu_pytorch_tanh",
31
+ "hidden_size": 1024,
32
+ "hidden_size_per_layer_input": 0,
33
+ "id2label": {
34
+ "0": "LABEL_0",
35
+ "1": "LABEL_1"
36
+ },
37
+ "initializer_range": 0.02,
38
+ "intermediate_size": 8192,
39
+ "is_encoder_decoder": false,
40
+ "label2id": {
41
+ "LABEL_0": 0,
42
+ "LABEL_1": 1
43
+ },
44
+ "layer_types": [
45
+ "sliding_attention",
46
+ "sliding_attention",
47
+ "sliding_attention",
48
+ "full_attention"
49
+ ],
50
+ "max_position_embeddings": 262144,
51
+ "model_type": "gemma4_unified_text",
52
+ "moe_intermediate_size": null,
53
+ "num_attention_heads": 16,
54
+ "num_experts": null,
55
+ "num_global_key_value_heads": 1,
56
+ "num_hidden_layers": 4,
57
+ "num_key_value_heads": 8,
58
+ "num_kv_shared_layers": 4,
59
+ "output_attentions": false,
60
+ "output_hidden_states": false,
61
+ "pad_token_id": 0,
62
+ "problem_type": null,
63
+ "return_dict": true,
64
+ "rms_norm_eps": 1e-06,
65
+ "rope_parameters": {
66
+ "full_attention": {
67
+ "partial_rotary_factor": 0.25,
68
+ "rope_theta": 1000000.0,
69
+ "rope_type": "proportional"
70
+ },
71
+ "sliding_attention": {
72
+ "rope_theta": 10000.0,
73
+ "rope_type": "default"
74
+ }
75
+ },
76
+ "sliding_window": 1024,
77
+ "tie_word_embeddings": true,
78
+ "top_k_experts": null,
79
+ "use_bidirectional_attention": "vision",
80
+ "use_cache": true,
81
+ "use_double_wide_mlp": false,
82
+ "vocab_size": 262144,
83
+ "vocab_size_per_layer_input": 0
84
+ },
85
+ "tie_word_embeddings": true,
86
+ "transformers_version": "5.10.0.dev0",
87
+ "use_ordered_embeddings": false
88
+ }
gemma-4-12B-it-assistant/drafter_manifest.json ADDED
@@ -0,0 +1,285 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "builder_runtime_identity": {
3
+ "binaries": {
4
+ "build/lib/_ttnncpp.so": "72186b132bf7be85bbb6e1b54bf6de125fbd7a557b12711c84efc6d3e2e08f8d",
5
+ "build/lib/libtt_metal.so": "0497308232f180a0973406e2afd290b9eafa9e310a16b09bf69acc4434eb0e78",
6
+ "ttnn/ttnn/_ttnn.so": "a6448405abd714d6ba310dbbc409c01842f42017b8c1d6bb1c30503fb194c7b2"
7
+ },
8
+ "gemma4_tree": {
9
+ "__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
10
+ "config.py": "ca64c3f783bd7135bec21e2c765004f311631539129ad7ab592b8ae15a2474e4",
11
+ "conftest.py": "5c9425f9367d48a2873dc83ae074bb0cf62fc433a5c8d140cf030930baf4c4ec",
12
+ "demo/__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
13
+ "demo/text_demo.py": "c95cee077db887fd748b4b32cf30f0751d21d14596cc0243048a4afafce98988",
14
+ "demo/text_demo_v2.py": "ad4785b8539e4fdeb32583932e6ac6e2d9bb342edb79d153f5d123d9b323a7d0",
15
+ "tt/__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
16
+ "tt/assistant/__init__.py": "229bb6de453e5f58c9c6bc80368cc417cc943ac0d6a6a268b845e91198ca2aa2",
17
+ "tt/assistant/model.py": "0e3247c34a359a73bd4c8ecc7365e028372084554eeeba41682d1fd818c296af",
18
+ "tt/attention/__init__.py": "5fef2a905f61e36f8a8ebfa19913460dac6845bb370f5688829bf514f0030400",
19
+ "tt/attention/config.py": "01b0f3e14c31aae5ed01104aee7f4cb03913007c0d0018f9d35d341d3fe049ef",
20
+ "tt/attention/decode.py": "d517d0f1df07be038658f77fd19d0e42c7ca7241efe8c59f3479576f7cb65042",
21
+ "tt/attention/kv_cache.py": "3a1e9f520cdde61dd49550dbfca1072d009ee34878bafe5923c655821c09bff7",
22
+ "tt/attention/kv_cache_hybrid.py": "9074a2d1de4b3a65d11271370f1fce39a78e5523b841752bcaa87dccb4b56470",
23
+ "tt/attention/operations.py": "f5cda2ade4031a66e631a732199f3b4b9a00b45ef2e252aee0d3138406bbcc43",
24
+ "tt/attention/prefill.py": "32c6e1c9423b9c22f14558c33d76d9164cb8a27c939cb5d1ccb6da039b7c831f",
25
+ "tt/attention/weights.py": "3ec94d60c6d1085bed0bf83fbc450aa9599a476341d3c6d77466b1b2e17c00d6",
26
+ "tt/ccl.py": "3580f2e25950381465e55dd2886d2832c33d8d6fc618b5a2b94088d0b1a91105",
27
+ "tt/common.py": "7e1a088a2f666cfef32e891dbe43ad143e15f1318a91ca895b4216bdfb68ca2e",
28
+ "tt/decode_mm.py": "978c6f7b111bfd652ac7e00f5c7ca90afc9699480ee9fc0be1614e1d43d3a1c1",
29
+ "tt/experts/__init__.py": "567aef533a30a692e4c0b28695c15d13e2cfc024ee621e4147c836b63220d5c3",
30
+ "tt/experts/config.py": "80bd7d18ffd0e0b6bd969d05a8e9972ff8473dc63631f1cb6de0bd619a50f33b",
31
+ "tt/experts/decode.py": "112b9df4b88e8395053fdbf69315e3a43e9f3f4877e699ca5a48462ccd8e804b",
32
+ "tt/experts/operations.py": "27db30001f4280dc46bd500d5d811f56ccd8a0896d4eb9f90a6ef50e9af3355b",
33
+ "tt/experts/prefill.py": "5f07a86da85d5ba64838e03d6627c7a78ccfcc06ebfbdaff39d8f7b52f84948d",
34
+ "tt/experts/weights.py": "dda56687ff669a92b8447dc098e4cad6007f94d66b642f1da48c17abd8568916",
35
+ "tt/gemma4_attention_config.py": "8e892a13f89acadd2415719c7c22d188ee906f5fabb1d326464b65afa06816be",
36
+ "tt/gemma4_expert_config.py": "e7b1037156bc3f20cacba6e7a59c5456ebaefa3ed772310f7476908b26aaf20a",
37
+ "tt/generator.py": "34914394ca36e2bc5942b28710ebc7362d7b610c4eb2db33745311b391ceee9e",
38
+ "tt/generator_trace.py": "d5c303e5fe936bf32d097d4013336dcfc587a3a2502cd71fd266c00ab6e2c9f3",
39
+ "tt/generator_vllm.py": "e3b9831358e44e6bcc1433e39ea13b82a7c15565826a38a00815a8a7640b65a4",
40
+ "tt/greedy_decode.py": "7604d9098ca2a36ba736eed21a2728dc72988c5ade03a1f6d8ce99c84d4ae5a4",
41
+ "tt/kernels/kv_rows_write.cpp": "f7419aa0ae5b10049546c4615780d53b6b54518c5f79ca76c708f1f76ce593ac",
42
+ "tt/layer.py": "e3261517b9328d460f84dd888f13694bb50047eb70240e6cc1967f16803599e0",
43
+ "tt/model.py": "9de68e12eb996f9f717d04f63dd42eb6b8cdbcc558d5c1789dc746ee25c7cc18",
44
+ "tt/model_config.py": "04ecd4a462f1ba726186e2c7fce7c3c86b78faf21091af9c9c098f325bb9556d",
45
+ "tt/moe.py": "37d5c5921fe9e5dc61f7b788bc0b3dfec349ce17915dc5450cf286a217f64e41",
46
+ "tt/precision.py": "588e9c3f9864c248175f0c8b2c8b45a5e45511568f71d448cf8ed1621c1d6571",
47
+ "tt/rms_norm.py": "cf0f9fb6191ae3d94191aa62a81686eeccb46b6c44635a1c75967d803886b213",
48
+ "tt/router.py": "61cfcfc0f6b3a7899f8d0b0a9de3738d8af933602e30d2fc4550c05ea9b19ff9",
49
+ "tt/shared_mlp.py": "31003feda56e095a8f871b210795306d1e988c0db3a860e799c4b025ee7f40db",
50
+ "tt/spec_decode.py": "8acf4609aca527c123b5d064236d4e817bca7e481702ad6580c0c417a30885e7",
51
+ "tt/spec_greedy.py": "1d3a4b1820f55eda775bcf4c56a93b24c13ecf8836f626185e191a6a95ebbd20",
52
+ "utils/__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
53
+ "utils/general_utils.py": "df7324ddc31c12d8177ae8ff007a220cb6e298b3ade498205ecaf5e93abe1eb9",
54
+ "utils/substate.py": "dd0c86aa2b1d6131d45f09ede19690e559543c51dc1166c858277a5c183963ac"
55
+ }
56
+ },
57
+ "dtype": "bfp8-lm4",
58
+ "files": {
59
+ "config.json": {
60
+ "bytes": 2346,
61
+ "sha256": "b6f19209588fcefe41f65b193fad6148446253c470d36e29441ecc5158a54e6d"
62
+ },
63
+ "generation_config.json": {
64
+ "bytes": 233,
65
+ "sha256": "02b56bd11e1cd1e363e701a85a2fd7fbaa2992ec3358c1cd7cc44ead7208f505"
66
+ },
67
+ "model.safetensors": {
68
+ "bytes": 845719296,
69
+ "sha256": "3279c173daddd7186e79d652ad94022415736d3a1370625696c898429b06d6df"
70
+ },
71
+ "tensors/assistant_tensor_cache_bfp8/final_norm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
72
+ "bytes": 2392,
73
+ "sha256": "70eb60efff28b3606bbc91b35a3311e7b4ccae3b70e0f3787c9c0b32cc4ebf3c"
74
+ },
75
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/input_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
76
+ "bytes": 2392,
77
+ "sha256": "1179fa4a06ab8e4bae470dce5c4848d5b3b5eca1a3c87d8a02001bfb1aacf02f"
78
+ },
79
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/mlp/down_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
80
+ "bytes": 8913264,
81
+ "sha256": "bc68d0c960e2782993ff0d354800e3f4d67dd80d1b31f0775dc5574114402787"
82
+ },
83
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/mlp/gate_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
84
+ "bytes": 8913264,
85
+ "sha256": "eab412d5d2c21880e9be0f06a8814e649f54995eacf78cdcc1f57b4c0914c97a"
86
+ },
87
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/mlp/up_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
88
+ "bytes": 8913264,
89
+ "sha256": "87737fd7e41bf956bf99c709389f994710633f5271f7b1a34899fc64097f3465"
90
+ },
91
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/post_attention_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
92
+ "bytes": 2392,
93
+ "sha256": "c7134e87abbbd99a3613b87356b26910c67c4dc88481883473a48a592e105ab9"
94
+ },
95
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/post_feedforward_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
96
+ "bytes": 2392,
97
+ "sha256": "f36ac602315f100f55e2ff60dd2d433f416e5088b7a1e132374d4fb5507d0ffa"
98
+ },
99
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/pre_feedforward_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
100
+ "bytes": 2392,
101
+ "sha256": "41226b5616d8526fa29f4b1d160a9aafab622c4520ad0aa5ac3d751ed105e87e"
102
+ },
103
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/self_attn/k_norm.weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
104
+ "bytes": 856,
105
+ "sha256": "196c1590df83891269c77d8e3c0f3e04644d5c7d30985bd4990b2056854fbd43"
106
+ },
107
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/self_attn/o_proj_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
108
+ "bytes": 4456816,
109
+ "sha256": "25138aecfea86617289b887c60e290cffc48d091f8dea8f3e6c7aa93a8145c13"
110
+ },
111
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/self_attn/q_norm.weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
112
+ "bytes": 856,
113
+ "sha256": "5f7bc4231c79615214a55cc24b23ed92fb8f280d8889294a1b9e510159b16c2a"
114
+ },
115
+ "tensors/assistant_tensor_cache_bfp8/layer_0/layer_0/self_attn/wqkv_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
116
+ "bytes": 8913264,
117
+ "sha256": "8ebb3b5a6fcef326e40a9beaff4b49628ee07a46ca9a8b54d1ed03eddfbee967"
118
+ },
119
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/input_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
120
+ "bytes": 2392,
121
+ "sha256": "16eeab1537e1770d8b129c67297def6e3792eadcc099663258f2da865eea7efd"
122
+ },
123
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/mlp/down_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
124
+ "bytes": 8913264,
125
+ "sha256": "60bd4ccea8173ae93aa96501d09a24507e30c1732e389550c8971b9197e7b412"
126
+ },
127
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/mlp/gate_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
128
+ "bytes": 8913264,
129
+ "sha256": "9005f3ef53275810d462d4792cd7e412f8fa95482bdf12f9c5e98ba857f476a7"
130
+ },
131
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/mlp/up_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
132
+ "bytes": 8913264,
133
+ "sha256": "ef2a8fe512348732555d51ef5e668fa399623bd8f9cd59389a358fab152cc4a9"
134
+ },
135
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/post_attention_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
136
+ "bytes": 2392,
137
+ "sha256": "2defec402158aa4c0b2c46a40e637480543f484b48d25b4a9c2517f6adb38fdc"
138
+ },
139
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/post_feedforward_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
140
+ "bytes": 2392,
141
+ "sha256": "5af05c890d4af050a8b752187e63796416404855d7f0f89e38c31dddcfd3e8c3"
142
+ },
143
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/pre_feedforward_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
144
+ "bytes": 2392,
145
+ "sha256": "dbb9eab778b1c9504fd5a1c189e02be9b901dee91ce108f83a0dc5121943e35b"
146
+ },
147
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/self_attn/k_norm.weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
148
+ "bytes": 856,
149
+ "sha256": "196c1590df83891269c77d8e3c0f3e04644d5c7d30985bd4990b2056854fbd43"
150
+ },
151
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/self_attn/o_proj_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
152
+ "bytes": 4456816,
153
+ "sha256": "b2c7d656e21b7fff14ee0b4345ef5cea6d8b6cb85cb2e1650674b162726652c9"
154
+ },
155
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/self_attn/q_norm.weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
156
+ "bytes": 856,
157
+ "sha256": "5f7bc4231c79615214a55cc24b23ed92fb8f280d8889294a1b9e510159b16c2a"
158
+ },
159
+ "tensors/assistant_tensor_cache_bfp8/layer_1/layer_1/self_attn/wqkv_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
160
+ "bytes": 8913264,
161
+ "sha256": "66e21ea442f9449f949a9368e0911f702dfa25d101dde38d83fda7becc36f39d"
162
+ },
163
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/input_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
164
+ "bytes": 2392,
165
+ "sha256": "c740974f83abf0bebe1368963ad09840575b2bdc9ce165abe36954ef932b2855"
166
+ },
167
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/mlp/down_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
168
+ "bytes": 8913264,
169
+ "sha256": "ed6a97a6b76c80cf73c0872b54090e3780eb1d708a50b31c5c1aeda5ec59463b"
170
+ },
171
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/mlp/gate_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
172
+ "bytes": 8913264,
173
+ "sha256": "349df26d7fe81c44bf157fc9efa6841e553b502355f2bd27695d2441c65d46c0"
174
+ },
175
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/mlp/up_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
176
+ "bytes": 8913264,
177
+ "sha256": "7b875ed7cdf3642c1a030aab5354e4cc9d98d6376365040eb8105d7ac9cf0e0e"
178
+ },
179
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/post_attention_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
180
+ "bytes": 2392,
181
+ "sha256": "de6fef7a1e07c44b18eddca38cf33bb5b6751b2400fedc258ec38bc257f3a3b7"
182
+ },
183
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/post_feedforward_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
184
+ "bytes": 2392,
185
+ "sha256": "77ec76e8b8766904eff7d03a34a950ad842bc8e49faadc5dc81049e17673d5a1"
186
+ },
187
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/pre_feedforward_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
188
+ "bytes": 2392,
189
+ "sha256": "cc57622fd05943bde0a5fb293f47e37fdf623ae8c546eecc71445a53788b465e"
190
+ },
191
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/self_attn/k_norm.weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
192
+ "bytes": 856,
193
+ "sha256": "196c1590df83891269c77d8e3c0f3e04644d5c7d30985bd4990b2056854fbd43"
194
+ },
195
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/self_attn/o_proj_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
196
+ "bytes": 4456816,
197
+ "sha256": "c64c03dffe767aea3cab9f727feceb23aaf26af97b297fb95258ce200a226ab2"
198
+ },
199
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/self_attn/q_norm.weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
200
+ "bytes": 856,
201
+ "sha256": "5f7bc4231c79615214a55cc24b23ed92fb8f280d8889294a1b9e510159b16c2a"
202
+ },
203
+ "tensors/assistant_tensor_cache_bfp8/layer_2/layer_2/self_attn/wqkv_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
204
+ "bytes": 8913264,
205
+ "sha256": "109e837cd97fac3baff29735a7473d2eb45c68ca0b28a603fd49090e37b82429"
206
+ },
207
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/input_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
208
+ "bytes": 2392,
209
+ "sha256": "4d1a717abb43f630dd58ab5815d07f5534bfdce1a718cd2fd5cf4f807a6739e9"
210
+ },
211
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/mlp/down_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
212
+ "bytes": 8913264,
213
+ "sha256": "54fd9c32751b389aa03d242b3a4fc02333d862dbc89eaf3cdeeae9f93cb534c8"
214
+ },
215
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/mlp/gate_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
216
+ "bytes": 8913264,
217
+ "sha256": "0061d7ac874d6e2bfe05d179963c706783a7441413f41cdc044bd7b88eba78f0"
218
+ },
219
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/mlp/up_proj.weight_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
220
+ "bytes": 8913264,
221
+ "sha256": "3d4275845c162a674c6ecb37b46f336774d3cd71dca3c14beccc23ba30d3dfd6"
222
+ },
223
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/post_attention_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
224
+ "bytes": 2392,
225
+ "sha256": "eefb4511a61815953b3edfe2ca1e2dbe43abf51548e95c6ac4bec8d39e466967"
226
+ },
227
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/post_feedforward_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
228
+ "bytes": 2392,
229
+ "sha256": "181044afe3caded023ab7b0b40bcab19a7656ec448034a3df8ad9f18758c7fcd"
230
+ },
231
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/pre_feedforward_layernorm/weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
232
+ "bytes": 2392,
233
+ "sha256": "9c99dfec9777e33091fe41333da4376f757bdd40f2130b4330aa3fff9a20a199"
234
+ },
235
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/self_attn/k_norm.weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
236
+ "bytes": 1368,
237
+ "sha256": "9ddd85636b12938302c387549428563d6b5509c1e0cd40353bb936fea0a64c07"
238
+ },
239
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/self_attn/o_proj_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
240
+ "bytes": 8913264,
241
+ "sha256": "3d7ff783e7d7318244d040d414a7eee4c3485fe437647a1cf507ee39f10ced76"
242
+ },
243
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/self_attn/q_norm.weight_dtype_BFLOAT16_layout_ROW_MAJOR.tensorbin": {
244
+ "bytes": 1368,
245
+ "sha256": "4512424c238a7935ee2a53d3b56028773aa1124af78a0b3cd76dbcd832e79f2c"
246
+ },
247
+ "tensors/assistant_tensor_cache_bfp8/layer_3/layer_3/self_attn/wqkv_bfp8_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
248
+ "bytes": 10027376,
249
+ "sha256": "e38b05bd4b4dc7eda8178bc223a979f08ac12452bad0ff898e6bc38855a42946"
250
+ },
251
+ "tensors/assistant_tensor_cache_bfp8/model_embed_tokens_weight_bfp4_dtype_BFLOAT4_B_layout_TILE.tensorbin": {
252
+ "bytes": 150995312,
253
+ "sha256": "fa589583b45be7c4fedda493517bd7150cc356f75632d34f9469daaec28a0b3b"
254
+ },
255
+ "tensors/assistant_tensor_cache_bfp8/post_projection_weight_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
256
+ "bytes": 4178288,
257
+ "sha256": "a926c8853869f484c379a9d5b3397488c04905b15d90082ac679ecf7e6466479"
258
+ },
259
+ "tensors/assistant_tensor_cache_bfp8/pre_projection_weight_dtype_BFLOAT8_B_layout_TILE.tensorbin": {
260
+ "bytes": 8356208,
261
+ "sha256": "ba02924ee433b4f9689d504ddadc0e3247d23972021668dc03bd9e27c17670dc"
262
+ }
263
+ },
264
+ "format": "gemma4-12b-it-assistant-ttnn-drafter-1x1-v1",
265
+ "source": {
266
+ "config_sha256": "b6f19209588fcefe41f65b193fad6148446253c470d36e29441ecc5158a54e6d",
267
+ "fetch_verification": {
268
+ "files": {
269
+ ".gitattributes": "ok",
270
+ "README.md": "ok",
271
+ "config.json": "ok",
272
+ "generation_config.json": "ok",
273
+ "model.safetensors": "ok",
274
+ "tokenizer.json": "ok",
275
+ "tokenizer_config.json": "ok"
276
+ },
277
+ "repo": "google/gemma-4-12B-it-assistant",
278
+ "revision": "46d4c6f13f0ac0ad827b915669b8df9b81c64c51"
279
+ },
280
+ "model_id": "google/gemma-4-12B-it-assistant",
281
+ "safetensors_sha256": "3279c173daddd7186e79d652ad94022415736d3a1370625696c898429b06d6df",
282
+ "source_dir": "/work/hf-assistant"
283
+ },
284
+ "target_manifest_sha256": "007d9f3da8d54363a75e8af84cf0c9b762eafd6edf23679de48b08c62d584d55"
285
+ }
gemma-4-12B-it-assistant/generation_config.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 2,
3
+ "do_sample": true,
4
+ "eos_token_id": 1,
5
+ "pad_token_id": 0,
6
+ "suppress_tokens": [
7
+ 258883,
8
+ 258882
9
+ ],
10
+ "temperature": 1.0,
11
+ "top_k": 64,
12
+ "top_p": 0.95,
13
+ "transformers_version": "5.10.0.dev0"
14
+ }
gemma-4-12B-it-assistant/spec_equivalence.json ADDED
@@ -0,0 +1,86 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "draft_len": 5,
3
+ "drafter_manifest_sha256": "2db4559e6ec51d81a7cd51b003743ca689fed5220d41b6a403b8d6947e1c74b7",
4
+ "max_new": 320,
5
+ "prompts_sha256": "12691a871d1878fcd5fd8fdaceacdf412a6300bc38ba3d61413a3cbe0ff76b3c",
6
+ "reference_metadata_sha256": "b1190856d991324991e0206cdbe0b2a3e5c9ab83351bf9f17cce6361aa6524fa",
7
+ "runtime_env": {
8
+ "GEMMA4_FUSE_GELU_MUL": "1",
9
+ "GEMMA4_VERIFY_SDPA": "batched"
10
+ },
11
+ "runtime_identity": {
12
+ "binaries": {
13
+ "build/lib/_ttnncpp.so": "72186b132bf7be85bbb6e1b54bf6de125fbd7a557b12711c84efc6d3e2e08f8d",
14
+ "build/lib/libtt_metal.so": "0497308232f180a0973406e2afd290b9eafa9e310a16b09bf69acc4434eb0e78",
15
+ "ttnn/ttnn/_ttnn.so": "a6448405abd714d6ba310dbbc409c01842f42017b8c1d6bb1c30503fb194c7b2"
16
+ },
17
+ "gemma4_tree": {
18
+ "__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
19
+ "config.py": "ca64c3f783bd7135bec21e2c765004f311631539129ad7ab592b8ae15a2474e4",
20
+ "conftest.py": "5c9425f9367d48a2873dc83ae074bb0cf62fc433a5c8d140cf030930baf4c4ec",
21
+ "demo/__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
22
+ "demo/text_demo.py": "c95cee077db887fd748b4b32cf30f0751d21d14596cc0243048a4afafce98988",
23
+ "demo/text_demo_v2.py": "ad4785b8539e4fdeb32583932e6ac6e2d9bb342edb79d153f5d123d9b323a7d0",
24
+ "tt/__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
25
+ "tt/assistant/__init__.py": "229bb6de453e5f58c9c6bc80368cc417cc943ac0d6a6a268b845e91198ca2aa2",
26
+ "tt/assistant/model.py": "0e3247c34a359a73bd4c8ecc7365e028372084554eeeba41682d1fd818c296af",
27
+ "tt/attention/__init__.py": "df24b2bf4311ffaee781692ff8f465dea0f415e2323cf87fba4ce22227edb8ae",
28
+ "tt/attention/config.py": "01b0f3e14c31aae5ed01104aee7f4cb03913007c0d0018f9d35d341d3fe049ef",
29
+ "tt/attention/decode.py": "0ec303ff75ded2a03f773c22aba27c719e2f055c71b9560fe18661d0272fd3db",
30
+ "tt/attention/kv_cache.py": "3a1e9f520cdde61dd49550dbfca1072d009ee34878bafe5923c655821c09bff7",
31
+ "tt/attention/kv_cache_hybrid.py": "9074a2d1de4b3a65d11271370f1fce39a78e5523b841752bcaa87dccb4b56470",
32
+ "tt/attention/operations.py": "f5cda2ade4031a66e631a732199f3b4b9a00b45ef2e252aee0d3138406bbcc43",
33
+ "tt/attention/prefill.py": "32c6e1c9423b9c22f14558c33d76d9164cb8a27c939cb5d1ccb6da039b7c831f",
34
+ "tt/attention/weights.py": "3ec94d60c6d1085bed0bf83fbc450aa9599a476341d3c6d77466b1b2e17c00d6",
35
+ "tt/ccl.py": "3580f2e25950381465e55dd2886d2832c33d8d6fc618b5a2b94088d0b1a91105",
36
+ "tt/common.py": "e4369b067c3c2e7930fb538a34ab04b237940048e771f2f8f0560776f03a578f",
37
+ "tt/decode_mm.py": "93ebedf2eb3b0c78ebbddfc00c630b13c726d9b13bb07fb53a073a207731bae4",
38
+ "tt/experts/__init__.py": "567aef533a30a692e4c0b28695c15d13e2cfc024ee621e4147c836b63220d5c3",
39
+ "tt/experts/config.py": "80bd7d18ffd0e0b6bd969d05a8e9972ff8473dc63631f1cb6de0bd619a50f33b",
40
+ "tt/experts/decode.py": "112b9df4b88e8395053fdbf69315e3a43e9f3f4877e699ca5a48462ccd8e804b",
41
+ "tt/experts/operations.py": "27db30001f4280dc46bd500d5d811f56ccd8a0896d4eb9f90a6ef50e9af3355b",
42
+ "tt/experts/prefill.py": "5f07a86da85d5ba64838e03d6627c7a78ccfcc06ebfbdaff39d8f7b52f84948d",
43
+ "tt/experts/weights.py": "dda56687ff669a92b8447dc098e4cad6007f94d66b642f1da48c17abd8568916",
44
+ "tt/gemma4_attention_config.py": "8e892a13f89acadd2415719c7c22d188ee906f5fabb1d326464b65afa06816be",
45
+ "tt/gemma4_expert_config.py": "e7b1037156bc3f20cacba6e7a59c5456ebaefa3ed772310f7476908b26aaf20a",
46
+ "tt/generator.py": "34914394ca36e2bc5942b28710ebc7362d7b610c4eb2db33745311b391ceee9e",
47
+ "tt/generator_trace.py": "d5c303e5fe936bf32d097d4013336dcfc587a3a2502cd71fd266c00ab6e2c9f3",
48
+ "tt/generator_vllm.py": "e3b9831358e44e6bcc1433e39ea13b82a7c15565826a38a00815a8a7640b65a4",
49
+ "tt/greedy_decode.py": "a03ac98749300d79ebf8e124b8fd9507df556fc4d518bf46776f7a7df08cf521",
50
+ "tt/kernels/kv_rows_write.cpp": "af410a5fe3dbeefc873eee2f3eaf5136d368fc8ad616730700519a1d6ce83b6d",
51
+ "tt/layer.py": "e3261517b9328d460f84dd888f13694bb50047eb70240e6cc1967f16803599e0",
52
+ "tt/model.py": "2075ea168f8c4fffe90195cadf3ea235bfafdcbaca6daf16081fd05e24106a6c",
53
+ "tt/model_config.py": "04ecd4a462f1ba726186e2c7fce7c3c86b78faf21091af9c9c098f325bb9556d",
54
+ "tt/moe.py": "37d5c5921fe9e5dc61f7b788bc0b3dfec349ce17915dc5450cf286a217f64e41",
55
+ "tt/precision.py": "588e9c3f9864c248175f0c8b2c8b45a5e45511568f71d448cf8ed1621c1d6571",
56
+ "tt/rms_norm.py": "cf0f9fb6191ae3d94191aa62a81686eeccb46b6c44635a1c75967d803886b213",
57
+ "tt/router.py": "61cfcfc0f6b3a7899f8d0b0a9de3738d8af933602e30d2fc4550c05ea9b19ff9",
58
+ "tt/serving_vllm.py": "93f5f00e4fa5a287254d3c7d24b12bccc76aa2e0930fe59f2ab60b47f728c2e4",
59
+ "tt/shared_mlp.py": "fee9764001ac7587810bbb253d80dc29b8997056a9e35540710b60292314e161",
60
+ "tt/single_user.py": "e1535718aeaaabb6b26074e66e3bf5a64f39be0e6dc44e17fcc3f70036b4b8df",
61
+ "tt/spec_decode.py": "8acf4609aca527c123b5d064236d4e817bca7e481702ad6580c0c417a30885e7",
62
+ "tt/spec_greedy.py": "efc6ef2f6df41b77b79003a1ba213c006b397e6e4c44baa723f81dda4ef3e956",
63
+ "utils/__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
64
+ "utils/general_utils.py": "df7324ddc31c12d8177ae8ff007a220cb6e298b3ade498205ecaf5e93abe1eb9",
65
+ "utils/substate.py": "dd0c86aa2b1d6131d45f09ede19690e559543c51dc1166c858277a5c183963ac"
66
+ }
67
+ },
68
+ "scope": "greedy token streams of speculative decoding equal non-speculative greedy decoding on the recorded prompts, same proven runtime",
69
+ "speculative_metadata_sha256": "8d73468115afbf703356394074e01a74a3cfcacac68650214bb52d19d0a3ed0d",
70
+ "streams_identical": true,
71
+ "streams_sha256": {
72
+ "code": "e81a12880f67e38b9409ceb85846ca731a0cdb9efe4046d458bae25c1be09417",
73
+ "code-c": "4755648cb1b633a42596b416aeae1d2e4d2519b29188d31d7be9f85a73eb05cc",
74
+ "code-lru": "2ddfbc830a509221a668dc6f430dd4c7f65d424f0a47ef0297fe9825ce64bcdd",
75
+ "code-sql": "36484b712be95b98a7fc1b0e06fd2b4a3d4f09017f3a6f791102529d803c37c8",
76
+ "explain": "c4c3e77a59cab64b75d4604f3dadcd75fa06cbcdfb77c4e6e3209a5f10691332",
77
+ "logic": "f28d9e2ee05d50b2225c186bfe2fdeb2c5464f5e4809ebe6fc4c70e7a72b62fa",
78
+ "long-code": "557d5b465ba6ba3b9260b33dd5b7d943c70599b8de06c669be70ec70178b50dd",
79
+ "long-summary": "f9a898b16dc9eed88e73e18d07c9abb019f937bea73287f1d101dea86272ff6e",
80
+ "math": "c109ed7a56f25c977d1e0ca80337a40083fe4c8172eed6a2dc8ad081f91244b2",
81
+ "story": "ded163eb533ec4a7e47cd4351644d35c209852c8418503e060b988ddf6765414"
82
+ },
83
+ "target_manifest_sha256": "007d9f3da8d54363a75e8af84cf0c9b762eafd6edf23679de48b08c62d584d55",
84
+ "target_proof_sha256": "1bd70995c42ce9d484c717d4169a000df36ce5181798c6f54b26138128daf296",
85
+ "tokens_compared": 3130
86
+ }
gemma-4-12B-it/chat_template.jinja ADDED
@@ -0,0 +1,390 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {#
2
+ Template: Google Gemma 4 Canonical Chat Template
3
+ Author: Google Gemma Engineering Team
4
+ Published: 2026-07-09
5
+ Context: Fixed tool-calling loops, turn closures, and thinking content-ordering.
6
+ #}
7
+ {%- macro format_parameters(properties, required, filter_keys=false) -%}
8
+ {%- set standard_keys = ['description', 'type', 'properties', 'required', 'nullable'] -%}
9
+ {%- set ns = namespace(found_first=false) -%}
10
+ {%- for key, value in properties | dictsort -%}
11
+ {%- set add_comma = false -%}
12
+ {%- if not filter_keys or key not in standard_keys -%}
13
+ {%- if ns.found_first %},{% endif -%}
14
+ {%- set ns.found_first = true -%}
15
+ {{ key }}:{
16
+ {%- if value['description'] -%}
17
+ description:<|"|>{{ value['description'] }}<|"|>
18
+ {%- set add_comma = true -%}
19
+ {%- endif -%}
20
+ {%- if value['type'] | upper == 'STRING' -%}
21
+ {%- if value['enum'] -%}
22
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
23
+ enum:{{ format_argument(value['enum']) }}
24
+ {%- endif -%}
25
+ {%- elif value['type'] | upper == 'ARRAY' -%}
26
+ {%- if value['items'] is mapping and value['items'] -%}
27
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
28
+ items:{
29
+ {%- set ns_items = namespace(found_first=false) -%}
30
+ {%- for item_key, item_value in value['items'] | dictsort -%}
31
+ {%- if item_value is not none -%}
32
+ {%- if ns_items.found_first %},{% endif -%}
33
+ {%- set ns_items.found_first = true -%}
34
+ {%- if item_key == 'properties' -%}
35
+ properties:{
36
+ {%- if item_value is mapping -%}
37
+ {{- format_parameters(item_value, value['items']['required'] | default([])) -}}
38
+ {%- endif -%}
39
+ }
40
+ {%- elif item_key == 'required' -%}
41
+ required:[
42
+ {%- for req_item in item_value -%}
43
+ <|"|>{{- req_item -}}<|"|>
44
+ {%- if not loop.last %},{% endif -%}
45
+ {%- endfor -%}
46
+ ]
47
+ {%- elif item_key == 'type' -%}
48
+ {%- if item_value is string -%}
49
+ type:{{ format_argument(item_value | upper) }}
50
+ {%- else -%}
51
+ type:{{ format_argument(item_value | map('upper') | list) }}
52
+ {%- endif -%}
53
+ {%- else -%}
54
+ {{ item_key }}:{{ format_argument(item_value) }}
55
+ {%- endif -%}
56
+ {%- endif -%}
57
+ {%- endfor -%}
58
+ }
59
+ {%- endif -%}
60
+ {%- endif -%}
61
+ {%- if value['nullable'] %}
62
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
63
+ nullable:true
64
+ {%- endif -%}
65
+ {%- if value['type'] | upper == 'OBJECT' -%}
66
+ {%- if value['properties'] is defined and value['properties'] is mapping -%}
67
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
68
+ properties:{
69
+ {{- format_parameters(value['properties'], value['required'] | default([])) -}}
70
+ }
71
+ {%- elif value is mapping -%}
72
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
73
+ properties:{
74
+ {{- format_parameters(value, value['required'] | default([]), filter_keys=true) -}}
75
+ }
76
+ {%- endif -%}
77
+ {%- if value['required'] -%}
78
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
79
+ required:[
80
+ {%- for item in value['required'] | default([]) -%}
81
+ <|"|>{{- item -}}<|"|>
82
+ {%- if not loop.last %},{% endif -%}
83
+ {%- endfor -%}
84
+ ]
85
+ {%- endif -%}
86
+ {%- endif -%}
87
+ {%- if add_comma %},{%- else -%} {%- set add_comma = true -%} {% endif -%}
88
+ type:<|"|>{{ value['type'] | upper }}<|"|>}
89
+ {%- endif -%}
90
+ {%- endfor -%}
91
+ {%- endmacro -%}
92
+ {%- macro format_function_declaration(tool_data) -%}
93
+ declaration:{{- tool_data['function']['name'] -}}{description:<|"|>{{- tool_data['function']['description'] -}}<|"|>
94
+ {%- set params = tool_data['function']['parameters'] -%}
95
+ {%- if params -%}
96
+ ,parameters:{
97
+ {%- if params['properties'] -%}
98
+ properties:{ {{- format_parameters(params['properties'], params['required']) -}} },
99
+ {%- endif -%}
100
+ {%- if params['required'] -%}
101
+ required:[
102
+ {%- for item in params['required'] -%}
103
+ <|"|>{{- item -}}<|"|>
104
+ {{- ',' if not loop.last -}}
105
+ {%- endfor -%}
106
+ ],
107
+ {%- endif -%}
108
+ {%- if params['type'] -%}
109
+ type:<|"|>{{- params['type'] | upper -}}<|"|>}
110
+ {%- endif -%}
111
+ {%- endif -%}
112
+ {%- if 'response' in tool_data['function'] -%}
113
+ {%- set response_declaration = tool_data['function']['response'] -%}
114
+ ,response:{
115
+ {%- if response_declaration['description'] -%}
116
+ description:<|"|>{{- response_declaration['description'] -}}<|"|>,
117
+ {%- endif -%}
118
+ {%- if response_declaration['type'] | upper == 'OBJECT' -%}
119
+ type:<|"|>{{- response_declaration['type'] | upper -}}<|"|>}
120
+ {%- endif -%}
121
+ {%- endif -%}
122
+ }
123
+ {%- endmacro -%}
124
+ {%- macro format_argument(argument, escape_keys=True) -%}
125
+ {%- if argument is none -%}
126
+ {{- 'null' -}}
127
+ {%- elif argument is string -%}
128
+ {{- '<|"|>' + argument + '<|"|>' -}}
129
+ {%- elif argument is boolean -%}
130
+ {{- 'true' if argument else 'false' -}}
131
+ {%- elif argument is mapping -%}
132
+ {{- '{' -}}
133
+ {%- set ns = namespace(found_first=false) -%}
134
+ {%- for key, value in argument | dictsort -%}
135
+ {%- if ns.found_first %},{% endif -%}
136
+ {%- set ns.found_first = true -%}
137
+ {%- if escape_keys -%}
138
+ {{- '<|"|>' + key + '<|"|>' -}}
139
+ {%- else -%}
140
+ {{- key -}}
141
+ {%- endif -%}
142
+ :{{- format_argument(value, escape_keys=escape_keys) -}}
143
+ {%- endfor -%}
144
+ {{- '}' -}}
145
+ {%- elif argument is sequence -%}
146
+ {{- '[' -}}
147
+ {%- for item in argument -%}
148
+ {{- format_argument(item, escape_keys=escape_keys) -}}
149
+ {%- if not loop.last %},{% endif -%}
150
+ {%- endfor -%}
151
+ {{- ']' -}}
152
+ {%- else -%}
153
+ {{- argument -}}
154
+ {%- endif -%}
155
+ {%- endmacro -%}
156
+ {%- macro strip_thinking(text) -%}
157
+ {%- set ns = namespace(result='') -%}
158
+ {%- for part in text.split('<channel|>') -%}
159
+ {%- if '<|channel>' in part -%}
160
+ {%- set ns.result = ns.result + part.split('<|channel>')[0] -%}
161
+ {%- else -%}
162
+ {%- set ns.result = ns.result + part -%}
163
+ {%- endif -%}
164
+ {%- endfor -%}
165
+ {{- ns.result | trim -}}
166
+ {%- endmacro -%}
167
+
168
+ {%- macro format_tool_response_block(tool_name, response) -%}
169
+ {{- '<|tool_response>' -}}
170
+ {%- if response is mapping -%}
171
+ {{- 'response:' + tool_name + '{' -}}
172
+ {%- for key, value in response | dictsort -%}
173
+ {{- key -}}:{{- format_argument(value, escape_keys=False) -}}
174
+ {%- if not loop.last %},{% endif -%}
175
+ {%- endfor -%}
176
+ {{- '}' -}}
177
+ {%- else -%}
178
+ {{- 'response:' + tool_name + '{value:' + format_argument(response, escape_keys=False) + '}' -}}
179
+ {%- endif -%}
180
+ {{- '<tool_response|>' -}}
181
+ {%- endmacro -%}
182
+
183
+ {#- ===== SETUP ===== -#}
184
+ {%- set ns = namespace(prev_message_type=None, prev_non_tool_role=None) -%}
185
+ {%- set loop_messages = messages -%}
186
+ {%- set enable_thinking = enable_thinking | default(false) -%}
187
+ {%- set preserve_thinking = preserve_thinking | default(false) -%}
188
+ {{- bos_token -}}
189
+ {#- Handle System/Tool Definitions Block -#}
190
+ {%- if enable_thinking or tools or (messages and messages[0]['role'] in ['system', 'developer']) -%}
191
+ {{- '<|turn>system\n' -}}
192
+ {#- Inject Thinking token at the very top of the FIRST system turn -#}
193
+ {%- if enable_thinking -%}
194
+ {{- '<|think|>\n' -}}
195
+ {%- set ns.prev_message_type = 'think' -%}
196
+ {%- endif -%}
197
+ {%- if messages and messages[0]['role'] in ['system', 'developer'] -%}
198
+ {%- if messages[0]['content'] is string -%}
199
+ {{- messages[0]['content'] | trim -}}
200
+ {%- elif messages[0]['content'] is sequence -%}
201
+ {%- for item in messages[0]['content'] -%}
202
+ {{- item['text'] | trim + ' '-}}
203
+ {%- endfor -%}
204
+ {%- endif -%}
205
+ {%- set loop_messages = messages[1:] -%}
206
+ {%- endif -%}
207
+ {%- if tools -%}
208
+ {%- for tool in tools %}
209
+ {{- '<|tool>' -}}
210
+ {{- format_function_declaration(tool) | trim -}}
211
+ {{- '<tool|>' -}}
212
+ {%- endfor %}
213
+ {%- set ns.prev_message_type = 'tool' -%}
214
+ {%- endif -%}
215
+ {{- '<turn|>\n' -}}
216
+ {%- endif %}
217
+
218
+ {#- Pre-scan: find last user message index for reasoning guard -#}
219
+ {%- set ns_turn = namespace(last_user_idx=-1) -%}
220
+ {%- for i in range(loop_messages | length) -%}
221
+ {%- if loop_messages[i]['role'] == 'user' -%}
222
+ {%- set ns_turn.last_user_idx = i -%}
223
+ {%- endif -%}
224
+ {%- endfor -%}
225
+
226
+ {#- Loop through messages -#}
227
+ {%- for message in loop_messages -%}
228
+ {%- if message['role'] != 'tool' -%}
229
+ {%- set ns.prev_message_type = None -%}
230
+ {%- set role = 'model' if message['role'] == 'assistant' else message['role'] -%}
231
+ {#- Detect continuation using tracked state — O(1) instead of O(n) backward scan -#}
232
+ {%- set continue_same_model_turn = (role == 'model' and ns.prev_non_tool_role == 'assistant') -%}
233
+ {%- if not continue_same_model_turn -%}
234
+ {{- '<|turn>' + role + '\n' }}
235
+
236
+ {%- endif -%}
237
+
238
+ {#- Render reasoning/reasoning_content as thinking channel -#}
239
+ {%- set thinking_text = message.get('reasoning') or message.get('reasoning_content') -%}
240
+ {%- set thinking_gate = (loop.index0 > ns_turn.last_user_idx) or (preserve_thinking and message.get('tool_calls')) -%}
241
+ {%- if thinking_text and thinking_gate -%}
242
+ {{- '<|channel>thought\n' + thinking_text + '\n<channel|>' -}}
243
+ {%- endif -%}
244
+
245
+ {%- if message.get('tool_calls') -%}
246
+ {%- for tool_call in message.get('tool_calls') -%}
247
+ {%- set function = tool_call['function'] -%}
248
+ {{- '<|tool_call>call:' + function['name'] + '{' -}}
249
+ {%- if function['arguments'] is mapping -%}
250
+ {%- set ns_args = namespace(found_first=false) -%}
251
+ {%- for key, value in function['arguments'] | dictsort -%}
252
+ {%- if ns_args.found_first %},{% endif -%}
253
+ {%- set ns_args.found_first = true -%}
254
+ {{- key -}}:{{- format_argument(value, escape_keys=False) -}}
255
+ {%- endfor -%}
256
+ {%- elif function['arguments'] is none -%}
257
+ {%- else -%}
258
+ {{- raise_exception(
259
+ "chat_template: tool_calls[].function.arguments must be a "
260
+ "JSON object (mapping), not a string. Deserialize arguments "
261
+ "before passing to the template."
262
+ ) -}}
263
+ {%- endif -%}
264
+ {{- '}<tool_call|>' -}}
265
+ {%- endfor -%}
266
+ {%- set ns.prev_message_type = 'tool_call' -%}
267
+ {%- endif -%}
268
+
269
+ {%- set ns_tr_out = namespace(flag=false) -%}
270
+ {%- if message.get('tool_responses') -%}
271
+ {#- Legacy: tool_responses embedded on the assistant message (Google/Gemma native) -#}
272
+ {%- for tool_response in message.get('tool_responses') -%}
273
+ {{- format_tool_response_block(tool_response['name'] | default('unknown', true), tool_response['response']) -}}
274
+ {%- set ns_tr_out.flag = true -%}
275
+ {%- set ns.prev_message_type = 'tool_response' -%}
276
+ {%- endfor -%}
277
+ {%- elif message.get('tool_calls') -%}
278
+ {#- OpenAI Chat Completions: forward-scan consecutive role:tool messages -#}
279
+ {%- set ns_tool_scan = namespace(stopped=false) -%}
280
+ {%- for k in range(loop.index0 + 1, loop_messages | length) -%}
281
+ {%- if ns_tool_scan.stopped -%}
282
+ {%- elif loop_messages[k]['role'] != 'tool' -%}
283
+ {%- set ns_tool_scan.stopped = true -%}
284
+ {%- else -%}
285
+ {%- set follow = loop_messages[k] -%}
286
+ {#- Resolve tool_call_id to function name -#}
287
+ {%- set ns_tname = namespace(name=follow.get('name') or 'unknown') -%}
288
+ {%- for tc in message.get('tool_calls') -%}
289
+ {%- if tc.get('id') == follow.get('tool_call_id') -%}
290
+ {%- set ns_tname.name = tc['function']['name'] -%}
291
+ {%- endif -%}
292
+ {%- endfor -%}
293
+ {#- Handle content as string or content-parts array -#}
294
+ {%- set tool_body = follow.get('content') -%}
295
+ {%- if tool_body is string -%}
296
+ {{- format_tool_response_block(ns_tname.name, tool_body) -}}
297
+ {%- elif tool_body is sequence and tool_body is not string -%}
298
+ {%- set ns_txt = namespace(s='') -%}
299
+ {%- for part in tool_body -%}
300
+ {%- if part.get('type') == 'text' -%}
301
+ {%- set ns_txt.s = ns_txt.s + (part.get('text') | default('')) -%}
302
+ {%- endif -%}
303
+ {%- endfor -%}
304
+ {{- format_tool_response_block(ns_tname.name, ns_txt.s) -}}
305
+ {%- for part in tool_body -%}
306
+ {%- if part.get('type') in ['image', 'image_url'] -%}
307
+ {{- '<|image|>' -}}
308
+ {%- elif part.get('type') in ['audio', 'input_audio'] -%}
309
+ {{- '<|audio|>' -}}
310
+ {%- elif part.get('type') == 'video' -%}
311
+ {{- '<|video|>' -}}
312
+ {%- endif -%}
313
+ {%- endfor -%}
314
+ {%- else -%}
315
+ {{- format_tool_response_block(ns_tname.name, tool_body) -}}
316
+ {%- endif -%}
317
+ {%- set ns_tr_out.flag = true -%}
318
+ {%- set ns.prev_message_type = 'tool_response' -%}
319
+ {%- endif -%}
320
+ {%- endfor -%}
321
+ {%- endif -%}
322
+
323
+ {%- set captured_content -%}
324
+ {%- if message.get('content') is string -%}
325
+ {%- if role == 'model' -%}
326
+ {{- strip_thinking(message['content']) -}}
327
+ {%- else -%}
328
+ {{- message['content'] | trim -}}
329
+ {%- endif -%}
330
+ {%- elif message.get('content') is sequence -%}
331
+ {%- for item in message['content'] -%}
332
+ {%- if item.get('type') == 'text' -%}
333
+ {%- if role == 'model' -%}
334
+ {{- strip_thinking(item['text']) -}}
335
+ {%- else -%}
336
+ {{- item['text'] | trim -}}
337
+ {%- endif -%}
338
+ {%- elif item.get('type') in ['image', 'image_url'] -%}
339
+ {{- '<|image|>' -}}
340
+ {%- elif item.get('type') in ['audio', 'input_audio'] -%}
341
+ {{- '<|audio|>' -}}
342
+ {%- elif item.get('type') == 'video' -%}
343
+ {{- '<|video|>' -}}
344
+ {%- endif -%}
345
+ {%- endfor -%}
346
+ {%- endif -%}
347
+ {%- endset -%}
348
+
349
+ {{- captured_content -}}
350
+ {%- set has_content = captured_content | trim | length > 0 -%}
351
+
352
+ {#- Forward-scan: find next non-tool message role for continuation detection -#}
353
+ {%- set next_nt = namespace(role=None, found=false) -%}
354
+ {%- for j in range(loop.index0 + 1, loop_messages | length) -%}
355
+ {%- if not next_nt.found -%}
356
+ {%- if loop_messages[j]['role'] != 'tool' -%}
357
+ {%- set next_nt.role = loop_messages[j]['role'] -%}
358
+ {%- set next_nt.found = true -%}
359
+ {%- endif -%}
360
+ {%- endif -%}
361
+ {%- endfor -%}
362
+
363
+ {%- set continues_into_next = (
364
+ role == 'model'
365
+ and next_nt.role == 'assistant'
366
+ and (not message.get('tool_calls') or ns_tr_out.flag)
367
+ ) -%}
368
+
369
+ {%- if ns.prev_message_type == 'tool_call' and not ns_tr_out.flag -%}
370
+ {{- '<|tool_response>' -}}
371
+ {%- elif continues_into_next -%}
372
+ {%- elif not (ns_tr_out.flag and not has_content and not next_nt.found) -%}
373
+ {{- '<turn|>\n' -}}
374
+ {%- endif -%}
375
+
376
+ {#- Track previous non-tool role for next iteration (avoids O(n) backward scan) -#}
377
+ {%- set ns.prev_non_tool_role = message['role'] -%}
378
+ {%- endif -%}
379
+ {%- endfor -%}
380
+
381
+ {%- if add_generation_prompt -%}
382
+ {%- if ns.prev_message_type != 'tool_response' and ns.prev_message_type != 'tool_call' -%}
383
+ {{- '<|turn>model\n' -}}
384
+ {%- if not enable_thinking -%}
385
+ {{- '<|channel>thought\n<channel|>' -}}
386
+ {%- endif -%}
387
+ {%- elif ns.prev_message_type == 'tool_response' and enable_thinking -%}
388
+ {{- '<|channel>thought\n' -}}
389
+ {%- endif -%}
390
+ {%- endif -%}
gemma-4-12B-it/config.json ADDED
@@ -0,0 +1,172 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Gemma4UnifiedForConditionalGeneration"
4
+ ],
5
+ "audio_config": {
6
+ "_name_or_path": "",
7
+ "architectures": null,
8
+ "audio_embed_dim": 640,
9
+ "audio_samples_per_token": 640,
10
+ "chunk_size_feed_forward": 0,
11
+ "dtype": null,
12
+ "hidden_size": 640,
13
+ "id2label": {
14
+ "0": "LABEL_0",
15
+ "1": "LABEL_1"
16
+ },
17
+ "initializer_range": 0.02,
18
+ "is_encoder_decoder": false,
19
+ "label2id": {
20
+ "LABEL_0": 0,
21
+ "LABEL_1": 1
22
+ },
23
+ "model_type": "gemma4_unified_audio",
24
+ "output_attentions": false,
25
+ "output_hidden_states": false,
26
+ "output_proj_dims": 640,
27
+ "problem_type": null,
28
+ "return_dict": true,
29
+ "rms_norm_eps": 1e-06
30
+ },
31
+ "audio_token_id": 258881,
32
+ "boa_token_id": 256000,
33
+ "boi_token_id": 255999,
34
+ "dtype": "bfloat16",
35
+ "eoa_token_index": 258883,
36
+ "eoi_token_id": 258882,
37
+ "eos_token_id": [
38
+ 1,
39
+ 106
40
+ ],
41
+ "image_token_id": 258880,
42
+ "initializer_range": 0.02,
43
+ "model_type": "gemma4_unified",
44
+ "text_config": {
45
+ "attention_bias": false,
46
+ "attention_dropout": 0.0,
47
+ "attention_k_eq_v": true,
48
+ "bos_token_id": 2,
49
+ "enable_moe_block": false,
50
+ "eos_token_id": 1,
51
+ "final_logit_softcapping": 30.0,
52
+ "global_head_dim": 512,
53
+ "head_dim": 256,
54
+ "hidden_activation": "gelu_pytorch_tanh",
55
+ "hidden_size": 3840,
56
+ "hidden_size_per_layer_input": 0,
57
+ "initializer_range": 0.02,
58
+ "intermediate_size": 15360,
59
+ "layer_types": [
60
+ "sliding_attention",
61
+ "sliding_attention",
62
+ "sliding_attention",
63
+ "sliding_attention",
64
+ "sliding_attention",
65
+ "full_attention",
66
+ "sliding_attention",
67
+ "sliding_attention",
68
+ "sliding_attention",
69
+ "sliding_attention",
70
+ "sliding_attention",
71
+ "full_attention",
72
+ "sliding_attention",
73
+ "sliding_attention",
74
+ "sliding_attention",
75
+ "sliding_attention",
76
+ "sliding_attention",
77
+ "full_attention",
78
+ "sliding_attention",
79
+ "sliding_attention",
80
+ "sliding_attention",
81
+ "sliding_attention",
82
+ "sliding_attention",
83
+ "full_attention",
84
+ "sliding_attention",
85
+ "sliding_attention",
86
+ "sliding_attention",
87
+ "sliding_attention",
88
+ "sliding_attention",
89
+ "full_attention",
90
+ "sliding_attention",
91
+ "sliding_attention",
92
+ "sliding_attention",
93
+ "sliding_attention",
94
+ "sliding_attention",
95
+ "full_attention",
96
+ "sliding_attention",
97
+ "sliding_attention",
98
+ "sliding_attention",
99
+ "sliding_attention",
100
+ "sliding_attention",
101
+ "full_attention",
102
+ "sliding_attention",
103
+ "sliding_attention",
104
+ "sliding_attention",
105
+ "sliding_attention",
106
+ "sliding_attention",
107
+ "full_attention"
108
+ ],
109
+ "max_position_embeddings": 262144,
110
+ "model_type": "gemma4_unified_text",
111
+ "moe_intermediate_size": null,
112
+ "num_attention_heads": 16,
113
+ "num_experts": null,
114
+ "num_global_key_value_heads": 1,
115
+ "num_hidden_layers": 48,
116
+ "num_key_value_heads": 8,
117
+ "num_kv_shared_layers": 0,
118
+ "pad_token_id": 0,
119
+ "rms_norm_eps": 1e-06,
120
+ "rope_parameters": {
121
+ "full_attention": {
122
+ "partial_rotary_factor": 0.25,
123
+ "rope_theta": 1000000.0,
124
+ "rope_type": "proportional"
125
+ },
126
+ "sliding_attention": {
127
+ "rope_theta": 10000.0,
128
+ "rope_type": "default"
129
+ }
130
+ },
131
+ "sliding_window": 1024,
132
+ "tie_word_embeddings": true,
133
+ "top_k_experts": null,
134
+ "use_bidirectional_attention": "vision",
135
+ "use_cache": true,
136
+ "use_double_wide_mlp": false,
137
+ "vocab_size": 262144,
138
+ "vocab_size_per_layer_input": 262144
139
+ },
140
+ "tie_word_embeddings": true,
141
+ "transformers_version": "5.10.0.dev0",
142
+ "video_token_id": 258884,
143
+ "vision_config": {
144
+ "_name_or_path": "",
145
+ "architectures": null,
146
+ "chunk_size_feed_forward": 0,
147
+ "dtype": null,
148
+ "id2label": {
149
+ "0": "LABEL_0",
150
+ "1": "LABEL_1"
151
+ },
152
+ "initializer_range": 0.02,
153
+ "is_encoder_decoder": false,
154
+ "label2id": {
155
+ "LABEL_0": 0,
156
+ "LABEL_1": 1
157
+ },
158
+ "mm_embed_dim": 3840,
159
+ "mm_posemb_size": 1120,
160
+ "model_patch_size": 48,
161
+ "model_type": "gemma4_unified_vision",
162
+ "num_soft_tokens": 280,
163
+ "output_attentions": false,
164
+ "output_hidden_states": false,
165
+ "output_proj_dims": 3840,
166
+ "patch_size": 16,
167
+ "pooling_kernel_size": 3,
168
+ "problem_type": null,
169
+ "return_dict": true,
170
+ "rms_norm_eps": 1e-06
171
+ }
172
+ }
gemma-4-12B-it/equivalence.json ADDED
@@ -0,0 +1,106 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "baseline_metadata_sha256": "903a232363321fef2ac2ba77417d1c0f7d86ae961136361f99cc614b97930c16",
3
+ "exact_logits_equal": true,
4
+ "manifest_sha256": "007d9f3da8d54363a75e8af84cf0c9b762eafd6edf23679de48b08c62d584d55",
5
+ "precision_plan": {
6
+ "default": {
7
+ "attention": "bfp8",
8
+ "embedding": "bf16",
9
+ "lm_head": "bfp8",
10
+ "shared_mlp": "bfp8"
11
+ }
12
+ },
13
+ "records_full_logits": [
14
+ {
15
+ "baseline_sha256": "6647144f4034bd40266b61df6a163edbdc9fce4e2a6b032fa2fa6e2c557a8eea",
16
+ "id": "chat-ultrachat-test-0-1k",
17
+ "restored_sha256": "6647144f4034bd40266b61df6a163edbdc9fce4e2a6b032fa2fa6e2c557a8eea",
18
+ "rows": 1023
19
+ },
20
+ {
21
+ "baseline_sha256": "9e24b2a7caa64b2ed24685d514365bdb63c5b6bfcf76a2656e9293562c46c8ad",
22
+ "id": "code-stdlib-_pydecimal-1k",
23
+ "restored_sha256": "9e24b2a7caa64b2ed24685d514365bdb63c5b6bfcf76a2656e9293562c46c8ad",
24
+ "rows": 1023
25
+ },
26
+ {
27
+ "baseline_sha256": "30dbb1527de2dc14ee58de257d088e4de66c89c34de3d15b5511d28605bcd5a1",
28
+ "id": "long-book-2k",
29
+ "restored_sha256": "30dbb1527de2dc14ee58de257d088e4de66c89c34de3d15b5511d28605bcd5a1",
30
+ "rows": 2047
31
+ }
32
+ ],
33
+ "records_nll_argmax_equal": [
34
+ "chat-ultrachat-test-0-1k",
35
+ "code-stdlib-_pydecimal-1k",
36
+ "long-book-2k"
37
+ ],
38
+ "restored_metadata_sha256": "673a8647733567367ad5ec14f9a86493dddf1dbf12cb0d892adca329a9c6ed92",
39
+ "runtime_env": {
40
+ "GEMMA4_FUSE_GELU_MUL": "1",
41
+ "GEMMA4_VERIFY_SDPA": "batched"
42
+ },
43
+ "runtime_identity": {
44
+ "binaries": {
45
+ "build/lib/_ttnncpp.so": "72186b132bf7be85bbb6e1b54bf6de125fbd7a557b12711c84efc6d3e2e08f8d",
46
+ "build/lib/libtt_metal.so": "0497308232f180a0973406e2afd290b9eafa9e310a16b09bf69acc4434eb0e78",
47
+ "ttnn/ttnn/_ttnn.so": "a6448405abd714d6ba310dbbc409c01842f42017b8c1d6bb1c30503fb194c7b2"
48
+ },
49
+ "gemma4_tree": {
50
+ "__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
51
+ "config.py": "ca64c3f783bd7135bec21e2c765004f311631539129ad7ab592b8ae15a2474e4",
52
+ "conftest.py": "5c9425f9367d48a2873dc83ae074bb0cf62fc433a5c8d140cf030930baf4c4ec",
53
+ "demo/__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
54
+ "demo/text_demo.py": "c95cee077db887fd748b4b32cf30f0751d21d14596cc0243048a4afafce98988",
55
+ "demo/text_demo_v2.py": "ad4785b8539e4fdeb32583932e6ac6e2d9bb342edb79d153f5d123d9b323a7d0",
56
+ "tt/__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
57
+ "tt/assistant/__init__.py": "229bb6de453e5f58c9c6bc80368cc417cc943ac0d6a6a268b845e91198ca2aa2",
58
+ "tt/assistant/model.py": "0e3247c34a359a73bd4c8ecc7365e028372084554eeeba41682d1fd818c296af",
59
+ "tt/attention/__init__.py": "df24b2bf4311ffaee781692ff8f465dea0f415e2323cf87fba4ce22227edb8ae",
60
+ "tt/attention/config.py": "01b0f3e14c31aae5ed01104aee7f4cb03913007c0d0018f9d35d341d3fe049ef",
61
+ "tt/attention/decode.py": "0ec303ff75ded2a03f773c22aba27c719e2f055c71b9560fe18661d0272fd3db",
62
+ "tt/attention/kv_cache.py": "3a1e9f520cdde61dd49550dbfca1072d009ee34878bafe5923c655821c09bff7",
63
+ "tt/attention/kv_cache_hybrid.py": "9074a2d1de4b3a65d11271370f1fce39a78e5523b841752bcaa87dccb4b56470",
64
+ "tt/attention/operations.py": "f5cda2ade4031a66e631a732199f3b4b9a00b45ef2e252aee0d3138406bbcc43",
65
+ "tt/attention/prefill.py": "32c6e1c9423b9c22f14558c33d76d9164cb8a27c939cb5d1ccb6da039b7c831f",
66
+ "tt/attention/weights.py": "3ec94d60c6d1085bed0bf83fbc450aa9599a476341d3c6d77466b1b2e17c00d6",
67
+ "tt/ccl.py": "3580f2e25950381465e55dd2886d2832c33d8d6fc618b5a2b94088d0b1a91105",
68
+ "tt/common.py": "e4369b067c3c2e7930fb538a34ab04b237940048e771f2f8f0560776f03a578f",
69
+ "tt/decode_mm.py": "93ebedf2eb3b0c78ebbddfc00c630b13c726d9b13bb07fb53a073a207731bae4",
70
+ "tt/experts/__init__.py": "567aef533a30a692e4c0b28695c15d13e2cfc024ee621e4147c836b63220d5c3",
71
+ "tt/experts/config.py": "80bd7d18ffd0e0b6bd969d05a8e9972ff8473dc63631f1cb6de0bd619a50f33b",
72
+ "tt/experts/decode.py": "112b9df4b88e8395053fdbf69315e3a43e9f3f4877e699ca5a48462ccd8e804b",
73
+ "tt/experts/operations.py": "27db30001f4280dc46bd500d5d811f56ccd8a0896d4eb9f90a6ef50e9af3355b",
74
+ "tt/experts/prefill.py": "5f07a86da85d5ba64838e03d6627c7a78ccfcc06ebfbdaff39d8f7b52f84948d",
75
+ "tt/experts/weights.py": "dda56687ff669a92b8447dc098e4cad6007f94d66b642f1da48c17abd8568916",
76
+ "tt/gemma4_attention_config.py": "8e892a13f89acadd2415719c7c22d188ee906f5fabb1d326464b65afa06816be",
77
+ "tt/gemma4_expert_config.py": "e7b1037156bc3f20cacba6e7a59c5456ebaefa3ed772310f7476908b26aaf20a",
78
+ "tt/generator.py": "34914394ca36e2bc5942b28710ebc7362d7b610c4eb2db33745311b391ceee9e",
79
+ "tt/generator_trace.py": "d5c303e5fe936bf32d097d4013336dcfc587a3a2502cd71fd266c00ab6e2c9f3",
80
+ "tt/generator_vllm.py": "e3b9831358e44e6bcc1433e39ea13b82a7c15565826a38a00815a8a7640b65a4",
81
+ "tt/greedy_decode.py": "a03ac98749300d79ebf8e124b8fd9507df556fc4d518bf46776f7a7df08cf521",
82
+ "tt/kernels/kv_rows_write.cpp": "af410a5fe3dbeefc873eee2f3eaf5136d368fc8ad616730700519a1d6ce83b6d",
83
+ "tt/layer.py": "e3261517b9328d460f84dd888f13694bb50047eb70240e6cc1967f16803599e0",
84
+ "tt/model.py": "2075ea168f8c4fffe90195cadf3ea235bfafdcbaca6daf16081fd05e24106a6c",
85
+ "tt/model_config.py": "04ecd4a462f1ba726186e2c7fce7c3c86b78faf21091af9c9c098f325bb9556d",
86
+ "tt/moe.py": "37d5c5921fe9e5dc61f7b788bc0b3dfec349ce17915dc5450cf286a217f64e41",
87
+ "tt/precision.py": "588e9c3f9864c248175f0c8b2c8b45a5e45511568f71d448cf8ed1621c1d6571",
88
+ "tt/rms_norm.py": "cf0f9fb6191ae3d94191aa62a81686eeccb46b6c44635a1c75967d803886b213",
89
+ "tt/router.py": "61cfcfc0f6b3a7899f8d0b0a9de3738d8af933602e30d2fc4550c05ea9b19ff9",
90
+ "tt/serving_vllm.py": "93f5f00e4fa5a287254d3c7d24b12bccc76aa2e0930fe59f2ab60b47f728c2e4",
91
+ "tt/shared_mlp.py": "fee9764001ac7587810bbb253d80dc29b8997056a9e35540710b60292314e161",
92
+ "tt/single_user.py": "e1535718aeaaabb6b26074e66e3bf5a64f39be0e6dc44e17fcc3f70036b4b8df",
93
+ "tt/spec_decode.py": "8acf4609aca527c123b5d064236d4e817bca7e481702ad6580c0c417a30885e7",
94
+ "tt/spec_greedy.py": "efc6ef2f6df41b77b79003a1ba213c006b397e6e4c44baa723f81dda4ef3e956",
95
+ "utils/__init__.py": "501d5b563ca412fa21db82bf39117b5480ecf99bcaf11996b9c3d5f1093bb9a8",
96
+ "utils/general_utils.py": "df7324ddc31c12d8177ae8ff007a220cb6e298b3ade498205ecaf5e93abe1eb9",
97
+ "utils/substate.py": "dd0c86aa2b1d6131d45f09ede19690e559543c51dc1166c858277a5c183963ac"
98
+ }
99
+ },
100
+ "runtime_run_sha256": {
101
+ "gemma4_tree_sha256": "ffae949a9a260c3a03a9b0a2ef891668bf79fb790df84f53e24e19048e9fe1fb",
102
+ "tool_sha256": "36485b68752f51c3642e3dd777f5da0cdb9de61ab819d5680f95402e1a24ede3"
103
+ },
104
+ "scope": "bit-exact full-vocabulary logits, original weights quantized at load vs native reload, same runtime; not a quality certification",
105
+ "tokens_compared": 4093
106
+ }
gemma-4-12B-it/generation_config.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 2,
3
+ "do_sample": true,
4
+ "eos_token_id": [
5
+ 1,
6
+ 106,
7
+ 50
8
+ ],
9
+ "pad_token_id": 0,
10
+ "suppress_tokens": [
11
+ 258883,
12
+ 258882
13
+ ],
14
+ "temperature": 1.0,
15
+ "top_k": 64,
16
+ "top_p": 0.95,
17
+ "transformers_version": "5.10.0.dev0"
18
+ }
gemma-4-12B-it/native_manifest.json ADDED
The diff for this file is too large to render. See raw diff
 
gemma-4-12B-it/precision_plan.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "default": {
3
+ "attention": "bfp8",
4
+ "shared_mlp": "bfp8",
5
+ "lm_head": "bfp8",
6
+ "embedding": "bf16"
7
+ }
8
+ }
gemma-4-12B-it/tokenizer_config.json ADDED
@@ -0,0 +1,120 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "audio_token": "<|audio|>",
3
+ "backend": "tokenizers",
4
+ "boa_token": "<|audio>",
5
+ "boi_token": "<|image>",
6
+ "bos_token": "<bos>",
7
+ "eoa_token": "<audio|>",
8
+ "eoc_token": "<channel|>",
9
+ "eoi_token": "<image|>",
10
+ "eos_token": "<eos>",
11
+ "eot_token": "<turn|>",
12
+ "escape_token": "<|\"|>",
13
+ "etc_token": "<tool_call|>",
14
+ "etd_token": "<tool|>",
15
+ "etr_token": "<tool_response|>",
16
+ "extra_special_tokens": [
17
+ "<|video|>"
18
+ ],
19
+ "image_token": "<|image|>",
20
+ "mask_token": "<mask>",
21
+ "model_max_length": 1000000000000000019884624838656,
22
+ "pad_token": "<pad>",
23
+ "padding_side": "left",
24
+ "processor_class": "Gemma4UnifiedProcessor",
25
+ "response_schema": {
26
+ "properties": {
27
+ "content": {
28
+ "type": "string"
29
+ },
30
+ "role": {
31
+ "const": "assistant"
32
+ },
33
+ "thinking": {
34
+ "type": "string"
35
+ },
36
+ "tool_calls": {
37
+ "items": {
38
+ "properties": {
39
+ "function": {
40
+ "properties": {
41
+ "arguments": {
42
+ "additionalProperties": {},
43
+ "type": "object",
44
+ "x-parser": "gemma4-tool-call"
45
+ },
46
+ "name": {
47
+ "type": "string"
48
+ }
49
+ },
50
+ "type": "object",
51
+ "x-regex": "call\\:(?P<name>\\w+)(?P<arguments>\\{.*\\})"
52
+ },
53
+ "type": {
54
+ "const": "function"
55
+ }
56
+ },
57
+ "type": "object"
58
+ },
59
+ "type": "array",
60
+ "x-regex-iterator": "<\\|tool_call>(.*?)<tool_call\\|>"
61
+ }
62
+ },
63
+ "type": "object",
64
+ "x-regex": "(\\<\\|channel\\>thought\\n(?P<thinking>.*?)\\<channel\\|\\>)?(?P<tool_calls>\\<\\|tool_call\\>.*\\<tool_call\\|\\>)?(?P<content>(?:(?!\\<turn\\|\\>)(?!\\<\\|tool_response\\>).)+)?(?:\\<turn\\|\\>|\\<\\|tool_response\\>)?"
65
+ },
66
+ "response_template": {
67
+ "defaults": {
68
+ "role": "assistant"
69
+ },
70
+ "fields": {
71
+ "content": {
72
+ "close": [
73
+ "<turn|>",
74
+ "<|tool_response>",
75
+ "<eos>"
76
+ ],
77
+ "content": "text"
78
+ },
79
+ "thinking": {
80
+ "close": "<channel|>",
81
+ "content": "text",
82
+ "open": "<|channel>thought\n"
83
+ },
84
+ "tool_calls": {
85
+ "close": "<tool_call|>",
86
+ "content": "json",
87
+ "content_args": {
88
+ "string_delims": [
89
+ [
90
+ "<|\"|>",
91
+ "<|\"|>"
92
+ ]
93
+ ],
94
+ "unquoted_keys": true
95
+ },
96
+ "open_pattern": "<\\|tool_call>call:(?P<name>\\w+)",
97
+ "repeats": true,
98
+ "transform": {
99
+ "function": {
100
+ "arguments": "{content}",
101
+ "name": "{name}"
102
+ },
103
+ "type": "function"
104
+ }
105
+ }
106
+ },
107
+ "start_anchor": [
108
+ "<|turn>model\n",
109
+ "<tool_response|>"
110
+ ]
111
+ },
112
+ "soc_token": "<|channel>",
113
+ "sot_token": "<|turn>",
114
+ "stc_token": "<|tool_call>",
115
+ "std_token": "<|tool>",
116
+ "str_token": "<|tool_response>",
117
+ "think_token": "<|think|>",
118
+ "tokenizer_class": "GemmaTokenizer",
119
+ "unk_token": "<unk>"
120
+ }
launch.py ADDED
@@ -0,0 +1,300 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Download and serve the public, verified Gemma 4 12B TT-native all-BFP8 P150 release.
3
+
4
+ Requires Python 3.11+, Docker, and an exclusively caller-owned P150 to serve.
5
+ No Hugging Face login, Python ML packages, or original model weights are needed.
6
+ Downloads are anonymous; serving is offline inside the ID-pinned container.
7
+ """
8
+
9
+ import argparse
10
+ import hashlib
11
+ import json
12
+ import os
13
+ from pathlib import Path
14
+ import re
15
+ import stat
16
+ import sys
17
+ import subprocess
18
+ import tempfile
19
+ from urllib.error import URLError
20
+ from urllib.parse import quote
21
+ from urllib.request import Request, urlopen
22
+
23
+
24
+ MODEL_REPO = "Lottolabs/gemma-4-12B-it-TT-BFP8-P150"
25
+ RELEASE_ID = "p3c-20260930"
26
+ IMAGE_TAG = "lottolabs/gemma4-12b-tt-p150:p3c"
27
+ CHECKPOINT_SUBDIR = "gemma-4-12B-it"
28
+ DRAFTER_SUBDIR = "gemma-4-12B-it-assistant"
29
+ HF_BASE = "https://huggingface.co"
30
+ SHA256 = re.compile(r"[0-9a-f]{64}")
31
+ CHUNK_SIZE = 1024 * 1024
32
+
33
+
34
+ def request(url):
35
+ # urllib does not load HF tokens, netrc credentials, or ML dependencies.
36
+ return urlopen(Request(url, headers={"User-Agent": "gemma4-12b-native-launch/1"}), timeout=120)
37
+
38
+
39
+ def fetch_json(url):
40
+ with request(url) as response:
41
+ raw = response.read(16 * CHUNK_SIZE + 1)
42
+ if len(raw) > 16 * CHUNK_SIZE:
43
+ raise ValueError("Remote JSON exceeds the 16 MiB metadata limit")
44
+ value = json.loads(raw)
45
+ if not isinstance(value, dict):
46
+ raise ValueError("Expected a JSON object")
47
+ return value, raw
48
+
49
+
50
+ def safe_relative(value):
51
+ if (not isinstance(value, str) or not value
52
+ or any(part in ("", ".", "..") for part in value.split("/"))
53
+ or "\\" in value or ":" in value
54
+ or any(ord(char) < 32 or ord(char) == 127 for char in value)):
55
+ raise ValueError(f"Unsafe release path: {value!r}")
56
+ return value
57
+
58
+
59
+ def file_records(value):
60
+ if not isinstance(value, dict) or not value:
61
+ raise ValueError("Manifest must contain a nonempty files object")
62
+ for relative, record in value.items():
63
+ safe_relative(relative)
64
+ if (not isinstance(record, dict)
65
+ or not isinstance(record.get("sha256"), str)
66
+ or SHA256.fullmatch(record["sha256"]) is None
67
+ or type(record.get("bytes")) is not int or record["bytes"] < 0):
68
+ raise ValueError(f"Invalid file checksum/size: {relative}")
69
+ parts = relative.split("/")
70
+ if any("/".join(parts[:index]) in value for index in range(1, len(parts))):
71
+ raise ValueError(f"File/directory collision: {relative}")
72
+ return value
73
+
74
+
75
+ def validate_release(release):
76
+ if (type(release.get("schema_version")) is not int or release["schema_version"] != 1
77
+ or release.get("model_repo") != MODEL_REPO
78
+ or release.get("release_id") != RELEASE_ID
79
+ or release.get("checkpoint_subdir") != CHECKPOINT_SUBDIR
80
+ or release.get("drafter_subdir") != DRAFTER_SUBDIR):
81
+ raise ValueError("Unsupported runtime release identity or schema")
82
+ image = release.get("image")
83
+ image_id = release.get("image_id")
84
+ image_archive = release.get("image_archive")
85
+ if (image != IMAGE_TAG or not isinstance(image_id, str)
86
+ or re.fullmatch(r"sha256:[0-9a-f]{64}", image_id) is None
87
+ or image_archive != "runtime-image.tar.gz"):
88
+ raise ValueError("Release must identify the checksum-pinned bundled runtime image")
89
+ files = file_records(release.get("files"))
90
+ target_manifest = f"{CHECKPOINT_SUBDIR}/native_manifest.json"
91
+ drafter_manifest = f"{DRAFTER_SUBDIR}/drafter_manifest.json"
92
+ required = {"serve_native.py", "native_checkpoint.py", target_manifest, f"{CHECKPOINT_SUBDIR}/equivalence.json",
93
+ drafter_manifest, f"{DRAFTER_SUBDIR}/spec_equivalence.json", image_archive}
94
+ if not required.issubset(files):
95
+ raise ValueError("Release is missing the server, native manifests, or equivalence proofs")
96
+ if "runtime-release.json" in files or "launch.py" in files:
97
+ raise ValueError("Release inventory must exclude runtime-release.json and launch.py")
98
+ if (release.get("checkpoint_manifest_sha256") != files[target_manifest]["sha256"]
99
+ or release.get("drafter_manifest_sha256") != files[drafter_manifest]["sha256"]):
100
+ raise ValueError("Release manifest checksums do not match its file inventory")
101
+ return files
102
+
103
+
104
+ def ensure_image(image, expected_id, archive):
105
+ inspect = subprocess.run(
106
+ ["docker", "image", "inspect", image, "--format", "{{.Id}}"],
107
+ text=True, capture_output=True)
108
+ if inspect.returncode != 0 or inspect.stdout.strip() != expected_id:
109
+ print(f"Loading {archive.name} into Docker", file=sys.stderr, flush=True)
110
+ loaded = subprocess.run(["docker", "load", "--input", str(archive)], text=True)
111
+ if loaded.returncode != 0:
112
+ raise ValueError("Docker could not load the verified runtime image archive")
113
+ inspect = subprocess.run(
114
+ ["docker", "image", "inspect", image, "--format", "{{.Id}}"],
115
+ text=True, capture_output=True)
116
+ if inspect.returncode != 0 or inspect.stdout.strip() != expected_id:
117
+ raise ValueError("Loaded runtime image identity does not match runtime-release.json")
118
+
119
+
120
+ def cache_path(root, relative, directory=False):
121
+ """Do not follow pre-existing links or special files within a snapshot."""
122
+ current = root
123
+ parts = safe_relative(relative).split("/")
124
+ for index, part in enumerate(parts):
125
+ current = current / part
126
+ if current.is_symlink():
127
+ raise ValueError(f"Refusing a symlink in the release cache: {current}")
128
+ if current.exists():
129
+ mode = current.stat().st_mode
130
+ expected = stat.S_ISDIR if directory or index < len(parts) - 1 else stat.S_ISREG
131
+ if not expected(mode):
132
+ raise ValueError(f"Unexpected cache entry type: {current}")
133
+ return current
134
+
135
+
136
+ def verified(path, record):
137
+ if not path.is_file() or path.stat().st_size != record["bytes"]:
138
+ return False
139
+ with path.open("rb") as source:
140
+ return hashlib.file_digest(source, "sha256").hexdigest() == record["sha256"]
141
+
142
+
143
+ def atomic_write(destination, chunks, record=None):
144
+ destination.parent.mkdir(parents=True, exist_ok=True)
145
+ temporary = None
146
+ try:
147
+ with tempfile.NamedTemporaryFile(prefix=".download-", dir=destination.parent, delete=False) as target:
148
+ temporary = Path(target.name)
149
+ digest = hashlib.sha256()
150
+ size = 0
151
+ for chunk in chunks:
152
+ size += len(chunk)
153
+ if record is not None and size > record["bytes"]:
154
+ raise ValueError(f"Download exceeds expected size: {destination.name}")
155
+ target.write(chunk)
156
+ digest.update(chunk)
157
+ if record is not None and (size != record["bytes"] or digest.hexdigest() != record["sha256"]):
158
+ raise ValueError(f"Download failed SHA256/size verification: {destination}")
159
+ target.flush()
160
+ os.fsync(target.fileno())
161
+ # The container must be able to read the read-only checkpoint mounts.
162
+ os.fchmod(target.fileno(), 0o644)
163
+ os.replace(temporary, destination)
164
+ finally:
165
+ if temporary is not None:
166
+ temporary.unlink(missing_ok=True)
167
+
168
+
169
+ def download(snapshot, base_url, relative, record):
170
+ destination = cache_path(snapshot, relative)
171
+ if verified(destination, record):
172
+ return
173
+ print(f"Downloading {relative} ({record['bytes']:,} bytes)", file=sys.stderr, flush=True)
174
+ with request(base_url + quote(relative, safe="/")) as response:
175
+ atomic_write(destination, iter(lambda: response.read(CHUNK_SIZE), b""), record)
176
+
177
+
178
+ def covered(snapshot, files, subdir, manifest_name, excluded):
179
+ """The published inventory must be exactly the manifest's file set plus the manifest and its proof."""
180
+ manifest = json.loads((snapshot / subdir / manifest_name).read_bytes())
181
+ if not isinstance(manifest, dict):
182
+ raise ValueError(f"{subdir}/{manifest_name} must be a JSON object")
183
+ listed = file_records(manifest.get("files"))
184
+ for relative, record in listed.items():
185
+ published = files.get(f"{subdir}/{relative}")
186
+ if published is None or any(published[key] != record[key] for key in ("sha256", "bytes")):
187
+ raise ValueError(f"Release inventory does not cover {subdir}: {relative}")
188
+ extra = {name for name in files if name.startswith(subdir + "/")} - {f"{subdir}/{name}" for name in (*listed, *excluded)}
189
+ if extra:
190
+ raise ValueError(f"Release inventory has files outside the {subdir} manifest: {sorted(extra)[:3]}")
191
+
192
+
193
+ def independent_kernel_cache(snapshot, kernel_cache):
194
+ if (kernel_cache == snapshot or kernel_cache.is_relative_to(snapshot)
195
+ or snapshot.is_relative_to(kernel_cache)):
196
+ raise ValueError("Kernel cache must be independent of the downloaded release directory")
197
+ if ":" in str(snapshot) or ":" in str(kernel_cache):
198
+ raise ValueError("Docker bind-mount paths cannot contain ':'")
199
+ kernel_cache.mkdir(parents=True, exist_ok=True)
200
+ with tempfile.TemporaryFile(dir=kernel_cache) as probe:
201
+ probe.write(b"writable")
202
+ probe.flush()
203
+
204
+
205
+ def parser():
206
+ cache_home = Path(os.environ.get("XDG_CACHE_HOME") or Path.home() / ".cache")
207
+ result = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter,
208
+ epilog="""Examples:
209
+ python3 launch.py --download-only --cache-root /large-disk/gemma4-12b
210
+ python3 launch.py --cache-root /large-disk/gemma4-12b --print-command
211
+ python3 launch.py --cache-root /large-disk/gemma4-12b --device-ownership-confirmed
212
+ python3 launch.py --speculation off --port 8001 --device-ownership-confirmed
213
+
214
+ --print-command downloads/verifies the release but does not run Docker.
215
+ --download-only downloads/verifies without requiring Docker or device ownership.
216
+ Serving never stops containers, resets the device, or takes ownership for you.
217
+ Stop your other accelerator workloads yourself before confirming ownership.
218
+ The API binds to 127.0.0.1 only. Keep the printed immutable revision to reproduce.
219
+ """)
220
+ result.add_argument("--cache-root", type=Path, default=cache_home / "gemma4-12b-tt-native",
221
+ help="Download/cache root (default: $XDG_CACHE_HOME/gemma4-12b-tt-native or ~/.cache/gemma4-12b-tt-native)")
222
+ result.add_argument("--kernel-cache", type=Path,
223
+ help="Independent writable TT kernel cache (default: CACHE_ROOT/kernel-cache)")
224
+ result.add_argument("--port", type=int, default=8000, help="Loopback API port (default: 8000)")
225
+ result.add_argument("--speculation", choices=("on", "off"), default="on",
226
+ help="Greedy speculative decoding with the proven assistant drafter, K=5 (default: on)")
227
+ result.add_argument("--max-model-len", type=int, default=262144,
228
+ help="Context limit, 1024–262144 (default: 262144)")
229
+ result.add_argument("--name", default="gemma4-12b-native-api",
230
+ help="Docker container name (default: gemma4-12b-native-api)")
231
+ result.add_argument("--revision", default="main", help="Public HF branch, tag, or commit (default: main)")
232
+ result.add_argument("--device-ownership-confirmed", action="store_true",
233
+ help="Confirm you exclusively own the P150 and have stopped your other workloads")
234
+ mode = result.add_mutually_exclusive_group()
235
+ mode.add_argument("--print-command", action="store_true", help="Download, verify, and print the Docker command without running it")
236
+ mode.add_argument("--download-only", action="store_true", help="Download and verify the native package without serving")
237
+ return result
238
+
239
+
240
+ def main():
241
+ cli = parser()
242
+ args = cli.parse_args()
243
+ if sys.version_info < (3, 11):
244
+ cli.error("Python 3.11 or newer is required")
245
+ if not 1 <= args.port <= 65535 or not 1024 <= args.max_model_len <= 262144:
246
+ cli.error("Port must be 1–65535 and max-model-len must be 1024–262144")
247
+ if re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9_.-]*", args.name) is None:
248
+ cli.error("Invalid Docker container name")
249
+ if not args.revision.strip():
250
+ cli.error("Revision must not be empty")
251
+ if not (args.print_command or args.download_only or args.device_ownership_confirmed):
252
+ cli.error("Confirm exclusive P150 ownership with --device-ownership-confirmed, or use --print-command/--download-only")
253
+ try:
254
+ info, _ = fetch_json(f"{HF_BASE}/api/models/{MODEL_REPO}/revision/{quote(args.revision, safe='')}")
255
+ commit = info.get("sha")
256
+ if not isinstance(commit, str) or re.fullmatch(r"[0-9a-f]{40}", commit) is None:
257
+ raise ValueError("Hugging Face did not resolve the revision to an immutable commit")
258
+ base_url = f"{HF_BASE}/{MODEL_REPO}/resolve/{commit}/"
259
+ release, release_raw = fetch_json(base_url + "runtime-release.json")
260
+ files = validate_release(release)
261
+ root = args.cache_root.expanduser().resolve()
262
+ root.mkdir(parents=True, exist_ok=True)
263
+ snapshot = cache_path(root, f"snapshots/{commit}", directory=True)
264
+ snapshot.mkdir(parents=True, exist_ok=True)
265
+ kernel_cache = (args.kernel_cache.expanduser() if args.kernel_cache else root / "kernel-cache").resolve()
266
+ independent_kernel_cache(snapshot, kernel_cache)
267
+ print(f"Revision: {commit}\nImage: {release['image']} ({release['image_id']})"
268
+ f"\nCheckpoint: {snapshot / CHECKPOINT_SUBDIR}\nDrafter: {snapshot / DRAFTER_SUBDIR}"
269
+ f"\nKernel cache: {kernel_cache}", file=sys.stderr, flush=True)
270
+ # Check both native inventories before downloading the multi-gigabyte tensors.
271
+ for subdir, manifest_name, proof in ((CHECKPOINT_SUBDIR, "native_manifest.json", "equivalence.json"),
272
+ (DRAFTER_SUBDIR, "drafter_manifest.json", "spec_equivalence.json")):
273
+ relative = f"{subdir}/{manifest_name}"
274
+ download(snapshot, base_url, relative, files[relative])
275
+ covered(snapshot, files, subdir, manifest_name, (manifest_name, proof))
276
+ for relative, record in files.items():
277
+ download(snapshot, base_url, relative, record)
278
+ atomic_write(cache_path(snapshot, "runtime-release.json"), (release_raw,))
279
+ print(f"All {len(files)} release files verified.", file=sys.stderr, flush=True)
280
+ if args.download_only:
281
+ return 0
282
+ ensure_image(release["image"], release["image_id"], snapshot / release["image_archive"])
283
+ command = [sys.executable, "-B", str(snapshot / "serve_native.py"),
284
+ "--cache-root", str(kernel_cache), "--image", release["image"],
285
+ "--port", str(args.port), "--name", args.name,
286
+ "--max-model-len", str(args.max_model_len)]
287
+ if args.speculation == "off":
288
+ command.append("--no-speculation")
289
+ if args.print_command:
290
+ command.append("--print-command")
291
+ if args.device_ownership_confirmed:
292
+ command.append("--device-ownership-confirmed")
293
+ os.execv(sys.executable, command)
294
+ except (OSError, URLError, ValueError) as error:
295
+ cli.exit(1, f"launch.py: {error}\n")
296
+ return 0
297
+
298
+
299
+ if __name__ == "__main__":
300
+ raise SystemExit(main())
native_checkpoint.py ADDED
@@ -0,0 +1,476 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """TT-native text-only checkpoint for google/gemma-4-12B-it on one P150 (MeshShape 1x1).
3
+
4
+ Layout of a checkpoint directory ``<root>`` (its basename must contain "12B" so the
5
+ runtime's per-model policy key resolves, e.g. ``.../gemma-4-12B-it``):
6
+
7
+ native_manifest.json format, source identity, precision plan, every file's sha256/bytes
8
+ precision_plan.json the Gemma4Precision plan the tensors were quantized with
9
+ config.json, tokenizer.json, tokenizer_config.json, chat_template.jinja, generation_config.json
10
+ small.safetensors every 1-D text tensor (norm weights, layer scalars), lossless BF16
11
+ tensors/tensor_cache_bf16/**.tensorbin
12
+ the exact TTNN host tensors (TILE bfp8/bfp4/bf16 matrices, row-major
13
+ bf16 embedding, norm tiles) that the runtime uploads to DRAM; produced by
14
+ the runtime's own ttnn.as_tensor cache writer from the ORIGINAL weights
15
+ equivalence.json written only by ``record`` from two real runs (never by hand)
16
+
17
+ Loading never touches the original HF safetensors: 2-D matrices are passed to the model
18
+ constructor as meta-device placeholders (shape only) and every ttnn.as_tensor call hits its
19
+ .tensorbin file; a missing/corrupt file makes the fallback conversion fail loudly on the meta
20
+ tensor. The loader refuses to serve without an equivalence proof that matches the manifest,
21
+ the runtime source tree and the TTNN binaries.
22
+
23
+ Subcommands (run inside the TT image):
24
+ finalize <root> --source <hf dir> --fetch-verification <json>
25
+ after the builder populated tensors/: write small.safetensors, copy config/tokenizer,
26
+ hash everything into native_manifest.json.
27
+ record <root> --baseline <teacher-run dir> --restored <teacher-run dir>
28
+ compare the two runs' full logits (original weights quantized at load vs native reload,
29
+ identical code/plan/input) and write equivalence.json iff bit-exact.
30
+ verify <root>
31
+ verify hashes + proof + runtime identity (what the loader does).
32
+
33
+ Speculative-decoding drafter (google/gemma-4-12B-it-assistant) directory ``<droot>``:
34
+ drafter_manifest.json source identity (fetch verification, config/safetensors sha256), draft weight
35
+ dtype, target checkpoint manifest, every file's sha256/bytes
36
+ config.json, model.safetensors (hardlink of the sha-verified HF file; small tensors are read from it)
37
+ tensors/assistant_tensor_cache_<dtype>/**.tensorbin TTNN drafter tensors written by the runtime
38
+ spec_equivalence.json written only by ``spec-record``: greedy streams of a speculative run equal the
39
+ non-speculative streams of the same proven target runtime
40
+ drafter-finalize <droot> --source <hf dir> --target <target root> --dtype bfp8
41
+ spec-record <droot> --target <root> --reference <generate run> --speculative <generate run>
42
+ """
43
+ import argparse
44
+ import hashlib
45
+ import json
46
+ import os
47
+ import shutil
48
+ from pathlib import Path
49
+
50
+ FORMAT = "gemma4-12b-it-ttnn-native-text-1x1-v1"
51
+ MODEL_ID = "google/gemma-4-12B-it"
52
+ MANIFEST = "native_manifest.json"
53
+ PROOF = "equivalence.json"
54
+ DRAFTER_FORMAT = "gemma4-12b-it-assistant-ttnn-drafter-1x1-v1"
55
+ DRAFTER_MANIFEST = "drafter_manifest.json"
56
+ SPEC_PROOF = "spec_equivalence.json"
57
+ META_FILES = ("config.json", "tokenizer.json", "tokenizer_config.json", "chat_template.jinja", "generation_config.json")
58
+ TT_ROOT = Path("/home/container_app_user/tt-metal")
59
+ TREE = TT_ROOT / "models/demos/gemma4"
60
+ BINARIES = ("build/lib/_ttnncpp.so", "build/lib/libtt_metal.so", "ttnn/ttnn/_ttnn.so")
61
+ TEXT_PREFIX = "model.language_model."
62
+
63
+
64
+ def digest(path):
65
+ h = hashlib.sha256()
66
+ with open(path, "rb") as f:
67
+ for b in iter(lambda: f.read(1 << 24), b""):
68
+ h.update(b)
69
+ return h.hexdigest()
70
+
71
+
72
+ def save_json(path, value):
73
+ path = Path(path)
74
+ tmp = path.with_suffix(path.suffix + ".partial")
75
+ tmp.write_text(json.dumps(value, indent=1, sort_keys=True) + "\n")
76
+ tmp.replace(path)
77
+
78
+
79
+ def validate_config(cfg):
80
+ t = cfg["text_config"]
81
+ expected = {
82
+ "hidden_size": 3840,
83
+ "num_hidden_layers": 48,
84
+ "intermediate_size": 15360,
85
+ "num_attention_heads": 16,
86
+ "num_key_value_heads": 8,
87
+ "num_global_key_value_heads": 1,
88
+ "head_dim": 256,
89
+ "global_head_dim": 512,
90
+ "vocab_size": 262144,
91
+ "sliding_window": 1024,
92
+ "attention_k_eq_v": True,
93
+ "final_logit_softcapping": 30.0,
94
+ "tie_word_embeddings": True,
95
+ "enable_moe_block": False,
96
+ "hidden_size_per_layer_input": 0,
97
+ "num_kv_shared_layers": 0,
98
+ }
99
+ if cfg.get("model_type") != "gemma4_unified":
100
+ raise ValueError("expected gemma4_unified config")
101
+ for k, v in expected.items():
102
+ if t.get(k) != v:
103
+ raise ValueError(f"unsupported config {k}={t.get(k)!r} (expected {v!r})")
104
+ return t
105
+
106
+
107
+ def text_shapes(cfg):
108
+ """Every text tensor of the source checkpoint (name -> shape)."""
109
+ t = validate_config(cfg)
110
+ H, I, V = t["hidden_size"], t["intermediate_size"], t["vocab_size"]
111
+ shapes = {TEXT_PREFIX + "embed_tokens.weight": [V, H], TEXT_PREFIX + "norm.weight": [H]}
112
+ for i, lt in enumerate(t["layer_types"]):
113
+ p = f"{TEXT_PREFIX}layers.{i}."
114
+ full = lt == "full_attention"
115
+ hd = t["global_head_dim"] if full else t["head_dim"]
116
+ nkv = t["num_global_key_value_heads"] if full else t["num_key_value_heads"]
117
+ shapes.update(
118
+ {
119
+ p + "input_layernorm.weight": [H],
120
+ p + "post_attention_layernorm.weight": [H],
121
+ p + "pre_feedforward_layernorm.weight": [H],
122
+ p + "post_feedforward_layernorm.weight": [H],
123
+ p + "layer_scalar": [1],
124
+ p + "mlp.gate_proj.weight": [I, H],
125
+ p + "mlp.up_proj.weight": [I, H],
126
+ p + "mlp.down_proj.weight": [H, I],
127
+ p + "self_attn.q_proj.weight": [t["num_attention_heads"] * hd, H],
128
+ p + "self_attn.k_proj.weight": [nkv * hd, H],
129
+ p + "self_attn.o_proj.weight": [H, t["num_attention_heads"] * hd],
130
+ p + "self_attn.q_norm.weight": [hd],
131
+ p + "self_attn.k_norm.weight": [hd],
132
+ }
133
+ )
134
+ if not full: # K=V tied on full-attention layers: no v_proj
135
+ shapes[p + "self_attn.v_proj.weight"] = [nkv * hd, H]
136
+ return shapes
137
+
138
+
139
+ def runtime_identity():
140
+ files = sorted(p for ext in ("*.py", "*.cpp", "*.hpp") for p in TREE.rglob(ext) if "__pycache__" not in p.parts and "tests" not in p.parts)
141
+ return {
142
+ "gemma4_tree": {str(p.relative_to(TREE)): digest(p) for p in files},
143
+ "binaries": {b: digest(TT_ROOT / b) for b in BINARIES if (TT_ROOT / b).exists()},
144
+ }
145
+
146
+
147
+ def read_manifest(root):
148
+ root = Path(root)
149
+ m = json.loads((root / MANIFEST).read_text())
150
+ if m.get("format") != FORMAT or m.get("scope") != "text-only, MeshShape 1x1 (single P150)":
151
+ raise ValueError("unsupported native checkpoint format/scope")
152
+ if m.get("source", {}).get("model_id") != MODEL_ID:
153
+ raise ValueError("native checkpoint model mismatch")
154
+ cfg = json.loads((root / "config.json").read_text())
155
+ shapes = text_shapes(cfg)
156
+ if m["text_tensor_shapes"] != shapes:
157
+ raise ValueError("manifest tensor inventory disagrees with config")
158
+ if json.loads((root / "precision_plan.json").read_text()) != m["precision_plan"]:
159
+ raise ValueError("precision_plan.json differs from manifest")
160
+ return m
161
+
162
+
163
+ def verify_files(root, m):
164
+ root = Path(root).resolve()
165
+ listed = set(m["files"])
166
+ present = {str(p.relative_to(root)) for p in root.rglob("*") if p.is_file()} - {MANIFEST, PROOF}
167
+ if present != listed:
168
+ raise ValueError(f"file set mismatch: extra={sorted(present - listed)[:5]} missing={sorted(listed - present)[:5]}")
169
+ for name, info in m["files"].items():
170
+ p = root / name
171
+ if p.stat().st_size != info["bytes"] or digest(p) != info["sha256"]:
172
+ raise ValueError(f"hash/size mismatch: {name}")
173
+
174
+
175
+ def require_proof(root, m):
176
+ root = Path(root)
177
+ proof = json.loads((root / PROOF).read_text())
178
+ if proof.get("manifest_sha256") != digest(root / MANIFEST):
179
+ raise ValueError("equivalence proof does not match this manifest")
180
+ if proof.get("exact_logits_equal") is not True or proof.get("tokens_compared", 0) < 1:
181
+ raise ValueError("equivalence proof is not an exact-logits proof")
182
+ if proof.get("precision_plan") != m["precision_plan"]:
183
+ raise ValueError("equivalence proof precision plan mismatch")
184
+ return proof
185
+
186
+
187
+ # GEMMA4_* variables that select paths / the serving entry point and never change numerics; every other
188
+ # GEMMA4_* variable (matmul accumulation, norms, GELU, ring / prefill chunk, ...) must match the proof.
189
+ # GEMMA4_DRAFT_LEN is checked against the drafter's speculative proof (draft_len) by the serving loader.
190
+ SERVING_SELECTORS = frozenset({"GEMMA4_NATIVE", "GEMMA4_SINGLE_USER", "GEMMA4_ASSISTANT_DIR", "GEMMA4_DRAFT_LEN",
191
+ "GEMMA4_TOOLS_DIR"})
192
+
193
+
194
+ def prepare_native_runtime(root, allow_unproven=False, verify_hashes=True):
195
+ """Loader entry: returns (model_path, state_dict, tt_cache_path, identity)."""
196
+ import torch
197
+ from safetensors.torch import load_file
198
+
199
+ root = Path(root).resolve()
200
+ m = read_manifest(root)
201
+ if verify_hashes:
202
+ verify_files(root, m)
203
+ identity = {"manifest_sha256": digest(root / MANIFEST), "proof": None}
204
+ if not allow_unproven:
205
+ proof = require_proof(root, m)
206
+ now = runtime_identity()
207
+ if now != proof["runtime_identity"]:
208
+ diff = [k for k in now["gemma4_tree"] if proof["runtime_identity"]["gemma4_tree"].get(k) != now["gemma4_tree"][k]]
209
+ raise ValueError(f"runtime differs from the one the equivalence proof covers: {diff[:8] or 'binaries'}")
210
+ env_now = {k: v for k, v in os.environ.items()
211
+ if k.startswith("GEMMA4_") and k != "GEMMA4_PRECISION_PLAN" and k not in SERVING_SELECTORS}
212
+ if env_now != proof["runtime_env"]:
213
+ raise ValueError(f"GEMMA4_* runtime environment {env_now} differs from the proof's {proof['runtime_env']}")
214
+ identity["proof"] = digest(root / PROOF)
215
+ small = load_file(str(root / "small.safetensors"))
216
+ state_dict = {}
217
+ for name, shape in m["text_tensor_shapes"].items():
218
+ if len(shape) == 2:
219
+ state_dict[name] = torch.empty(shape, dtype=torch.bfloat16, device="meta")
220
+ else:
221
+ state_dict[name] = small[name]
222
+ return str(root), state_dict, str(root / "tensors"), identity
223
+
224
+
225
+ def finalize(root, source, fetch_verification):
226
+ import torch
227
+ from safetensors import safe_open
228
+ from safetensors.torch import save_file
229
+
230
+ root, source = Path(root).resolve(), Path(source).resolve()
231
+ cfg = json.loads((source / "config.json").read_text())
232
+ shapes = text_shapes(cfg)
233
+ for f in META_FILES:
234
+ shutil.copyfile(source / f, root / f)
235
+ small = {}
236
+ with safe_open(str(source / "model.safetensors"), framework="pt") as fh:
237
+ keys = set(fh.keys())
238
+ missing = [k for k in shapes if k not in keys]
239
+ if missing:
240
+ raise ValueError(f"source lacks text tensors: {missing[:5]}")
241
+ for name, shape in shapes.items():
242
+ if list(fh.get_slice(name).get_shape()) != shape:
243
+ raise ValueError(f"source shape mismatch {name}")
244
+ if len(shape) != 2:
245
+ small[name] = fh.get_tensor(name).to(torch.bfloat16).contiguous()
246
+ save_file(small, str(root / "small.safetensors"))
247
+ tensor_files = sorted(p for p in (root / "tensors").rglob("*.tensorbin"))
248
+ if not tensor_files:
249
+ raise ValueError("builder produced no tensors")
250
+ files = {}
251
+ for p in sorted(q for q in root.rglob("*") if q.is_file()):
252
+ rel = str(p.relative_to(root))
253
+ if rel in (MANIFEST, PROOF):
254
+ continue
255
+ files[rel] = {"bytes": p.stat().st_size, "sha256": digest(p)}
256
+ fv = json.loads(Path(fetch_verification).read_text())
257
+ manifest = {
258
+ "format": FORMAT,
259
+ "scope": "text-only, MeshShape 1x1 (single P150)",
260
+ "source": {
261
+ "model_id": MODEL_ID,
262
+ "fetch_verification": fv,
263
+ "source_dir": str(source),
264
+ "config_sha256": digest(source / "config.json"),
265
+ },
266
+ "precision_plan": json.loads((root / "precision_plan.json").read_text()),
267
+ "text_tensor_shapes": shapes,
268
+ "tensor_files": len(tensor_files),
269
+ "tensor_bytes": sum(p.stat().st_size for p in tensor_files),
270
+ "files": files,
271
+ "builder_runtime_identity": runtime_identity(),
272
+ }
273
+ save_json(root / MANIFEST, manifest)
274
+ return {"files": len(files), "tensor_files": len(tensor_files), "tensor_GB": manifest["tensor_bytes"] / 1e9}
275
+
276
+
277
+ def record(root, baseline, restored):
278
+ import numpy as np
279
+
280
+ root, baseline, restored = Path(root).resolve(), Path(baseline), Path(restored)
281
+ m = read_manifest(root)
282
+ verify_files(root, m)
283
+ a = json.loads((baseline / "metadata.json").read_text())
284
+ b = json.loads((restored / "metadata.json").read_text())
285
+ for meta in (a, b):
286
+ if meta.get("status") != "complete" or meta.get("mode") != "teacher" or not meta.get("saved_logits"):
287
+ raise ValueError("both runs must be completed teacher runs with saved full logits")
288
+ for key in ("input_sha256", "records", "saved_logits"):
289
+ if a[key] != b[key]:
290
+ raise ValueError(f"run configuration mismatch: {key}")
291
+ if a["runtime"] != b["runtime"]:
292
+ raise ValueError("runtime (source tree/tool/env/plan) differs between the two runs")
293
+ if a["runtime"]["precision_plan"] != m["precision_plan"]:
294
+ raise ValueError("baseline plan differs from the checkpoint plan")
295
+ if a["weights"]["native"] or not a["weights"]["model_dir"]:
296
+ raise ValueError("baseline must load the original HF weights")
297
+ if Path(a["weights"]["cache"]).resolve() == (root / "tensors").resolve():
298
+ raise ValueError("baseline must quantize from the original weights, not reuse the checkpoint tensors")
299
+ if not b["weights"]["native"] or b["native_identity"].get("manifest_sha256") != digest(root / MANIFEST):
300
+ raise ValueError("restored run is not bound to this native checkpoint")
301
+ tokens, recs = 0, []
302
+ for rid in a["saved_logits"]:
303
+ x = np.load(baseline / f"{rid}.logits.npy", mmap_mode="r")
304
+ y = np.load(restored / f"{rid}.logits.npy", mmap_mode="r")
305
+ if x.shape != y.shape or x.shape[1] != 262144:
306
+ raise ValueError(f"logit shape mismatch for {rid}")
307
+ for s in range(0, x.shape[0], 64):
308
+ if not np.array_equal(np.asarray(x[s : s + 64]), np.asarray(y[s : s + 64])):
309
+ raise ValueError(f"native reload logits differ for {rid} at row {s}")
310
+ tokens += x.shape[0]
311
+ recs.append(
312
+ {
313
+ "id": rid,
314
+ "rows": int(x.shape[0]),
315
+ "baseline_sha256": digest(baseline / f"{rid}.logits.npy"),
316
+ "restored_sha256": digest(restored / f"{rid}.logits.npy"),
317
+ }
318
+ )
319
+ for rid in a["records"]:
320
+ za, zb = np.load(baseline / f"{rid}.npz"), np.load(restored / f"{rid}.npz")
321
+ if not (np.array_equal(za["nll"], zb["nll"]) and np.array_equal(za["argmax"], zb["argmax"])):
322
+ raise ValueError(f"teacher-forced NLL/argmax differ for {rid}")
323
+ proof = {
324
+ "manifest_sha256": digest(root / MANIFEST),
325
+ "exact_logits_equal": True,
326
+ "tokens_compared": tokens,
327
+ "records_full_logits": recs,
328
+ "records_nll_argmax_equal": a["records"],
329
+ "precision_plan": m["precision_plan"],
330
+ "runtime_identity": runtime_identity(),
331
+ "runtime_run_sha256": {"gemma4_tree_sha256": a["runtime"]["gemma4_tree_sha256"], "tool_sha256": a["runtime"]["tool_sha256"]},
332
+ "runtime_env": a["runtime"]["env"],
333
+ "baseline_metadata_sha256": digest(baseline / "metadata.json"),
334
+ "restored_metadata_sha256": digest(restored / "metadata.json"),
335
+ "scope": "bit-exact full-vocabulary logits, original weights quantized at load vs native reload, same runtime; not a quality certification",
336
+ }
337
+ save_json(root / PROOF, proof)
338
+ return {k: proof[k] for k in ("exact_logits_equal", "tokens_compared", "manifest_sha256")}
339
+
340
+
341
+ def _drafter_files(droot):
342
+ return {
343
+ str(p.relative_to(droot)): {"bytes": p.stat().st_size, "sha256": digest(p)}
344
+ for p in sorted(q for q in droot.rglob("*") if q.is_file())
345
+ if str(p.relative_to(droot)) not in (DRAFTER_MANIFEST, SPEC_PROOF)
346
+ }
347
+
348
+
349
+ def drafter_finalize(droot, source, target, dtype):
350
+ droot, source, target = Path(droot).resolve(), Path(source).resolve(), Path(target).resolve()
351
+ fv = json.loads((source / "fetch-verification.json").read_text())
352
+ if not list((droot / "tensors").rglob("*.tensorbin")):
353
+ raise ValueError("drafter tensors/ is empty: run the runtime once with --assistant-cache <droot>/tensors")
354
+ for name in ("config.json", "model.safetensors"):
355
+ if digest(droot / name) != digest(source / name):
356
+ raise ValueError(f"{name} differs from the source")
357
+ manifest = {
358
+ "format": DRAFTER_FORMAT,
359
+ "source": {"model_id": "google/gemma-4-12B-it-assistant", "fetch_verification": fv, "source_dir": str(source),
360
+ "config_sha256": digest(source / "config.json"), "safetensors_sha256": digest(source / "model.safetensors")},
361
+ "dtype": dtype,
362
+ "target_manifest_sha256": digest(target / MANIFEST),
363
+ "files": _drafter_files(droot),
364
+ "builder_runtime_identity": runtime_identity(),
365
+ }
366
+ save_json(droot / DRAFTER_MANIFEST, manifest)
367
+ return {"files": len(manifest["files"]), "dtype": dtype}
368
+
369
+
370
+ def prepare_drafter(droot, target_identity, allow_unproven=False):
371
+ """Loader entry for the drafter: verify files (+ speculative proof); returns (dir, cache path, identity)."""
372
+ droot = Path(droot).resolve()
373
+ m = json.loads((droot / DRAFTER_MANIFEST).read_text())
374
+ if m.get("format") != DRAFTER_FORMAT:
375
+ raise ValueError("unsupported drafter format")
376
+ if _drafter_files(droot) != m["files"]:
377
+ raise ValueError("drafter file set / hash mismatch")
378
+ if target_identity.get("manifest_sha256") != m["target_manifest_sha256"]:
379
+ raise ValueError("drafter was built for a different target checkpoint")
380
+ identity = {"drafter_manifest_sha256": digest(droot / DRAFTER_MANIFEST), "spec_proof": None, "dtype": m["dtype"]}
381
+ if not allow_unproven:
382
+ proof = json.loads((droot / SPEC_PROOF).read_text())
383
+ if proof.get("drafter_manifest_sha256") != identity["drafter_manifest_sha256"] or proof.get("streams_identical") is not True:
384
+ raise ValueError("speculative proof does not cover this drafter")
385
+ if proof.get("target_proof_sha256") != target_identity.get("proof"):
386
+ raise ValueError("speculative proof was made against a different target proof")
387
+ if proof["runtime_identity"] != runtime_identity():
388
+ raise ValueError("runtime differs from the one the speculative proof covers")
389
+ identity["spec_proof"] = digest(droot / SPEC_PROOF)
390
+ identity["draft_len"] = proof["draft_len"]
391
+ return str(droot), str(droot / "tensors"), identity
392
+
393
+
394
+ def spec_record(droot, target, reference, speculative):
395
+ droot, target = Path(droot).resolve(), Path(target).resolve()
396
+ a = json.loads((Path(reference) / "metadata.json").read_text())
397
+ b = json.loads((Path(speculative) / "metadata.json").read_text())
398
+ for meta in (a, b):
399
+ if meta.get("status") != "complete" or meta.get("mode") != "generate":
400
+ raise ValueError("both runs must be completed generate runs")
401
+ ident = meta.get("native_identity") or {}
402
+ if ident.get("manifest_sha256") != digest(target / MANIFEST) or not ident.get("proof"):
403
+ raise ValueError("runs must load the proven target checkpoint (no --allow-unproven)")
404
+ if a["runtime"] != b["runtime"] or a["native_identity"] != b["native_identity"]:
405
+ raise ValueError("reference and speculative runs differ in runtime / target identity")
406
+ if a.get("spec") or not b.get("spec"):
407
+ raise ValueError("reference must be non-speculative and the second run speculative")
408
+ if b["spec"]["drafter"]["drafter_manifest_sha256"] != digest(droot / DRAFTER_MANIFEST):
409
+ raise ValueError("speculative run did not use this drafter")
410
+ if a["prompts_sha256"] != b["prompts_sha256"] or a["max_new"] != b["max_new"]:
411
+ raise ValueError("runs used different prompts / lengths")
412
+ ra = {r["id"]: r["gen_ids"] for r in map(json.loads, open(Path(reference) / "generations.jsonl"))}
413
+ rb = {r["id"]: r["gen_ids"] for r in map(json.loads, open(Path(speculative) / "generations.jsonl"))}
414
+ if ra.keys() != rb.keys() or any(ra[k] != rb[k] for k in ra):
415
+ raise ValueError("speculative greedy streams differ from the non-speculative ones")
416
+ proof = {
417
+ "drafter_manifest_sha256": digest(droot / DRAFTER_MANIFEST),
418
+ "target_manifest_sha256": digest(target / MANIFEST),
419
+ "target_proof_sha256": a["native_identity"]["proof"],
420
+ "streams_identical": True,
421
+ "prompts_sha256": a["prompts_sha256"],
422
+ "max_new": a["max_new"],
423
+ "draft_len": b["spec"]["draft_len"],
424
+ "tokens_compared": sum(len(v) for v in ra.values()),
425
+ "streams_sha256": {k: hashlib.sha256(json.dumps(v).encode()).hexdigest() for k, v in sorted(ra.items())},
426
+ "runtime_identity": runtime_identity(),
427
+ "runtime_env": a["runtime"]["env"],
428
+ "reference_metadata_sha256": digest(Path(reference) / "metadata.json"),
429
+ "speculative_metadata_sha256": digest(Path(speculative) / "metadata.json"),
430
+ "scope": "greedy token streams of speculative decoding equal non-speculative greedy decoding on the recorded prompts, same proven runtime",
431
+ }
432
+ save_json(droot / SPEC_PROOF, proof)
433
+ return {k: proof[k] for k in ("streams_identical", "tokens_compared", "draft_len")}
434
+
435
+
436
+ def main():
437
+ ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
438
+ sub = ap.add_subparsers(dest="cmd", required=True)
439
+ f = sub.add_parser("finalize")
440
+ f.add_argument("root")
441
+ f.add_argument("--source", required=True)
442
+ f.add_argument("--fetch-verification", required=True)
443
+ r = sub.add_parser("record")
444
+ r.add_argument("root")
445
+ r.add_argument("--baseline", required=True)
446
+ r.add_argument("--restored", required=True)
447
+ v = sub.add_parser("verify")
448
+ v.add_argument("root")
449
+ df = sub.add_parser("drafter-finalize")
450
+ df.add_argument("root")
451
+ df.add_argument("--source", required=True)
452
+ df.add_argument("--target", required=True)
453
+ df.add_argument("--dtype", default="bfp8")
454
+ sr = sub.add_parser("spec-record")
455
+ sr.add_argument("root")
456
+ sr.add_argument("--target", required=True)
457
+ sr.add_argument("--reference", required=True)
458
+ sr.add_argument("--speculative", required=True)
459
+ a = ap.parse_args()
460
+ if a.cmd == "drafter-finalize":
461
+ print(json.dumps(drafter_finalize(a.root, a.source, a.target, a.dtype), indent=1))
462
+ return
463
+ if a.cmd == "spec-record":
464
+ print(json.dumps(spec_record(a.root, a.target, a.reference, a.speculative), indent=1))
465
+ return
466
+ if a.cmd == "finalize":
467
+ print(json.dumps(finalize(a.root, a.source, a.fetch_verification), indent=1))
468
+ elif a.cmd == "record":
469
+ print(json.dumps(record(a.root, a.baseline, a.restored), indent=1))
470
+ else:
471
+ _, sd, cache, ident = prepare_native_runtime(a.root)
472
+ print(json.dumps({"verified": True, "tensors": len(sd), "cache": cache, "identity": ident}, indent=1))
473
+
474
+
475
+ if __name__ == "__main__":
476
+ main()
provenance/build_native_checkpoint.py ADDED
@@ -0,0 +1,54 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Build a Gemma-4-12B TT-native checkpoint from the ORIGINAL HF weights.
3
+
4
+ build_native_checkpoint.py --source <hf model dir named gemma-4-12B-it> --plan <plan.json>
5
+ --fetch-verification <fetch-verification.json> --out <root>/gemma-4-12B-it
6
+
7
+ The runtime's own weight path (create_tt_model -> ttnn.as_tensor) quantizes every tensor
8
+ per the plan and serializes the exact host TTNN tensors into <out>/tensors; then
9
+ native_checkpoint.finalize writes the lossless small tensors, metadata and manifest.
10
+ Needs the device (model construction uploads to DRAM). No proof is written here: run the
11
+ proof recipe (prove_native.sh) afterwards.
12
+ """
13
+ import argparse
14
+ import json
15
+ import os
16
+ import shutil
17
+ import sys
18
+ from pathlib import Path
19
+
20
+ sys.path.insert(0, str(Path(__file__).resolve().parent))
21
+ import native_checkpoint as nc # noqa: E402
22
+
23
+
24
+ def main():
25
+ ap = argparse.ArgumentParser()
26
+ ap.add_argument("--source", required=True)
27
+ ap.add_argument("--plan", required=True)
28
+ ap.add_argument("--fetch-verification", required=True)
29
+ ap.add_argument("--out", required=True)
30
+ a = ap.parse_args()
31
+ out = Path(a.out)
32
+ if "12b" not in out.name.lower():
33
+ raise SystemExit("checkpoint dir basename must contain 12B (runtime policy key)")
34
+ if out.exists():
35
+ raise SystemExit(f"{out} exists; refusing to overwrite")
36
+ (out / "tensors").mkdir(parents=True)
37
+ shutil.copyfile(a.plan, out / "precision_plan.json")
38
+ os.environ["TT_CACHE_PATH"] = str(out / "tensors")
39
+ os.environ["GEMMA4_PRECISION_PLAN"] = str(out / "precision_plan.json")
40
+
41
+ import ttnn
42
+ from models.demos.gemma4.tt.common import create_tt_model
43
+
44
+ mesh = ttnn.open_mesh_device(mesh_shape=ttnn.MeshShape(1, 1))
45
+ try:
46
+ create_tt_model(mesh_device=mesh, max_batch_size=1, max_seq_len=2048, model_path=a.source, create_kv_cache=False)
47
+ finally:
48
+ ttnn.close_mesh_device(mesh)
49
+ print(json.dumps(nc.finalize(out, a.source, a.fetch_verification), indent=1))
50
+ print("BUILD_DONE")
51
+
52
+
53
+ if __name__ == "__main__":
54
+ main()
release-manifest.json ADDED
The diff for this file is too large to render. See raw diff
 
reproduction.json ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 1,
3
+ "release_id": "p3c-20260930",
4
+ "model_repo": "Lottolabs/gemma-4-12B-it-TT-BFP8-P150",
5
+ "source_model": {
6
+ "model_id": "google/gemma-4-12B-it",
7
+ "revision": "707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7"
8
+ },
9
+ "source_drafter": {
10
+ "model_id": "google/gemma-4-12B-it-assistant",
11
+ "revision": "46d4c6f13f0ac0ad827b915669b8df9b81c64c51"
12
+ },
13
+ "runtime_image_tag": "lottolabs/gemma4-12b-tt-p150:p3c",
14
+ "runtime_image": "sha256:53428e7f60a705b75b51b5166f44a912b19c3032981013fb788dd1ce9757aa73",
15
+ "runtime_source": "runtime/source.json",
16
+ "checkpoint_build": {
17
+ "builder": "provenance/build_native_checkpoint.py",
18
+ "manifest_tool": "native_checkpoint.py",
19
+ "note": "Tensors are the runtime's own ttnn.as_tensor host tensors written from the original BF16 weights with precision_plan.json (all matrices bfloat8_b, embedding bf16)."
20
+ },
21
+ "serve": [
22
+ "python3",
23
+ "launch.py",
24
+ "--cache-root",
25
+ "/large-disk/gemma4-12b-tt-native",
26
+ "--device-ownership-confirmed"
27
+ ],
28
+ "non_speculative": [
29
+ "python3",
30
+ "launch.py",
31
+ "--cache-root",
32
+ "/large-disk/gemma4-12b-tt-native",
33
+ "--speculation",
34
+ "off",
35
+ "--device-ownership-confirmed"
36
+ ],
37
+ "serving_policy": {
38
+ "max_num_seqs": 1,
39
+ "max_model_len": 262144,
40
+ "block_size": 64,
41
+ "prefill_chunk_tokens": 16384,
42
+ "prefix_caching": false,
43
+ "sliding_kv_ring_tokens": 18432,
44
+ "speculative": {
45
+ "method": "gemma4_assistant",
46
+ "num_speculative_tokens": 5,
47
+ "applies_to": "greedy (temperature 0) requests"
48
+ }
49
+ },
50
+ "proofs": {
51
+ "target": "gemma-4-12B-it/equivalence.json",
52
+ "drafter": "gemma-4-12B-it-assistant/spec_equivalence.json"
53
+ },
54
+ "evidence": {
55
+ "quality": "evidence/quality/",
56
+ "http": "evidence/http-p3c/",
57
+ "localmaxxing": "evidence/localmaxxing/speed-test.json",
58
+ "public_download": "evidence/public-download-verification.json",
59
+ "public_serving": "evidence/public-serving-smoke.json"
60
+ },
61
+ "runtime_distribution": {
62
+ "path": "runtime-image.tar.gz",
63
+ "contract": "launch.py verifies the archive checksum, loads it with Docker, and requires the exact recorded image ID."
64
+ }
65
+ }
runtime-release.json ADDED
The diff for this file is too large to render. See raw diff
 
runtime/Dockerfile.fast1 ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ FROM gemma4-12b:p2c
2
+ USER root
3
+ COPY --chown=1000:1000 gemma4/ /home/container_app_user/tt-metal/models/demos/gemma4/
4
+ COPY --chown=1000:1000 tools/ /home/container_app_user/gemma4-tools/
5
+ ENV GEMMA4_VERIFY_SDPA=batched
6
+ ENV GEMMA4_FUSE_GELU_MUL=1
7
+ USER container_app_user
runtime/Dockerfile.p2b ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ FROM ghcr.io/tenstorrent/tt-inference-server/vllm-tt-metal-src-release-ubuntu-22.04-amd64:0.20.0-de59f8a-03fa3af
2
+ USER root
3
+ COPY --chown=1000:1000 ttnn/ /home/container_app_user/tt-metal/ttnn/
4
+ USER container_app_user
5
+ WORKDIR /home/container_app_user/tt-metal
6
+ RUN cmake --build build_Release --target ttnn --parallel 24 \
7
+ && cp build_Release/ttnn/_ttnncpp.so build/lib/_ttnncpp.so \
8
+ && cp build_Release/ttnn/_ttnn.so build/lib/_ttnn.so \
9
+ && cp build_Release/ttnn/_ttnn.so ttnn/ttnn/_ttnn.so
10
+ USER root
11
+ WORKDIR /home/container_app_user/app/src
12
+ COPY --chown=1000:1000 gemma4/ /home/container_app_user/tt-metal/models/demos/gemma4/
13
+ COPY --chown=1000:1000 tools/ /home/container_app_user/gemma4-tools/
14
+ USER container_app_user
runtime/Dockerfile.p2c ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ FROM gemma4-12b:p2b
2
+ USER root
3
+
4
+ COPY --chown=1000:1000 gemma4/ /home/container_app_user/tt-metal/models/demos/gemma4/
5
+ COPY --chown=1000:1000 tools/ /home/container_app_user/gemma4-tools/
6
+ USER container_app_user
runtime/Dockerfile.p3c ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ FROM gemma4-12b:fast1
2
+ USER root
3
+
4
+ COPY --chown=1000:1000 gemma4/ /home/container_app_user/tt-metal/models/demos/gemma4/
5
+ COPY --chown=1000:1000 tools/ /home/container_app_user/gemma4-tools/
6
+ COPY --chown=1000:1000 serving-overlay/ /
7
+ USER container_app_user
runtime/build.sh ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/bin/bash
2
+ # Build the phase-2 runtime image gemma4-12b:<tag>: upstream TT image + this gemma4 tree (Python + JIT
3
+ # kernel sources) + the runner tools. With TTNN_OVERLAY=1 (p2b on) the TTNN sources under ttnn/overlay
4
+ # (in0-multicast 1D matmul factory + in1 reader kernel) are copied into the tree and the TTNN host
5
+ # libraries are rebuilt inside the image (incremental, ~1 min).
6
+ # BASE_IMAGE=<image> builds on a previous image (e.g. gemma4-12b:p2b to reuse its TTNN build).
7
+ # SERVING_OVERLAY=1 (p3 on) also copies serving/overlay/ (vLLM core + TT plugin files for the single-user
8
+ # gemma-4 server: speculative method gemma4_assistant, multi-token stop handling) over the image root.
9
+ # Usage: [TTNN_OVERLAY=1] [SERVING_OVERLAY=1] [BASE_IMAGE=...] build.sh TAG
10
+ set -euo pipefail
11
+ TAG=${1:?tag}
12
+ W=/var/tmp/gemma4-12b
13
+ BASE=${BASE_IMAGE:-ghcr.io/tenstorrent/tt-inference-server/vllm-tt-metal-src-release-ubuntu-22.04-amd64:0.20.0-de59f8a-03fa3af}
14
+ B=$W/build/$TAG
15
+ rm -rf "$B" && mkdir -p "$B/tools"
16
+ rsync -a --exclude __pycache__ --exclude .git --exclude .gitignore "$W/src/gemma4/" "$B/gemma4/"
17
+ cp "$W/tools/tt_eval.py" "$W/tools/native_checkpoint.py" "$W/tools/compare_gen.py" "$B/tools/"
18
+ SERVING_STEP=""
19
+ if [ "${SERVING_OVERLAY:-0}" = 1 ]; then
20
+ rsync -a --exclude __pycache__ "$W/serving/overlay/" "$B/serving-overlay/"
21
+ SERVING_STEP='COPY --chown=1000:1000 serving-overlay/ /'
22
+ fi
23
+ TTNN_STEP=""
24
+ if [ "${TTNN_OVERLAY:-0}" = 1 ]; then
25
+ rsync -a --exclude .git "$W/ttnn/overlay/ttnn/" "$B/ttnn/"
26
+ TTNN_STEP='COPY --chown=1000:1000 ttnn/ /home/container_app_user/tt-metal/ttnn/
27
+ USER container_app_user
28
+ WORKDIR /home/container_app_user/tt-metal
29
+ RUN cmake --build build_Release --target ttnn --parallel 24 \
30
+ && cp build_Release/ttnn/_ttnncpp.so build/lib/_ttnncpp.so \
31
+ && cp build_Release/ttnn/_ttnn.so build/lib/_ttnn.so \
32
+ && cp build_Release/ttnn/_ttnn.so ttnn/ttnn/_ttnn.so
33
+ USER root
34
+ WORKDIR /home/container_app_user/app/src'
35
+ fi
36
+ cat > "$B/Dockerfile" <<D
37
+ FROM $BASE
38
+ USER root
39
+ $TTNN_STEP
40
+ COPY --chown=1000:1000 gemma4/ /home/container_app_user/tt-metal/models/demos/gemma4/
41
+ COPY --chown=1000:1000 tools/ /home/container_app_user/gemma4-tools/
42
+ $SERVING_STEP
43
+ USER container_app_user
44
+ D
45
+ docker build -q -t "gemma4-12b:$TAG" "$B"
46
+ docker image inspect "gemma4-12b:$TAG" --format '{{.Id}}'
runtime/source.json ADDED
@@ -0,0 +1,59 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 1,
3
+ "runtime_image": "sha256:53428e7f60a705b75b51b5166f44a912b19c3032981013fb788dd1ce9757aa73",
4
+ "runtime_image_tag": "lottolabs/gemma4-12b-tt-p150:p3c",
5
+ "gemma4_tree": {
6
+ "path": "runtime/gemma4/",
7
+ "installed_at": "/home/container_app_user/tt-metal/models/demos/gemma4/",
8
+ "git_branch": "master",
9
+ "git_commit": "ce67960382fffed8550b38717cedf574e821da06",
10
+ "note": "Byte-identical to the tree in the runtime image (compared file by file against the image); .git and __pycache__ excluded."
11
+ },
12
+ "tools": {
13
+ "path": "runtime/tools/",
14
+ "installed_at": "/home/container_app_user/gemma4-tools/"
15
+ },
16
+ "serving_overlay": {
17
+ "path": "runtime/serving-overlay/",
18
+ "installed_at": "/",
19
+ "content": "vLLM core and vllm-tt-plugin files for the single-user Gemma 4 server (speculative method gemma4_assistant, multi-token stop handling)"
20
+ },
21
+ "ttnn_overlay": {
22
+ "path": "runtime/ttnn-overlay/",
23
+ "installed_at": "/home/container_app_user/tt-metal/ttnn/",
24
+ "content": "in0-multicast 1D matmul program factory; TTNN host libraries rebuilt in layer p2b"
25
+ },
26
+ "build_chain": [
27
+ {
28
+ "layer": "base",
29
+ "image": "ghcr.io/tenstorrent/tt-inference-server/vllm-tt-metal-src-release-ubuntu-22.04-amd64:0.20.0-de59f8a-03fa3af",
30
+ "image_id": "sha256:710de20012a5417ec632e9d1db33b45f81fa9d9961da882a94868d15c5f99735"
31
+ },
32
+ {
33
+ "layer": "p2b",
34
+ "dockerfile": "runtime/Dockerfile.p2b",
35
+ "image_id": "sha256:9fc0b344293da678ef107d570753fea90205968a9cced44ec81831f10163cb3b"
36
+ },
37
+ {
38
+ "layer": "p2c",
39
+ "dockerfile": "runtime/Dockerfile.p2c",
40
+ "image_id": "sha256:b91c7e92ec16c507c916229ba7d675093379fe18ab05c18732eedcd3b91156f2"
41
+ },
42
+ {
43
+ "layer": "fast1",
44
+ "dockerfile": "runtime/Dockerfile.fast1",
45
+ "image_id": "sha256:aade01a9328d028c1979070df60aa17ade1ce729e936c74815b67802a086fff9"
46
+ },
47
+ {
48
+ "layer": "p3c",
49
+ "dockerfile": "runtime/Dockerfile.p3c",
50
+ "image_id": "sha256:53428e7f60a705b75b51b5166f44a912b19c3032981013fb788dd1ce9757aa73"
51
+ }
52
+ ],
53
+ "build_script": "runtime/build.sh",
54
+ "rebuild_note": "Provenance only: intermediate layers copied earlier gemma4 trees that the p3c layer overwrites; the Dockerfiles reference local intermediate tags. Serve the checksum-pinned runtime-image.tar.gz.",
55
+ "licenses": {
56
+ "tt-metal": "runtime/licenses/tt-metal/",
57
+ "vllm": "runtime/licenses/vllm/"
58
+ }
59
+ }
serve_native.py ADDED
@@ -0,0 +1,128 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Serve the TT-native all-BFP8 google/gemma-4-12B-it release (text) on one Tenstorrent P150 through
3
+ vLLM (TT plugin) at its native 262,144-token context, single user, with exact greedy speculative decoding.
4
+
5
+ Normally invoked by launch.py from a verified snapshot. Before launching it verifies, from this release
6
+ directory: the native checkpoint (manifest, every file hash, exact-logits equivalence proof bound to the
7
+ manifest), the drafter (manifest, hashes, speculative-equivalence proof bound to the target proof and
8
+ draft length) and the runtime image ID pinned in runtime-release.json. Inside the container the loader
9
+ re-verifies the checkpoint and both proofs against the image's runtime identity before loading.
10
+
11
+ Serving policy (models/demos/gemma4/tt/serving_vllm.py): max_num_seqs 1, block 64, chunked prefill in
12
+ 16,384-token chunks, prefix caching off, sliding-window KV as 18,432-token rings, full-attention KV
13
+ paged through vLLM's block table. Greedy requests (temperature 0) decode speculatively (drafter K=5,
14
+ exact per-row verify); other sampling decodes one token per step with vLLM's host sampler. Text only.
15
+ Standard library only.
16
+ """
17
+ import argparse
18
+ import hashlib
19
+ import json
20
+ import os
21
+ import shlex
22
+ import subprocess
23
+ import sys
24
+ import time
25
+ from pathlib import Path
26
+
27
+ HERE = Path(__file__).resolve().parent
28
+ sys.path.insert(0, str(HERE))
29
+ from native_checkpoint import ( # noqa: E402
30
+ DRAFTER_MANIFEST, MANIFEST, MODEL_ID, SPEC_PROOF, _drafter_files, digest, read_manifest, require_proof,
31
+ verify_files,
32
+ )
33
+
34
+ RELEASE = json.loads((HERE / "runtime-release.json").read_text())
35
+ p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
36
+ p.add_argument("--checkpoint", type=Path, default=HERE / RELEASE["checkpoint_subdir"])
37
+ p.add_argument("--drafter", type=Path, default=HERE / RELEASE["drafter_subdir"],
38
+ help="it-assistant drafter dir with spec_equivalence.json")
39
+ p.add_argument("--no-speculation", action="store_true", help="serve without the drafter (one token per step)")
40
+ p.add_argument("--cache-root", type=Path, required=True, help="independent writable TT kernel cache directory")
41
+ p.add_argument("--image", default=RELEASE["image"])
42
+ p.add_argument("--port", type=int, default=8000)
43
+ p.add_argument("--name", default="gemma4-12b-native-api")
44
+ p.add_argument("--max-model-len", type=int, default=262144)
45
+ p.add_argument("--device-ownership-confirmed", action="store_true")
46
+ p.add_argument("--print-command", action="store_true")
47
+ a = p.parse_args()
48
+
49
+ checkpoint, cache = a.checkpoint.resolve(), a.cache_root.resolve()
50
+ try:
51
+ manifest = read_manifest(checkpoint)
52
+ verify_files(checkpoint, manifest)
53
+ proof = require_proof(checkpoint, manifest)
54
+ except (OSError, ValueError, KeyError) as error:
55
+ p.error(f"native checkpoint verification failed: {error}")
56
+ if digest(checkpoint / MANIFEST) != RELEASE["checkpoint_manifest_sha256"]:
57
+ p.error("checkpoint manifest is not the one pinned by runtime-release.json")
58
+ image_id = subprocess.run(["docker", "image", "inspect", a.image, "--format", "{{.Id}}"],
59
+ capture_output=True, text=True).stdout.strip()
60
+ if image_id != RELEASE["image_id"]:
61
+ p.error(f"image {a.image} is {image_id or 'missing'}, not the pinned runtime {RELEASE['image_id']}")
62
+ if not 1024 <= a.max_model_len <= 262144 or not 1 <= a.port <= 65535:
63
+ p.error("invalid context length (1024-262144) or port")
64
+ for mount in (checkpoint, cache):
65
+ if ":" in str(mount):
66
+ p.error("Docker bind-mount paths cannot contain ':'")
67
+ if cache == checkpoint or cache.is_relative_to(checkpoint) or checkpoint.is_relative_to(cache):
68
+ p.error("use independent checkpoint and writable cache directories")
69
+
70
+ draft_len = None
71
+ if not a.no_speculation:
72
+ drafter = a.drafter.resolve()
73
+ if ":" in str(drafter):
74
+ p.error("Docker bind-mount paths cannot contain ':'")
75
+ dm = json.loads((drafter / DRAFTER_MANIFEST).read_text())
76
+ spec = json.loads((drafter / SPEC_PROOF).read_text())
77
+ if _drafter_files(drafter) != dm["files"]:
78
+ p.error("drafter file set / hash mismatch")
79
+ if digest(drafter / DRAFTER_MANIFEST) != RELEASE["drafter_manifest_sha256"]:
80
+ p.error("drafter manifest is not the one pinned by runtime-release.json")
81
+ if dm["target_manifest_sha256"] != digest(checkpoint / MANIFEST):
82
+ p.error("drafter was built for a different target checkpoint")
83
+ if (spec.get("drafter_manifest_sha256") != digest(drafter / DRAFTER_MANIFEST) or spec.get("streams_identical") is not True
84
+ or spec.get("target_proof_sha256") != digest(checkpoint / "equivalence.json")):
85
+ p.error("drafter lacks a speculative-equivalence proof for this target")
86
+ draft_len = int(spec["draft_len"])
87
+
88
+ # Kernel cache keyed by what compiles into it (image + checkpoint precision + serving policy).
89
+ cache_identity = hashlib.sha256(json.dumps({"image": image_id, "manifest": proof["manifest_sha256"],
90
+ "max_model_len": a.max_model_len}, sort_keys=True).encode()).hexdigest()[:16]
91
+ env = {
92
+ "GEMMA4_SINGLE_USER": "1", "GEMMA4_NATIVE": "/model", "HF_MODEL": "/model", "HF_HUB_OFFLINE": "1",
93
+ "TRANSFORMERS_OFFLINE": "1", "VLLM_TARGET_DEVICE": "tt", "TORCHDYNAMO_DISABLE": "1",
94
+ "TT_METAL_CACHE": f"/cache/{cache_identity}",
95
+ **proof.get("runtime_env", {}),
96
+ }
97
+ command = ["docker", "run", "--rm", "--name", a.name, "--publish", f"127.0.0.1:{a.port}:8000",
98
+ "--device", "/dev/tenstorrent:/dev/tenstorrent",
99
+ "-v", "/dev/hugepages:/dev/hugepages", "-v", "/dev/hugepages-1G:/dev/hugepages-1G",
100
+ "-v", f"{checkpoint}:/model:ro", "-v", f"{cache}:/cache",
101
+ "--workdir", "/home/container_app_user/tt-metal",
102
+ "--entrypoint", "/home/container_app_user/tt-metal/python_env/bin/python"]
103
+ if draft_len is not None:
104
+ command += ["-v", f"{drafter}:/drafter:ro"]
105
+ for key, value in sorted(env.items()):
106
+ command += ["-e", f"{key}={value}"]
107
+ command += [a.image, "-m", "vllm.entrypoints.openai.api_server", "--model", "/model",
108
+ "--served-model-name", MODEL_ID, "--host", "0.0.0.0", "--port", "8000",
109
+ "--max-model-len", str(a.max_model_len), "--max-num-seqs", "1", "--block-size", "64",
110
+ "--max-num-batched-tokens", "16384", "--enable-chunked-prefill", "--no-enable-prefix-caching",
111
+ "--seed", "0", "--enable-auto-tool-choice", "--tool-call-parser", "gemma4",
112
+ "--reasoning-parser", "gemma4",
113
+ "--additional-config", json.dumps({"tt": {"trace_mode": "all", "enable_model_warmup": True,
114
+ "trace_region_size": 200000000}})]
115
+ if draft_len is not None:
116
+ command += ["--speculative-config", json.dumps({"method": "gemma4_assistant", "model": "/drafter",
117
+ "num_speculative_tokens": draft_len})]
118
+ print(shlex.join(command), flush=True)
119
+ if not a.print_command:
120
+ if not a.device_ownership_confirmed:
121
+ p.error("stop other P150 workloads and pass --device-ownership-confirmed; this launcher never stops them")
122
+ cache.mkdir(parents=True, exist_ok=True)
123
+ # Never open the P150 while another process holds it (procfs lists its users; the file size is 0).
124
+ pids = Path("/proc/driver/tenstorrent/0/pids")
125
+ while pids.exists() and pids.read_text().strip():
126
+ print("waiting for the P150 to be released", flush=True)
127
+ time.sleep(20)
128
+ os.execvp(command[0], command)