Gemma 4 12B IT โ TT-native all-BFP8, single P150, 262K context, speculative decoding
A hardware-specific, text-only checkpoint and custom runtime derived from google/gemma-4-12B-it, revision 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7, for one Tenstorrent Blackhole P150. The drafter for speculative decoding is derived from google/gemma-4-12B-it-assistant, revision 46d4c6f13f0ac0ad827b915669b8df9b81c64c51.
This is a community experimental release by Lottolabs, not an official Google or Tenstorrent release. It is not GGUF, AWQ, GPTQ, bitsandbytes, or an ordinary Transformers checkpoint. transformers.from_pretrained(), llama.cpp and hosted Inference Providers cannot load it. It needs the bundled runtime.
Quantization
All-BFP8: every text matrix (attention, MLP, LM head) is stored as TTNN bfloat8_b. The embedding is BF16; 1-D tensors (norms, layer scalars) are lossless BF16 in small.safetensors. See gemma-4-12B-it/precision_plan.json and gemma-4-12B-it/native_manifest.json.
tensors/tensor_cache_bf16/*.tensorbin holds the exact TTNN host tensors that the runtime uploads to device DRAM. The directory name comes from the runtime's tensor-cache writer. The files are not a regenerable cache: they are the checkpoint. Every file is hashed in the manifest, and the loader refuses to start if any file is missing or changed. The drafter (gemma-4-12B-it-assistant/) stores its tensors the same way. Its recorded dtype is bfp8-lm4: BFP8 weights and a BFP4 drafter LM head.
The loader requires two proofs, both bound to the manifests and recorded on the current runtime (p3d, see Runtime):
gemma-4-12B-it/equivalence.json: the full logits of the native reload are bit-exact to quantizing the original weights at load time, over 4,093 teacher-forced tokens (exact_logits_equal: true).gemma-4-12B-it-assistant/spec_equivalence.json: greedy speculative decoding with draft length 5 produced token streams identical to non-speculative greedy decoding on the same proven runtime (streams_identical: true, 3,130 tokens compared).
Quality versus the original BF16 model
This table is the precision-selection study. It is a teacher-forced comparison against the original model run in BF16 on CPU, measured on the exact (unfused) TT build (evidence/quality/quant_table.json). For the shipped runtime's own gate, see Runtime. dPPL is the perplexity change relative to BF16, top-1 is argmax agreement with BF16, and KL is the mean KL divergence on the short-prompt set:
| Plan | Weight read | chat dPPL / top-1 | code dPPL / top-1 | 16K book dPPL / top-1 | KL | GSM8K-50 |
|---|---|---|---|---|---|---|
| all-BFP8 (this release) | 12.7 GB | โ0.3% / 97.1% | +1.0% / 98.3% | +1.6% / 97.2% | 0.009 | 50/50 |
| BFP8 + BFP4 LM head | 12.2 GB | +0.6% / 95.5% | +2.0% / 97.0% | +2.6% / 96.1% | 0.019 | 49/50 |
| MLP in BFP4 | 8.4 GB | +5.2% / 87.1% | +3.9% / 90.8% | +9.9% / 87.7% | 0.175 | 49/50 |
| all-BFP4 | 7.2 GB | +7.2% / 82.7% | +8.3% / 87.4% | +19.2% / 80.1% | 0.296 | โ |
Why BFP8 and not BFP4: with the MLP in BFP4, perplexity rises about 5% and top-1 agreement falls to about 87%. All-BFP4 is worse still. BFP8 stays within about 1โ2% of BF16 perplexity.
Runtime
The current runtime is p3d: image lottolabs/gemma4-12b-tt-p150:p3d, ID sha256:331e79b24fa2575c436158be994c6bc0e184d776d15305e82538d2e4d646f544. Its Gemma 4 tree is at git commit 061e48d. On top of the earlier runtime it adds batched speculative verification, a fused GELUรup kernel, a decode KV write through a rows kernel, and dual-NoC weight reads for the decode matmuls (the TTNN in1 reader overlay, dn branch efc465f).
These fast paths are numerics switches, and the proofs bind them (runtime_env: GEMMA4_VERIFY_SDPA=batched, GEMMA4_FUSE_GELU_MUL=1, GEMMA4_DECODE_KV_ROWS=1, GEMMA4_DECODE_MM_DN=1). They change floating-point rounding relative to the previous runtime p3c, so greedy outputs are not token-identical to p3c's. Within p3d, speculative output is identical to non-speculative output, and HTTP output is identical to the direct runtime (see below). The checkpoint tensors are byte-identical to the p3c release. Only equivalence.json and spec_equivalence.json were re-recorded on p3d.
Quality gate for p3d. Decode-path teacher forcing was compared against the original model in BF16 on CPU (evidence/quality/p3d-gate/):
| Domain | Positions | dPPL vs BF16 | top-1 vs BF16 | argmax agreement with exact build |
|---|---|---|---|---|
| chat | 5,888 | +1.00% | 98.2% | 98.5% |
| code | 3,328 | +0.61% | 98.7% | 98.9% |
| long book | 2,048 | +1.15% | 96.5% | 97.9% |
GSM8K-100 scored 97/100 on p3d. The exact (unfused) build and the BF16 CPU reference also scored 97/100.
Previous runtime: p3c, image lottolabs/gemma4-12b-tt-p150:p3c, ID sha256:53428e7f60a705b75b51b5166f44a912b19c3032981013fb788dd1ce9757aa73. It was published at repo revision 3d9001d5, and its evidence is under evidence/previous-p3c/. On that runtime, HTTP decode ran at 50.6 / 49.3 / 46.9 / 45.1 / 29.7 / 21.8 tok/s at the prompt lengths in the table below. To serve p3c, use the launch.py from that revision; each launcher pins its own runtime image.
Performance
These numbers were measured on runtime p3d over HTTP on one P150 with greedy decoding, 512-token outputs, chat prompts, one user and speculative decoding on (evidence/http-p3d/):
| Prompt tokens | 128 | 2K | 8K | 32K | 131K | 262K (261,632) |
|---|---|---|---|---|---|---|
| Decode tok/s | 63.1 | 52.4 | 53.9 | 43.4 | 30.9 | 24.6 |
| TTFT (s) | 0.09 | 0.65 | 3.2 | 17.0 | 124 | 395 |
Decode speed with speculative decoding depends on how many drafted tokens the content lets the model accept. That is why it does not fall monotonically with prompt length.
LocalMaxxing, 2K prompt (p3d). The server reached 53.9 output tok/s (median of 5), 3,063 tok/s prefill and a 651 ms TTFT, with 1,996 prompt and 256 output tokens. The run was submitted as cmuosddj10kd1lq010lf2ahkj and approved as an unverified run (evidence/localmaxxing/p3d-2k/speed-test.json; the prompt is included).
Serving correctness (p3d)
- HTTP matches the direct runtime token for token. All 28 of 28 checks passed (
evidence/http-p3d/exact.json). They cover chat and completion outputs on 10 prompts, request isolation, sampled requests followed by greedy ones, cancellation mid-prefill and mid-decode, and rejection of image input. The 512-token outputs at 32K and 261,632 prompt tokens were also identical to the direct runtime (evidence/http-p3d/tput_vs_direct.json). - Long context works to the native limit. Passkeys at the beginning, middle and end of the prompt were retrieved at 32,768, 131,072 and 262,016 prompt tokens. A request totalling 262,145 tokens was rejected with HTTP 400 before generation started (
evidence/http-p3d/long.json). - The public package was verified end to end. It was downloaded fresh, anonymously, and served on a P150 (
evidence/public-download-verification.json,evidence/public-serving-smoke.json).
Download and serve
Prerequisites:
- Linux with a supported P150 driver and hugepage mounts at
/dev/hugepagesand/dev/hugepages-1G. - Docker and Python 3.11 or newer.
- Exclusive ownership of one P150.
- Disk space for about 21.3 GB of downloads: a 14.7 GB checkpoint, a 1.2 GB drafter and a 5.4 GB compressed runtime archive. You also need room for the loaded image and a writable kernel cache.
Download the standard-library launcher:
curl --fail --location --output launch.py \
https://huggingface.co/Lottolabs/gemma-4-12B-it-TT-BFP8-P150/resolve/main/launch.py
python3 launch.py \
--cache-root "$HOME/.cache/gemma4-12b-tt-native" \
--device-ownership-confirmed
The launcher does the following before it serves anything:
- Resolves the requested Hugging Face revision to an immutable commit.
- Downloads only the serving package (the checkpoint, drafter, server script and runtime archive) and verifies every byte against
runtime-release.json. - Checks that the published file sets match both native manifests.
- Loads the checksum-pinned runtime archive into Docker and checks the resulting image ID.
serve_native.py then verifies the manifests, file hashes and both proofs again, checks the image ID once more, and starts vLLM. Inside the container, the loader repeats the checks against the image's runtime identity.
The original BF16 weights are not needed. The API is OpenAI-compatible, binds to 127.0.0.1:8000, and serves the model name google/gemma-4-12B-it. The first launch compiles kernels into the kernel cache, so it takes longer than later launches.
Useful modes:
# Download and verify without Docker or a device
python3 launch.py --cache-root /large-disk/gemma4-12b --download-only
# Print the exact serving command without starting it
python3 launch.py --cache-root /large-disk/gemma4-12b --print-command
# Non-speculative serving (no drafter)
python3 launch.py --cache-root /large-disk/gemma4-12b --speculation off --device-ownership-confirmed
Save the immutable revision that the launcher prints, and pass it with --revision to reproduce the same setup exactly.
curl http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "google/gemma-4-12B-it", "temperature": 0, "max_tokens": 256,
"messages": [{"role": "user", "content": "Explain database transactions in three sentences."}]}'
Limits
- One user at a time (
max_num_seqs 1), one P150 (MeshShape 1x1), no multi-device support. - Text only. Requests with images are rejected.
- Context up to 262,144 tokens in total. Sliding-window layers keep their KV cache as 18,432-token rings. Prefill runs in 16,384-token chunks, and prefix caching is off.
- Only greedy requests (
temperature: 0) decode speculatively (drafter K=5, exact per-row verification). Requests with other sampling settings are served but decode one token per step with vLLM's host sampler, so they are slower. - Long prompts have a long time to first token: about 17 s at 32K and about 6.6 min at 262K.
Integrity and provenance
runtime-release.json: the serving inventory the launcher downloads, the runtime archive checksum, the expected Docker image ID and the upstream revisions.release-manifest.json: the complete repository inventory with size and sha256 for every file.SHA256SUMS: file checksums.reproduction.json: source, runtime, serving and evidence contract.runtime/: source of the runtime image layers, recorded inruntime/source.json. It contains:- the Gemma 4 TT model tree (
runtime/gemma4/, git commit061e48ddf6a0eab5dac1f958731fcb70768231d1, byte-identical to the tree in the image); - the TTNN overlay (
runtime/ttnn-overlay/: 1D matmul factory and the dual-NoC in1 reader,dncommitefc465f4b1a44025ae4a33e4dbc7d371b3f9266c, byte-identical to the image); - the vLLM/TT-plugin serving overlay;
- the Dockerfile chain (p2b โ p2c โ ttnn-dn1 โ fast2 โ p3d) and the build script.
This is provenance: serving uses the checksum-pinned
runtime-image.tar.gz.- the Gemma 4 TT model tree (
native_checkpoint.pyandprovenance/build_native_checkpoint.py: the manifest, proof and verification tool, and the checkpoint builder.
License and attribution
Derived from google/gemma-4-12B-it and google/gemma-4-12B-it-assistant. Apache-2.0; see LICENSE. The runtime includes tt-metal and vLLM, which keep their upstream notices; see runtime/licenses/.