Prismyra decision FP8, 36 layers, for Qwen3.6-35B-A3B-FP8

A full-weight FP8 checkpoint truncated to 36 of the base model's 40 layers, with a LoRA trained specifically for this depth -- distilled from the 40-layer checkpoint's own output distribution -- folded into the FP8 weights and requantized per 128x128 block. It is for Prismyra: an engine that answers typed questions (booleans, choices) about a document by reading the probability of each declared option's token in a single forward pass, rather than generating an answer -- see the Prismyra repository for what that trades off and where it does not apply.

Of the three FP8 checkpoints in this family (40, 36, 32 layers), this is the one recommended by default: on the five sets below it ties or clears every bar the 40-layer checkpoint does, while running measurably faster.

This revision raises the folded-in LoRA's rank from 16 to 64 (alpha scaled from 32 to 128 to keep the same effective update magnitude) and retrains it by knowledge distillation from this checkpoint's own 36-layer forward pass. The previous revision (rank-16 LoRA) stays available at commit e24a3c89e7185f28f02f8ebca878804bbffe768d; every absolute number quoted below for "the previous revision" is that commit's own published numbers, not re-measured here.

Use

pip install "prismyra[server,fast] @ git+https://github.com/littlemex/Prismyra@v0.3.0"
from prismyra import Prismyra, Boolean

engine = Prismyra("littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l")
result = engine.ask(document, [Boolean(id="q", prompt="Is this a two-sided agreement?")])
prismyra-serve --model littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l --require-kernels

--require-kernels refuses to start rather than serve at a fraction of the speed. Prismyra's fused kernels are matched against the checkpoint by per-layer tensor shape, not by layer count or adapter rank, so this checkpoint gets every kernel the base model gets regardless of which revision's LoRA is folded in -- there is no flag to set for this.

Training

The adapter targets 250 modules (attention and gated-delta-net projections, the shared expert) at rank 64 instead of the previous revision's 16 (76.8M parameters instead of 19.2M), initialised by zero-padding the previous revision's own trained rank-16 adapter up to rank 64 so training continues from where that adapter already was rather than from zero. Training data is the same pool used for the previous revision's own adapter (RACE, SciQ, BoolQ-family and twelve others, 31,418 rows); the distillation target is this checkpoint's own 36-layer forward pass, with the gold label on about 7% of rows -- the rows where the pool's own labeled answer and a frontier LLM's (Kimi K3) answer disagreed -- taken from that frontier LLM's answer letter rather than the pool's own label. One epoch, learning rate 2.5e-5, three L40S GPUs, about 19 GPU-hours.

On 2,500 held-out training-distribution rows, this reduced the KL divergence to the distillation target by 51.2%, against 26.8% for the previous revision's own rank-16 adapter measured the same way -- the only point in this project's rank progression (16, 64, 128) where raising the rank measurably helped on this axis; rank 128, tried next, did not improve on this result on any of three held-out checks and was not published.

Evaluation

Five sets, served by Prismyra on one L40S, compared against the previous revision (rank-16 LoRA, commit e24a3c89), the same protocol as that revision's own published numbers: RACE (579 questions), BoolQ (400), bury7k and bury10k (the same RACE questions buried in about 7,000 and 10,000 tokens of unrelated articles, 157 questions each), and Kev (a 764-question transfer test, of which about 30% comes from families this checkpoint never trained on in any form).

set point difference vs. the previous revision 95% interval sign-test p
RACE (579) +0.17 [-0.69, +1.04] --
BoolQ (400) +0.25 [-0.75, +1.25] --
bury7k (157) 0.00 [-3.18, +3.18] --
bury10k (157) +1.27 [-1.27, +4.46] --
Kev (764) +0.92 [-0.39, +2.36] 0.265

This is the first point in this project's rank progression with a point estimate at or above the previous revision on every one of these five sets. It is also the first to clear this project's weaker, 2,050-question held-out set (disjoint from every set above and from this checkpoint's own training data) by a margin a sign test calls significant: +1.51 points, 95% interval [+0.49, +2.59], p = 0.0062. Read together with the row above: the margin is positive and significant on a broad held-out slice, but on this project's specific production bar -- Kev, the 764-question set, sign-test p < 0.05 -- it is positive but not significant (p = 0.265).

bury7k is the plainest result to read on its own: zero point difference on 157 questions buried in roughly 7,000 tokens of unrelated text, so the higher-rank adapter shows no advantage on that specific long-context slice even where bury10k, 3,000 tokens longer, shows one.

This checkpoint is also measured against JevBench (a 231-question benchmark none of these checkpoints trained on) and against four general-purpose LLMs on the same five sets above, in the training recipe's README. The short version, measured on the previous (rank-16) revision and not re-measured for this one: this checkpoint's family is not significantly different from the 40-layer checkpoint on JevBench, but its calibration there (Brier and expected calibration error) is worse than the 40-layer checkpoint's; and against general-purpose LLMs it loses on accuracy overall, most clearly on Kev, at one to a few orders of magnitude lower estimated GPU-time cost depending on how long the context is.

Latency

Two conditions. First, this revision against the previous revision directly, same L40S, same script, measured twice in alternation to separate a real difference from GPU/driver warm-up state:

1 question 16 questions 64 questions
previous revision (rank-16 LoRA) 233.7 ms 348.4 ms 623.1 ms
this revision (rank-64 LoRA) 233.3 ms 348.4 ms 622.8 ms

Within 0.2 ms at every width -- folding in a higher-rank adapter costs nothing measurable, which is expected: the adapter changes weight values, not the checkpoint's shape or layer count, so it gets the same kernels at the same speed.

Second, measured on the previous revision against the 40- and 32-layer checkpoints and the untrained base, same script and GPU, on the same ~5,300-token document repeated across calls (repeating a document, once warmed up, costs the same as making it fresh on every call -- see docs/PERFORMANCE.md in the repository). Not re-measured for this revision, but the comparison above shows adapter rank does not move these numbers:

checkpoint 1 question 16 questions 64 questions
base, untrained, 40 layers 276.8 ms 408.7 ms 733.4 ms
40-layer 269.9 ms 399.4 ms 716.8 ms
36-layer (this family) 244.3 ms 362.1 ms 650.3 ms
32-layer 218.0 ms 324.5 ms 584.7 ms

The speed difference across the three published checkpoints tracks their layer count, which is what fused kernels applying at every depth looks like -- a checkpoint that fell back to the unfused path would be roughly three times slower than the figures above, at any depth.

How it relates to the other checkpoints

This project publishes five repositories:

Licence and data terms

Apache-2.0, the base model's licence, and this repository carries the base model's LICENSE file unchanged. Some of the training sources behind this checkpoint's LoRA carry their own terms -- RACE and SciQ are distributed for non-commercial research use -- so check those terms before commercial use of this checkpoint. About 7% of training rows (this revision only; the previous revision's adapter did not use this) take their gold label from a frontier LLM (Kimi K3, called through its provider's API) rather than from the pool's own label, used only as a single-letter answer where it disagreed with the pool; no generated rationale or free text from that model is reproduced in this checkpoint's weights or in this card.

Downloads last month
325
Safetensors
Model size
33B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l

Adapter
(8)
this model