Prismyra decision NVFP4, 32 layers, for Qwen3.6-35B-A3B-FP8

A 32-of-40-layer checkpoint of Prismyra with its 256-per-layer routed experts (most of this architecture's weights) in NVFP4 (4-bit floating point weights, block scales of 16 elements, one FP32 scale per expert matrix) and everything else -- attention, the gated-delta-net projections, the shared expert, the multi-token-prediction (MTP) head -- left in the FP8 this project's other checkpoints use. Quantizing just the experts takes them from about 26 GB (FP8) to about 14.5 GB (NVFP4); the repository as a whole is about 19.5 GB, roughly two thirds the size of the FP8 32-layer checkpoint (about 29 GB) it is built from. The LoRA folded into the non-expert weights is a rank-64 adapter trained for this exact checkpoint (32 layers, routed experts already in NVFP4), distilled against the un-truncated 40-layer FP8 checkpoint's own forward pass -- see Training below.

This checkpoint is about 29.2 billion parameters, the same as the FP8 32-layer checkpoint it is built from, not the much smaller total the Hub's automatic parameter counter will show for this repository, for the same structural reason given in the 36-layer NVFP4 card: that counter does not count the packed 4-bit codes and per-group scales inside nvfp4_experts.safetensors as model weights.

This checkpoint requires an NVIDIA Blackwell GPU with native 4-bit floating-point tensor cores; verified only on the RTX PRO 4500 (compute capability 12.0, "sm_120"). It will not run its NVFP4 path on Ada or Hopper cards (L40S, H100), and datacenter Blackwell (B200/GB200, "sm_100") has not been tested. vLLM's NVFP4 MoE kernel this checkpoint depends on is Blackwell-only; on a card or vLLM build without it, loading with PRISMYRA_EXPERTS=nvfp4 raises rather than silently falling back to FP8 (untested outside Blackwell as of this card).

This checkpoint is faster than the FP8 32-layer checkpoint it is built from and scores higher on this project's 5-set production average, but it is not this project's most accurate NVFP4 checkpoint -- the 36-layer NVFP4 checkpoint still holds that position. The margin over FP8-32L is small (91.73 vs 91.39, no significance test run, and the two were measured on different GPUs -- see Evaluation), and this checkpoint is behind FP8-32L on one of the five sets. Choose this one when the FP8 32-layer checkpoint is the one you would otherwise run and 32-layer NVFP4's speed matters; choose the 36-layer NVFP4 checkpoint when accuracy matters most and its slightly higher latency is acceptable.

Status: requires prismyra >= 0.3.1

prismyra.kernels.nvfp4 (reads the routed experts, runs them via vLLM's CUTLASS/FlashInfer NVFP4 kernels, leaves the shared expert and router untouched) is wired into prismyra-serve as of v0.3.1: engine.py converts the routed experts before the rest of the fast-kernel adapter runs, and the adapter (prismyra.kernels.qwen3_moe) recognises the already-converted experts instead of re-wrapping them in FP8. pip install "prismyra[server,fast] @ git+https://github.com/littlemex/Prismyra@v0.3.1" gets you this path. Verified on one RTX PRO 4500: prismyra-serve --require-kernels started against this exact repository (every fast kernel, including the NVFP4 routed experts, applied) and answered two smoke-test questions correctly through the HTTP /ask endpoint; see Evaluation for accuracy and speed measured on this same checkpoint.

Contents

  • layers-*.safetensors, config.json, tokenizer files: the 32 transformer layers with their routed-expert tensors removed from the shards and from model.safetensors.index.json. mtp.safetensors and outside.safetensors are unchanged and keep their own routed-expert tensors (the MTP head's experts and the embeddings/lm-head/final-norm weights are not part of this conversion).
  • nvfp4_experts.safetensors: the 256 routed experts of each of the 32 transformer layers, pre-converted to NVFP4 (packed 4-bit codes, 16-element block scales, one FP32 scale per expert matrix), in this project's own on-disk format (prismyra.kernels.nvfp4.prepare_layer's output) -- not the compressed-tensors layout vLLM's own AutoModel-style quantization loader expects.
  • nvfp4_calib.json: the per-layer activation maxima (measured on 128 training rows) that the NVFP4 forward pass uses to scale its activations.

Use (verified on one RTX PRO 4500 against this exact repository -- see Status)

pip install "prismyra[server,fast] @ git+https://github.com/littlemex/Prismyra@v0.3.1"
PRISMYRA_EXPERTS=nvfp4 \
PRISMYRA_NVFP4_EXPERTS=/path/to/nvfp4_experts.safetensors \
PRISMYRA_NVFP4_CALIB=/path/to/nvfp4_calib.json \
prismyra-serve --model /path/to/this/checkpoint/directory --require-kernels

The checkpoint directory and the two side files are downloaded from this repository as a whole (huggingface_hub.snapshot_download); PRISMYRA_NVFP4_EXPERTS and PRISMYRA_NVFP4_CALIB must point at the two files once downloaded, since the loader reads them from a path, not from the checkpoint directory automatically.

Training

The LoRA is rank 64, alpha 128, targeting attention and gated-delta-net projections and the shared expert, not the routed experts -- the same split this project's other NVFP4 checkpoint's own LoRA uses. The NVFP4 conversion quantizes the routed experts from the un-adapted FP8 weights, so the LoRA and the conversion touch disjoint weights and can be applied independently; the Evaluation table below measures the combined checkpoint directly. The student starts from the FP8 32-layer checkpoint with its routed experts already converted to NVFP4; the teacher is the un-truncated 40-layer FP8 checkpoint's own forward pass (not the 36-layer NVFP4 checkpoint). Training data: 31,418 rows, the same set this project's 36-layer NVFP4 checkpoint's own LoRA trains on (23,585 rows carrying a single-letter answer from Kimi K3 with reasoning disabled, the remaining 7,833 carrying distillation signal, recomputed from the 40-layer teacher for this run). One epoch, learning rate 2.5e-5, gradient accumulation 4, batches capped at 12,500 tokens, trained from a fresh initialisation, random seed 0. After training, the adapter was merged into the base weights; 200 of 200 adapted modules changed, maximum re-quantization error 1.85%. nvfp4_calib.json's activation maxima were measured on the 32-layer FP8 checkpoint before this LoRA existed and were not re-measured after merging; the Evaluation results below are measured with this calibration and the LoRA together.

Evaluation

This project's five production sets (RACE, BoolQ, bury7k, bury10k, Kev; see the 36-layer NVFP4 card for what each measures), scored with the same tool and the same flags on both sides (e6.py --perm 1 --group 32), FP8 sets measured on an L40S and NVFP4 sets measured on an RTX PRO 4500:

set FP8 32-layer (published; this checkpoint's base) NVFP4 experts, before this LoRA this checkpoint (NVFP4 experts + rank-64 LoRA) NVFP4 36-layer (published, most accurate NVFP4 checkpoint)
RACE (579) 95.68 95.34 95.68 96.03
BoolQ (400) 89.25 88.75 89.75 90.25
bury7k (157) 94.27 95.54 94.90 94.90
bury10k (157) 94.90 93.63 93.63 92.36
Kev (764) 82.85 81.41 84.69 85.73
5-set average 91.39 90.93 91.73 91.86

Against the FP8 32-layer checkpoint it is built from, this checkpoint scores 0.34 points higher on the 5-set average: ahead on three sets (BoolQ, bury7k, Kev), tied on RACE, and behind by two questions on bury10k (93.63 vs 94.90). The margin is small (about 15 questions out of 2,057 pooled across the five sets), no significance test was run, and the FP8 and NVFP4 columns were measured on different GPUs (L40S and RTX PRO 4500 respectively), so this comparison is directional rather than a tested claim of superiority. Against the NVFP4 experts alone (before this LoRA), the LoRA recovers +0.80 points, with the largest single-set gain on Kev (+3.28 points). Against the published 36-layer NVFP4 checkpoint, this checkpoint is 0.13 points lower on the 5-set average -- it does not reach that checkpoint's accuracy, and this card does not claim that it does.

This project also maintains a 5,200-question confirmation set and a held-out question pool; neither was run for this checkpoint, and no result in this card comes from them.

Loading this checkpoint's weights in two separate processes on the same RTX PRO 4500 and asking the same 65 questions (16 RACE documents) against each gave bit-identical probabilities (maximum difference 0.0) and no changed answers between the two runs.

Speed, measured on this repository's own files (not inferred from an earlier variant) with this project's profiling script -- an in-process forward pass through the prismyra engine directly, not through the HTTP server -- --group 32, median of 5 runs after 2 warm-up runs, one race-comprehension document of about 5,300 tokens, "N questions" meaning N questions asked about that one document in a single call, on one RTX PRO 4500 (32 GB):

checkpoint 1 question 16 questions 64 questions
this checkpoint (NVFP4 32-layer) 255.2 ms 376.5 ms 678.6 ms
FP8 32-layer (its base, measured the same way, same card, same session) 273.6 ms out of device memory not completable (already out of memory at 16 questions)

This checkpoint answers one question with 6.7% lower latency than the FP8 32-layer checkpoint it is built from (255.2 ms vs 273.6 ms), and -- because the NVFP4 experts take less device memory than the FP8 experts they replace -- it completes the 16- and 64-question measurements that the FP8 checkpoint ran out of memory on at this context length, on the same card. These numbers are close to (within about 1%) a separate measurement of the same NVFP4 32-layer model without this LoRA (254.8 / 374.3 / 674.4 ms), which is expected: the LoRA changes which answer is more likely, not how many FLOPs or how much memory the forward pass takes. This card does not compare these numbers against the 36-layer NVFP4 checkpoint's own published speed figures, since those were measured through a different serving path (see that checkpoint's own card for its numbers).

How it relates to the other checkpoints

This project's FP8 checkpoints are: 40 layers, 36 layers (recommended by default), and 32 layers (the checkpoint this one's weights are built from). It also publishes one other NVFP4 checkpoint, 36 layers, which is more accurate than this one and is the NVFP4 checkpoint to prefer unless this checkpoint's speed and smaller memory footprint at 32 layers specifically matter for your use. Both NVFP4 checkpoints need a Blackwell GPU; none of the FP8 checkpoints do.

Licence and data terms

Apache-2.0, the base model's licence, and this repository carries the base model's LICENSE file unchanged. Some of the training sources behind this checkpoint's LoRA carry their own terms -- RACE and SciQ are distributed for non-commercial research use -- so check those terms before commercial use of this checkpoint, same as for this project's other checkpoints. The majority of this LoRA's training rows carry a label from a frontier LLM (Kimi K3, called through its provider's API) rather than from the training pool's own original label. Of Kimi K3's output, only the single-letter answer is used; no generated rationale or free text from it is reflected in this checkpoint's weights or in this card.

Downloads last month
14
Safetensors
Model size
3B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for littlemex/prismyra-decision-qwen3.6-35b-a3b-nvfp4-32l

Adapter
(7)
this model