Qwen3.8-Flash-Next EXL3 K4.25 PLE FP8 v1

A byte-copy hybrid of wrldsuksgo2mars/Qwen3.8-Flash-Next-EXL3-K4.25-v1 and the FP8 PLE n-gram table from nvidia/Qwen3.8-Flash-Next-NVFP4.

All EXL3 tensors are preserved exactly except the 128 BF16 n-gram table parts, which are replaced by NVIDIA's 128 FP8 E4M3 parts and their shared BF16 scale. The routed experts remain calibrated mixed K4/K5 EXL3, averaging 4.25 bits per expert weight, including the EXL3 MTP experts. The other PLE parameters, hash buffers, dense weights, vision weights and tokenizer/processor assets come from the EXL3 release. No NVIDIA NVFP4 expert weights are imported.

The result contains 302,617 tensors and 128,166,418,426 bytes of tensor payload, saving 51,200,245,758 payload bytes versus the BF16-PLE parent.

Data integrity

Construction copies existing tensor byte ranges without decoding, casting, requantization or scale calculation. PLE geometry and hash buffers match; rewritten tensors are verified by readback hashes. Unchanged shards are byte-identical to the EXL3 source. hybrid-provenance.json records source revisions and the rewritten tensor audit; release-manifest.json records file digests. No KL, perplexity or other quality evaluation was performed on this hybrid. Parent benchmark results are not measurements of this checkpoint.

Loading

This is EXL3 with FP8 PLE storage, not an NVFP4 checkpoint. A loader must support both EXL3 and NVIDIA's FP8 n-gram table with its shared scale tensor: model.language_model.layers.1.ple.ple_embedding.ngram_embedding.weight_scale. The EXL3 quantization configuration is preserved with an additional meta.qflashrt_ple annotation describing the copied FP8 table. Stock-loader compatibility is not implied. QFlashRT exports native experts and PLE artifacts separately. Expert activations and KV-cache precision are runtime choices, not changes made by this data transform.

Provenance and licenses

  • EXL3 source: wrldsuksgo2mars/Qwen3.8-Flash-Next-EXL3-K4.25-v1@73a050c27b8c488c65acd6d1c74e45ff02be5fab.
  • FP8 PLE source: nvidia/Qwen3.8-Flash-Next-NVFP4@2061e0b0c5d92bdf7c8fbd4241bbc2af239d7e2d.
  • Original model: Qwen/Qwen3.8-Flash-Next@de4b8e4d43b917e7706784d8bb445c9af86a3540.

The Qwen Community License is included in LICENSE. NVIDIA's copied component is Licensed by NVIDIA Corporation under the NVIDIA Open Model License; see LICENSE-NVIDIA.pdf and NOTICE.txt.

Downloads last month
117
Safetensors
Model size
90B params
Tensor type
BF16
·
F16
·
I16
·
I64
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wrldsuksgo2mars/Qwen3.8-Flash-Next-EXL3-K4.25-PLE-FP8-v1