DeepSeek-V4.1-EXL3-K3.25-v1

Routed-expert-only mixed EXL3 K3/K4 quantization of deepseek-ai/DeepSeek-V4.1-Flash, source revision dba1be0a40aa45a94ad051997016db3960a90277. Built with the custom V4.1 support in our GPTQModel fork and the reproducible ds41rt quantization workflow.

Quantization recipe

All routed gate, up and down projections are quantized: 40 main-model blocks with 384 experts each, plus three dSpark blocks with 128 experts each. The 47,232 projections comprise 35,424 K3 and 11,808 K4 projections. K4 assignments rank base-K3 reconstruction error weighted by natural squared router-gate mass, with gate:up:down allocation quotas of 3:5:8 per block. The routed matrix average is 3.25 bits per weight; this is not the whole-model storage rate and excludes scales and other packing overhead.

Calibration uses the unchanged GLM-5.3 EXL3 NEXT corpus: 1,441 original records, 1,056,269 tokens under this checkpoint's tokenizer. dSpark uses 327,680 fixed stratified anchors (seed 20260809), grouped jointly by original prompt. Calibration propagates the selected mixed weights layer by layer. Activation generation runs on two RTX GPUs; trellis search uses those GPUs and four Spark workers.

Storage and loading

Weights use standard safetensors files and a Hugging Face weight index, with checkpoint-native tensor names. quantize_config.json records per-projection EXL3 storage and preserves the source's native quantization configuration. Non-routed tensors retain their source representation. Each PLE table and its scales occupy an isolated shard group, separate from the other table and all other tensors. A PLE group can span multiple files; tensors are not split. These groups can be reused through hard links when constructing later variants.

Standard file formatting does not imply stock Transformers or an existing inference engine can execute this mixed V4.1 checkpoint. Loading requires support for V4.1, native non-routed weights, EXL3 routed projections, and dSpark as needed. The bundled source reference inference code is preserved for architectural reference and is not an EXL3 serving implementation.

Validation and limitations

The publication workflow checks tensor inventories, packed-buffer geometry, per-block bit quotas, shard/index/config consistency and PLE file isolation. It does not perform a final full-model replay or retain calibration data for one. Behavioral quality, normal prompts and tool calls remain unvalidated pending inference-engine integration. No benchmark results or stock-loader compatibility are claimed. Source-model capabilities described in README.source.md are not validation results for this quantization. See LICENSE for the source license.

Downloads last month
62
Safetensors
Model size
320B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wrldsuksgo2mars/DeepSeek-V4.1-EXL3-K3.25-v1

Quantized
(101)
this model
Quantizations
1 model