Upload NVFP4 (ModelOpt) quantization
Browse files- .gitattributes +0 -1
- README.md +5 -1
.gitattributes
CHANGED
|
@@ -34,4 +34,3 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 37 |
-
model.safetensors.index.json filter=lfs diff=lfs merge=lfs -text
|
|
|
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
|
|
README.md
CHANGED
|
@@ -32,7 +32,7 @@ _Before/after sample generation was skipped for this run (`SKIP_GENERATE=1`)._
|
|
| 32 |
|
| 33 |
### Quantized MTP/NEXTN draft head (research probe)
|
| 34 |
|
| 35 |
-
Unlike a BF16-draft-head build, this checkpoint NVFP4-quantizes the MTP/NEXTN draft head (a full-attention + shared-expert-MoE block) and bakes FP8 KV scales onto its attention, so the SGLang p51 KV-scale loader exercises the MTP head. Because transformers drops the MTP layer at load and the multinode FSDP2 export carries no mtp.* at all, the head is produced by a post-export surgical
|
| 36 |
|
| 37 |
### MTP scales are heuristic, not calibrated
|
| 38 |
|
|
@@ -46,4 +46,8 @@ Qwen3.6-35B-A3B carries a 27-layer vision tower. Calibration is text-only, so th
|
|
| 46 |
|
| 47 |
Serving needs the mamba scheduler knobs to avoid a spec-v2-vs-radix-cache boot crash: mamba_scheduler_strategy extra_buffer plus SGLANG_ENABLE_SPEC_V2=1, along with the flashinfer attention backend and flashinfer_cutlass MoE/FP4-GEMM backends. NEXTN speculative decoding at num_steps=2 was the validated throughput sweet spot; higher step counts (3/4/5) decayed.
|
| 48 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 49 |
<!-- END quant:remarks -->
|
|
|
|
| 32 |
|
| 33 |
### Quantized MTP/NEXTN draft head (research probe)
|
| 34 |
|
| 35 |
+
Unlike a BF16-draft-head build, this checkpoint NVFP4-quantizes the MTP/NEXTN draft head (a full-attention + shared-expert-MoE block) and bakes FP8 KV scales onto its attention, so the SGLang p51 KV-scale loader exercises the MTP head. Because transformers drops the MTP layer at load and the multinode FSDP2 export carries no mtp.* at all, the head is produced by a post-export surgical requant that sources the raw BF16 mtp.* from the original checkpoint and NVFP4-packs them, mirroring each module's main-model analogue. Treat as a research probe until a smoke-serve + GSM8K + acceptance-rate check clears it.
|
| 36 |
|
| 37 |
### MTP scales are heuristic, not calibrated
|
| 38 |
|
|
|
|
| 46 |
|
| 47 |
Serving needs the mamba scheduler knobs to avoid a spec-v2-vs-radix-cache boot crash: mamba_scheduler_strategy extra_buffer plus SGLANG_ENABLE_SPEC_V2=1, along with the flashinfer attention backend and flashinfer_cutlass MoE/FP4-GEMM backends. NEXTN speculative decoding at num_steps=2 was the validated throughput sweet spot; higher step counts (3/4/5) decayed.
|
| 48 |
|
| 49 |
+
### Serving requires SGLang runtime patches (stock SGLang does not load this)
|
| 50 |
+
|
| 51 |
+
Stock SGLang 0.5.15 crashes loading this checkpoint on GB10/sm121; it needs five launch-time source patches, shipped in the public dgxarley repo (https://github.com/vroomfondel/dgxarley) under roles/k8s_dgx/files/sglang_patches/ and applied by sglang_launch.sh: (1) p50 -- allow QUANTIZED attention for modelopt_fp4 (this is a uniform-W4A4 export that also quantizes the Gated-DeltaNet linear-attn, unlike NVIDIA's MoE-only NVFP4); (2) p51 -- load the main model's baked FP8 KV scales onto RadixAttention (remap ...self_attn.k_proj.k_scale -> ...attn.k_scale), else the full-attn layers default to scale 1.0; (3) p43 -- keep quant_config for the QUANTIZED MTP draft head, else SGLang assumes a BF16 MTP, builds it unquantized, and the fused-MoE loader narrows the unpacked intermediate dim over the NVFP4-packed expert weight ("start(0)+length(512) exceeds dimension size(256)"); (4) p44 -- apply the same KV-scale mapper in the MTP load path (qwen3_5_mtp.py has its own load_weights) so the draft attention loads its baked FP8 scales instead of the 1.0 sentinel; (5) the arch-independent sm121 CUTLASS-FP4 mma patch. p43/p44 are specific to a quantized MTP and are inert for a BF16-MTP checkpoint. Independently, config.json pins torch_dtype: bfloat16 (the Qwen3.6 source config omits it) so SGLang's --dtype auto resolves bf16; without it the unquantized in_proj_ba loads as fp16 and crashes at the CUDA-graph capture ("mat1 and mat2 ... BFloat16 != Half"). Validated end to end on GB10/sm121 (image 0.5.15.post1-sm121): loads register-free, NEXTN speculation active, coherent output, MTP draft attn loads its borrowed KV scales (k_scale 0.0398, byte-identical to the last full-attn layer, the requant borrow source).
|
| 52 |
+
|
| 53 |
<!-- END quant:remarks -->
|