vroomfondel commited on
Commit
7524d39
·
verified ·
1 Parent(s): 167f543

Upload NVFP4 (ModelOpt) quantization

Browse files
Files changed (2) hide show
  1. .gitattributes +0 -1
  2. README.md +5 -1
.gitattributes CHANGED
@@ -34,4 +34,3 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
- model.safetensors.index.json filter=lfs diff=lfs merge=lfs -text
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
README.md CHANGED
@@ -32,7 +32,7 @@ _Before/after sample generation was skipped for this run (`SKIP_GENERATE=1`)._
32
 
33
  ### Quantized MTP/NEXTN draft head (research probe)
34
 
35
- Unlike a BF16-draft-head build, this checkpoint NVFP4-quantizes the MTP/NEXTN draft head (a full-attention + shared-expert-MoE block) and bakes FP8 KV scales onto its attention, so the SGLang p51 KV-scale loader exercises the MTP head. Because transformers drops the MTP layer at load and the multinode FSDP2 export carries no mtp.* at all, the head is produced by a post-export surgical pass (quantizer/requant_mtp_nvfp4.py) that sources the raw BF16 mtp.* from the original checkpoint and NVFP4-packs them, mirroring each module's main-model analogue. Treat as a research probe until a smoke-serve + GSM8K + acceptance-rate check clears it.
36
 
37
  ### MTP scales are heuristic, not calibrated
38
 
@@ -46,4 +46,8 @@ Qwen3.6-35B-A3B carries a 27-layer vision tower. Calibration is text-only, so th
46
 
47
  Serving needs the mamba scheduler knobs to avoid a spec-v2-vs-radix-cache boot crash: mamba_scheduler_strategy extra_buffer plus SGLANG_ENABLE_SPEC_V2=1, along with the flashinfer attention backend and flashinfer_cutlass MoE/FP4-GEMM backends. NEXTN speculative decoding at num_steps=2 was the validated throughput sweet spot; higher step counts (3/4/5) decayed.
48
 
 
 
 
 
49
  <!-- END quant:remarks -->
 
32
 
33
  ### Quantized MTP/NEXTN draft head (research probe)
34
 
35
+ Unlike a BF16-draft-head build, this checkpoint NVFP4-quantizes the MTP/NEXTN draft head (a full-attention + shared-expert-MoE block) and bakes FP8 KV scales onto its attention, so the SGLang p51 KV-scale loader exercises the MTP head. Because transformers drops the MTP layer at load and the multinode FSDP2 export carries no mtp.* at all, the head is produced by a post-export surgical requant that sources the raw BF16 mtp.* from the original checkpoint and NVFP4-packs them, mirroring each module's main-model analogue. Treat as a research probe until a smoke-serve + GSM8K + acceptance-rate check clears it.
36
 
37
  ### MTP scales are heuristic, not calibrated
38
 
 
46
 
47
  Serving needs the mamba scheduler knobs to avoid a spec-v2-vs-radix-cache boot crash: mamba_scheduler_strategy extra_buffer plus SGLANG_ENABLE_SPEC_V2=1, along with the flashinfer attention backend and flashinfer_cutlass MoE/FP4-GEMM backends. NEXTN speculative decoding at num_steps=2 was the validated throughput sweet spot; higher step counts (3/4/5) decayed.
48
 
49
+ ### Serving requires SGLang runtime patches (stock SGLang does not load this)
50
+
51
+ Stock SGLang 0.5.15 crashes loading this checkpoint on GB10/sm121; it needs five launch-time source patches, shipped in the public dgxarley repo (https://github.com/vroomfondel/dgxarley) under roles/k8s_dgx/files/sglang_patches/ and applied by sglang_launch.sh: (1) p50 -- allow QUANTIZED attention for modelopt_fp4 (this is a uniform-W4A4 export that also quantizes the Gated-DeltaNet linear-attn, unlike NVIDIA's MoE-only NVFP4); (2) p51 -- load the main model's baked FP8 KV scales onto RadixAttention (remap ...self_attn.k_proj.k_scale -> ...attn.k_scale), else the full-attn layers default to scale 1.0; (3) p43 -- keep quant_config for the QUANTIZED MTP draft head, else SGLang assumes a BF16 MTP, builds it unquantized, and the fused-MoE loader narrows the unpacked intermediate dim over the NVFP4-packed expert weight ("start(0)+length(512) exceeds dimension size(256)"); (4) p44 -- apply the same KV-scale mapper in the MTP load path (qwen3_5_mtp.py has its own load_weights) so the draft attention loads its baked FP8 scales instead of the 1.0 sentinel; (5) the arch-independent sm121 CUTLASS-FP4 mma patch. p43/p44 are specific to a quantized MTP and are inert for a BF16-MTP checkpoint. Independently, config.json pins torch_dtype: bfloat16 (the Qwen3.6 source config omits it) so SGLang's --dtype auto resolves bf16; without it the unquantized in_proj_ba loads as fp16 and crashes at the CUDA-graph capture ("mat1 and mat2 ... BFloat16 != Half"). Validated end to end on GB10/sm121 (image 0.5.15.post1-sm121): loads register-free, NEXTN speculation active, coherent output, MTP draft attn loads its borrowed KV scales (k_scale 0.0398, byte-identical to the last full-attn layer, the requant borrow source).
52
+
53
  <!-- END quant:remarks -->