vroomfondel's picture
Upload NVFP4 (ModelOpt) quantization
5da9eb4 verified
|
Raw History Blame
4.75 kB
metadata
library_name: transformers
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
pipeline_tag: text-generation
tags:
  - modelopt
  - nvfp4
  - fp4

qwen3.8-27b-nvfp4-modelopt

NVFP4 quantization of Qwen/Qwen3.8-27B via NVIDIA ModelOpt.

Quantization details (auto-generated)

  • source model: Qwen/Qwen3.8-27B
  • qformat: nvfp4 kv_cache: fp8
  • calibration: ? samples from ?
  • producer: NVIDIA ModelOpt ?
  • generated: ?

Before/after sample generation was skipped for this run (SKIP_GENERATE=1).

Notes

Validation status -- structurally verified and smoke-tested coherent

Quantized and served on a single DGX Spark (GB10/sm121) on 2026-08-15. The export passed every structural check: the vision tower is BF16 (333 tensors, zero scale tensors, dtypes and shapes identical to the source), quant_algo is NVFP4 rather than MIXED_PRECISION, the FP8 KV scales sit on exactly the 16 full-attention layers (indices 3, 7, ..., 63), in_proj_qkv is NVFP4-packed on all 48 Gated-DeltaNet layers while conv1d/in_proj_a/in_proj_b stayed BF16, and all 15 mtp.* tensors are present and unquantized. Serving on SGLang 0.5.17 then produced coherent output on all four probe types (German two-sentence explanation, a multi-step train word problem, a five-sentence historical paragraph, a translation), each finishing with finish_reason=stop. The word problem was solved correctly (17:00, with the right derivation), which is the more informative signal: a broken Gated-DeltaNet path degrades into word salad rather than into arithmetic mistakes. This is a four-prompt smoke test, not an evaluation -- GSM8K or a comparable suite is still owed before any quality claim.

Uniform W4A4 on a Gated-DeltaNet hybrid

Unlike NVIDIA's Qwen3.5/3.6 NVFP4 releases, which quantize the FFN/expert path and leave attention in BF16, this build quantizes attention as well, including the Gated-DeltaNet linear-attention path that carries 48 of the model's 64 layers. That is the deliberate point of the profile and it is also where the risk sits: an inadequately loaded scale on the fused linear_attn.in_proj_qkv degrades this architecture into complete word salad rather than into a measurable accuracy drop. Judge the build on generated output, not on the fact that it loads.

Serving requires the qwen3_5 attention-quant and KV-scale loader changes

Because attention is quantized and FP8 KV scales are baked into the checkpoint, SGLang needs the qwen3_5 attention-quant override and baked-KV-scale loader changes (sgl-project/sglang PR #31220) plus the NVFP4 scalar-scale fix for merged and fused linears (PR #29151, merged upstream 2026-07-13) that the fused Gated-DeltaNet in_proj_qkv depends on. Use the flashinfer attention backend: the triton backend hits a forward-time crash in RadixLinearAttention.forward for this exact configuration (sgl-project/sglang#29577, still open). A reliable check that the KV scales actually loaded is that the server logs "Using FP8 KV cache but no scaling factors provided" zero times.

Multimodal, vision tower kept BF16

Calibration is text-only, so the 27-layer vision tower, its merger and the embeddings are excluded from quantization and stay BF16, avoiding the amax=0 degenerate-quant failure mode. The language-side FFNs that consume projected image tokens ARE quantized, and they were calibrated on text alone, so the image path is the least-validated surface of this build and should be checked against the BF16 source before being relied on.

BF16 MTP head, speculative decoding off by default

The NEXTN/MTP draft head (mtp_num_hidden_layers=1) ships unquantized: transformers drops mtp.* at load for every Qwen3.5 architecture, so ModelOpt never sees it and passes the tensors through as BF16. Speculative decoding is therefore available but is left disabled in the shipped serving profile until bare serving is verified coherent, so that a quality regression can never be confused with a drafter problem.

Quantize on a single Spark with offload, not force-on-GPU

Quantize with SEQ_DEVICE_MAP=0 (offload / device_map=auto). GPU and CPU share the same ~121 GB unified memory pool on a DGX Spark, so offloading the 55.6 GB source costs nothing here. Do NOT use SEQ_DEVICE_MAP=1 with a high GPU_MAX_MEM_PCT on a single Spark: that reserves most of the shared pool as GPU and OOM-kills weight loading regardless of the percentage. Force-on-GPU is a multi-GPU (4x H200) setting only.