Download quantization/VALIDATION.md from ukisai/Swift-Qwen3.8-27B-NVFP4: direct link, hf CLI and curl.
- Browser
- Download file 2.67 kB
-
https://huggingface.co/ukisai/Swift-Qwen3.8-27B-NVFP4/resolve/main/quantization/VALIDATION.md
- Command line
-
hf download hf://ukisai/Swift-Qwen3.8-27B-NVFP4/quantization/VALIDATION.md
-
curl -L -o VALIDATION.md https://huggingface.co/ukisai/Swift-Qwen3.8-27B-NVFP4/resolve/main/quantization/VALIDATION.md
Fresh validation passed on both H100s. The detached suite completed at 17:31 UTC on 2026-09-15, with exit code 0.
| Metric | Old quant | Clean ModelOpt |
|---|---|---|
| Checkpoint size (GB) | 28.572 | 21.945 |
| Weight VRAM (GiB) | 26.146 | 19.995 |
| Free VRAM after load (GiB) | 52.059 | 58.270 |
| Cache memory (GiB) — constrained | 3.084 | 9.294 |
| Reported cache tokens — constrained | 47,824 | 149,744 |
| Tested actual context — constrained | 47,040 | 148,960 |
| Cache memory (GiB) — full | 42.674 | 48.885 |
| Reported cache tokens — full | 691,036 | 791,861 |
| Tested actual context — full | 262,144 | 262,144 |
| GSM8K | 124/128 | 124/128 |
| MTP arithmetic | 32/32 | 32/32 |
| KV cache quantized? | No | No |
Each paired inference scenario used identical settings, on the same physical H100: 40% memory utilization for constrained tests and 90% for full-budget tests. GPUs 6 and 7 ran independent scenarios concurrently.
Stronger context checks passed for both models in both budgets: retrieve all three independent values placed at 10%, 50%, and 90% of the prompt, then perform a real 256-token decode at the capacity boundary. Tested context above is actual prompt plus generated tokens, not unused generation allowance. Both full-budget tests reach the unchanged 262,144-token architecture ceiling.
Fresh strict main/MTP loading found no missing or unexpected parameters. MTP arithmetic and image checks passed. Actual cache-entry specifications are identical, with BF16 full-attention KV and unchanged default DeltaNet states. All 193 NVFP4 / 208 FP8 projection assignments match NVIDIA’s pinned reference and the source architecture. All 798 retained BF16 tensors, including 15 MTP tensors, remain byte-identical to the source. All 208 FP8 weights and 995 scale tensors passed numerical checks. Complete checkpoint file hashes match before and after testing.
Calibration audit verified 2,048 examples, 6,459,330 tokens, 64 disjoint held-out examples, pinned dataset revisions, the exact recipe hash, two recorded calibration passes, and coverage of all 193 NVFP4 modules. Calibration and likelihood were not rerun during this validation; their saved records were audited. The fresh GPU quality test is GSM8K above.
One legacy metadata discrepancy was found: the old checkpoint index overstates its tensor payload by 251,008 bytes. Its tensor names/shards and loading are valid; measured checkpoint sizes use actual files. The new checkpoint index is correct.
Limit: H100 uses Marlin weight-only NVFP4 inference. Native Blackwell W4A4 execution remains untested; these are sanity tests, not a complete production benchmark suite.