ukisai's picture
initial release
17d5f9d
|
Raw History Blame Contribute Delete
2.67 kB

Fresh validation passed on both H100s. The detached suite completed at 17:31 UTC on 2026-09-15, with exit code 0.

Metric Old quant Clean ModelOpt
Checkpoint size (GB) 28.572 21.945
Weight VRAM (GiB) 26.146 19.995
Free VRAM after load (GiB) 52.059 58.270
Cache memory (GiB) — constrained 3.084 9.294
Reported cache tokens — constrained 47,824 149,744
Tested actual context — constrained 47,040 148,960
Cache memory (GiB) — full 42.674 48.885
Reported cache tokens — full 691,036 791,861
Tested actual context — full 262,144 262,144
GSM8K 124/128 124/128
MTP arithmetic 32/32 32/32
KV cache quantized? No No

Each paired inference scenario used identical settings, on the same physical H100: 40% memory utilization for constrained tests and 90% for full-budget tests. GPUs 6 and 7 ran independent scenarios concurrently.

Stronger context checks passed for both models in both budgets: retrieve all three independent values placed at 10%, 50%, and 90% of the prompt, then perform a real 256-token decode at the capacity boundary. Tested context above is actual prompt plus generated tokens, not unused generation allowance. Both full-budget tests reach the unchanged 262,144-token architecture ceiling.

Fresh strict main/MTP loading found no missing or unexpected parameters. MTP arithmetic and image checks passed. Actual cache-entry specifications are identical, with BF16 full-attention KV and unchanged default DeltaNet states. All 193 NVFP4 / 208 FP8 projection assignments match NVIDIA’s pinned reference and the source architecture. All 798 retained BF16 tensors, including 15 MTP tensors, remain byte-identical to the source. All 208 FP8 weights and 995 scale tensors passed numerical checks. Complete checkpoint file hashes match before and after testing.

Calibration audit verified 2,048 examples, 6,459,330 tokens, 64 disjoint held-out examples, pinned dataset revisions, the exact recipe hash, two recorded calibration passes, and coverage of all 193 NVFP4 modules. Calibration and likelihood were not rerun during this validation; their saved records were audited. The fresh GPU quality test is GSM8K above.

One legacy metadata discrepancy was found: the old checkpoint index overstates its tensor payload by 251,008 bytes. Its tensor names/shards and loading are valid; measured checkpoint sizes use actual files. The new checkpoint index is correct.

Limit: H100 uses Marlin weight-only NVFP4 inference. Native Blackwell W4A4 execution remains untested; these are sanity tests, not a complete production benchmark suite.