V100-Validated Quantizations
Models built for and runtime-validated on NVIDIA Tesla V100/SM70, with frozen recipes, checksums, benchmarks, and explicit quality limits.
Image-Text-to-Text • 125B • Updated • 112Note Checkpoint size: 82.08 GB (76.45 GiB). Activation-aware selective AWQ W4A16; validated on 4x V100 32GB TP4 with no CPU weight offload.
leoncca/Laguna-S-2.1-Selective-AWQ
Text Generation • 118B • Updated • 375Note Checkpoint size: 94.30 GB (87.82 GiB). Selective AWQ W4A16 for 4x V100 32GB TP4. Quantization/runtime validated; known Laguna-family output-quality caveats remain, and this artifact does not claim to fix them.
leoncca/Qwen3.8-27B-AEON-Mixed-FP8
Image-Text-to-Text • 28B • Updated • 318Note Checkpoint size: 32.92 GB (30.66 GiB). High-quality mixed E4M3 block-128 FP8 for 4x V100 32GB TP4. Vision, native MTP, full-attention Q/K/V/O, and state-sensitive tensors remain BF16; validated with FP16 KV through 246K context.
leoncca/Qwen3.8-27B-Huihui-Mixed-FP8
Image-Text-to-Text • 28B • Updated • 928Note Checkpoint size: 32.92 GB (30.66 GiB). Versioned mixed E4M3 block-128 FP8 repository. The V100-validated release is v1-d42ca897: 4x V100 32GB TP4 with FP16 KV through 246K context. Current main/v2-739e3c5b follows the updated Huihui mother and has complete artifact/static audit; exact V100/SM70 runtime validation is pending, so v1 quality, speed, long-context, and production results must not be transferred to v2.
leoncca/Qwen3.8-27B-Huihui-AWQ
Image-Text-to-Text • 28B • Updated • 396Note Checkpoint size: 19.55 GB (18.21 GiB). Versioned AWQ W4A16 g128 repository. The V100-validated release is v1-d42ca897: 4x V100 32GB TP4 with FP16 compute/KV; historical 4096-to-512 median decode MTP0/3/7 = 60.922/94.105/86.840 tok/s and AIME 2026 medium/xhigh = 23/30 and 29/30. Current main/v2-739e3c5b follows the updated Huihui mother and has artifact audit plus RTX 3090 native semantic smoke; exact V100 runtime validation is pending, so v1 metrics must not be transferred to v2.
leoncca/Qwen3.8-Flash-Next-NVFP4-QSA-FP8-E4M3-KV-Scales
UpdatedNote Scale payload: 2.64 kB (2.58 KiB); no model weights. Checkpoint-bound 24-tensor calibrated FP8 E4M3 QSA KV scale pack for RadixArk/Qwen3.8-Flash-Next-NVFP4@7b719225. V100/SM70 TP4 MTP0 validated through 128K with 1.836x KV capacity at the same memory budget. Requires draft 1Cat-vLLM #447; the performance matrix used optional sibling drafts #452 and #453. This is not a standalone model.
leoncca/Qwen3.8-Flash-Next-AWQ-g32
Image-Text-to-Text • 181B • Updated • 62 • 1Note Formal tag v1-formal-r6-da62ba7b. Model files: 138.13 GB (128.65 GiB). Native per-expert asymmetric AWQ W4A16 g32; 512/512 expert IDs and 24,568/24,576 layer-expert pairs were naturally activated, with all 93 below-128-token pairs selectively augmented. Bundles checkpoint-bound calibrated FP8 E4M3 QSA KV scales (main QSA KV E4M3, indexer FP16). Validated on 4x V100 32GB with TP4/MTP0 through 128K, including image and video; measured KV capacity 1.8345x versus FP16 KV under the accepted contract.