--- base_model: ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4 tags: - qwen - qwen3.8 - flash-next - swift - nvfp4 - fp8 - sglang - pennyroyal - blackwell - local-llm --- # Swift-1.5 Qwen3.8 Flash-Next NVFP4 — FP8 PLE Production-oriented derivative of: `ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4` The model weights remain in their original NVFP4 layout while the large PLE embedding table has been converted from BF16 to FP8 E4M3. ## FP8 PLE conversion Original BF16 PLE: - 95.37 GiB Converted FP8 PLE: - 47.68 GiB - FP8 E4M3 - shared BF16 scale This substantially reduces the storage and memory footprint of the PLE component. ## Validated configuration Production validation was performed with Pennyroyal / SGLang on an NVIDIA RTX PRO 6000 Blackwell. Configuration: - Qwen3.8 Flash-Next - native NEXTN - FR-Spec disabled - native context: 262144 tokens - FP8 KV cache - FP8 PLE - NVMe SSD-streamed PLE - online MXFP8 enabled at runtime - max running requests: 4 The production deployment uses Pennyroyal's prepared NVMe PLE overlay. The NVMe overlay is **not included** here because it is a runtime-specific, regenerable artifact. This repository contains the portable checkpoint from which the overlay is created. Online MXFP8 is also a runtime optimization and is not baked into this checkpoint. ## Validation results Selected measurements on RTX PRO 6000 Blackwell: - MMLU-Pro: 226 / 280 = 80.71% - historical agentic workload: 174.68 effective tok/s - fixed 4096 generation median: 133.07 tok/s - native 262144-token context: PASS - retrieval at ~261.7K input tokens: PASS - agentic tool/workflow smoke: 7/7 Performance numbers are runtime- and hardware-specific. ## PLE conversion validation The FP8 PLE derivative was validated against the source model before production deployment. ## Runtime Validated with Pennyroyal / SGLang on NVIDIA Blackwell. Other runtimes may require support for the model's NVFP4 checkpoint format and FP8 PLE representation. ## License and provenance This is a derivative of the upstream Swift/Qwen model. Please review the upstream model card and license files included in the repository before redistribution or deployment.