Download README.md from d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE: direct link, hf CLI and curl.
- Browser
- Download file 2.19 kB
-
https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE/resolve/main/README.md
- Command line
-
hf download hf://d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE/README.md
-
curl -L -o README.md https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE/resolve/main/README.md
base_model: ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4
tags:
- qwen
- qwen3.8
- flash-next
- swift
- nvfp4
- fp8
- sglang
- pennyroyal
- blackwell
- local-llm
Swift-1.5 Qwen3.8 Flash-Next NVFP4 — FP8 PLE
Production-oriented derivative of:
ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4
The model weights remain in their original NVFP4 layout while the large PLE embedding table has been converted from BF16 to FP8 E4M3.
FP8 PLE conversion
Original BF16 PLE:
- 95.37 GiB
Converted FP8 PLE:
- 47.68 GiB
- FP8 E4M3
- shared BF16 scale
This substantially reduces the storage and memory footprint of the PLE component.
Validated configuration
Production validation was performed with Pennyroyal / SGLang on an NVIDIA RTX PRO 6000 Blackwell.
Configuration:
- Qwen3.8 Flash-Next
- native NEXTN
- FR-Spec disabled
- native context: 262144 tokens
- FP8 KV cache
- FP8 PLE
- NVMe SSD-streamed PLE
- online MXFP8 enabled at runtime
- max running requests: 4
The production deployment uses Pennyroyal's prepared NVMe PLE overlay.
The NVMe overlay is not included here because it is a runtime-specific, regenerable artifact. This repository contains the portable checkpoint from which the overlay is created.
Online MXFP8 is also a runtime optimization and is not baked into this checkpoint.
Validation results
Selected measurements on RTX PRO 6000 Blackwell:
- MMLU-Pro: 226 / 280 = 80.71%
- historical agentic workload: 174.68 effective tok/s
- fixed 4096 generation median: 133.07 tok/s
- native 262144-token context: PASS
- retrieval at ~261.7K input tokens: PASS
- agentic tool/workflow smoke: 7/7
Performance numbers are runtime- and hardware-specific.
PLE conversion validation
The FP8 PLE derivative was validated against the source model before production deployment.
Runtime
Validated with Pennyroyal / SGLang on NVIDIA Blackwell.
Other runtimes may require support for the model's NVFP4 checkpoint format and FP8 PLE representation.
License and provenance
This is a derivative of the upstream Swift/Qwen model.
Please review the upstream model card and license files included in the repository before redistribution or deployment.