Swift-1.5 Qwen3.8 Flash-Next NVFP4 — FP8 PLE

Production-oriented derivative of:

ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4

The model weights remain in their original NVFP4 layout while the large PLE embedding table has been converted from BF16 to FP8 E4M3.

FP8 PLE conversion

Original BF16 PLE:

  • 95.37 GiB

Converted FP8 PLE:

  • 47.68 GiB
  • FP8 E4M3
  • shared BF16 scale

This substantially reduces the storage and memory footprint of the PLE component.

Validated configuration

Production validation was performed with Pennyroyal / SGLang on an NVIDIA RTX PRO 6000 Blackwell.

Configuration:

  • Qwen3.8 Flash-Next
  • native NEXTN
  • FR-Spec disabled
  • native context: 262144 tokens
  • FP8 KV cache
  • FP8 PLE
  • NVMe SSD-streamed PLE
  • online MXFP8 enabled at runtime
  • max running requests: 4

The production deployment uses Pennyroyal's prepared NVMe PLE overlay.

The NVMe overlay is not included here because it is a runtime-specific, regenerable artifact. This repository contains the portable checkpoint from which the overlay is created.

Online MXFP8 is also a runtime optimization and is not baked into this checkpoint.

Validation results

Selected measurements on RTX PRO 6000 Blackwell:

  • MMLU-Pro: 226 / 280 = 80.71%
  • historical agentic workload: 174.68 effective tok/s
  • fixed 4096 generation median: 133.07 tok/s
  • native 262144-token context: PASS
  • retrieval at ~261.7K input tokens: PASS
  • agentic tool/workflow smoke: 7/7

Performance numbers are runtime- and hardware-specific.

PLE conversion validation

The FP8 PLE derivative was validated against the source model before production deployment.

Runtime

Validated with Pennyroyal / SGLang on NVIDIA Blackwell.

Other runtimes may require support for the model's NVFP4 checkpoint format and FP8 PLE representation.

License and provenance

This is a derivative of the upstream Swift/Qwen model.

Please review the upstream model card and license files included in the repository before redistribution or deployment.

Downloads last month
17
Safetensors
Model size
120B params
Tensor type
F8_E4M3
·
BF16
·
U8
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE