d0xin's picture
Add production FP8-PLE model card
eab5a3f verified
|
Raw History Blame Contribute Delete
2.19 kB
metadata
base_model: ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4
tags:
  - qwen
  - qwen3.8
  - flash-next
  - swift
  - nvfp4
  - fp8
  - sglang
  - pennyroyal
  - blackwell
  - local-llm

Swift-1.5 Qwen3.8 Flash-Next NVFP4 — FP8 PLE

Production-oriented derivative of:

ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4

The model weights remain in their original NVFP4 layout while the large PLE embedding table has been converted from BF16 to FP8 E4M3.

FP8 PLE conversion

Original BF16 PLE:

  • 95.37 GiB

Converted FP8 PLE:

  • 47.68 GiB
  • FP8 E4M3
  • shared BF16 scale

This substantially reduces the storage and memory footprint of the PLE component.

Validated configuration

Production validation was performed with Pennyroyal / SGLang on an NVIDIA RTX PRO 6000 Blackwell.

Configuration:

  • Qwen3.8 Flash-Next
  • native NEXTN
  • FR-Spec disabled
  • native context: 262144 tokens
  • FP8 KV cache
  • FP8 PLE
  • NVMe SSD-streamed PLE
  • online MXFP8 enabled at runtime
  • max running requests: 4

The production deployment uses Pennyroyal's prepared NVMe PLE overlay.

The NVMe overlay is not included here because it is a runtime-specific, regenerable artifact. This repository contains the portable checkpoint from which the overlay is created.

Online MXFP8 is also a runtime optimization and is not baked into this checkpoint.

Validation results

Selected measurements on RTX PRO 6000 Blackwell:

  • MMLU-Pro: 226 / 280 = 80.71%
  • historical agentic workload: 174.68 effective tok/s
  • fixed 4096 generation median: 133.07 tok/s
  • native 262144-token context: PASS
  • retrieval at ~261.7K input tokens: PASS
  • agentic tool/workflow smoke: 7/7

Performance numbers are runtime- and hardware-specific.

PLE conversion validation

The FP8 PLE derivative was validated against the source model before production deployment.

Runtime

Validated with Pennyroyal / SGLang on NVIDIA Blackwell.

Other runtimes may require support for the model's NVFP4 checkpoint format and FP8 PLE representation.

License and provenance

This is a derivative of the upstream Swift/Qwen model.

Please review the upstream model card and license files included in the repository before redistribution or deployment.