|
Download README.md from d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE: direct link, hf CLI and curl.
- Browser
- Download file 9.74 kB
-
https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE/resolve/main/README.md
- Command line
-
hf download hf://d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE/README.md
-
curl -L -o README.md https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE/resolve/main/README.md
9.74 kB
| base_model: ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4 | |
| tags: | |
| - qwen | |
| - qwen3.8 | |
| - flash-next | |
| - swift | |
| - nvfp4 | |
| - fp8 | |
| - sglang | |
| - pennyroyal | |
| - blackwell | |
| - local-llm | |
| <!-- D0XIN_RELEASE_NAV_START --> | |
| > **Related models:** [all models](https://huggingface.co/d0xin) | |
| > | |
| > [Swift-1.5-Qwen3.8-27B-Uncensored-BF16](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16) · [Swift-1.5-Qwen3.8-27B-Uncensored-FP8](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8) · [Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer) · **[Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE)** · [Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch) · [Swift-Qwen3.8-27B-FP8](https://huggingface.co/d0xin/Swift-Qwen3.8-27B-FP8) · [Swift-Qwen3.8-27B-Uncensored-BF16](https://huggingface.co/d0xin/Swift-Qwen3.8-27B-Uncensored-BF16) · [Swift-Qwen3.8-27B-Uncensored-FP8](https://huggingface.co/d0xin/Swift-Qwen3.8-27B-Uncensored-FP8) · [Swift-Qwen3.8-27B-Uncensored-NVFP4-LocalHessian-ActivationHeadroom-NInfer](https://huggingface.co/d0xin/Swift-Qwen3.8-27B-Uncensored-NVFP4-LocalHessian-ActivationHeadroom-NInfer) | |
| <!-- D0XIN_RELEASE_NAV_END --> | |
| # Swift-1.5 Qwen3.8 Flash-Next NVFP4 — FP8 PLE | |
| Production-oriented derivative of: | |
| `ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4` | |
| The model weights remain in their original NVFP4 layout while the large | |
| PLE embedding table has been converted from BF16 to FP8 E4M3. | |
| ## FP8 PLE conversion | |
| This checkpoint keeps the original Flash-Next NVFP4 model weights and converts the large PLE embedding table from BF16 to FP8 E4M3. | |
| | Component | Representation | Size | | |
| |---|---|---:| | |
| | Original PLE | BF16 | 95.37 GiB | | |
| | This release | FP8 E4M3 + shared BF16 scale | **47.68 GiB** | | |
| This cuts the PLE storage footprint by approximately 50% while retaining the original NVFP4 weight layout. | |
| The repository contains the **portable checkpoint**. The Pennyroyal NVMe PLE overlay is deliberately not published because it is a regenerable, machine-local runtime artifact. | |
| ## Quick start — RTX PRO 6000 96 GB | |
| The production-qualified path uses Pennyroyal / SGLang v2.5.3 on a single NVIDIA RTX PRO 6000 Blackwell 96 GB GPU. | |
| Reference image: | |
| ```text | |
| ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3 | |
| ``` | |
| ### 1. Download the checkpoint | |
| ```bash | |
| hf download \ | |
| d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE \ | |
| --local-dir /srv/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE | |
| ``` | |
| ### 2. Get Pennyroyal v2.5.3 setup files | |
| ```bash | |
| git clone \ | |
| --depth 1 \ | |
| --branch pennyroyal-v2.5.3-setup1 \ | |
| https://github.com/jpezzulli/sglang-rtxpro6000.git \ | |
| pennyroyal | |
| cd pennyroyal/docker/pennyroyal | |
| cp .env.example .env | |
| ``` | |
| ### 3. Configure the validated profile | |
| A practical starting `.env` matching the validated production profile: | |
| ```dotenv | |
| PENNYROYAL_PROFILE=next-plain | |
| PENNYROYAL_IMAGE=ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3 | |
| PENNYROYAL_PORT=8001 | |
| HOST_MODELS_ROOT=/srv/models | |
| HOST_CACHE_BASE=/var/cache/pennyroyal | |
| HOST_NIXL_STORAGE_BASE=/srv/pennyroyal-nixl | |
| USER_ID=1000 | |
| GROUP_ID=1000 | |
| NVIDIA_GPU=0 | |
| TARGET_MODEL=/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE | |
| PENNY_PLE_BACKEND=nvme | |
| PENNY_PLE_NVME_MODEL=/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-nvme | |
| SGLANG_SM120_ONLINE_MXFP8=true | |
| SGLANG_MM_PREPROCESS_DEVICE=cpu | |
| MAX_RUNNING_REQUESTS=4 | |
| MAX_MAMBA_CACHE_SIZE=24 | |
| ``` | |
| Create the writable cache directories with the same UID/GID used by the container: | |
| ```bash | |
| sudo install -d \ | |
| -o 1000 \ | |
| -g 1000 \ | |
| /var/cache/pennyroyal \ | |
| /srv/pennyroyal-nixl | |
| ``` | |
| ### 4. Prepare the NVMe PLE overlay once | |
| Pennyroyal's normal `/models` mount is read-only, so the one-time preparation uses a writable override: | |
| ```bash | |
| docker compose run --rm --no-deps \ | |
| -v /srv/models:/models \ | |
| pennyroyal exec .venv/bin/python \ | |
| scripts/pennyroyal/prepare_ple_nvme.py \ | |
| --source /models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE \ | |
| --output /models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-nvme | |
| ``` | |
| Keep both the source checkpoint and prepared overlay immutable while serving. | |
| The prepared FP8 PLE table occupies approximately **47.68 GiB** on local SSD/NVMe. NVMe mode replaces its fixed pinned-RAM residency with bounded streaming buffers and reclaimable filesystem page cache. | |
| ### 5. Start | |
| ```bash | |
| docker compose config --quiet | |
| docker compose pull | |
| docker compose up -d | |
| docker compose logs -f pennyroyal | |
| ``` | |
| Startup can take several minutes while the checkpoint is validated and loaded, kernels are prepared, CUDA graphs are captured, and the prepared PLE table checksum is verified. | |
| Check readiness: | |
| ```bash | |
| curl -fsS http://127.0.0.1:8001/health | |
| ``` | |
| OpenAI-compatible API: | |
| ```text | |
| http://127.0.0.1:8001/v1 | |
| ``` | |
| Use `pennyroyal` as the served model name. | |
| Smoke test: | |
| ```bash | |
| curl -s \ | |
| http://127.0.0.1:8001/v1/chat/completions \ | |
| -H 'Content-Type: application/json' \ | |
| -d '{ | |
| "model": "pennyroyal", | |
| "messages": [ | |
| {"role": "user", "content": "Reply with exactly: READY"} | |
| ], | |
| "temperature": 0, | |
| "max_tokens": 64 | |
| }' | |
| ``` | |
| ## Validated production configuration | |
| Production qualification was performed on a single NVIDIA RTX PRO 6000 Blackwell 96 GB GPU with: | |
| - Qwen3.8 Flash-Next | |
| - Swift 1.5 derivative | |
| - NVFP4 target weights | |
| - FP8 E4M3 PLE | |
| - FP8 KV cache | |
| - native NEXTN / MTP speculative decoding | |
| - `next-plain` / FR-Spec disabled | |
| - 262,144-token configured context | |
| - NVMe-backed PLE | |
| - online MXFP8 enabled | |
| - TP=1 | |
| - maximum running requests: 4 | |
| - Pennyroyal / SGLang v2.5.3 in the current production deployment | |
| Current Pennyroyal releases also support other profiles, including FR-Spec and larger context/pool configurations. The measurements below refer to this specific validated profile and should not be compared as if they came from the same benchmark regime. | |
| ## Validation results | |
| ### Quality | |
| | Test | Result | | |
| |---|---:| | |
| | MMLU-Pro | **226 / 280 = 80.71%** | | |
| This was the model-quality check used during qualification of the FP8-PLE derivative. | |
| ### Long context | |
| | Test | Result | | |
| |---|---:| | |
| | Configured native context | **262,144 tokens — PASS** | | |
| | Retrieval near context limit | **~261.7K input tokens — PASS** | | |
| The near-limit retrieval check verifies that the model can prefill and recover the expected information close to the configured 262K context boundary rather than merely accepting a large `max_model_len` value. | |
| ### Agentic / tool behavior | |
| | Test | Result | | |
| |---|---:| | |
| | Agentic tool/workflow smoke | **7 / 7 PASS** | | |
| | Direct tool-call regression on v2.5.3 | **PASS** | | |
| | Hybrid Auto end-to-end routing | **PASS** | | |
| The v2.5.3 production regression test returned the expected `write_file` tool call with correctly formed arguments and `finish_reason=tool_calls`. | |
| The same production target also passed an end-to-end request through the `hybrid_auto` routing layer. | |
| ### Throughput | |
| | Measurement | Result | | |
| |---|---:| | |
| | Historical agentic workload | **174.68 effective tok/s** | | |
| | Fixed 4096-token generation | **133.07 tok/s median** | | |
| | Current v2.5.3 production smoke | **~161.52 E2E tok/s** | | |
| The current production smoke generated 512 tokens in 3.169812 seconds, corresponding to approximately 161.52 end-to-end tokens/s. | |
| These numbers come from different workloads and harnesses. They are reported separately and should not be interpreted as directly comparable benchmark samples. | |
| ### PLE/runtime validation | |
| The FP8 PLE derivative was validated before production deployment, and the deployed NVMe path additionally exercises Pennyroyal's source/config/index/header identity checks and prepared-table checksum verification at startup. | |
| The current production container reports healthr and serves the model through the normal OpenAI-compatible endpoint. | |
| ## What the tests establish | |
| The validation suite was designed to cover different failure modes rather than a single synthetic throughput score: | |
| - **MMLU-Pro** — quality retention after the FP8-PLE conversion. | |
| - **262K / ~261.7K retrieval** — actual near-limit long-context operation. | |
| - **7/7 agentic smoke** — tool/workflow behavior. | |
| - **direct tool regression** — structured function-call correctness on the deployed v2.5.3 runtime. | |
| - **Hybrid Auto check** — successful integration through the production routing layer. | |
| - **throughput runs** — practical single-GPU serving performance. | |
| - **startup integrity checks** — consistency between the portable checkpoint and prepared NVMe PLE derivative. | |
| ## Runtime notes | |
| Pennyroyal v2.5.3 is the recommended reference runtime for this checkpoint. | |
| The v2.5.3 container includes the matching NVMe PLE reader. A separate reader installation is therefore not required for the Docker deployment, but the NVMe overlay must still be prepared once from this checkpoint. | |
| Pennyroyal documentation: | |
| https://github.com/jpezzulli/sglang-rtxpro6000 | |
| NVMe PLE documentation: | |
| https://github.com/jpezzulli/sglang-rtxpro6000/blob/pennyroyal-main-sm120-final/NVME-PLE.md | |
| ## Rank-2 abliteration patch | |
| A reproducible Rank-2 refusal-direction ablation / uncensoring patch for this exact checkpoint is published separately: | |
| https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch | |
| The patch repository contains the projection basis, transformation plan, validated offline baker and reference shard hashes instead of duplicating this full checkpoint. | |
| ## License and provenance | |
| This is a derivative of the upstream Swift/Qwen model. | |
| Please review the upstream model card and license files included in the | |
| repository before redistribution or deployment. | |