d0xin's picture
Add production quick start and validation details
1a3c606 verified
|
Raw History Blame Contribute Delete
9.74 kB
---
base_model: ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4
tags:
- qwen
- qwen3.8
- flash-next
- swift
- nvfp4
- fp8
- sglang
- pennyroyal
- blackwell
- local-llm
---
<!-- D0XIN_RELEASE_NAV_START -->
> **Related models:** [all models](https://huggingface.co/d0xin)
>
> [Swift-1.5-Qwen3.8-27B-Uncensored-BF16](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16) · [Swift-1.5-Qwen3.8-27B-Uncensored-FP8](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8) · [Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer) · **[Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE)** · [Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch](https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch) · [Swift-Qwen3.8-27B-FP8](https://huggingface.co/d0xin/Swift-Qwen3.8-27B-FP8) · [Swift-Qwen3.8-27B-Uncensored-BF16](https://huggingface.co/d0xin/Swift-Qwen3.8-27B-Uncensored-BF16) · [Swift-Qwen3.8-27B-Uncensored-FP8](https://huggingface.co/d0xin/Swift-Qwen3.8-27B-Uncensored-FP8) · [Swift-Qwen3.8-27B-Uncensored-NVFP4-LocalHessian-ActivationHeadroom-NInfer](https://huggingface.co/d0xin/Swift-Qwen3.8-27B-Uncensored-NVFP4-LocalHessian-ActivationHeadroom-NInfer)
<!-- D0XIN_RELEASE_NAV_END -->
# Swift-1.5 Qwen3.8 Flash-Next NVFP4 — FP8 PLE
Production-oriented derivative of:
`ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4`
The model weights remain in their original NVFP4 layout while the large
PLE embedding table has been converted from BF16 to FP8 E4M3.
## FP8 PLE conversion
This checkpoint keeps the original Flash-Next NVFP4 model weights and converts the large PLE embedding table from BF16 to FP8 E4M3.
| Component | Representation | Size |
|---|---|---:|
| Original PLE | BF16 | 95.37 GiB |
| This release | FP8 E4M3 + shared BF16 scale | **47.68 GiB** |
This cuts the PLE storage footprint by approximately 50% while retaining the original NVFP4 weight layout.
The repository contains the **portable checkpoint**. The Pennyroyal NVMe PLE overlay is deliberately not published because it is a regenerable, machine-local runtime artifact.
## Quick start — RTX PRO 6000 96 GB
The production-qualified path uses Pennyroyal / SGLang v2.5.3 on a single NVIDIA RTX PRO 6000 Blackwell 96 GB GPU.
Reference image:
```text
ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3
```
### 1. Download the checkpoint
```bash
hf download \
d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE \
--local-dir /srv/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE
```
### 2. Get Pennyroyal v2.5.3 setup files
```bash
git clone \
--depth 1 \
--branch pennyroyal-v2.5.3-setup1 \
https://github.com/jpezzulli/sglang-rtxpro6000.git \
pennyroyal
cd pennyroyal/docker/pennyroyal
cp .env.example .env
```
### 3. Configure the validated profile
A practical starting `.env` matching the validated production profile:
```dotenv
PENNYROYAL_PROFILE=next-plain
PENNYROYAL_IMAGE=ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3
PENNYROYAL_PORT=8001
HOST_MODELS_ROOT=/srv/models
HOST_CACHE_BASE=/var/cache/pennyroyal
HOST_NIXL_STORAGE_BASE=/srv/pennyroyal-nixl
USER_ID=1000
GROUP_ID=1000
NVIDIA_GPU=0
TARGET_MODEL=/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE
PENNY_PLE_BACKEND=nvme
PENNY_PLE_NVME_MODEL=/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-nvme
SGLANG_SM120_ONLINE_MXFP8=true
SGLANG_MM_PREPROCESS_DEVICE=cpu
MAX_RUNNING_REQUESTS=4
MAX_MAMBA_CACHE_SIZE=24
```
Create the writable cache directories with the same UID/GID used by the container:
```bash
sudo install -d \
-o 1000 \
-g 1000 \
/var/cache/pennyroyal \
/srv/pennyroyal-nixl
```
### 4. Prepare the NVMe PLE overlay once
Pennyroyal's normal `/models` mount is read-only, so the one-time preparation uses a writable override:
```bash
docker compose run --rm --no-deps \
-v /srv/models:/models \
pennyroyal exec .venv/bin/python \
scripts/pennyroyal/prepare_ple_nvme.py \
--source /models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE \
--output /models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-nvme
```
Keep both the source checkpoint and prepared overlay immutable while serving.
The prepared FP8 PLE table occupies approximately **47.68 GiB** on local SSD/NVMe. NVMe mode replaces its fixed pinned-RAM residency with bounded streaming buffers and reclaimable filesystem page cache.
### 5. Start
```bash
docker compose config --quiet
docker compose pull
docker compose up -d
docker compose logs -f pennyroyal
```
Startup can take several minutes while the checkpoint is validated and loaded, kernels are prepared, CUDA graphs are captured, and the prepared PLE table checksum is verified.
Check readiness:
```bash
curl -fsS http://127.0.0.1:8001/health
```
OpenAI-compatible API:
```text
http://127.0.0.1:8001/v1
```
Use `pennyroyal` as the served model name.
Smoke test:
```bash
curl -s \
http://127.0.0.1:8001/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "pennyroyal",
"messages": [
{"role": "user", "content": "Reply with exactly: READY"}
],
"temperature": 0,
"max_tokens": 64
}'
```
## Validated production configuration
Production qualification was performed on a single NVIDIA RTX PRO 6000 Blackwell 96 GB GPU with:
- Qwen3.8 Flash-Next
- Swift 1.5 derivative
- NVFP4 target weights
- FP8 E4M3 PLE
- FP8 KV cache
- native NEXTN / MTP speculative decoding
- `next-plain` / FR-Spec disabled
- 262,144-token configured context
- NVMe-backed PLE
- online MXFP8 enabled
- TP=1
- maximum running requests: 4
- Pennyroyal / SGLang v2.5.3 in the current production deployment
Current Pennyroyal releases also support other profiles, including FR-Spec and larger context/pool configurations. The measurements below refer to this specific validated profile and should not be compared as if they came from the same benchmark regime.
## Validation results
### Quality
| Test | Result |
|---|---:|
| MMLU-Pro | **226 / 280 = 80.71%** |
This was the model-quality check used during qualification of the FP8-PLE derivative.
### Long context
| Test | Result |
|---|---:|
| Configured native context | **262,144 tokens — PASS** |
| Retrieval near context limit | **~261.7K input tokens — PASS** |
The near-limit retrieval check verifies that the model can prefill and recover the expected information close to the configured 262K context boundary rather than merely accepting a large `max_model_len` value.
### Agentic / tool behavior
| Test | Result |
|---|---:|
| Agentic tool/workflow smoke | **7 / 7 PASS** |
| Direct tool-call regression on v2.5.3 | **PASS** |
| Hybrid Auto end-to-end routing | **PASS** |
The v2.5.3 production regression test returned the expected `write_file` tool call with correctly formed arguments and `finish_reason=tool_calls`.
The same production target also passed an end-to-end request through the `hybrid_auto` routing layer.
### Throughput
| Measurement | Result |
|---|---:|
| Historical agentic workload | **174.68 effective tok/s** |
| Fixed 4096-token generation | **133.07 tok/s median** |
| Current v2.5.3 production smoke | **~161.52 E2E tok/s** |
The current production smoke generated 512 tokens in 3.169812 seconds, corresponding to approximately 161.52 end-to-end tokens/s.
These numbers come from different workloads and harnesses. They are reported separately and should not be interpreted as directly comparable benchmark samples.
### PLE/runtime validation
The FP8 PLE derivative was validated before production deployment, and the deployed NVMe path additionally exercises Pennyroyal's source/config/index/header identity checks and prepared-table checksum verification at startup.
The current production container reports healthr and serves the model through the normal OpenAI-compatible endpoint.
## What the tests establish
The validation suite was designed to cover different failure modes rather than a single synthetic throughput score:
- **MMLU-Pro** — quality retention after the FP8-PLE conversion.
- **262K / ~261.7K retrieval** — actual near-limit long-context operation.
- **7/7 agentic smoke** — tool/workflow behavior.
- **direct tool regression** — structured function-call correctness on the deployed v2.5.3 runtime.
- **Hybrid Auto check** — successful integration through the production routing layer.
- **throughput runs** — practical single-GPU serving performance.
- **startup integrity checks** — consistency between the portable checkpoint and prepared NVMe PLE derivative.
## Runtime notes
Pennyroyal v2.5.3 is the recommended reference runtime for this checkpoint.
The v2.5.3 container includes the matching NVMe PLE reader. A separate reader installation is therefore not required for the Docker deployment, but the NVMe overlay must still be prepared once from this checkpoint.
Pennyroyal documentation:
https://github.com/jpezzulli/sglang-rtxpro6000
NVMe PLE documentation:
https://github.com/jpezzulli/sglang-rtxpro6000/blob/pennyroyal-main-sm120-final/NVME-PLE.md
## Rank-2 abliteration patch
A reproducible Rank-2 refusal-direction ablation / uncensoring patch for this exact checkpoint is published separately:
https://huggingface.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch
The patch repository contains the projection basis, transformation plan, validated offline baker and reference shard hashes instead of duplicating this full checkpoint.
## License and provenance
This is a derivative of the upstream Swift/Qwen model.
Please review the upstream model card and license files included in the
repository before redistribution or deployment.