File size: 3,644 Bytes
b51cd45 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 | ---
license: mit
base_model: deepreinforce-ai/Ornith-1.0-35B
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- compressed-tensors
- fp8
- moe
- vllm
- multimodal
---
# Ornith-1.0-35B-FP8
**FP8_DYNAMIC** W8A8 quantisation of [`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B) — a 35B Mixture-of-Experts (256 routed experts, 8 active, + 1 shared expert) multimodal model with a hybrid linear+full attention stack and a vision tower. Produced with `llm-compressor`.
## Quantisation
- **Scheme**: **FP8_DYNAMIC** W8A8 — 8-bit FP8 (E4M3) weights with dynamic per-token FP8 activation quantisation (data-free; activation scales computed at inference).
- **Tooling**: `llm-compressor` 0.12.0 + `compressed-tensors` 0.17.1, transformers 5.8.1, on a Blackwell RTX PRO 6000 (SM120).
- **Quantised**: the 256 routed `mlp.experts.*` FFNs, the `mlp.shared_expert` FFN, and the full-attention `self_attn` projections.
- **Kept in bf16** (`quantization_config.ignore`): `lm_head`, all `model.visual.*` (vision tower), the entire `linear_attn` Mamba/SSM block (`in_proj_*`, `out_proj`, `conv1d`), the MoE router `mlp.gate`, and `mlp.shared_expert_gate`.
- **No MTP head**: the Ornith base ships **0** `mtp.*` tensors (despite `mtp_num_hidden_layers: 1` in its config), so there is no speculative-decoding head to preserve or graft.
- **Recipe**: [`recipe.yaml`](recipe.yaml).
## Quality
KLD is per-token KL divergence vs the bf16 base over 8 [`neuralmagic/calibration`](https://huggingface.co/datasets/neuralmagic/calibration) samples (`max_seq=1024`). Because transformers cannot execute the fused-MoE compressed checkpoint (see below), the checkpoint was decompressed to bf16 — applying the quant→dequant round-trip so precision loss is captured — and evaluated there. PPL is wikitext-2-raw (test), non-overlapping 2048-token chunks, identical tokenizer to the base.
| Model | KLD vs base (nats) | PPL (wikitext-2-raw) | ΔPPL |
|---|---|---|---|
| [`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B) (bf16) | 0 | 6.7539 | — |
| **this** (FP8) | **0.0161** | **6.7547** | **+0.01%** |
> The FP8 KLD below is **weight-only**: the bf16-decompressed eval does not model FP8_DYNAMIC's per-token activation quant, so the served KLD is marginally higher.
## Inference — vLLM / SGLang (not transformers)
> ⚠️ **This compressed MoE checkpoint does not run under 🤗 transformers.** `Qwen3_5MoeForConditionalGeneration` fuses the experts into a 3D `torch._grouped_mm` kernel that rejects fp8/fp4 weights (`RuntimeError: Expected mat_a to be Float32, BFloat16 or Float16 matrix, got Float8_e4m3fn`), and the per-expert scales load as `UNEXPECTED` keys. Serve it with **vLLM** or **SGLang**, which have the compressed-tensors fused-MoE kernels.
```bash
vllm serve huginnfork/Ornith-1.0-35B-FP8 \
--trust-remote-code \
--gpu-memory-utilization 0.85 \
--max-model-len 8192 \
--quantization compressed-tensors
```
**On Blackwell (RTX PRO 6000 / B200, SM120/SM100):** FlashInfer's JIT arch probe currently mis-detects SM120 and aborts engine startup (`FlashInfer requires GPUs with sm75 or higher`). Disable it:
```bash
TORCH_CUDA_ARCH_LIST=12.0+PTX VLLM_USE_FLASHINFER_SAMPLER=0 VLLM_USE_FLASHINFER=0 \
vllm serve huginnfork/Ornith-1.0-35B-FP8 --trust-remote-code --quantization compressed-tensors --max-model-len 8192
```
Validated on an RTX PRO 6000 (SM120) with vLLM 0.24.0: both the FP8 and NVFP4A16 quants load and generate coherently (e.g. "17 times 24 is 408", correct multi-step reasoning traces).
|