Hemmingway-1 FP8
FP8 build of Altworld/Hemmingway-1 (Qwen3.8-27B fine-tune for everyday writing) for vLLM, made from darrellbest/Hemmingway-1-VL (the same weights in the stock Qwen3.8 layout, vision restored). 35 GB instead of 55.6 GB. All credit for the model goes to Altworld. An uncensored version is at darrellbest/Hemmingway-1-Heretic-FP8.
What is quantized
| Part | Precision |
|---|---|
| MLP linear layers, and the attention projections of the full-attention layers | FP8 E4M3 weights with per-channel scales, dynamic per-token FP8 activations (W8A8) |
Gated DeltaNet (linear_attn) layers, vision tower, MTP head, embeddings, lm_head, norms |
bf16, unchanged |
Made with llm-compressor 0.13.0 (scheme="FP8_DYNAMIC"). The
multi-token-prediction weights, which the quantized save drops, were copied back unchanged.
Tested (vLLM 0.30, RTX PRO 6000)
| bf16 | FP8 | |
|---|---|---|
| Same top token as bf16 (held-out UltraChat, 24.6K positions) | 100% | 96.9% |
| Perplexity, same text | 5.21 | 5.22 |
| MATH-500 levels 4-5, thinking mode (40) | 32/40 | 32/40 |
get_weather tool call |
correct | correct |
No measurable loss on reasoning or tool calling.
vllm serve darrellbest/Hemmingway-1-FP8
Licence: CC BY-NC 4.0, as Hemmingway-1: non-commercial use, with credit to Hemmingway-1 / Altworld. Commercial use: luka@hemmingway.io.
- Downloads last month
- -