Hemmingway-1-FP8 / README.md
dbest's picture
Hemmingway-1 FP8 (tested)
ce550a0 verified
|
Raw History Blame
1.89 kB
---
license: cc-by-nc-4.0
base_model:
- darrellbest/Hemmingway-1-VL
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- fp8
- compressed-tensors
- vllm
- qwen3.8
- hemmingway
---
# Hemmingway-1 FP8
FP8 build of [Altworld/Hemmingway-1](https://huggingface.co/Altworld/Hemmingway-1) (Qwen3.8-27B fine-tune for
everyday writing) for **vLLM**, made from
[darrellbest/Hemmingway-1-VL](https://huggingface.co/darrellbest/Hemmingway-1-VL) (the same weights in the stock
Qwen3.8 layout, vision restored). 35 GB instead of 55.6 GB. All credit for the model goes to
[Altworld](https://hemmingway.io). An uncensored version is at
[darrellbest/Hemmingway-1-Heretic-FP8](https://huggingface.co/darrellbest/Hemmingway-1-Heretic-FP8).
## What is quantized
| Part | Precision |
|---|---|
| MLP linear layers, and the attention projections of the full-attention layers | **FP8 E4M3 weights with per-channel scales, dynamic per-token FP8 activations (W8A8)** |
| Gated DeltaNet (`linear_attn`) layers, vision tower, MTP head, embeddings, `lm_head`, norms | bf16, unchanged |
Made with [llm-compressor](https://github.com/vllm-project/llm-compressor) 0.13.0 (`scheme="FP8_DYNAMIC"`). The
multi-token-prediction weights, which the quantized save drops, were copied back unchanged.
## Tested (vLLM 0.30, RTX PRO 6000)
| | bf16 | **FP8** |
|---|---:|---:|
| Same top token as bf16 (held-out UltraChat, 24.6K positions) | 100% | **96.9%** |
| Perplexity, same text | 5.21 | 5.22 |
| MATH-500 levels 4-5, thinking mode (40) | 32/40 | **32/40** |
| `get_weather` tool call | correct | correct |
No measurable loss on reasoning or tool calling.
```sh
vllm serve darrellbest/Hemmingway-1-FP8
```
Licence: [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/), as Hemmingway-1: non-commercial use, with
credit to Hemmingway-1 / Altworld. Commercial use: luka@hemmingway.io.