Hemmingway-1-FP8 / README.md
dbest's picture
Hemmingway-1 FP8 (tested)
ce550a0 verified
|
Raw History Blame
1.89 kB
metadata
license: cc-by-nc-4.0
base_model:
  - darrellbest/Hemmingway-1-VL
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
  - fp8
  - compressed-tensors
  - vllm
  - qwen3.8
  - hemmingway

Hemmingway-1 FP8

FP8 build of Altworld/Hemmingway-1 (Qwen3.8-27B fine-tune for everyday writing) for vLLM, made from darrellbest/Hemmingway-1-VL (the same weights in the stock Qwen3.8 layout, vision restored). 35 GB instead of 55.6 GB. All credit for the model goes to Altworld. An uncensored version is at darrellbest/Hemmingway-1-Heretic-FP8.

What is quantized

Part Precision
MLP linear layers, and the attention projections of the full-attention layers FP8 E4M3 weights with per-channel scales, dynamic per-token FP8 activations (W8A8)
Gated DeltaNet (linear_attn) layers, vision tower, MTP head, embeddings, lm_head, norms bf16, unchanged

Made with llm-compressor 0.13.0 (scheme="FP8_DYNAMIC"). The multi-token-prediction weights, which the quantized save drops, were copied back unchanged.

Tested (vLLM 0.30, RTX PRO 6000)

bf16 FP8
Same top token as bf16 (held-out UltraChat, 24.6K positions) 100% 96.9%
Perplexity, same text 5.21 5.22
MATH-500 levels 4-5, thinking mode (40) 32/40 32/40
get_weather tool call correct correct

No measurable loss on reasoning or tool calling.

vllm serve darrellbest/Hemmingway-1-FP8

Licence: CC BY-NC 4.0, as Hemmingway-1: non-commercial use, with credit to Hemmingway-1 / Altworld. Commercial use: luka@hemmingway.io.