--- license: cc-by-nc-4.0 base_model: - darrellbest/Hemmingway-1-VL base_model_relation: quantized pipeline_tag: image-text-to-text tags: - fp8 - compressed-tensors - vllm - qwen3.8 - hemmingway --- # Hemmingway-1 FP8 FP8 build of [Altworld/Hemmingway-1](https://huggingface.co/Altworld/Hemmingway-1) (Qwen3.8-27B fine-tune for everyday writing) for **vLLM**, made from [darrellbest/Hemmingway-1-VL](https://huggingface.co/darrellbest/Hemmingway-1-VL) (the same weights in the stock Qwen3.8 layout, vision restored). 35 GB instead of 55.6 GB. All credit for the model goes to [Altworld](https://hemmingway.io). An uncensored version is at [darrellbest/Hemmingway-1-Heretic-FP8](https://huggingface.co/darrellbest/Hemmingway-1-Heretic-FP8). ## What is quantized | Part | Precision | |---|---| | MLP linear layers, and the attention projections of the full-attention layers | **FP8 E4M3 weights with per-channel scales, dynamic per-token FP8 activations (W8A8)** | | Gated DeltaNet (`linear_attn`) layers, vision tower, MTP head, embeddings, `lm_head`, norms | bf16, unchanged | Made with [llm-compressor](https://github.com/vllm-project/llm-compressor) 0.13.0 (`scheme="FP8_DYNAMIC"`). The multi-token-prediction weights, which the quantized save drops, were copied back unchanged. ## Tested (vLLM 0.30, RTX PRO 6000) | | bf16 | **FP8** | |---|---:|---:| | Same top token as bf16 (held-out UltraChat, 24.6K positions) | 100% | **96.9%** | | Perplexity, same text | 5.21 | 5.22 | | MATH-500 levels 4-5, thinking mode (40) | 32/40 | **32/40** | | `get_weather` tool call | correct | correct | No measurable loss on reasoning or tool calling. ```sh vllm serve darrellbest/Hemmingway-1-FP8 ``` Licence: [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/), as Hemmingway-1: non-commercial use, with credit to Hemmingway-1 / Altworld. Commercial use: luka@hemmingway.io.