File size: 5,077 Bytes
b08f8ab
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
---
license: apache-2.0
base_model: amalia-llm/AMALIA-9B-0626-DPO
base_model_relation: quantized
language:
  - pt
library_name: transformers
tags:
  - portuguese
  - european-portuguese
  - fp8
  - vllm
  - compressed-tensors
  - quantized
pipeline_tag: text-generation
---

# AMALIA-9B-0626-DPO-FP8

FP8 quantization of [amalia-llm/AMALIA-9B-0626-DPO](https://huggingface.co/amalia-llm/AMALIA-9B-0626-DPO),
the European Portuguese language model, for serving with vLLM.

To the best of our knowledge this is the first FP8 quantization of AMALIA. Every other
quantization published on the Hub targets llama.cpp (GGUF) or Apple Silicon (MLX).

## Why FP8

The goal was a version that runs well on server GPUs **without measurable quality loss**.
FP8 was chosen over 4-bit formats for one reason that matters specifically for a
language-specialised model: **FP8 dynamic quantization requires no calibration data**.
Weight scales are computed per channel and activation scales at runtime.

4-bit schemes need a calibration pass, and the default calibration datasets are English.
Calibrating a European Portuguese model on English text would degrade exactly the property
that makes the model worth using.

## Results

Measured against the BF16 original on an NVIDIA DGX Spark (GB10 Grace-Blackwell, 121 GB
unified memory, sm_121a).

**Perplexity** over 401 tokens of European Portuguese across four registers. The colloquial
passage deliberately contains *comboio*, *pequeno-almoço*, *casa de banho* and *telemóvel*,
so that any degradation of specifically European vocabulary would show up.

| Register | BF16 | FP8 | Delta |
|---|---|---|---|
| Literary | 15.6962 | 16.0568 | +2.30% |
| Administrative | 4.7820 | 4.8075 | +0.53% |
| Colloquial | 9.8653 | 10.1136 | +2.52% |
| Technical | 18.1637 | 17.8324 | -1.82% |
| **Global** | **11.0987** | **11.2033** | **+0.94%** |

**Serving**

| | BF16 | FP8 |
|---|---|---|
| Size on disk | 18 GB | **9.6 GB** |
| Throughput (single stream, DGX Spark) | 12.2 tok/s | **22.3 tok/s** |
| KV cache at the same memory fraction | 140,912 tokens | **190,112 tokens** |
| Startup time | 270s | 180s |

**Language markers.** Across 12 generation prompts, both variants produced 7 European
Portuguese lexical markers and **zero** Brazilian ones, with the European gerund
construction (*estou a fazer*) and never the Brazilian one (*estou fazendo*).

## Caveat

Sub-1% perplexity difference does not mean identical answers. Greedy decoding amplifies
tiny logit differences: across 12 prompts, none of the FP8 answers matched the BF16 answer
token for token, even though the aggregate quality is equivalent. On one geography question
the FP8 answer placed Madeira near the Strait of Gibraltar, which is wrong, where the BF16
answer correctly placed it west of Morocco.

Statistical equivalence is not answer-by-answer equivalence. Use accordingly.

## Usage

```bash
vllm serve <this-repo> \
  --served-model-name amalia \
  --max-model-len 32768 \
  --host 0.0.0.0 --port 8000
```

vLLM detects `compressed-tensors` automatically. No extra flags are needed.

## How it was made

```python
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier

recipe = QuantizationModifier(
    targets="Linear",
    scheme="FP8_DYNAMIC",
    ignore=["lm_head"],
)
oneshot(model=model, recipe=recipe)
```

Total time: 209 seconds, of which **2 seconds** are the quantization itself. The rest is
reading the model from disk and writing it back. `llm-compressor` selects a `DataFreePipeline`
because the recipe needs no calibration data.

`lm_head` is kept in original precision.

### A note on weight-only quantization

We also produced a weight-only variant (W8A16, `input_activations=None`) expecting it to be
more accurate, since activation quantization is usually where the error lives. It was
**worse**: +2.31% perplexity against +0.94% for W8A8.

The configuration files show that weights are quantized identically in both variants, so
the entire difference comes from the compute path. W8A8 runs on the Blackwell native FP8
tensor cores with FP32 accumulation; W8A16 upconverts the weights back to BF16. On an
accelerator with native FP8 support, using that support turned out to be more accurate than
avoiding it.

## Em português

Esta é uma quantização FP8 do AMALIA-9B, o modelo de língua de português europeu, para ser
servida em vLLM. Foi escolhido FP8 em vez de 4 bits porque **não precisa de calibração**, e
calibrar um modelo de português com textos ingleses degradaria precisamente aquilo que o
distingue.

Resultado: metade do tamanho, quase o dobro do débito, e menos de 1% de degradação de
perplexidade medida sobre quatro registos de português europeu.

## Attribution

Base model: [amalia-llm/AMALIA-9B-0626-DPO](https://huggingface.co/amalia-llm/AMALIA-9B-0626-DPO),
Apache 2.0. All credit for the model itself belongs to the AMALIA team. This repository
contains only a quantization.

Quantization method: [llm-compressor](https://github.com/vllm-project/llm-compressor).