File size: 3,117 Bytes
29c8d08
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
---
base_model: mistralai/Voxtral-Mini-4B-Realtime-2602
language:
  - ar
  - de
  - en
  - es
  - fr
  - hi
  - it
  - nl
  - pt
  - zh
  - ja
  - ko
  - ru
library_name: mlx
license: apache-2.0
pipeline_tag: automatic-speech-recognition
tags:
  - mlx
  - mlx-audio
  - speech-to-text
  - streaming
  - realtime
  - voxtral
  - fp16
---

# Voxtral Mini 4B Realtime 4bit (float16)

This is a **4-bit quantized, float16-base** [MLX](https://github.com/ml-explore/mlx) conversion of [mistralai/Voxtral-Mini-4B-Realtime-2602](https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602).

## Which variant should you pick?

| Chip           | Recommended                                                                                            | Why                                                                                                                      |
|----------------|--------------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------|
| **M1 / M2**    | **This repo (`-4bit-fp16`)**                                                                           | Metal on M1/M2 has no native `bfloat16` ALU; bf16 ops fall back to a slower path. `float16` stays on the fast GPU path.  |
| **M3 / M4+**   | [`iris-sfg/Voxtral-Mini-4B-Realtime-2602-4bit`](https://huggingface.co/iris-sfg/Voxtral-Mini-4B-Realtime-2602-4bit) (bf16) | bf16 is natively supported and gives the same speed as fp16 with a wider dynamic range (slightly safer numerics).       |

Only the non-quantized weights differ between the two repos (norms, biases, scales, some embeddings). The quantized mat-mul weights are bit-identical. Transcription output is byte-identical on a 20 s French clip at temperature 0 (verified locally).

## Conversion

Source model:

- `mistralai/Voxtral-Mini-4B-Realtime-2602`

Local conversion command:

```bash
python -m mlx_audio.convert \
  --hf-path mistralai/Voxtral-Mini-4B-Realtime-2602 \
  --mlx-path /path/to/Voxtral-Mini-4B-Realtime-2602-4bit-fp16 \
  --quantize \
  --q-group-size 64 \
  --q-bits 4 \
  --dtype float16 \
  --model-domain stt
```

Quantization config:

- bits: `4`
- group size: `64`
- mode: `affine`
- non-quant dtype: `float16`

## Files

Only the MLX runtime artifacts needed for inference:

- `model.safetensors`
- `model.safetensors.index.json`
- `config.json`
- `generation_config.json`
- `params.json`
- `processor_config.json`
- `tekken.json`

## Usage

```bash
pip install "mlx-audio[stt]"
```

```python
from mlx_audio.stt.utils import load_model

model = load_model("iris-sfg/Voxtral-Mini-4B-Realtime-2602-4bit-fp16")
result = model.generate("audio.wav")
print(result.text)
```

## Notes

- Base model license remains Apache 2.0.
- On M3/M4, prefer the `-4bit` (bf16) repo; there is no speed benefit to fp16 there and bf16's wider exponent range is slightly more robust.
- Transcription quality was verified identical to the bf16 variant at `temperature=0` on a 20 s French parliamentary audio clip.