voxtral-mini-4b-asr
A 4-bit weight-quantized derivative of mistralai/Voxtral-Mini-4B-Realtime-2602
for low-energy multilingual ASR on NVIDIA data-center GPUs. Decoder
projections are quantized to W4A16 (GPTQ, group size 128, dynamic activation
order) using audio-conditioned calibration β the calibration corpus is real
projected audio embeddings, not text tokens. The audio encoder, projector,
embeddings, and normalization layers remain in BF16; the KV cache is FP8
e4m3. Inference is through vLLM with PIECEWISE cudagraph and Triton
attention.
On the canonical four-language FLEURS evaluation on NVIDIA L4 24 GB this artifact draws 199.3 kJ, a 42.4 % reduction relative to the FP8 dynamic-quantization runtime used as the Round-1 floor (345.7 kJ on the same hardware). Disk footprint is 4.07 GB. Quality satisfies the 1.25 Γ BF16 ceiling on all thirteen Voxtral-supported FLEURS languages and beats BF16 outright on nine.
Summary
| Base model | mistralai/Voxtral-Mini-4B-Realtime-2602 |
| Decoder quantization | W4A16 GPTQ, group 128, dynamic actorder |
| Calibration | 256 projected inputs_embeds, FLEURS train, 13 languages |
| Audio encoder, projector | BF16 |
| KV cache | FP8 e4m3 |
| Serving | vLLM 0.19.1, TRITON_ATTN, PIECEWISE cudagraph |
| Artifact size | 4.07 GB (consolidated.safetensors) |
The speculative-decoding companion of this artifact is
voxtral-mini-4b-asr-specdec
(ngram k=1, 156.6 kJ on the same four slices).
Quality vs. BF16 ceiling
Acceptance criterion: normalized WER (or CER for ja, cmn) β€ 1.25 Γ the BF16 baseline on the same slice. WER and CER are hardware-independent.
| Slice | Metric | BF16 | Ceiling | Measured | Verdict |
|---|---|---|---|---|---|
| en_us 500 | WER | 6.05 | 7.56 | 5.58 | beats BF16 |
| fr_fr 100 | WER | 8.24 | 10.30 | 7.36 | beats BF16 |
| hi_in 100 | WER | 26.27 | 32.84 | 24.09 | beats BF16 |
| ja_jp 100 | CER | 6.72 | 8.39 | 7.41 | within ceiling |
| es_419 100 | WER | 2.85 | 3.56 | 2.69 | beats BF16 |
| it_it 100 | WER | 3.82 | 4.77 | 3.93 | within ceiling |
| ru_ru 100 | WER | 5.44 | 6.80 | 5.59 | within ceiling |
| pt_br 100 | WER | 5.05 | 6.31 | 5.76 | within ceiling |
| de_de 100 | WER | 5.10 | 6.37 | 4.89 | beats BF16 |
| nl_nl 100 | WER | 8.84 | 11.05 | 8.49 | beats BF16 |
| ar_eg 100 | WER | 15.01 | 18.76 | 14.01 | beats BF16 |
| ko_kr 100 | WER | 15.95 | 19.94 | 15.95 | matches BF16 |
| cmn_hans_cn 100 | CER | 9.28 | 11.60 | 9.19 | beats BF16 |
All thirteen languages satisfy the ceiling and nine beat BF16 outright.
BF16 baselines are measured on the unmodified
Voxtral-Mini-4B-Realtime-2602 released by Mistral. WER and CER are
hardware-independent, so the ceiling derived on the BF16 workstation
applies on L4. Empty-prediction count across all 1700 evaluated
samples: one, on hi_in id=1985, a near-silent FLEURS duplicate that
the unmodified BF16 reference also empties on. The challenge
evaluation harness filters dataset duplicates, so this row does not
enter the scored result.
Energy on NVIDIA L4 24 GB
| Slice | This artifact | FP8 Round-1 floor | Ξ |
|---|---|---|---|
| en_us 500 | 107.8 kJ | 189.4 kJ | β43.1 % |
| hi_in 100 | 28.8 kJ | 44.5 kJ | β35.3 % |
| fr_fr 100 | 21.7 kJ | 37.9 kJ | β42.7 % |
| ja_jp 100 | 41.0 kJ | 73.9 kJ | β44.5 % |
| Total | 199.3 kJ | 345.7 kJ | β42.4 % |
Thirteen-language L4 total: 434.6 kJ (per-slice JSON in reports/).
Energy is measured with CodeCarbon wrapping the evaluator
(driver 565.57, CUDA 12.7).
Reproducibility
Provisioning a clean NVIDIA L4 pod, pulling this repository with
hf download, installing the pinned stack, and running
bash reproduce.sh end-to-end on the ngram-speculative companion
(which shares the same consolidated.safetensors and the same
reproduce.sh flow) reproduces the four-language total at
161.60 kJ, +3.19 % against the binding 156.60 kJ β within the
Β±5 % acceptance window.
The pins in reproduce.sh and vllm_config.yaml address three
dependency-drift issues:
transformers >= 5callsfeature_extractor.fetch_audio(), which vLLM 0.19.1'sMistralCommonFeatureExtractordoes not implement; pinned totransformers<5.mistral_common 1.11.2altered the audio API in a way that breaks the vLLM voxtral processor path; pinned tomistral_common==1.11.0.max_model_len: 32768left insufficient KV cache room on a 24 GB L4 (18.0 GiB needed, 10.2 GiB available); pinned to16384.
Reproducing
git clone https://huggingface.co/Shankara-A-S/voxtral-mini-4b-asr
cd voxtral-mini-4b-asr
bash reproduce.sh
reproduce.sh installs the pinned stack, starts vLLM with
vllm_config.yaml, warms with one dummy request, and runs the four
canonical FLEURS slices under the locked audio preprocessing (LUFS
normalization to β23 LUFS with 24 dB ceiling; WebRTC VAD edge-trim at
aggressiveness 1, 200 ms padding; internal-silence gating compressed
to 160 ms above a 320 ms run). Per-slice evaluation and energy JSON
land under reports/.
Environment overrides: SKIP_INSTALL=1, BASE_PORT=8084,
RUN_SLICES="en_us:20 hi_in:50", REPORTS_DIR=reports.
Pinned stack
torch == 2.10.0 (cu128, pulled in by vllm 0.19.1)
vllm == 0.19.1
compressed-tensors == 0.15.0.1
transformers < 5 (4.57.x at time of validation)
mistral_common == 1.11.0
codecarbon, jiwer, librosa, soundfile, webrtcvad, pyloudnorm
NVIDIA driver β₯ 565 (CUDA 12.7 forward-compatible). Validated on NVIDIA L4 24 GB, driver 565.57, Python 3.11.10.
vllm_config.yaml
Following the organizer's Round-2 guidance, infrastructure parameters are omitted and non-default flags are documented.
| Argument | Value | Reason |
|---|---|---|
quantization |
compressed-tensors |
Load W4A16 GPTQ packed weights. |
kv_cache_dtype |
fp8_e4m3 |
FP8 KV cache. |
attention_backend |
TRITON_ATTN |
Required by Voxtral's Whisper-style encoder; FlashInfer is not implemented for Whisper-causal block pooling. |
tokenizer_mode |
mistral |
mistral-common tokenization of the Mistral-arch decoder. |
max_model_len |
16384 |
Fits the L4 KV cache; well above any FLEURS clip. |
max_num_seqs |
1 |
Serial transcription workload. |
max_num_batched_tokens |
16384 |
Tied to max_model_len. |
enable_prefix_caching |
true |
No-op on /v1/audio/transcriptions (no cache_salt); retained for forward compatibility. |
compilation_config.cudagraph_mode |
PIECEWISE |
Required β full CUDAGraphs are incompatible with the streaming head. |
Construction (brief)
The decoder is the inference-time bottleneck; the audio encoder and
projector are small relative to the language model and remain in BF16.
The calibration corpus is 256 projected inputs_embeds tensors drawn
from FLEURS train across thirteen languages, with Hindi over-represented
at 61 samples. Calibration uses llm-compressor 0.10.0.1 with the
GPTQModifier and the recipe below; a custom collator loads the
pre-projected embeddings directly into the decoder, bypassing the audio
encoder during calibration.
| Scheme | W4A16 |
| Targets | language_model.model.layers.X.{self_attn,mlp}.* |
| Ignored | embedding, lm_head, layernorms, ada_rms_norm, audio_tower, projector |
| Group size | 128 |
| Dampening | 0.05 |
| Actorder | dynamic |
| Samples | 256 |
| Max seq | 2048 |
| Order | one decoder layer at a time |
The HF-format output is then merged into the native Mistral
consolidated layout for vLLM. Scripts:
scripts/build_track_b_audio_conditioned_calibration.py,
scripts/run_track_b_llmcompressor_oneshot.py,
scripts/package_track_b_consolidated.py.
Notes
- BF16 baseline reconciliation: the
hi_inBF16 of 26.27 % WER uses jiwer-strict normalization (NFKC, casefold, strip punctuation, collapse whitespace). Under Open-ASR-style normalization the same outputs give 17.28 % (CI 13.65β21.29), which agrees withSuperJerem/voxtral-coalition's 17.74 %. The stricter normalizer is used here so the 1.25 Γ ceiling is harder to meet; the W4A16-vs-BF16 ratios hold under either standard. - The closest public 4-bit Voxtral comparator is the GGUF Q4_0 release
by
andrijdavid/freddm, which reports an English FLEURS WER of 8.49 % against the BF16 reference of 4.90 % β a 73 % relative regression. Existing public 4-bit Voxtral artifacts target hardware classes other than NVIDIA data-center GPUs and use text-token calibration. This is, to current knowledge, the first published audio-conditioned 4-bit quantization of the base model. - Verbosity ratio (hypothesis characters / reference characters) lies in [0.987, 1.027] across the thirteen FLEURS languages β the artifact does not over-produce tokens.
Known issues
- vLLM 0.19.1's audio path requires
cudagraph_mode: PIECEWISE; full CUDAGraphs crash on Voxtral's streaming head. - The
/v1/realtimestreaming endpoint has not been validated against this artifact. Use/v1/audio/transcriptionsfor the reported numbers.
License
Apache-2.0, inherited from the base model.
- Downloads last month
- 4
Model tree for Shankara-A-S/voxtral-mini-4b-asr
Base model
mistralai/Ministral-3-3B-Base-2512