You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

voxtral-mini-4b-asr

A 4-bit weight-quantized derivative of mistralai/Voxtral-Mini-4B-Realtime-2602 for low-energy multilingual ASR on NVIDIA data-center GPUs. Decoder projections are quantized to W4A16 (GPTQ, group size 128, dynamic activation order) using audio-conditioned calibration β€” the calibration corpus is real projected audio embeddings, not text tokens. The audio encoder, projector, embeddings, and normalization layers remain in BF16; the KV cache is FP8 e4m3. Inference is through vLLM with PIECEWISE cudagraph and Triton attention.

On the canonical four-language FLEURS evaluation on NVIDIA L4 24 GB this artifact draws 199.3 kJ, a 42.4 % reduction relative to the FP8 dynamic-quantization runtime used as the Round-1 floor (345.7 kJ on the same hardware). Disk footprint is 4.07 GB. Quality satisfies the 1.25 Γ— BF16 ceiling on all thirteen Voxtral-supported FLEURS languages and beats BF16 outright on nine.

Summary

Base model mistralai/Voxtral-Mini-4B-Realtime-2602
Decoder quantization W4A16 GPTQ, group 128, dynamic actorder
Calibration 256 projected inputs_embeds, FLEURS train, 13 languages
Audio encoder, projector BF16
KV cache FP8 e4m3
Serving vLLM 0.19.1, TRITON_ATTN, PIECEWISE cudagraph
Artifact size 4.07 GB (consolidated.safetensors)

The speculative-decoding companion of this artifact is voxtral-mini-4b-asr-specdec (ngram k=1, 156.6 kJ on the same four slices).

Quality vs. BF16 ceiling

Acceptance criterion: normalized WER (or CER for ja, cmn) ≀ 1.25 Γ— the BF16 baseline on the same slice. WER and CER are hardware-independent.

Slice Metric BF16 Ceiling Measured Verdict
en_us 500 WER 6.05 7.56 5.58 beats BF16
fr_fr 100 WER 8.24 10.30 7.36 beats BF16
hi_in 100 WER 26.27 32.84 24.09 beats BF16
ja_jp 100 CER 6.72 8.39 7.41 within ceiling
es_419 100 WER 2.85 3.56 2.69 beats BF16
it_it 100 WER 3.82 4.77 3.93 within ceiling
ru_ru 100 WER 5.44 6.80 5.59 within ceiling
pt_br 100 WER 5.05 6.31 5.76 within ceiling
de_de 100 WER 5.10 6.37 4.89 beats BF16
nl_nl 100 WER 8.84 11.05 8.49 beats BF16
ar_eg 100 WER 15.01 18.76 14.01 beats BF16
ko_kr 100 WER 15.95 19.94 15.95 matches BF16
cmn_hans_cn 100 CER 9.28 11.60 9.19 beats BF16

All thirteen languages satisfy the ceiling and nine beat BF16 outright. BF16 baselines are measured on the unmodified Voxtral-Mini-4B-Realtime-2602 released by Mistral. WER and CER are hardware-independent, so the ceiling derived on the BF16 workstation applies on L4. Empty-prediction count across all 1700 evaluated samples: one, on hi_in id=1985, a near-silent FLEURS duplicate that the unmodified BF16 reference also empties on. The challenge evaluation harness filters dataset duplicates, so this row does not enter the scored result.

Energy on NVIDIA L4 24 GB

Slice This artifact FP8 Round-1 floor Ξ”
en_us 500 107.8 kJ 189.4 kJ βˆ’43.1 %
hi_in 100 28.8 kJ 44.5 kJ βˆ’35.3 %
fr_fr 100 21.7 kJ 37.9 kJ βˆ’42.7 %
ja_jp 100 41.0 kJ 73.9 kJ βˆ’44.5 %
Total 199.3 kJ 345.7 kJ βˆ’42.4 %

Thirteen-language L4 total: 434.6 kJ (per-slice JSON in reports/). Energy is measured with CodeCarbon wrapping the evaluator (driver 565.57, CUDA 12.7).

Reproducibility

Provisioning a clean NVIDIA L4 pod, pulling this repository with hf download, installing the pinned stack, and running bash reproduce.sh end-to-end on the ngram-speculative companion (which shares the same consolidated.safetensors and the same reproduce.sh flow) reproduces the four-language total at 161.60 kJ, +3.19 % against the binding 156.60 kJ β€” within the Β±5 % acceptance window.

The pins in reproduce.sh and vllm_config.yaml address three dependency-drift issues:

  1. transformers >= 5 calls feature_extractor.fetch_audio(), which vLLM 0.19.1's MistralCommonFeatureExtractor does not implement; pinned to transformers<5.
  2. mistral_common 1.11.2 altered the audio API in a way that breaks the vLLM voxtral processor path; pinned to mistral_common==1.11.0.
  3. max_model_len: 32768 left insufficient KV cache room on a 24 GB L4 (18.0 GiB needed, 10.2 GiB available); pinned to 16384.

Reproducing

git clone https://huggingface.co/Shankara-A-S/voxtral-mini-4b-asr
cd voxtral-mini-4b-asr
bash reproduce.sh

reproduce.sh installs the pinned stack, starts vLLM with vllm_config.yaml, warms with one dummy request, and runs the four canonical FLEURS slices under the locked audio preprocessing (LUFS normalization to βˆ’23 LUFS with 24 dB ceiling; WebRTC VAD edge-trim at aggressiveness 1, 200 ms padding; internal-silence gating compressed to 160 ms above a 320 ms run). Per-slice evaluation and energy JSON land under reports/.

Environment overrides: SKIP_INSTALL=1, BASE_PORT=8084, RUN_SLICES="en_us:20 hi_in:50", REPORTS_DIR=reports.

Pinned stack

torch              == 2.10.0  (cu128, pulled in by vllm 0.19.1)
vllm               == 0.19.1
compressed-tensors == 0.15.0.1
transformers       <  5      (4.57.x at time of validation)
mistral_common     == 1.11.0
codecarbon, jiwer, librosa, soundfile, webrtcvad, pyloudnorm

NVIDIA driver β‰₯ 565 (CUDA 12.7 forward-compatible). Validated on NVIDIA L4 24 GB, driver 565.57, Python 3.11.10.

vllm_config.yaml

Following the organizer's Round-2 guidance, infrastructure parameters are omitted and non-default flags are documented.

Argument Value Reason
quantization compressed-tensors Load W4A16 GPTQ packed weights.
kv_cache_dtype fp8_e4m3 FP8 KV cache.
attention_backend TRITON_ATTN Required by Voxtral's Whisper-style encoder; FlashInfer is not implemented for Whisper-causal block pooling.
tokenizer_mode mistral mistral-common tokenization of the Mistral-arch decoder.
max_model_len 16384 Fits the L4 KV cache; well above any FLEURS clip.
max_num_seqs 1 Serial transcription workload.
max_num_batched_tokens 16384 Tied to max_model_len.
enable_prefix_caching true No-op on /v1/audio/transcriptions (no cache_salt); retained for forward compatibility.
compilation_config.cudagraph_mode PIECEWISE Required β€” full CUDAGraphs are incompatible with the streaming head.

Construction (brief)

The decoder is the inference-time bottleneck; the audio encoder and projector are small relative to the language model and remain in BF16. The calibration corpus is 256 projected inputs_embeds tensors drawn from FLEURS train across thirteen languages, with Hindi over-represented at 61 samples. Calibration uses llm-compressor 0.10.0.1 with the GPTQModifier and the recipe below; a custom collator loads the pre-projected embeddings directly into the decoder, bypassing the audio encoder during calibration.

Scheme W4A16
Targets language_model.model.layers.X.{self_attn,mlp}.*
Ignored embedding, lm_head, layernorms, ada_rms_norm, audio_tower, projector
Group size 128
Dampening 0.05
Actorder dynamic
Samples 256
Max seq 2048
Order one decoder layer at a time

The HF-format output is then merged into the native Mistral consolidated layout for vLLM. Scripts: scripts/build_track_b_audio_conditioned_calibration.py, scripts/run_track_b_llmcompressor_oneshot.py, scripts/package_track_b_consolidated.py.

Notes

  • BF16 baseline reconciliation: the hi_in BF16 of 26.27 % WER uses jiwer-strict normalization (NFKC, casefold, strip punctuation, collapse whitespace). Under Open-ASR-style normalization the same outputs give 17.28 % (CI 13.65–21.29), which agrees with SuperJerem/voxtral-coalition's 17.74 %. The stricter normalizer is used here so the 1.25 Γ— ceiling is harder to meet; the W4A16-vs-BF16 ratios hold under either standard.
  • The closest public 4-bit Voxtral comparator is the GGUF Q4_0 release by andrijdavid / freddm, which reports an English FLEURS WER of 8.49 % against the BF16 reference of 4.90 % β€” a 73 % relative regression. Existing public 4-bit Voxtral artifacts target hardware classes other than NVIDIA data-center GPUs and use text-token calibration. This is, to current knowledge, the first published audio-conditioned 4-bit quantization of the base model.
  • Verbosity ratio (hypothesis characters / reference characters) lies in [0.987, 1.027] across the thirteen FLEURS languages β€” the artifact does not over-produce tokens.

Known issues

  • vLLM 0.19.1's audio path requires cudagraph_mode: PIECEWISE; full CUDAGraphs crash on Voxtral's streaming head.
  • The /v1/realtime streaming endpoint has not been validated against this artifact. Use /v1/audio/transcriptions for the reported numbers.

License

Apache-2.0, inherited from the base model.

Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Shankara-A-S/voxtral-mini-4b-asr