You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

voxtral-mini-4b-asr-specdec

The W4A16 audio-conditioned GPTQ decoder of voxtral-mini-4b-asr served with prompt-lookup ngram speculative decoding inside vLLM. The weights are unchanged; this artifact differs from its parent only in the speculative_config block of vllm_config.yaml. Rejection sampling preserves the output distribution, so served quality is identical to the verifier alone; the energy reduction comes from fewer verifier forward passes per generated token.

On the canonical four-language FLEURS evaluation on NVIDIA L4 24 GB this configuration draws 156.6 kJ — a 21.4 % reduction relative to the W4A16 verifier alone (199.3 kJ) and a 54.7 % reduction relative to the FP8 Round-1 floor (345.7 kJ).

Summary

Base model mistralai/Voxtral-Mini-4B-Realtime-2602
Decoder weights W4A16 GPTQ, group 128, dynamic actorder (audio-conditioned)
Audio encoder, projector BF16
KV cache FP8 e4m3
Speculative drafter ngram (prompt lookup), k = 1
Serving vLLM 0.19.1, TRITON_ATTN, PIECEWISE cudagraph
Artifact size 4.07 GB (same consolidated.safetensors as the parent)

Energy on NVIDIA L4 — canonical four-language

Slice FP8 Round-1 floor W4A16 alone W4A16 + ngram (this)
en_us 500 6.15 % · 189.4 kJ 5.58 % · 107.8 kJ 5.54 % · 85.7 kJ
hi_in 100 25.43 % · 44.5 kJ 24.09 % · 28.8 kJ 24.09 % · 23.0 kJ
fr_fr 100 8.45 % · 37.9 kJ 7.36 % · 21.7 kJ 7.43 % · 17.7 kJ
ja_jp 100 7.09 % CER · 73.9 kJ 7.41 % · 41.0 kJ 6.77 % · 30.2 kJ
Total 345.7 kJ 199.3 kJ 156.6 kJ
vs FP8 floor baseline −42.4 % −54.7 %
vs W4A16 alone — baseline −21.4 %

WER is normalized; CER is reported for ja, cmn. Energy is measured with CodeCarbon wrapping the evaluator on NVIDIA L4 24 GB (driver 565.57, CUDA 12.7).

Quality vs. BF16 ceiling

Acceptance criterion: normalized WER (or CER) ≤ 1.25 × the BF16 baseline on the same slice.

Slice Metric BF16 Ceiling Measured Verdict
en_us 500 WER 6.05 7.56 5.56 beats BF16
fr_fr 100 WER 8.24 10.30 7.43 beats BF16
hi_in 100 WER 26.27 32.84 24.01 beats BF16
ja_jp 100 CER 6.72 8.39 6.74 within ceiling
es_419 100 WER 2.85 3.56 2.73 beats BF16
it_it 100 WER 3.82 4.77 3.97 within ceiling
ru_ru 100 WER 5.44 6.80 5.70 within ceiling
pt_br 100 WER 5.05 6.31 5.79 within ceiling
de_de 100 WER 5.10 6.37 4.89 beats BF16
nl_nl 100 WER 8.84 11.05 8.36 beats BF16
ar_eg 100 WER 15.01 18.76 14.16 beats BF16
ko_kr 100 WER 15.95 19.94 16.16 within ceiling
cmn_hans_cn 100 CER 9.28 11.60 9.31 within ceiling

All thirteen languages satisfy the ceiling and seven beat BF16 outright. The thirteen-language L4 energy total is 339.04 kJ, against 434.6 kJ for the W4A16 verifier alone — a 21.99 % reduction that holds roughly uniformly across Latin, Indic, CJK, Semitic, and Slavic families. Empty-prediction count across 1700 evaluated samples: one, on hi_in id=1985, a near-silent FLEURS duplicate that the unmodified BF16 reference also empties on. The challenge evaluation harness filters dataset duplicates, so this row does not enter the scored result.

Reproducibility

Provisioning a clean NVIDIA L4 pod, pulling this repository with hf download, installing the pinned stack, and running bash reproduce.sh end-to-end on the four canonical slices reproduces the four-language total at 161.60 kJ, +3.19 % against the binding 156.60 kJ — within the ±5 % acceptance window.

The pins in reproduce.sh and vllm_config.yaml address three dependency-drift issues:

  1. transformers >= 5 calls feature_extractor.fetch_audio(), which vLLM 0.19.1's MistralCommonFeatureExtractor does not implement; pinned to transformers<5.
  2. mistral_common 1.11.2 altered the audio API in a way that breaks the vLLM voxtral processor path; pinned to mistral_common==1.11.0.
  3. max_model_len: 32768 left insufficient KV cache room on a 24 GB L4 (18.0 GiB needed, 10.2 GiB available); pinned to 16384.

How speculative decoding helps here

The verifier is autoregressive: each token costs one forward pass. Prompt-lookup ngram speculation maintains a hash table of n-grams observed earlier in the same generation; when the recent context matches an indexed prefix, the next k = 1 token is proposed from the lookup and verified together with the verifier's own prediction in a single batched forward pass. Accepted tokens skip a verifier forward; rejected tokens fall back. Rejection sampling preserves the output distribution exactly.

For FLEURS transcription this catches function-word bigrams across utterances, padding tokens from the streaming head, and CJK sentence-closing punctuation. Empirical acceptance is 0.4–0.5 tokens per draft call, which compounds into a 1.2–1.5× wall-clock speedup and a proportional energy reduction.

Higher speculative depths (k = 2, 3, 4) crash vLLM 0.19.1's /v1/audio/transcriptions path with a tensor-shape mismatch in the audio-token-prefix copy (an upstream limitation). k = 1 is the deepest stable configuration on this stack.

Reproducing

git clone https://huggingface.co/Shankara-A-S/voxtral-mini-4b-asr-specdec
cd voxtral-mini-4b-asr-specdec
bash reproduce.sh

reproduce.sh installs the pinned stack, starts vLLM with vllm_config.yaml, warms with one dummy request, and runs the four canonical FLEURS slices under the locked audio preprocessing chain (LUFS normalization to −23 LUFS with 24 dB ceiling; WebRTC VAD edge-trim at aggressiveness 1, 200 ms padding; internal-silence gating compressed to 160 ms above a 320 ms run). Per-slice JSON lands under reports/.

Environment overrides: SKIP_INSTALL=1, BASE_PORT=8084, RUN_SLICES="en_us:20", REPORTS_DIR=reports.

Pinned stack

torch              == 2.10.0  (cu128, pulled in by vllm 0.19.1)
vllm               == 0.19.1
compressed-tensors == 0.15.0.1
transformers       <  5      (4.57.x at time of validation)
mistral_common     == 1.11.0
codecarbon, jiwer, librosa, soundfile, webrtcvad, pyloudnorm

NVIDIA driver ≥ 565 (CUDA 12.7 forward-compatible). Validated on NVIDIA L4 24 GB, driver 565.57, Python 3.11.10.

vllm_config.yaml

Infrastructure parameters omitted per Round-2 guidance; non-default flags documented.

Argument Value Reason
quantization compressed-tensors Load W4A16 GPTQ packed weights.
kv_cache_dtype fp8_e4m3 FP8 KV cache.
attention_backend TRITON_ATTN Required by Voxtral's Whisper-style encoder.
tokenizer_mode mistral mistral-common tokenization.
max_model_len 16384 Fits the L4 KV cache; well above any FLEURS clip.
max_num_seqs 1 Serial transcription.
max_num_batched_tokens 16384 Tied to max_model_len.
enable_prefix_caching true No-op on the audio endpoint; retained.
compilation_config.cudagraph_mode PIECEWISE Required — full CUDAGraphs crash on the streaming head.
speculative_config.method ngram Prompt-lookup drafter.
speculative_config.num_speculative_tokens 1 Maximum stable depth on vLLM 0.19.1's audio path.
speculative_config.prompt_lookup_max 2 Maximum lookup prefix length.
speculative_config.prompt_lookup_min 1 Minimum lookup prefix length.

EAGLE-3 drafter

A 1-layer Llama-architecture EAGLE-3 draft head with a 32 k draft vocabulary was trained on audio-conditioned trajectories across all thirteen languages. Offline acceptance: 53.84 % top-1, 69.69 % top-5. On the L4 binding it draws 183.3 kJ — 17 % higher than ngram in this artifact — and degrades served quality on multiple languages, which indicates that vLLM 0.19.1's EAGLE + Voxtral integration breaks the rejection-sampling guarantee in practice. The trained draft head remains available at voxtral-mini-4b-asr-eagle3-specdec for follow-up work; ngram is the production speculative configuration.

Known issues

  • vLLM 0.19.1's audio path requires cudagraph_mode: PIECEWISE.
  • num_speculative_tokens >= 2 crashes the audio endpoint (upstream vLLM bug); k = 1 is fixed for this artifact.
  • /v1/realtime streaming has not been validated with speculative configuration — use /v1/audio/transcriptions.

License

Apache-2.0, inherited from the base model.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shankara-A-S/voxtral-mini-4b-asr-specdec