voxtral-mini-4b-asr-specdec
The W4A16 audio-conditioned GPTQ decoder of
voxtral-mini-4b-asr
served with prompt-lookup ngram speculative decoding inside vLLM. The
weights are unchanged; this artifact differs from its parent only in the
speculative_config block of vllm_config.yaml. Rejection sampling
preserves the output distribution, so served quality is identical to the
verifier alone; the energy reduction comes from fewer verifier forward
passes per generated token.
On the canonical four-language FLEURS evaluation on NVIDIA L4 24 GB this configuration draws 156.6 kJ — a 21.4 % reduction relative to the W4A16 verifier alone (199.3 kJ) and a 54.7 % reduction relative to the FP8 Round-1 floor (345.7 kJ).
Summary
| Base model | mistralai/Voxtral-Mini-4B-Realtime-2602 |
| Decoder weights | W4A16 GPTQ, group 128, dynamic actorder (audio-conditioned) |
| Audio encoder, projector | BF16 |
| KV cache | FP8 e4m3 |
| Speculative drafter | ngram (prompt lookup), k = 1 |
| Serving | vLLM 0.19.1, TRITON_ATTN, PIECEWISE cudagraph |
| Artifact size | 4.07 GB (same consolidated.safetensors as the parent) |
Energy on NVIDIA L4 — canonical four-language
| Slice | FP8 Round-1 floor | W4A16 alone | W4A16 + ngram (this) |
|---|---|---|---|
| en_us 500 | 6.15 % · 189.4 kJ | 5.58 % · 107.8 kJ | 5.54 % · 85.7 kJ |
| hi_in 100 | 25.43 % · 44.5 kJ | 24.09 % · 28.8 kJ | 24.09 % · 23.0 kJ |
| fr_fr 100 | 8.45 % · 37.9 kJ | 7.36 % · 21.7 kJ | 7.43 % · 17.7 kJ |
| ja_jp 100 | 7.09 % CER · 73.9 kJ | 7.41 % · 41.0 kJ | 6.77 % · 30.2 kJ |
| Total | 345.7 kJ | 199.3 kJ | 156.6 kJ |
| vs FP8 floor | baseline | −42.4 % | −54.7 % |
| vs W4A16 alone | — | baseline | −21.4 % |
WER is normalized; CER is reported for ja, cmn. Energy is measured with CodeCarbon wrapping the evaluator on NVIDIA L4 24 GB (driver 565.57, CUDA 12.7).
Quality vs. BF16 ceiling
Acceptance criterion: normalized WER (or CER) ≤ 1.25 × the BF16 baseline on the same slice.
| Slice | Metric | BF16 | Ceiling | Measured | Verdict |
|---|---|---|---|---|---|
| en_us 500 | WER | 6.05 | 7.56 | 5.56 | beats BF16 |
| fr_fr 100 | WER | 8.24 | 10.30 | 7.43 | beats BF16 |
| hi_in 100 | WER | 26.27 | 32.84 | 24.01 | beats BF16 |
| ja_jp 100 | CER | 6.72 | 8.39 | 6.74 | within ceiling |
| es_419 100 | WER | 2.85 | 3.56 | 2.73 | beats BF16 |
| it_it 100 | WER | 3.82 | 4.77 | 3.97 | within ceiling |
| ru_ru 100 | WER | 5.44 | 6.80 | 5.70 | within ceiling |
| pt_br 100 | WER | 5.05 | 6.31 | 5.79 | within ceiling |
| de_de 100 | WER | 5.10 | 6.37 | 4.89 | beats BF16 |
| nl_nl 100 | WER | 8.84 | 11.05 | 8.36 | beats BF16 |
| ar_eg 100 | WER | 15.01 | 18.76 | 14.16 | beats BF16 |
| ko_kr 100 | WER | 15.95 | 19.94 | 16.16 | within ceiling |
| cmn_hans_cn 100 | CER | 9.28 | 11.60 | 9.31 | within ceiling |
All thirteen languages satisfy the ceiling and seven beat BF16
outright. The thirteen-language L4 energy total is 339.04 kJ, against
434.6 kJ for the W4A16 verifier alone — a 21.99 % reduction that holds
roughly uniformly across Latin, Indic, CJK, Semitic, and Slavic
families. Empty-prediction count across 1700 evaluated samples: one,
on hi_in id=1985, a near-silent FLEURS duplicate that the unmodified
BF16 reference also empties on. The challenge evaluation harness
filters dataset duplicates, so this row does not enter the scored
result.
Reproducibility
Provisioning a clean NVIDIA L4 pod, pulling this repository with
hf download, installing the pinned stack, and running
bash reproduce.sh end-to-end on the four canonical slices reproduces
the four-language total at 161.60 kJ, +3.19 % against the binding
156.60 kJ — within the ±5 % acceptance window.
The pins in reproduce.sh and vllm_config.yaml address three
dependency-drift issues:
transformers >= 5callsfeature_extractor.fetch_audio(), which vLLM 0.19.1'sMistralCommonFeatureExtractordoes not implement; pinned totransformers<5.mistral_common 1.11.2altered the audio API in a way that breaks the vLLM voxtral processor path; pinned tomistral_common==1.11.0.max_model_len: 32768left insufficient KV cache room on a 24 GB L4 (18.0 GiB needed, 10.2 GiB available); pinned to16384.
How speculative decoding helps here
The verifier is autoregressive: each token costs one forward pass.
Prompt-lookup ngram speculation maintains a hash table of n-grams
observed earlier in the same generation; when the recent context matches
an indexed prefix, the next k = 1 token is proposed from the lookup
and verified together with the verifier's own prediction in a single
batched forward pass. Accepted tokens skip a verifier forward; rejected
tokens fall back. Rejection sampling preserves the output distribution
exactly.
For FLEURS transcription this catches function-word bigrams across utterances, padding tokens from the streaming head, and CJK sentence-closing punctuation. Empirical acceptance is 0.4–0.5 tokens per draft call, which compounds into a 1.2–1.5× wall-clock speedup and a proportional energy reduction.
Higher speculative depths (k = 2, 3, 4) crash vLLM 0.19.1's
/v1/audio/transcriptions path with a tensor-shape mismatch in the
audio-token-prefix copy (an upstream limitation). k = 1 is the
deepest stable configuration on this stack.
Reproducing
git clone https://huggingface.co/Shankara-A-S/voxtral-mini-4b-asr-specdec
cd voxtral-mini-4b-asr-specdec
bash reproduce.sh
reproduce.sh installs the pinned stack, starts vLLM with
vllm_config.yaml, warms with one dummy request, and runs the four
canonical FLEURS slices under the locked audio preprocessing chain
(LUFS normalization to −23 LUFS with 24 dB ceiling; WebRTC VAD
edge-trim at aggressiveness 1, 200 ms padding; internal-silence gating
compressed to 160 ms above a 320 ms run). Per-slice JSON lands under
reports/.
Environment overrides: SKIP_INSTALL=1, BASE_PORT=8084,
RUN_SLICES="en_us:20", REPORTS_DIR=reports.
Pinned stack
torch == 2.10.0 (cu128, pulled in by vllm 0.19.1)
vllm == 0.19.1
compressed-tensors == 0.15.0.1
transformers < 5 (4.57.x at time of validation)
mistral_common == 1.11.0
codecarbon, jiwer, librosa, soundfile, webrtcvad, pyloudnorm
NVIDIA driver ≥ 565 (CUDA 12.7 forward-compatible). Validated on NVIDIA L4 24 GB, driver 565.57, Python 3.11.10.
vllm_config.yaml
Infrastructure parameters omitted per Round-2 guidance; non-default flags documented.
| Argument | Value | Reason |
|---|---|---|
quantization |
compressed-tensors |
Load W4A16 GPTQ packed weights. |
kv_cache_dtype |
fp8_e4m3 |
FP8 KV cache. |
attention_backend |
TRITON_ATTN |
Required by Voxtral's Whisper-style encoder. |
tokenizer_mode |
mistral |
mistral-common tokenization. |
max_model_len |
16384 |
Fits the L4 KV cache; well above any FLEURS clip. |
max_num_seqs |
1 |
Serial transcription. |
max_num_batched_tokens |
16384 |
Tied to max_model_len. |
enable_prefix_caching |
true |
No-op on the audio endpoint; retained. |
compilation_config.cudagraph_mode |
PIECEWISE |
Required — full CUDAGraphs crash on the streaming head. |
speculative_config.method |
ngram |
Prompt-lookup drafter. |
speculative_config.num_speculative_tokens |
1 |
Maximum stable depth on vLLM 0.19.1's audio path. |
speculative_config.prompt_lookup_max |
2 |
Maximum lookup prefix length. |
speculative_config.prompt_lookup_min |
1 |
Minimum lookup prefix length. |
EAGLE-3 drafter
A 1-layer Llama-architecture EAGLE-3 draft head with a 32 k draft
vocabulary was trained on audio-conditioned trajectories across all
thirteen languages. Offline acceptance: 53.84 % top-1, 69.69 % top-5.
On the L4 binding it draws 183.3 kJ — 17 % higher than ngram in this
artifact — and degrades served quality on multiple languages, which
indicates that vLLM 0.19.1's EAGLE + Voxtral integration breaks the
rejection-sampling guarantee in practice. The trained draft head
remains available at
voxtral-mini-4b-asr-eagle3-specdec
for follow-up work; ngram is the production speculative configuration.
Known issues
- vLLM 0.19.1's audio path requires
cudagraph_mode: PIECEWISE. num_speculative_tokens >= 2crashes the audio endpoint (upstream vLLM bug);k = 1is fixed for this artifact./v1/realtimestreaming has not been validated with speculative configuration — use/v1/audio/transcriptions.
License
Apache-2.0, inherited from the base model.
- Downloads last month
- -
Model tree for Shankara-A-S/voxtral-mini-4b-asr-specdec
Base model
mistralai/Ministral-3-3B-Base-2512