voxtral-mini-4b-realtime-2602-p300x2
Voxtral-Mini-4B-Realtime-2602, Mistral AI's realtime speech-to-text model (a 970M-parameter causal audio encoder feeding a 3.4B Ministral text decoder; 13 languages; one text token per 80 ms of audio), served on four Blackhole chips (2 x p300c, 1x4 ring, TP=4) through a TTNN autoport and the Tenstorrent vLLM plugin as an OpenAI-compatible /v1/audio/transcriptions endpoint.
Runs on p300x2 (mesh P300x2) β 131,072-token context, up to 32 concurrent sequences.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
At a glance
| Architecture | 4B realtime ASR (970M causal audio encoder + 3.4B Ministral-3B decoder) |
| Hardware | p300x2 |
| Context | 131,072 tokens |
| License | apache-2.0 |
| Status | Experimental community bring-up (tt-model-bringup autoport, 11 stages, locally committed) |
Intended use
Direct use: Speech-to-text transcription of mono speech audio (wav, flac, mp3 and other formats soundfile decodes; resampled to 16 kHz) through the OpenAI audio transcriptions API, one interactive user or up to 32 concurrent requests, in the 13 languages the checkpoint covers.
Out-of-scope use: Chat, text generation, translation, speaker diarization, or audio question answering: the checkpoint is transcription-only and the server exposes only that task.
Quickstart
uv tool install tenstorrent # once β the Tenstorrent CLI, `tt`
tt model pull ndaly/Voxtral-Mini-4B-Realtime-2602-tt-p300x2
tt serve ndaly/Voxtral-Mini-4B-Realtime-2602-tt-p300x2
tt model pull (or tt-model pull --with-weights) downloads the Docker image and the mistralai/Voxtral-Mini-4B-Realtime-2602 weights at 2769294da9567371363522aac9bbcfdd19447add (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.
Without tt-cli β tt-model alone does the whole job:
tt-model pull ndaly/Voxtral-Mini-4B-Realtime-2602-tt-p300x2 --with-weights
tt-model serve ndaly/Voxtral-Mini-4B-Realtime-2602-tt-p300x2
What to expect
- Hardware: a TT-QuietBox 2 or any host with 2 x p300c (4 Blackhole chips, 32 GB GDDR6 each), Docker, 1G hugepages.
- Weights: one 8.9 GB safetensors file (plus tokenizer and config files) is downloaded into your Hugging Face cache on first
serve. - First boot compiles kernels and captures the decode trace plus 32 per-slot prefill traces (several minutes); later boots reuse the
kernel cache under
~/.cache/tt-model/voxtral-mini-4b-realtime-2602-p300x2/. - Measured (stage 10 serving run, 2026-09-23): 37.5 ms served TTFT and ~50 tokens/s per user for one user; 606 tokens/s aggregate
for 32 concurrent requests. Full evidence:
doc/optimized_vllm/README.mdanddoc/tti_release/RUN_NOTES.mdin the tt-metal autoport tree.
Use with your client
Point an OpenAI-compatible client at http://127.0.0.1:20000 (or the port tt-model serve printed) and call the audio
transcriptions API with model id mistralai/Voxtral-Mini-4B-Realtime-2602 (see Using it below).
Using it
The server speaks the OpenAI API at http://127.0.0.1:20000/v1 (or whichever port your serve command reported; chat completions, completions, and /v1/models). Pass "model": "mistralai/Voxtral-Mini-4B-Realtime-2602" β the weights id, not this package's name.
The endpoint is POST /v1/audio/transcriptions (multipart form, OpenAI shape). With the server on the port tt-model serve printed
(20000 by default):
curl -s http://127.0.0.1:20000/v1/audio/transcriptions \
-F model=mistralai/Voxtral-Mini-4B-Realtime-2602 -F file=@clip.wav
# streaming (server-sent events, text deltas as they are decoded):
curl -sN http://127.0.0.1:20000/v1/audio/transcriptions \
-F model=mistralai/Voxtral-Mini-4B-Realtime-2602 -F file=@clip.wav -F stream=true
Any OpenAI client works the same way (client.audio.transcriptions.create(model=..., file=...)). tt-model curl sends a chat
completion and gets a 404 from this model: use the request above instead.
Expected performance
Measured on p300x2 through the served endpoint (stage 10 of the bring-up, 2026-09-23): single user, one 5.28 s clip (78 completion tokens, greedy, streaming): 37.5 ms served time-to-first-token (the adapter's prefill wall), 20.14 ms per decode step (49.7 tokens/s per user), 1.59 s end to end; 32 concurrent requests: 606 tokens/s aggregate, 3.63 s median end to end, 32/32 complete. The generator's own traced decode is 19.55 ms per token. Accuracy: 5.64 % WER on LibriSpeech test-other (2,939 utterances, greedy, tt-inference-server release run through lmms-eval); transcripts identical to the Hugging Face reference on 8 of 10 qualitative clips (the other 2 differ by near-tie words); top-1 token agreement 0.995 against the bf16 reference at the selected precision (bfp4 + LoFi inner projections, bfp8 LM head and KV cache, 6.8 GB per chip).
Limitations
Transcription only: the checkpoint supports no chat or text completion, so the server mounts /v1/audio/transcriptions (and /v1/models) but not /v1/chat/completions; tt-model curl cannot talk to it (see Using it). One TP=4 replica on exactly four chips (p300x2); no data parallelism, no other board validated. At most 32 concurrent sequences; 131,072-token context contract (the checkpoint's advertised context); block_size is fixed at 32 by the generator. Requires --tokenizer-mode mistral (tekken.json). Chunked prefill and prefix caching are disabled by the TT backend. Validated on utterances up to about 35 s (LibriSpeech) and a 30 s benchmark payload; longer audio runs through the same streaming path but was not measured. The vLLM stack is the tenstorrent/vllm fork's empty-target wheel plus its in-tree plugin with local commits, not the standalone vllm-tt-plugin.
Risks and safety considerations
The bfp4 + LoFi precision policy drops one word onset on 1 of 2,939 LibriSpeech test-other utterances where the bf16 reference does not. Under concurrency, 1 of 2,939 utterances came back empty from a rotation-dependent near-tie onset (the Hugging Face reference also returns empty on the evaluator's exact encoding). The release accuracy gate compared against a placeholder reference because no published LibriSpeech test-other WER exists for this checkpoint. Transcripts are greedy by default; sampled decoding is supported but not evaluated for accuracy.
Licensing
Weights under Mistral AI's Apache-2.0 licence; the TTNN port and serving code in code/ are Apache-2.0 (tt-metal), as are the vLLM fork and its plugin inside the image.
Feedback
Questions or problems with this package: open a discussion at https://huggingface.co/ndaly/Voxtral-Mini-4B-Realtime-2602-tt-p300x2/discussions β that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.
Provenance
The exact sources the image was built from β code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | a local checkout β commit not published (dirty tree β the image includes uncommitted changes) |
| vLLM | vllm-0.1.dev14195+g8c28fcecb.d20260916.empty-py3-none-any.whl β a wheel the author built |
| vllm-tt-plugin | a local checkout β commit not published (dirty tree β the image includes uncommitted changes) |
code/ digest |
b26d652115a867ab (sha256, first 16 hex digits) |
| built | 2026-09-25T16:20:42+00:00 by tt-model 0.1.0 |
Model tree for ndaly/Voxtral-Mini-4B-Realtime-2602-tt-p300x2
Base model
mistralai/Ministral-3-3B-Base-2512