Qwen3-ASR-0.6B ONNX β int4 encoder
andrewleech/qwen3-asr-0.6b-onnx
with the audio encoder requantized to int4. The decoders, embedding table,
config, and tokenizer are that repo's published int4 files, unchanged.
Ultimately derived from Qwen/Qwen3-ASR-0.6B.
Apache-2.0 throughout.
Why
The upstream repo's encoder.int4.onnx is misleadingly named: it holds FP32
weights (746 MB). Quantizing it for real shrinks the package, and on ARM it cuts
latency too.
| upstream | this | |
|---|---|---|
| encoder | 746 MB (FP32) | 121 MB (int4) |
| total on disk | 2.0 GB | 1.4 GB |
| Raspberry Pi 5, 3.5 s utterance, 4 threads | 1.577 s | 1.36β1.49 s |
| Ryzen 9 5950X, 3.5 s utterance, 16 threads | 0.408 s | 0.411 s |
| Pi 5 peak RSS, short utterance | β | 1.7 GB |
The encoder is compute-bound, so the speedup is ARM-only; on x86 this is purely a memory saving. Transcripts matched the FP32 encoder on 9 of 10 Home Assistant test clips β the tenth was an entity name both get wrong without a context prompt and both get right with one.
Files
| File | Description |
|---|---|
encoder.int4.onnx (+ .data) |
Audio encoder, int4 MatMulNBits |
decoder_init.int4.onnx |
Decoder prefill; takes input_ids, emits logits + KV cache |
decoder_step.int4.onnx |
Autoregressive step; takes input_embeds + KV cache |
decoder_weights.int4.data |
Shared external weights for both decoders |
embed_tokens.bin |
Token embeddings [151936, 1024], float16 |
config.json, tokenizer.json |
Architecture config, special tokens, mel params, tokenizer |
Keep embed_tokens.bin in fp16 and cast one row per generated token; casting the
whole table at load costs ~300 MB of RSS for nothing.
Inference
- Log-mel spectrogram (Whisper parameters: 16 kHz, 128 bins, n_fft 400, hop 160, Hann, Slaney mel, 0β8 kHz)
encoder.int4.onnxβ audio features- Build prompt ids, then
decoder_init.int4.onnxwithinput_ids,position_ids,audio_features,audio_offset - Greedy loop on
decoder_step.int4.onnxuntil<|im_end|>or<|endoftext|> - Drop everything up to and including
<asr_text>(the language preamble)
The prompt is the Qwen chat template:
<|im_start|>system\n{context}<|im_end|>\n
<|im_start|>user\n<|audio_start|>{audio_pad Γ N}<|audio_end|><|im_end|>\n
<|im_start|>assistant\n{language {Name}<asr_text>}
Note the token ids for the words system and user are 8948 and 872.
The upstream reference src/prompt.py hardcodes 9125 and 882, which decode to
" Current" and " time".
Context biasing
Free-form text in the system turn biases decoding toward specific spellings β useful for smart-home entity names:
system: Vocabulary: Ecobee, office lamp.
"What's the temperature of the incubator?" β "What's the temperature of the Ecobee?"
Verified clean with a 77-token entity list: no dropped outputs, no prompt echo.
(The int8 export at OpenVoiceOS/qwen3-asr-0.6b-onnx degenerates under the same
load β empty output, or echoing the vocabulary list back as the transcript. int4
MatMulNBits' per-group scales handle the decoder's outlier weights that per-tensor
int8 does not.)
Reproducing the encoder
import onnx
from onnxruntime.quantization.matmul_nbits_quantizer import (
MatMulNBitsQuantizer, RTNWeightOnlyQuantConfig)
from onnxruntime.quantization.quant_utils import QuantFormat
q = MatMulNBitsQuantizer(
model=onnx.load("encoder.int4.onnx"), # the FP32-weighted file from upstream
block_size=64, is_symmetric=False, accuracy_level=4,
algo_config=RTNWeightOnlyQuantConfig(quant_format=QuantFormat.QOperator),
)
q.process()
q.model.save_model_to_file("encoder.int4.onnx", use_external_data_format=True)
block_size=64 / accuracy_level=4 match the recipe the decoders were built
with. Don't change them casually β upstream measured block_size=32 at the same
accuracy level producing 99.98% WER, and requantizing the decoder with these
settings produced empty output on both x86 and ARM.
- Downloads last month
- 101