Qwen3-ASR-0.6B ONNX β€” int4 encoder

andrewleech/qwen3-asr-0.6b-onnx with the audio encoder requantized to int4. The decoders, embedding table, config, and tokenizer are that repo's published int4 files, unchanged.

Ultimately derived from Qwen/Qwen3-ASR-0.6B. Apache-2.0 throughout.

Why

The upstream repo's encoder.int4.onnx is misleadingly named: it holds FP32 weights (746 MB). Quantizing it for real shrinks the package, and on ARM it cuts latency too.

upstream this
encoder 746 MB (FP32) 121 MB (int4)
total on disk 2.0 GB 1.4 GB
Raspberry Pi 5, 3.5 s utterance, 4 threads 1.577 s 1.36–1.49 s
Ryzen 9 5950X, 3.5 s utterance, 16 threads 0.408 s 0.411 s
Pi 5 peak RSS, short utterance β€” 1.7 GB

The encoder is compute-bound, so the speedup is ARM-only; on x86 this is purely a memory saving. Transcripts matched the FP32 encoder on 9 of 10 Home Assistant test clips β€” the tenth was an entity name both get wrong without a context prompt and both get right with one.

Files

File Description
encoder.int4.onnx (+ .data) Audio encoder, int4 MatMulNBits
decoder_init.int4.onnx Decoder prefill; takes input_ids, emits logits + KV cache
decoder_step.int4.onnx Autoregressive step; takes input_embeds + KV cache
decoder_weights.int4.data Shared external weights for both decoders
embed_tokens.bin Token embeddings [151936, 1024], float16
config.json, tokenizer.json Architecture config, special tokens, mel params, tokenizer

Keep embed_tokens.bin in fp16 and cast one row per generated token; casting the whole table at load costs ~300 MB of RSS for nothing.

Inference

  1. Log-mel spectrogram (Whisper parameters: 16 kHz, 128 bins, n_fft 400, hop 160, Hann, Slaney mel, 0–8 kHz)
  2. encoder.int4.onnx β†’ audio features
  3. Build prompt ids, then decoder_init.int4.onnx with input_ids, position_ids, audio_features, audio_offset
  4. Greedy loop on decoder_step.int4.onnx until <|im_end|> or <|endoftext|>
  5. Drop everything up to and including <asr_text> (the language preamble)

The prompt is the Qwen chat template:

<|im_start|>system\n{context}<|im_end|>\n
<|im_start|>user\n<|audio_start|>{audio_pad Γ— N}<|audio_end|><|im_end|>\n
<|im_start|>assistant\n{language {Name}<asr_text>}

Note the token ids for the words system and user are 8948 and 872. The upstream reference src/prompt.py hardcodes 9125 and 882, which decode to " Current" and " time".

Context biasing

Free-form text in the system turn biases decoding toward specific spellings β€” useful for smart-home entity names:

system: Vocabulary: Ecobee, office lamp.

"What's the temperature of the incubator?" β†’ "What's the temperature of the Ecobee?"

Verified clean with a 77-token entity list: no dropped outputs, no prompt echo. (The int8 export at OpenVoiceOS/qwen3-asr-0.6b-onnx degenerates under the same load β€” empty output, or echoing the vocabulary list back as the transcript. int4 MatMulNBits' per-group scales handle the decoder's outlier weights that per-tensor int8 does not.)

Reproducing the encoder

import onnx
from onnxruntime.quantization.matmul_nbits_quantizer import (
    MatMulNBitsQuantizer, RTNWeightOnlyQuantConfig)
from onnxruntime.quantization.quant_utils import QuantFormat

q = MatMulNBitsQuantizer(
    model=onnx.load("encoder.int4.onnx"),   # the FP32-weighted file from upstream
    block_size=64, is_symmetric=False, accuracy_level=4,
    algo_config=RTNWeightOnlyQuantConfig(quant_format=QuantFormat.QOperator),
)
q.process()
q.model.save_model_to_file("encoder.int4.onnx", use_external_data_format=True)

block_size=64 / accuracy_level=4 match the recipe the decoders were built with. Don't change them casually β€” upstream measured block_size=32 at the same accuracy level producing 99.98% WER, and requantizing the decoder with these settings produced empty output on both x86 and ARM.

Downloads last month
101
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for rhasspy/qwen3-asr-0.6b-onnx-int4

Quantized
(2)
this model