tiro-qwen3-asr-1.7b-v2

Tiro is a far-field fine-tune of Qwen3-ASR-1.7B, with better performance on reverberant and noisy speech, scaling as conditions get harder, while maintaining general ASR capabilities. See benchmarks below for more details.

Currently #1 on the FFASR leaderboard.

English only, bf16. Same architecture and parameter count as the base.

Benchmarks

Lower is better. βœ… marks the better result.

FFASR Leaderboard

Condition Base (WER) Tiro (WER)
Average 13.41 12.01 βœ…
Near field speech 3.76 βœ… 3.80
High SNR 6.74 6.22 βœ…
Mid SNR 13.89 12.00 βœ…
Low SNR 29.26 26.02 βœ…

Other benchmarks

Benchmark Condition Base (WER) Tiro (WER)
VOiCES real rooms, replayed speech 8.88 8.30 βœ…
CHiME-4 real noisy environments 4.71 4.19 βœ…
CHiME-4 simulated arm 7.00 6.63 βœ…
VITW-Bench 8 perturbations, English 7.09 6.40 βœ…
NOIZEUS additive noise, 0-15 dB 9.23 7.85 βœ…
Treble10-Speech reverberation only 3.29 3.19 βœ…
AMI distant mic (SDM) 19.73 19.33 βœ…
AMI headset control (IHM) 8.24 βœ… 8.28
Open ASR Leaderboard mean of 8 public English splits 4.311 4.310
Per-condition breakdown
Benchmark Condition Base (WER) Tiro (WER)
VOiCES near microphone 3.09 2.85 βœ…
VOiCES far microphone 14.68 13.76 βœ…
VOiCES no distractor 3.27 3.07 βœ…
VOiCES music 8.06 7.08 βœ…
VOiCES television 8.88 8.29 βœ…
VOiCES babble 15.33 14.77 βœ…
CHiME-4 real bus 5.98 5.42 βœ…
CHiME-4 real cafe 4.84 4.09 βœ…
CHiME-4 real pedestrian 4.18 3.89 βœ…
CHiME-4 real street 3.82 3.35 βœ…
NOIZEUS 0 dB SNR 23.24 20.76 βœ…
NOIZEUS 5 dB SNR 8.42 6.51 βœ…
NOIZEUS 10 dB SNR 3.20 2.74 βœ…
NOIZEUS 15 dB SNR 2.07 1.39 βœ…
VITW-Bench real recordings 5.30 4.84 βœ…
VITW-Bench simulated 7.80 7.03 βœ…
Open ASR Leaderboard

Base column is the published Qwen3-ASR-1.7B result on the 8 public English splits. The leaderboard's headline average also includes two private sets that are not publicly runnable, so the figure shown there is higher than this one.

Split Base (WER) Tiro (WER)
LibriSpeech test-clean 1.26 1.25 βœ…
LibriSpeech test-other 2.94 βœ… 2.95
VoxPopuli 2.85 2.62 βœ…
AMI 8.31 βœ… 8.45
GigaSpeech 7.22 βœ… 7.23
Earnings22 5.84 5.70 βœ…
SPGISpeech 2.57 βœ… 2.77
Monsoon 3.50 βœ… 3.51
mean 4.311 4.310

Usage

Identical to the base model.

from transformers import AutoProcessor, Qwen3ASRForConditionalGeneration

model_id = "rohansheth/tiro-qwen3-asr-1.7b-v2"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3ASRForConditionalGeneration.from_pretrained(model_id, device_map="auto")

Citations

Trained by KL distillation from the frozen base model on noise-free audio, following:

@inproceedings{li2017large,
  title     = {Large-Scale Domain Adaptation via Teacher-Student Learning},
  author    = {Li, Jinyu and Seltzer, Michael L. and Wang, Xi and Zhao, Rui and Gong, Yifan},
  booktitle = {Interspeech},
  year      = {2017},
  eprint    = {1708.05466}
}

@inproceedings{mosner2019improving,
  title     = {Improving Noise Robustness of Automatic Speech Recognition via
               Parallel Data and Teacher-Student Learning},
  author    = {Mo{\v{s}}ner, Ladislav and Wu, Minhua and Raju, Anirudh and
               Parthasarathi, Sree Hari Krishnan and Kumatani, Kenichi and
               Sundaram, Shiva and Maas, Roland and Hoffmeister, Bj{\"o}rn},
  booktitle = {ICASSP},
  year      = {2019},
  eprint    = {1901.02348}
}

Base model:

@misc{qwen3asr,
  title  = {Qwen3-ASR},
  author = {Qwen Team},
  year   = {2026},
  url    = {https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf}
}
Downloads last month
214
Safetensors
Model size
2B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for rohansheth/tiro-qwen3-asr-1.7b-v2

Finetuned
(10)
this model

Papers for rohansheth/tiro-qwen3-asr-1.7b-v2