Improving noise robustness of automatic speech recognition via parallel data and teacher-student learning
Paper β’ 1901.02348 β’ Published
Tiro is a far-field fine-tune of Qwen3-ASR-1.7B, with better performance on reverberant and noisy speech, scaling as conditions get harder, while maintaining general ASR capabilities. See benchmarks below for more details.
Currently #1 on the FFASR leaderboard.
English only, bf16. Same architecture and parameter count as the base.
Lower is better. β marks the better result.
| Condition | Base (WER) | Tiro (WER) |
|---|---|---|
| Average | 13.41 | 12.01 β |
| Near field speech | 3.76 β | 3.80 |
| High SNR | 6.74 | 6.22 β |
| Mid SNR | 13.89 | 12.00 β |
| Low SNR | 29.26 | 26.02 β |
| Benchmark | Condition | Base (WER) | Tiro (WER) |
|---|---|---|---|
| VOiCES | real rooms, replayed speech | 8.88 | 8.30 β |
| CHiME-4 | real noisy environments | 4.71 | 4.19 β |
| CHiME-4 | simulated arm | 7.00 | 6.63 β |
| VITW-Bench | 8 perturbations, English | 7.09 | 6.40 β |
| NOIZEUS | additive noise, 0-15 dB | 9.23 | 7.85 β |
| Treble10-Speech | reverberation only | 3.29 | 3.19 β |
| AMI | distant mic (SDM) | 19.73 | 19.33 β |
| AMI | headset control (IHM) | 8.24 β | 8.28 |
| Open ASR Leaderboard | mean of 8 public English splits | 4.311 | 4.310 |
| Benchmark | Condition | Base (WER) | Tiro (WER) |
|---|---|---|---|
| VOiCES | near microphone | 3.09 | 2.85 β |
| VOiCES | far microphone | 14.68 | 13.76 β |
| VOiCES | no distractor | 3.27 | 3.07 β |
| VOiCES | music | 8.06 | 7.08 β |
| VOiCES | television | 8.88 | 8.29 β |
| VOiCES | babble | 15.33 | 14.77 β |
| CHiME-4 real | bus | 5.98 | 5.42 β |
| CHiME-4 real | cafe | 4.84 | 4.09 β |
| CHiME-4 real | pedestrian | 4.18 | 3.89 β |
| CHiME-4 real | street | 3.82 | 3.35 β |
| NOIZEUS | 0 dB SNR | 23.24 | 20.76 β |
| NOIZEUS | 5 dB SNR | 8.42 | 6.51 β |
| NOIZEUS | 10 dB SNR | 3.20 | 2.74 β |
| NOIZEUS | 15 dB SNR | 2.07 | 1.39 β |
| VITW-Bench | real recordings | 5.30 | 4.84 β |
| VITW-Bench | simulated | 7.80 | 7.03 β |
Base column is the published Qwen3-ASR-1.7B result on the 8 public English splits. The leaderboard's headline average also includes two private sets that are not publicly runnable, so the figure shown there is higher than this one.
| Split | Base (WER) | Tiro (WER) |
|---|---|---|
| LibriSpeech test-clean | 1.26 | 1.25 β |
| LibriSpeech test-other | 2.94 β | 2.95 |
| VoxPopuli | 2.85 | 2.62 β |
| AMI | 8.31 β | 8.45 |
| GigaSpeech | 7.22 β | 7.23 |
| Earnings22 | 5.84 | 5.70 β |
| SPGISpeech | 2.57 β | 2.77 |
| Monsoon | 3.50 β | 3.51 |
| mean | 4.311 | 4.310 |
Identical to the base model.
from transformers import AutoProcessor, Qwen3ASRForConditionalGeneration
model_id = "rohansheth/tiro-qwen3-asr-1.7b-v2"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3ASRForConditionalGeneration.from_pretrained(model_id, device_map="auto")
Trained by KL distillation from the frozen base model on noise-free audio, following:
@inproceedings{li2017large,
title = {Large-Scale Domain Adaptation via Teacher-Student Learning},
author = {Li, Jinyu and Seltzer, Michael L. and Wang, Xi and Zhao, Rui and Gong, Yifan},
booktitle = {Interspeech},
year = {2017},
eprint = {1708.05466}
}
@inproceedings{mosner2019improving,
title = {Improving Noise Robustness of Automatic Speech Recognition via
Parallel Data and Teacher-Student Learning},
author = {Mo{\v{s}}ner, Ladislav and Wu, Minhua and Raju, Anirudh and
Parthasarathi, Sree Hari Krishnan and Kumatani, Kenichi and
Sundaram, Shiva and Maas, Roland and Hoffmeister, Bj{\"o}rn},
booktitle = {ICASSP},
year = {2019},
eprint = {1901.02348}
}
Base model:
@misc{qwen3asr,
title = {Qwen3-ASR},
author = {Qwen Team},
year = {2026},
url = {https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf}
}
Base model
Qwen/Qwen3-ASR-1.7B-hf