bupalinyu commited on
Commit
1b0b5e7
·
verified ·
1 Parent(s): 9c9ed0e

Update README

Browse files
Files changed (1) hide show
  1. README.md +4 -4
README.md CHANGED
@@ -36,7 +36,7 @@ repository: https://github.com/AutoArk/open-audio-opd
36
 
37
  <div align="center">
38
 
39
- # ARK-ASR-3B: State-of-the-Art Multilingual ASR with Online Policy Distillation
40
 
41
  [![GitHub](https://img.shields.io/badge/GitHub-AutoArk%2Fopen--audio--opd-blue?logo=github)](https://github.com/AutoArk/open-audio-opd)
42
  [![arXiv](https://img.shields.io/badge/arXiv-2605.28139-b31b1b?logo=arxiv)](https://arxiv.org/abs/2605.28139)
@@ -44,13 +44,13 @@ repository: https://github.com/AutoArk/open-audio-opd
44
 
45
  </div>
46
 
47
- > **TL;DR** ARK-ASR-3B is an automatic speech recognition model trained with teacher-data adaptation and on-policy distillation. It achieves current state-of-the-art results on the Hugging Face Open ASR Leaderboard English short-form benchmark, with an average WER of **5.13%** across AMI, Earnings22, GigaSpeech, LibriSpeech, SPGISpeech, and VoxPopuli. The accompanying training, inference, and evaluation code is available at [AutoArk/open-audio-opd](https://github.com/AutoArk/open-audio-opd).
48
 
49
  ## Abstract
50
 
51
- ARK-ASR is an audio ASR student model optimized with the **teacher-data adaptation + online policy distillation (TD + OPD)** recipe from `open-audio-opd`.
52
 
53
- Instead of relying only on static supervised transcripts, OPD lets the student generate transcripts online and trains it against token-level teacher scores on the student's own generated behavior. This checkpoint is the 3B-scale ARK-ASR model trained with the TD + OPD recipe.
54
 
55
  ARK-ASR currently supports Chinese, English, German, Japanese, French, Korean, Spanish, Polish, Italian, Romanian, Hungarian, Czech, Dutch, Finnish, Croatian, Slovak, Slovene, Estonian, and Lithuanian ASR.
56
 
 
36
 
37
  <div align="center">
38
 
39
+ # ARK-ASR-3B: State-of-the-Art Multilingual ASR
40
 
41
  [![GitHub](https://img.shields.io/badge/GitHub-AutoArk%2Fopen--audio--opd-blue?logo=github)](https://github.com/AutoArk/open-audio-opd)
42
  [![arXiv](https://img.shields.io/badge/arXiv-2605.28139-b31b1b?logo=arxiv)](https://arxiv.org/abs/2605.28139)
 
44
 
45
  </div>
46
 
47
+ > **TL;DR** ARK-ASR-3B is a multilingual automatic speech recognition model. It achieves current state-of-the-art results on the Hugging Face Open ASR Leaderboard English short-form benchmark, with an average WER of **5.13%** across AMI, Earnings22, GigaSpeech, LibriSpeech, SPGISpeech, and VoxPopuli. The accompanying training, inference, and evaluation code is available at [AutoArk/open-audio-opd](https://github.com/AutoArk/open-audio-opd).
48
 
49
  ## Abstract
50
 
51
+ ARK-ASR-3B is a 3B-scale audio-capable autoregressive Transformers model for automatic speech recognition.
52
 
53
+ It combines a Whisper-style audio encoder, an MLP adapter, and a Qwen decoder with custom `arkasr` remote code.
54
 
55
  ARK-ASR currently supports Chinese, English, German, Japanese, French, Korean, Spanish, Polish, Italian, Romanian, Hungarian, Czech, Dutch, Finnish, Croatian, Slovak, Slovene, Estonian, and Lithuanian ASR.
56