Automatic Speech Recognition
Transformers
Safetensors
PyTorch
arkasr
text-generation
speech
audio
vllm
ark-asr
custom_code
Eval Results
Instructions to use Edge0/ARK-ASR-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Edge0/ARK-ASR-3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="Edge0/ARK-ASR-3B", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Edge0/ARK-ASR-3B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README
Browse files
README.md
CHANGED
|
@@ -36,7 +36,7 @@ repository: https://github.com/AutoArk/open-audio-opd
|
|
| 36 |
|
| 37 |
<div align="center">
|
| 38 |
|
| 39 |
-
# ARK-ASR-3B: State-of-the-Art Multilingual ASR
|
| 40 |
|
| 41 |
[](https://github.com/AutoArk/open-audio-opd)
|
| 42 |
[](https://arxiv.org/abs/2605.28139)
|
|
@@ -44,13 +44,13 @@ repository: https://github.com/AutoArk/open-audio-opd
|
|
| 44 |
|
| 45 |
</div>
|
| 46 |
|
| 47 |
-
> **TL;DR** ARK-ASR-3B is
|
| 48 |
|
| 49 |
## Abstract
|
| 50 |
|
| 51 |
-
ARK-ASR is
|
| 52 |
|
| 53 |
-
|
| 54 |
|
| 55 |
ARK-ASR currently supports Chinese, English, German, Japanese, French, Korean, Spanish, Polish, Italian, Romanian, Hungarian, Czech, Dutch, Finnish, Croatian, Slovak, Slovene, Estonian, and Lithuanian ASR.
|
| 56 |
|
|
|
|
| 36 |
|
| 37 |
<div align="center">
|
| 38 |
|
| 39 |
+
# ARK-ASR-3B: State-of-the-Art Multilingual ASR
|
| 40 |
|
| 41 |
[](https://github.com/AutoArk/open-audio-opd)
|
| 42 |
[](https://arxiv.org/abs/2605.28139)
|
|
|
|
| 44 |
|
| 45 |
</div>
|
| 46 |
|
| 47 |
+
> **TL;DR** ARK-ASR-3B is a multilingual automatic speech recognition model. It achieves current state-of-the-art results on the Hugging Face Open ASR Leaderboard English short-form benchmark, with an average WER of **5.13%** across AMI, Earnings22, GigaSpeech, LibriSpeech, SPGISpeech, and VoxPopuli. The accompanying training, inference, and evaluation code is available at [AutoArk/open-audio-opd](https://github.com/AutoArk/open-audio-opd).
|
| 48 |
|
| 49 |
## Abstract
|
| 50 |
|
| 51 |
+
ARK-ASR-3B is a 3B-scale audio-capable autoregressive Transformers model for automatic speech recognition.
|
| 52 |
|
| 53 |
+
It combines a Whisper-style audio encoder, an MLP adapter, and a Qwen decoder with custom `arkasr` remote code.
|
| 54 |
|
| 55 |
ARK-ASR currently supports Chinese, English, German, Japanese, French, Korean, Spanish, Polish, Italian, Romanian, Hungarian, Czech, Dutch, Finnish, Croatian, Slovak, Slovene, Estonian, and Lithuanian ASR.
|
| 56 |
|