Automatic Speech Recognition
Transformers
Safetensors
Chinese
English
audio8_asr_infinite
text-generation
streaming
realtime
speech-recognition
audio
custom_code
Instructions to use Edge0/Audio8-ASR-Infinite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Edge0/Audio8-ASR-Infinite with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="Edge0/Audio8-ASR-Infinite", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Edge0/Audio8-ASR-Infinite", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 7,825 Bytes
41d6d0b ed50ac2 41d6d0b ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de 7476824 b4413de 7476824 ed50ac2 7476824 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 b4413de ed50ac2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 | ---
license: apache-2.0
language:
- zh
- en
library_name: transformers
pipeline_tag: automatic-speech-recognition
tags:
- streaming
- realtime
- speech-recognition
- audio
---
<div align="center">
# Audio8 ASR Infinite
[](https://huggingface.co/Edge0/Audio8-ASR-Infinite)
[](https://github.com/Edge0-AI/Audio8-ASR-Infinite)
[](https://github.com/Edge0-AI/Audio8-ASR-Infinite)
[](https://github.com/Edge0-AI/Audio8-ASR-Infinite/blob/main/LICENSE)
</div>
**Audio8 ASR Infinite** is a native streaming speech recognition model built to be
as responsive as possible. It offers a selectable audio clock (80/120/160 ms) and
a transcription delay (240β560 ms).
With our adapted vLLM build it transcribes unlimited-length audio **24/7** without drifting.
## Highlights
- **Super responsive** β the native streaming architecture decodes 12.5 times per second.
- **Unlimited-length transcription** β a rolling KV Cache keeps both **memory and
latency constant**, even in **24/7 operation**.
- **Selectable streaming clock** β one text token per clock step
(12.5 / 8.3 / 6.25 decisions per second), balancing perception granularity and resource cost.
- **Configurable transcription delay** β set how much delay to trade for accuracy.
- **Semantic VAD** β distinguishes thinking pauses, stuttering and real end of turn, where traditional acoustic VAD usually fails.
- **Bilingual** β Chinese and English.
## See Audio8-ASR-Infinite in action
The checkpoint has a native context of 30 seconds. But with Rolling KV Cache, it can transcribe 24/7 nonstop.
<video controls playsinline width="100%" preload="metadata"
src="https://huggingface.co/Edge0/Audio8-ASR-Infinite/resolve/main/Audio8-Asr-Infinite-Demo.mp4"></video>
## Optimized operation points
The following combinations of frame length and delay are post-trained. Other combinations can be used but performance may not be optimum.
| audio clock | `frame_len` | `streaming_n_left_pad_tokens` | selectable `target_delay_ms` |
| --- | --- | --- | --- |
| 80 ms | 4 | 18 | 240 / 320 / 480 / 560 |
| 120 ms | 6 | 12 | 240 / 480 |
| 160 ms | 8 | 9 | 320 / 480 |
`target_delay_ms` must be an integer multiple of the selected clock, so longer
delays stay available at every clock even when they are not listed above.
## Architecture
Inherits the Voxtral realtime audio architecture and DSM-style streaming.
| Component | Initial weights | Trained |
| --- | --- | --- |
| Causal Audio Tower | Voxtral Realtime 4B | β
|
| Audio Projector | random initialisation | β
|
| Frame Length Embedding | random initialisation | β
|
| Decoder | Qwen2.5-3B-Instruct | β
|
| LM Head | Qwen2.5-3B-Instruct | β
|
Checkpoint specification:
| | |
| --- | --- |
| audio tower | 32 layers, hidden 1280, 128 mel bins, sliding window 750 |
| text decoder | 36 layers, hidden 2048, 16 query heads / 2 KV heads |
| projector | max frame len 8 β projection size 10240, gelu |
| frame-length conditioning | enabled (`use_frame_len_embedding: true`) |
| semantic VAD heads | `semantic_vad_heads.safetensors`, 8 classes, horizons 0.5 / 1.0 / 2.0 / 3.0 s |
| vocab size | 151936 |
| dtype | bfloat16 |
| weights | 8.17 GB `model.safetensors` (+ `semantic_vad_heads.safetensors`) |
## Roadmap
This is the **preview release**: it delivers the transcription base. Realtime
semantic perception is being built on the same frame grid and the same acoustic
forward pass.
| Stage | Status | Scope |
| --- | --- | --- |
| **Preview β ASR base** | β
done | Streaming Chinese/English transcription: selectable 80/120/160 ms clock, configurable `target_delay_ms`, unlimited-length rolling KV window |
| **Formal release** | πin progress | Frame-level semantic perception on the same grid, beyond transcription |
## Evaluation
### 480 ms Delay, 80ms frame length
| test set | metric | Audio8 ASR Infinite | Voxtral-Mini-4B-Realtime-2602 | nemotron-3.5-asr-streaming-0.6b |
| --- | --- | --- | --- | --- |
| aishell1/test | CER | **1.750** | 16.795 | 12.927@560ms |
| aishell4/test | CER | **2.893** | 16.456 | 14.677@560ms |
| librispeech test.clean | WER | 3.042 | **2.210** | 3.353@560ms |
| librispeech test.other | WER | 6.808 | **5.552** | 7.140@560ms |
| **average** | | **3.623** | 10.253 (2 sets) | 9.524 |
Greedy decode with EOS suppressed, at the 80 ms audio clock with
`target_delay_ms = 480` (6 delay tokens). Error rates in percent. No repetition
loops and no dropped trailing words.
## Usage
Programmatic simulated-streaming decode with the embedded remote code:
```python
import numpy as np
import torch
from transformers import AutoFeatureExtractor, AutoTokenizer
from audio8_asr_infinite.modeling.modeling_audio8_asr_infinite import (
Audio8ASRInfiniteForConditionalGeneration,
resolve_qwen_language_token_id,
resolve_qwen_streaming_special_token_ids,
)
from audio8_asr_infinite.streaming_inference import simulated_streaming_greedy_decode_batch
checkpoint = "Edge0/Audio8-ASR-Infinite"
tokenizer = AutoTokenizer.from_pretrained(checkpoint, trust_remote_code=True)
feature_extractor = AutoFeatureExtractor.from_pretrained(checkpoint, trust_remote_code=True)
model = Audio8ASRInfiniteForConditionalGeneration.from_pretrained(
checkpoint, trust_remote_code=True, torch_dtype=torch.bfloat16
).eval().cuda()
class AudioConfig: # duck-typed: raw_audio_samples_per_token / streaming_n_left_pad_tokens / sampling_rate
raw_audio_samples_per_token = 1280 # 80 ms @ 16 kHz
streaming_n_left_pad_tokens = 18
sampling_rate = 16000
waveform = np.load("sample.npy", allow_pickle=False).astype(np.float32) # [-1, 1], 16 kHz mono
results = simulated_streaming_greedy_decode_batch(
model=model,
tokenizer=tokenizer,
feature_extractor=feature_extractor,
waveforms=[waveform],
language_token_ids=[resolve_qwen_language_token_id(tokenizer, "zh")],
special_ids=resolve_qwen_streaming_special_token_ids(tokenizer),
audio_config=AudioConfig(),
num_delay_tokens=[480 // 80],
right_pad_text_tokens=10,
dtype=torch.bfloat16,
device=next(model.parameters()).device,
max_new_tokens=512,
)
print(results[0]["final_text"])
```
Only a full merged weight directory is supported (this repository as-is);
adapter-style or partially converted weights are not.
## 24/7 inference with vLLM
Docker compose is the canonical deployment path; it also serves the web demo:
```bash
cd docker
AUDIO8_MODEL_DIR=/path/to/checkpoint docker compose up -d
```
Verify with the web client shipped in the same stack:
```
http://localhost:8080/ # plain HTTP
https://localhost:8443/ # TLS proxy; accept the self-signed certificate
```
The same socket can be driven from a terminal:
```bash
python -m audio8_asr_infinite.examples.vllm_realtime_client \
--ws-url ws://127.0.0.1:18191/v1/realtime \
--audio sample.wav --language zh --target-delay-ms 480 --pace
```
`18191` is the host port published by `docker/docker-compose.yml`; the service
itself listens on `18190` inside the compose network. The rolling KV window is
30 s with exact RoPE re-basing, which is what keeps memory and latency bounded
over 24/7 operation.
## Torch inference (simulated streaming decode)
```bash
python -m audio8_asr_infinite.examples.torch_streaming_decode \
--checkpoint /path/to/checkpoint \
--audio sample.wav --language zh --transcription-delay-ms 480
```
|