Automatic Speech Recognition
Safetensors
VibeVoice
vibevoice.c
speculative-decoding
dflash
awq
compressed-tensors
speech-recognition
4-bit precision
Instructions to use Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- VibeVoice
How to use Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 with VibeVoice:
import torch, soundfile as sf, librosa, numpy as np from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor from vibevoice.modular.modeling_vibevoice_inference import VibeVoiceForConditionalGenerationInference # Load voice sample (should be 24kHz mono) voice, sr = sf.read("path/to/voice_sample.wav") if voice.ndim > 1: voice = voice.mean(axis=1) if sr != 24000: voice = librosa.resample(voice, sr, 24000) processor = VibeVoiceProcessor.from_pretrained("Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2") model = VibeVoiceForConditionalGenerationInference.from_pretrained( "Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2", torch_dtype=torch.bfloat16 ).to("cuda").eval() model.set_ddpm_inference_steps(5) inputs = processor(text=["Speaker 0: Hello!\nSpeaker 1: Hi there!"], voice_samples=[[voice]], return_tensors="pt") audio = model.generate(**inputs, cfg_scale=1.3, tokenizer=processor.tokenizer).speech_outputs[0] sf.write("output.wav", audio.cpu().numpy().squeeze(), 24000) - Notebooks
- Google Colab
- Kaggle
File size: 2,600 Bytes
463a3ca | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 | ---
license: mit
library_name: vibevoice.c
base_model: Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM
tags:
- speculative-decoding
- dflash
- awq
- compressed-tensors
- speech-recognition
- vibevoice
pipeline_tag: automatic-speech-recognition
---
# VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2
**VibeVoice-ASR-Streaming-7B**, AWQ W4A16 (asymmetric, groups of 128), **with its
[DFlash 2](https://inco.ai/blog/dflash2/) drafter bundled** in `drafter/`:
one download, and [vibevoice.c](https://github.com/Ar4ikov/vibevoice.c)
decodes with speculative decoding -- the drafter proposes 8 tokens in one
pass, the model checks them in one pass and keeps the ones it agrees with.
**The check is exact**: every checked row is computed with the arithmetic of
the model's own decode step, so the transcript is byte-for-byte the one
without the drafter.
## Use
Needs vibevoice.c with DFlash 2 support: branch `dflash2`
([PR #48](https://github.com/Ar4ikov/vibevoice.c/pull/48)), in the next
release. A model directory's `drafter/` is used without asking:
```bash
vv_cli --model ./VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 --audio talk.wav # with the drafter
vv_cli --model ./VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 --audio talk.wav --draft none # plain decoding
vv_cli serve --model ./VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 --slots 4 # streaming sessions (WebSocket, SSE) too
```
## Results
vibevoice.c `b72be15` (branch `dflash2`), RTX 3090, greedy decoding, decode tokens per second:
| | plain | drafted | speedup | tokens per block | same transcript |
|---|---|---|---|---|---|
| 20 held-out clips, 8 rows | 149 tok/s | 364 tok/s | **2.44x** | 3.57 | 20/20 |
Streaming sessions (22 + 4 frames a chunk). Plain = the same model with `--draft none`; `--draft-check exact`.
## Inside
* The model: the files of [Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM)
at revision `1cc2b627`, unchanged (7.01 GB) -- its card has the
quantization, the calibration and the WER.
* `drafter/`: [Ar4ikov/VibeVoice-ASR-Streaming-7B-DFlash2-Drafter-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/VibeVoice-ASR-Streaming-7B-DFlash2-Drafter-AWQ-W4A16-ASYM) at revision
`65e803b7` (0.55 GB): 5 Qwen3-style layers reading the model's
layers 1/7/13/19/25, a candidate selector, a 32768-id draft vocabulary; its
projections stored as INT4 (compressed-tensors `pack-quantized`). Its card
has the architecture and the training.
## License
MIT, like VibeVoice.
|