File size: 2,600 Bytes
463a3ca
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
---
license: mit
library_name: vibevoice.c
base_model: Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM
tags:
- speculative-decoding
- dflash
- awq
- compressed-tensors
- speech-recognition
- vibevoice
pipeline_tag: automatic-speech-recognition
---

# VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2

**VibeVoice-ASR-Streaming-7B**, AWQ W4A16 (asymmetric, groups of 128), **with its
[DFlash 2](https://inco.ai/blog/dflash2/) drafter bundled** in `drafter/`:
one download, and [vibevoice.c](https://github.com/Ar4ikov/vibevoice.c)
decodes with speculative decoding -- the drafter proposes 8 tokens in one
pass, the model checks them in one pass and keeps the ones it agrees with.
**The check is exact**: every checked row is computed with the arithmetic of
the model's own decode step, so the transcript is byte-for-byte the one
without the drafter.

## Use

Needs vibevoice.c with DFlash 2 support: branch `dflash2`
([PR #48](https://github.com/Ar4ikov/vibevoice.c/pull/48)), in the next
release. A model directory's `drafter/` is used without asking:

```bash
vv_cli --model ./VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 --audio talk.wav                  # with the drafter
vv_cli --model ./VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 --audio talk.wav --draft none     # plain decoding
vv_cli serve --model ./VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 --slots 4                  # streaming sessions (WebSocket, SSE) too
```

## Results

vibevoice.c `b72be15` (branch `dflash2`), RTX 3090, greedy decoding, decode tokens per second:

|  | plain | drafted | speedup | tokens per block | same transcript |
|---|---|---|---|---|---|
| 20 held-out clips, 8 rows | 149 tok/s | 364 tok/s | **2.44x** | 3.57 | 20/20 |

Streaming sessions (22 + 4 frames a chunk). Plain = the same model with `--draft none`; `--draft-check exact`.

## Inside

* The model: the files of [Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM)
  at revision `1cc2b627`, unchanged (7.01 GB) -- its card has the
  quantization, the calibration and the WER.
* `drafter/`: [Ar4ikov/VibeVoice-ASR-Streaming-7B-DFlash2-Drafter-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/VibeVoice-ASR-Streaming-7B-DFlash2-Drafter-AWQ-W4A16-ASYM) at revision
  `65e803b7` (0.55 GB): 5 Qwen3-style layers reading the model's
  layers 1/7/13/19/25, a candidate selector, a 32768-id draft vocabulary; its
  projections stored as INT4 (compressed-tensors `pack-quantized`). Its card
  has the architecture and the training.

## License

MIT, like VibeVoice.