--- license: mit library_name: vibevoice.c base_model: Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM tags: - speculative-decoding - dflash - awq - compressed-tensors - speech-recognition - vibevoice pipeline_tag: automatic-speech-recognition --- # VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 **VibeVoice-ASR-Streaming-7B**, AWQ W4A16 (asymmetric, groups of 128), **with its [DFlash 2](https://inco.ai/blog/dflash2/) drafter bundled** in `drafter/`: one download, and [vibevoice.c](https://github.com/Ar4ikov/vibevoice.c) decodes with speculative decoding -- the drafter proposes 8 tokens in one pass, the model checks them in one pass and keeps the ones it agrees with. **The check is exact**: every checked row is computed with the arithmetic of the model's own decode step, so the transcript is byte-for-byte the one without the drafter. ## Use Needs vibevoice.c with DFlash 2 support: branch `dflash2` ([PR #48](https://github.com/Ar4ikov/vibevoice.c/pull/48)), in the next release. A model directory's `drafter/` is used without asking: ```bash vv_cli --model ./VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 --audio talk.wav # with the drafter vv_cli --model ./VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 --audio talk.wav --draft none # plain decoding vv_cli serve --model ./VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM-DFlash2 --slots 4 # streaming sessions (WebSocket, SSE) too ``` ## Results vibevoice.c `b72be15` (branch `dflash2`), RTX 3090, greedy decoding, decode tokens per second: | | plain | drafted | speedup | tokens per block | same transcript | |---|---|---|---|---|---| | 20 held-out clips, 8 rows | 149 tok/s | 364 tok/s | **2.44x** | 3.57 | 20/20 | Streaming sessions (22 + 4 frames a chunk). Plain = the same model with `--draft none`; `--draft-check exact`. ## Inside * The model: the files of [Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/VibeVoice-ASR-Streaming-7B-AWQ-W4A16-ASYM) at revision `1cc2b627`, unchanged (7.01 GB) -- its card has the quantization, the calibration and the WER. * `drafter/`: [Ar4ikov/VibeVoice-ASR-Streaming-7B-DFlash2-Drafter-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/VibeVoice-ASR-Streaming-7B-DFlash2-Drafter-AWQ-W4A16-ASYM) at revision `65e803b7` (0.55 GB): 5 Qwen3-style layers reading the model's layers 1/7/13/19/25, a candidate selector, a 32768-id draft vocabulary; its projections stored as INT4 (compressed-tensors `pack-quantized`). Its card has the architecture and the training. ## License MIT, like VibeVoice.