VibeVoice-ASR-Streaming-1.5B-AWQ-W4A16-ASYM-DFlash2

VibeVoice-ASR-Streaming-1.5B, AWQ W4A16 (asymmetric, groups of 128), with its DFlash 2 drafter bundled in drafter/: one download, and vibevoice.c decodes with speculative decoding -- the drafter proposes 8 tokens in one pass, the model checks them in one pass and keeps the ones it agrees with. The check is exact: every checked row is computed with the arithmetic of the model's own decode step, so the transcript is byte-for-byte the one without the drafter.

Use

Needs vibevoice.c with DFlash 2 support: branch dflash2 (PR #48), in the next release. A model directory's drafter/ is used without asking:

vv_cli --model ./VibeVoice-ASR-Streaming-1.5B-AWQ-W4A16-ASYM-DFlash2 --audio talk.wav                  # with the drafter
vv_cli --model ./VibeVoice-ASR-Streaming-1.5B-AWQ-W4A16-ASYM-DFlash2 --audio talk.wav --draft none     # plain decoding
vv_cli serve --model ./VibeVoice-ASR-Streaming-1.5B-AWQ-W4A16-ASYM-DFlash2 --slots 4                  # streaming sessions (WebSocket, SSE) too

Results

vibevoice.c 7d7d43e (branch dflash2), RTX 3090, greedy decoding, decode tokens per second:

plain drafted speedup tokens per block same transcript
20 held-out clips, 8 rows 409 tok/s 866 tok/s 2.12x 3.53 20/20
2-minute file, 8 rows 416 tok/s 1047 tok/s 2.52x 4.03 yes
32-minute file, 8 rows 341 tok/s 722 tok/s 2.12x 4.03 yes

Streaming sessions (22 + 4 frames a chunk). Plain = the same model with --draft none; --draft-check exact.

Inside

License

MIT, like VibeVoice.

Downloads last month
19
Safetensors
Model size
2B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ar4ikov/VibeVoice-ASR-Streaming-1.5B-AWQ-W4A16-ASYM-DFlash2

Collection including Ar4ikov/VibeVoice-ASR-Streaming-1.5B-AWQ-W4A16-ASYM-DFlash2