Text-to-Speech
VibeVoice
ONNX
English
audio
realtime
iamvaar's picture
Upload README.md with huggingface_hub
2401747 verified
|
Raw History Blame
1.56 kB
---
language:
- en
license: mit
tags:
- onnx
- audio
- text-to-speech
- realtime
- vibevoice
datasets:
- WenetHQ/LibriTTS
- WenetHQ/GigaTTS
---
# VibeVoice-Realtime-0.5B (ONNX Export)
This repository contains an ONNX-exported version of the `microsoft/VibeVoice-Realtime-0.5B` model.
This export was manually created to allow cross-platform inference in environments like ONNX Runtime Web (JavaScript) and Flutter (Dart).
## πŸ† Credits
All credit for the original model architecture, training, and base weights goes to the **Microsoft VibeVoice Team**.
Please see their original repository for full details and research:
- [Original Model Card](https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B)
- [VibeVoice GitHub Repository](https://github.com/microsoft/VibeVoice)
The original weights and software are licensed under the **MIT License**.
## πŸ“¦ What's included?
Due to the streaming nature of VibeVoice, the ONNX export is modularized into the following specific components (with accompanying `.data` files for external weights):
- `language_model.onnx`
- `tts_language_model.onnx`
- `tts_eos_classifier.onnx`
- `acoustic_tokenizer.onnx`
*Note: The `ir_version` for these models has been set to 9 to natively support standard Flutter `onnxruntime` bindings.*
## πŸš€ Usage
These models are optimized for [ONNX Runtime](https://onnxruntime.ai/). They can be loaded directly into client-side applications instead of maintaining heavy PyTorch backends. Check the corresponding JS and Flutter demo applications for integration guidance!