Any-to-Any
Transformers
ONNX
Safetensors
English
Chinese
multimodal
audio
video
speech
streaming
full-duplex
long-video
custom-code
Instructions to use inclusionAI/Realtime-Venus with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use inclusionAI/Realtime-Venus with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("inclusionAI/Realtime-Venus", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Center README headings and refresh Chinese translation
Browse files- README.md +11 -14
- README_zh.md +81 -136
README.md
CHANGED
|
@@ -16,26 +16,24 @@ tags:
|
|
| 16 |
- custom-code
|
| 17 |
---
|
| 18 |
|
| 19 |
-
<
|
| 20 |
<img src="./assets/venus-logo-white.gif" alt="Realtime-Venus logo" width="180">
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
# Realtime-Venus
|
| 24 |
|
| 25 |
-
|
| 26 |
|
| 27 |
-
|
| 28 |
|
| 29 |
-
<
|
| 30 |
|
|
|
|
| 31 |
<a href="https://realtime-venus.github.io/"><img src="https://img.shields.io/badge/Project_Page-4c9aff.svg?logo=googlechrome&logoColor=white" alt="Project Page"></a>
|
| 32 |
<a href="https://huggingface.co/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/Hugging_Face-Realtime--Venus-FFD21E.svg?logo=huggingface&logoColor=000" alt="Realtime-Venus on Hugging Face"></a>
|
| 33 |
<a href="https://www.modelscope.cn/models/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/ModelScope-Realtime--Venus-624AFF.svg?logo=modelscope&logoColor=white" alt="Realtime-Venus on ModelScope"></a>
|
| 34 |
<a href="https://arxiv.org/abs/2609.13814"><img src="https://img.shields.io/badge/arXiv-2609.13814-b31b1b.svg?logo=arxiv&logoColor=white" alt="arXiv"></a>
|
| 35 |
<a href="https://github.com/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/GitHub-Realtime--Venus-181717.svg?logo=github&logoColor=white" alt="GitHub"></a>
|
| 36 |
-
<a href="./LICENSE"><img src="https://img.shields.io/badge/License-Apache_2.0-0b7285.svg?logo=apache&logoColor=white" alt="Apache License 2.0"></a>
|
| 37 |
-
|
| 38 |
-
</div>
|
| 39 |
|
| 40 |
<p align="center">
|
| 41 |
<a href="https://arxiv.org/abs/2609.13814">
|
|
@@ -141,10 +139,9 @@ huggingface-cli download inclusionAI/Realtime-Venus --local-dir .
|
|
| 141 |
python -m pip install -r Realtime-Venus-Omni/requirements.txt
|
| 142 |
```
|
| 143 |
|
| 144 |
-
All example paths below are relative to
|
| 145 |
-
|
| 146 |
-
|
| 147 |
-
paths after the download.
|
| 148 |
|
| 149 |
## 7. 🎙️ Realtime-Venus-Omni Usages
|
| 150 |
|
|
|
|
| 16 |
- custom-code
|
| 17 |
---
|
| 18 |
|
| 19 |
+
<p align="center">
|
| 20 |
<img src="./assets/venus-logo-white.gif" alt="Realtime-Venus logo" width="180">
|
| 21 |
+
</p>
|
|
|
|
|
|
|
| 22 |
|
| 23 |
+
<h1 align="center" style="text-align: center;">Realtime-Venus</h1>
|
| 24 |
|
| 25 |
+
<p align="center" style="text-align: center;"><strong>A full-duplex interaction system with asynchronous delegation</strong></p>
|
| 26 |
|
| 27 |
+
<p align="center" style="text-align: center;"><strong>English</strong> | <a href="./README_zh.md">简体中文</a></p>
|
| 28 |
|
| 29 |
+
<p align="center">
|
| 30 |
<a href="https://realtime-venus.github.io/"><img src="https://img.shields.io/badge/Project_Page-4c9aff.svg?logo=googlechrome&logoColor=white" alt="Project Page"></a>
|
| 31 |
<a href="https://huggingface.co/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/Hugging_Face-Realtime--Venus-FFD21E.svg?logo=huggingface&logoColor=000" alt="Realtime-Venus on Hugging Face"></a>
|
| 32 |
<a href="https://www.modelscope.cn/models/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/ModelScope-Realtime--Venus-624AFF.svg?logo=modelscope&logoColor=white" alt="Realtime-Venus on ModelScope"></a>
|
| 33 |
<a href="https://arxiv.org/abs/2609.13814"><img src="https://img.shields.io/badge/arXiv-2609.13814-b31b1b.svg?logo=arxiv&logoColor=white" alt="arXiv"></a>
|
| 34 |
<a href="https://github.com/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/GitHub-Realtime--Venus-181717.svg?logo=github&logoColor=white" alt="GitHub"></a>
|
| 35 |
+
<a href="https://huggingface.co/inclusionAI/Realtime-Venus/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache_2.0-0b7285.svg?logo=apache&logoColor=white" alt="Apache License 2.0"></a>
|
| 36 |
+
</p>
|
|
|
|
| 37 |
|
| 38 |
<p align="center">
|
| 39 |
<a href="https://arxiv.org/abs/2609.13814">
|
|
|
|
| 139 |
python -m pip install -r Realtime-Venus-Omni/requirements.txt
|
| 140 |
```
|
| 141 |
|
| 142 |
+
All example paths below are relative to the downloaded repository's root
|
| 143 |
+
directory. The examples load the local `Realtime-Venus-Omni/` and
|
| 144 |
+
`Realtime-Venus-Audio/` checkpoints through Hugging Face Transformers.
|
|
|
|
| 145 |
|
| 146 |
## 7. 🎙️ Realtime-Venus-Omni Usages
|
| 147 |
|
README_zh.md
CHANGED
|
@@ -1,76 +1,45 @@
|
|
| 1 |
-
-
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
tags:
|
| 9 |
-
- multimodal
|
| 10 |
-
- audio
|
| 11 |
-
- video
|
| 12 |
-
- speech
|
| 13 |
-
- streaming
|
| 14 |
-
- full-duplex
|
| 15 |
-
- long-video
|
| 16 |
-
- custom-code
|
| 17 |
-
---
|
| 18 |
-
|
| 19 |
-
<div align="center">
|
| 20 |
-
<img src="./assets/venus-logo-white.gif" alt="Realtime-Venus logo" width="180">
|
| 21 |
-
<br>
|
| 22 |
-
|
| 23 |
-
# Realtime-Venus
|
| 24 |
-
|
| 25 |
-
**支持异步委派的全双工交互系统**
|
| 26 |
-
|
| 27 |
-
[English](./README.md) | **简体中文**
|
| 28 |
-
|
| 29 |
-
<br>
|
| 30 |
|
|
|
|
|
|
|
|
|
|
| 31 |
<a href="https://realtime-venus.github.io/"><img src="https://img.shields.io/badge/Project_Page-4c9aff.svg?logo=googlechrome&logoColor=white" alt="项目主页"></a>
|
| 32 |
<a href="https://huggingface.co/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/Hugging_Face-Realtime--Venus-FFD21E.svg?logo=huggingface&logoColor=000" alt="Hugging Face 上的 Realtime-Venus"></a>
|
| 33 |
<a href="https://www.modelscope.cn/models/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/ModelScope-Realtime--Venus-624AFF.svg?logo=modelscope&logoColor=white" alt="ModelScope 上的 Realtime-Venus"></a>
|
| 34 |
-
<a href="https://arxiv.org/abs/2609.13814"><img src="https://img.shields.io/badge/arXiv-2609.13814-b31b1b.svg?logo=arxiv&logoColor=white" alt="arXiv"></a>
|
| 35 |
-
<a href="https://github.com/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/GitHub-Realtime--Venus-181717.svg?logo=github&logoColor=white" alt="GitHub"></a>
|
| 36 |
-
<a href="./LICENSE"><img src="https://img.shields.io/badge/License-Apache_2.0-0b7285.svg?logo=apache&logoColor=white" alt="Apache License 2.0"></a>
|
| 37 |
-
|
| 38 |
-
</div>
|
| 39 |
|
| 40 |
<p align="center">
|
| 41 |
<a href="https://arxiv.org/abs/2609.13814">
|
| 42 |
-
<img src="https://arxiv.org/html/2609.13814v1/case.png" alt="Realtime-Venus 的主动交互、异步委派与全双工交互示例" width="100%">
|
| 43 |
</a>
|
| 44 |
</p>
|
| 45 |
-
<p align="center"><em>Realtime-Venus 支持主动
|
| 46 |
|
| 47 |
## 1. 🧭 概览
|
| 48 |
|
| 49 |
-
本仓库
|
| 50 |
|
| 51 |
-
- **Realtime-Venus-Omni**(`Realtime-Venus-Omni/`):9B 音视频交互模型。模型
|
| 52 |
-
|
| 53 |
-
MiniCPM-o 4.5 改造,支持主动式交互、语义级打断处理,以及免训练的长视频 Memory。
|
| 54 |
-
- **Realtime-Venus-Audio**(`Realtime-Venus-Audio/`):基于同一流式骨干的音频模型,
|
| 55 |
-
用于不包含视觉输入的音频理解与音频对话,输出文本或语音。
|
| 56 |
|
| 57 |
-
两个目录
|
| 58 |
-
Realtime-Venus-Harness 及其外部工具集成见
|
| 59 |
-
[GitHub 仓库](https://github.com/inclusionAI/Realtime-Venus)。
|
| 60 |
|
| 61 |
## 2. ✨ 核心亮点
|
| 62 |
|
| 63 |
-
- **原生全双工对话:** 在说话
|
| 64 |
-
- **Omni-Proactive 主动交互:** 持续处理时间对齐的视频和音频,在
|
| 65 |
-
|
| 66 |
-
- **
|
| 67 |
-
|
| 68 |
-
Realtime-Venus-Harness 运行时,见
|
| 69 |
-
[GitHub 仓库](https://github.com/inclusionAI/Realtime-Venus)。)
|
| 70 |
-
- **免训练长视频 Memory:** 归档视觉信息丰富的时刻,检索与问题相关且不冗余的证据,
|
| 71 |
-
并重新组装对应的音视频上下文,无需额外训练模型。
|
| 72 |
-
- **文本与语音输出:** 通过随仓库提供的 Token2wav 资源和参考音色,同时生成回复文本
|
| 73 |
-
与原生语音。
|
| 74 |
|
| 75 |
## 3. 📋 模型信息
|
| 76 |
|
|
@@ -81,19 +50,19 @@ Realtime-Venus-Harness 及其外部工具集成见
|
|
| 81 |
| 视觉编码器 | SigLIP2 | 推理时不使用 |
|
| 82 |
| 音频编码器 | Whisper-Medium | Whisper-Medium |
|
| 83 |
| 语言模型骨干 | Qwen3-8B | Qwen3-8B |
|
| 84 |
-
| 语音生成 | 离散 S3 语音 token
|
| 85 |
| 输入 | 视频/图像、音频和文本 | 音频和文本 |
|
| 86 |
-
| 输出 | 文本及可选语音波形 | 文本和语音波形 |
|
| 87 |
| 上下文长度 | 40,960 tokens | 40,960 tokens |
|
| 88 |
| 权重精度 | BF16 | BF16 |
|
| 89 |
|
| 90 |
## 4. 📊 评测
|
| 91 |
|
| 92 |
-
下
|
| 93 |
|
| 94 |
-
<p align="center"><img src="assets/paper-understanding.svg" width="100%" alt="Omni 视频理解与 Audio 音频理解的雷达图
|
| 95 |
|
| 96 |
-
<p align="center"><img src="assets/paper-duplex.svg" width="100%" alt="不同
|
| 97 |
|
| 98 |
## 5. 🗂️ 仓库结构
|
| 99 |
|
|
@@ -102,26 +71,25 @@ Realtime-Venus-Harness 及其外部工具集成见
|
|
| 102 |
├── Realtime-Venus-Omni/ # 音视频全双工模型
|
| 103 |
│ ├── model-*.safetensors # 分片模型权重
|
| 104 |
│ ├── config.json, *.py # 模型配置与自定义 Transformers 代码
|
| 105 |
-
│ ├── realtime_venus_omni_memory.py #
|
| 106 |
-
│ ├── memory_adapter/ #
|
| 107 |
│ ├── assets/ # 参考音色、Token2wav、演示视频
|
| 108 |
│ └── requirements.txt
|
| 109 |
-
├── Realtime-Venus-Audio/ # 音频模型
|
| 110 |
│ ├── model-*.safetensors # 分片模型权重
|
| 111 |
│ ├── config.json, *.py # 模型配置与自定义 Transformers 代码
|
| 112 |
│ └── assets/ # 参考音色、Token2wav、演示音频
|
| 113 |
-
├── assets/ #
|
| 114 |
├── README.md
|
| 115 |
├── README_zh.md
|
| 116 |
└── LICENSE
|
| 117 |
```
|
| 118 |
|
| 119 |
-
|
| 120 |
|
| 121 |
## 6. 🛠️ 安装
|
| 122 |
|
| 123 |
-
需要 Python 3.10、CUDA 和 FFmpeg。先下载仓库(两个模型分别位于
|
| 124 |
-
再安装 Python 依赖:
|
| 125 |
|
| 126 |
```bash
|
| 127 |
huggingface-cli download inclusionAI/Realtime-Venus --local-dir .
|
|
@@ -130,18 +98,15 @@ huggingface-cli download inclusionAI/Realtime-Venus --local-dir .
|
|
| 130 |
python -m pip install -r Realtime-Venus-Omni/requirements.txt
|
| 131 |
```
|
| 132 |
|
| 133 |
-
下文所有示例路径均相对于
|
| 134 |
-
仓库的子目录,因此示例在下载完成后指向本地的 `Realtime-Venus-Omni/` 与
|
| 135 |
-
`Realtime-Venus-Audio/` 路径。
|
| 136 |
|
| 137 |
## 7. 🎙️ Realtime-Venus-Omni 使用方法
|
| 138 |
|
| 139 |
-
这些示例的
|
| 140 |
-
[Omni cookbook](https://github.com/inclusionAI/Realtime-Venus/tree/main/frontend/Realtime-Venus-Omni)。
|
| 141 |
|
| 142 |
### 7.1 🧱 模型初始化
|
| 143 |
|
| 144 |
-
|
| 145 |
|
| 146 |
<details>
|
| 147 |
<summary>点击展开 Omni 模型加载代码。</summary>
|
|
@@ -156,7 +121,7 @@ Path("output").mkdir(exist_ok=True)
|
|
| 156 |
set_seed(42)
|
| 157 |
print("Loading model ...")
|
| 158 |
model = AutoModel.from_pretrained(
|
| 159 |
-
"./Realtime-Venus-Omni", #
|
| 160 |
trust_remote_code=True,
|
| 161 |
local_files_only=True,
|
| 162 |
attn_implementation="sdpa",
|
|
@@ -170,14 +135,9 @@ print("Model loaded.")
|
|
| 170 |
|
| 171 |
### 7.2 🔊 全双工 Omni 模式
|
| 172 |
|
| 173 |
-
`model = model.as_duplex()`
|
| 174 |
-
随后每秒输入由一次 `streaming_prefill()` 与 `streaming_generate()` 处理;
|
| 175 |
-
`as_simplex()` 切换回离线模式。必须在导入 `minicpmo.utils` 前设置
|
| 176 |
-
`MAX_NUM_FRAMES`,否则超过 64 秒的视频会被默认帧数上限截断。
|
| 177 |
|
| 178 |
-
字幕字体说明:
|
| 179 |
-
fontconfig 解析。若系统没有支持中日韩字符的字体,中文等非拉丁文字会显示为空白方框。
|
| 180 |
-
在任意 Linux 发行版中均可无需 root 权限安装字体并刷新缓存:
|
| 181 |
|
| 182 |
```bash
|
| 183 |
mkdir -p ~/.local/share/fonts
|
|
@@ -187,17 +147,14 @@ curl --fail --location --retry 3 \
|
|
| 187 |
fc-cache -f
|
| 188 |
```
|
| 189 |
|
| 190 |
-
也可以通过包管理器安装:Debian/Ubuntu 使用 `apt install -y fonts-noto-cjk`;
|
| 191 |
-
RHEL/Alibaba Cloud Linux 使用
|
| 192 |
-
`yum install -y cjkuni-ukai-fonts cjkuni-uming-fonts`。无需修改模型代码。
|
| 193 |
|
| 194 |
-
#### 7.2.1 全双工
|
| 195 |
|
| 196 |
-
逐秒流式输入演示视频,并在 `question_times` 指定的
|
| 197 |
-
文本问题。模型会持续聆听,并在回答时生成语音。
|
| 198 |
|
| 199 |
<details>
|
| 200 |
-
<summary>点击展开全双工
|
| 201 |
|
| 202 |
```python
|
| 203 |
import os
|
|
@@ -206,11 +163,11 @@ os.environ["MAX_NUM_FRAMES"] = "100000"
|
|
| 206 |
|
| 207 |
from minicpmo.utils import get_video_frame_audio_segments, generate_duplex_video
|
| 208 |
|
| 209 |
-
model = model.as_duplex() #
|
| 210 |
model.prepare()
|
| 211 |
|
| 212 |
video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
|
| 213 |
-
#
|
| 214 |
question_times = [60, 128]
|
| 215 |
questions = [
|
| 216 |
"What do you see in the video so far?",
|
|
@@ -251,13 +208,12 @@ generate_duplex_video(
|
|
| 251 |
|
| 252 |
</details>
|
| 253 |
|
| 254 |
-
#### 7.2.2 语音输入全双工
|
| 255 |
|
| 256 |
-
|
| 257 |
-
因此不注入文本,模型必须直接听取问题。
|
| 258 |
|
| 259 |
<details>
|
| 260 |
-
<summary>点击展开语音输入全双工
|
| 261 |
|
| 262 |
```python
|
| 263 |
import os
|
|
@@ -266,7 +222,7 @@ os.environ["MAX_NUM_FRAMES"] = "100000"
|
|
| 266 |
|
| 267 |
from minicpmo.utils import get_video_frame_audio_segments, generate_duplex_video
|
| 268 |
|
| 269 |
-
model = model.as_duplex() #
|
| 270 |
model.prepare()
|
| 271 |
|
| 272 |
video_path = "Realtime-Venus-Omni/assets/speech_in.mp4"
|
|
@@ -303,12 +259,12 @@ generate_duplex_video(
|
|
| 303 |
|
| 304 |
</details>
|
| 305 |
|
| 306 |
-
#### 7.2.3
|
| 307 |
|
| 308 |
-
在进入
|
| 309 |
|
| 310 |
<details>
|
| 311 |
-
<summary>点击展开
|
| 312 |
|
| 313 |
```python
|
| 314 |
import os
|
|
@@ -317,8 +273,8 @@ os.environ["MAX_NUM_FRAMES"] = "100000"
|
|
| 317 |
|
| 318 |
from minicpmo.utils import get_video_frame_audio_segments, generate_duplex_video
|
| 319 |
|
| 320 |
-
model.use_memory(memory_minutes=40) #
|
| 321 |
-
model = model.as_duplex() #
|
| 322 |
model.prepare()
|
| 323 |
|
| 324 |
video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
|
|
@@ -359,16 +315,14 @@ generate_duplex_video(
|
|
| 359 |
|
| 360 |
### 7.3 💬 半双工 Omni 模式
|
| 361 |
|
| 362 |
-
`model.chat(...)` 以完整视频为输入,
|
| 363 |
|
| 364 |
-
#### 7.3.1 离线
|
| 365 |
|
| 366 |
-
采样后的视频帧、逐秒音频和问题
|
| 367 |
-
(`MAX_NUM_FRAMES`)用于控制视觉输入规模,`max_inp_length=32768` 用于设置输入
|
| 368 |
-
token 预算。完整音频仍会保留,因此即使进行视频帧采样,超长视频仍可能超过该预算。
|
| 369 |
|
| 370 |
<details>
|
| 371 |
-
<summary>点击展开离线
|
| 372 |
|
| 373 |
```python
|
| 374 |
import os
|
|
@@ -377,7 +331,7 @@ os.environ.setdefault("MAX_NUM_FRAMES", "128")
|
|
| 377 |
|
| 378 |
from minicpmo.utils import get_video_frame_audio_segments
|
| 379 |
|
| 380 |
-
model.init_tts() #
|
| 381 |
|
| 382 |
video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
|
| 383 |
question = "What is the color of the cooler labeled PRIME near the team bench?"
|
|
@@ -412,13 +366,12 @@ print(response)
|
|
| 412 |
|
| 413 |
</details>
|
| 414 |
|
| 415 |
-
#### 7.3.2
|
| 416 |
|
| 417 |
-
在调用
|
| 418 |
-
4 个近期帧,并为每个选中帧保留前后各 1 秒的音频。
|
| 419 |
|
| 420 |
<details>
|
| 421 |
-
<summary>点击展开
|
| 422 |
|
| 423 |
```python
|
| 424 |
import os
|
|
@@ -427,8 +380,8 @@ os.environ["MAX_NUM_FRAMES"] = "100000"
|
|
| 427 |
|
| 428 |
from minicpmo.utils import get_video_frame_audio_segments
|
| 429 |
|
| 430 |
-
model.use_memory() #
|
| 431 |
-
model.init_tts() #
|
| 432 |
|
| 433 |
video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
|
| 434 |
question = "What is the color of the cooler labeled PRIME near the team bench?"
|
|
@@ -465,17 +418,13 @@ print(response)
|
|
| 465 |
|
| 466 |
## 8. 🎧 Realtime-Venus-Audio 使用方法
|
| 467 |
|
| 468 |
-
这些示例的
|
| 469 |
-
[Audio cookbook](https://github.com/inclusionAI/Realtime-Venus/tree/main/frontend/Realtime-Venus-Audio)。
|
| 470 |
|
| 471 |
-
Audio 模型
|
| 472 |
-
全双工流式 API(语音回复)。输入统一解码为 16 kHz 单声道音频,可来自任意
|
| 473 |
-
音频或视频文件。
|
| 474 |
|
| 475 |
### 8.1 🧱 模型初始化
|
| 476 |
|
| 477 |
-
下方代码
|
| 478 |
-
纯文本对话可改用 `init_tts=False`,加载更快。
|
| 479 |
|
| 480 |
<details>
|
| 481 |
<summary>点击展开 Audio 模型加载代码。</summary>
|
|
@@ -499,22 +448,21 @@ model = AutoModel.from_pretrained(
|
|
| 499 |
local_files_only=True,
|
| 500 |
attn_implementation="sdpa",
|
| 501 |
torch_dtype=torch.bfloat16,
|
| 502 |
-
init_vision=False, #
|
| 503 |
init_audio=True,
|
| 504 |
-
init_tts=True, #
|
| 505 |
).eval().cuda()
|
| 506 |
print("Model loaded.")
|
| 507 |
```
|
| 508 |
|
| 509 |
</details>
|
| 510 |
|
| 511 |
-
### 8.2 💭 离线
|
| 512 |
|
| 513 |
-
对完整音频输入进行一轮确定性
|
| 514 |
-
`model.chat()` 调用传入。
|
| 515 |
|
| 516 |
<details>
|
| 517 |
-
<summary>点击展开离线
|
| 518 |
|
| 519 |
```python
|
| 520 |
import librosa
|
|
@@ -543,21 +491,19 @@ print(answer)
|
|
| 543 |
|
| 544 |
</details>
|
| 545 |
|
| 546 |
-
### 8.3 🎙️ 全双工
|
| 547 |
|
| 548 |
-
`model.as_duplex(generate_audio=True)` 切换为全双工流式模式:音频
|
| 549 |
-
模型持续聆听、在回答时开口说话。示例在输入末尾追加 10 秒静音,让模型在
|
| 550 |
-
输入结束后说完回复,生成的语音写入 `output/audio_full_duplex.wav`。
|
| 551 |
|
| 552 |
<details>
|
| 553 |
-
<summary>点击展开全双工
|
| 554 |
|
| 555 |
```python
|
| 556 |
import librosa
|
| 557 |
import numpy as np
|
| 558 |
import soundfile as sf
|
| 559 |
|
| 560 |
-
duplex = model.as_duplex(generate_audio=True) #
|
| 561 |
duplex.prepare(prompt_wav_path="Realtime-Venus-Audio/assets/HT_ref_audio.wav")
|
| 562 |
|
| 563 |
audio, _ = librosa.load(
|
|
@@ -586,7 +532,7 @@ for chunk_index in range(total_chunks):
|
|
| 586 |
if result["audio_waveform"] is not None and not result["is_listen"]:
|
| 587 |
timed_audio.append((chunk_index, result["audio_waveform"]))
|
| 588 |
|
| 589 |
-
#
|
| 590 |
sample_rate = 24000
|
| 591 |
total_samples = max(
|
| 592 |
t * sample_rate + len(np.asarray(w, dtype=np.float32).squeeze())
|
|
@@ -617,5 +563,4 @@ print("Saved generated speech to output/audio_full_duplex.wav")
|
|
| 617 |
|
| 618 |
## 10. 📄 许可证
|
| 619 |
|
| 620 |
-
本仓库
|
| 621 |
-
所涉及数据的许可证与可接受使用条款。
|
|
|
|
| 1 |
+
<p align="center" style="text-align: center;">
|
| 2 |
+
<img src="./assets/venus-logo-white.gif" alt="Realtime-Venus 标志" width="180">
|
| 3 |
+
</p>
|
| 4 |
+
|
| 5 |
+
<h1 align="center" style="text-align: center;">Realtime-Venus</h1>
|
| 6 |
+
|
| 7 |
+
<p align="center" style="text-align: center;"><strong>支持异步任务委派的全双工交互系统</strong></p>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
|
| 9 |
+
<p align="center" style="text-align: center;"><a href="./README.md">English</a> | <strong>简体中文</strong></p>
|
| 10 |
+
|
| 11 |
+
<p align="center" style="text-align: center;">
|
| 12 |
<a href="https://realtime-venus.github.io/"><img src="https://img.shields.io/badge/Project_Page-4c9aff.svg?logo=googlechrome&logoColor=white" alt="项目主页"></a>
|
| 13 |
<a href="https://huggingface.co/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/Hugging_Face-Realtime--Venus-FFD21E.svg?logo=huggingface&logoColor=000" alt="Hugging Face 上的 Realtime-Venus"></a>
|
| 14 |
<a href="https://www.modelscope.cn/models/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/ModelScope-Realtime--Venus-624AFF.svg?logo=modelscope&logoColor=white" alt="ModelScope 上的 Realtime-Venus"></a>
|
| 15 |
+
<a href="https://arxiv.org/abs/2609.13814"><img src="https://img.shields.io/badge/arXiv-2609.13814-b31b1b.svg?logo=arxiv&logoColor=white" alt="arXiv 论文"></a>
|
| 16 |
+
<a href="https://github.com/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/GitHub-Realtime--Venus-181717.svg?logo=github&logoColor=white" alt="GitHub 代码仓库"></a>
|
| 17 |
+
<a href="https://huggingface.co/inclusionAI/Realtime-Venus/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache_2.0-0b7285.svg?logo=apache&logoColor=white" alt="Apache License 2.0"></a>
|
| 18 |
+
</p>
|
|
|
|
| 19 |
|
| 20 |
<p align="center">
|
| 21 |
<a href="https://arxiv.org/abs/2609.13814">
|
| 22 |
+
<img src="https://arxiv.org/html/2609.13814v1/case.png" alt="Realtime-Venus 的主动交互、异步任务委派与全双工交互示例" width="100%">
|
| 23 |
</a>
|
| 24 |
</p>
|
| 25 |
+
<p align="center"><em>Realtime-Venus 支持主动音视频交互、异步任务委派,以及能够感知打断的全双工对话。</em></p>
|
| 26 |
|
| 27 |
## 1. 🧭 概览
|
| 28 |
|
| 29 |
+
本仓库提供 [Realtime-Venus](https://realtime-venus.github.io/) 系统的两个模型检查点:
|
| 30 |
|
| 31 |
+
- **Realtime-Venus-Omni**(`Realtime-Venus-Omni/`):9B 音视频交互模型。它持续观看和聆听,判断是否响应以及何时响应,并在共享的因果时间线上生成文本和语音。该模型基于 MiniCPM-o 4.5 改造,支持主动交互、语义级打断处理,以及无需额外训练的长视频记忆(Memory)。
|
| 32 |
+
- **Realtime-Venus-Audio**(`Realtime-Venus-Audio/`):专注于音频的模型检查点,采用相同的流式骨干网络,用于音频理解和音频驱动的对话,可输出文本或语音。
|
|
|
|
|
|
|
|
|
|
| 33 |
|
| 34 |
+
两个目录均包含模型权重与自定义 Hugging Face Transformers 代码。异步运行的 Realtime-Venus-Harness 及其外部工具集成位于 [GitHub 仓库](https://github.com/inclusionAI/Realtime-Venus)。
|
|
|
|
|
|
|
| 35 |
|
| 36 |
## 2. ✨ 核心亮点
|
| 37 |
|
| 38 |
+
- **原生全双工对话:** 在说话时持续感知输入,能够区分简短附和、打断、纠正和话题转向。
|
| 39 |
+
- **Omni-Proactive 主动交互:** 持续处理时间对齐的视频和音频,在事件需要响应时主动发言,无需等待用户提示。
|
| 40 |
+
- **异步任务委派:** 在共享的因果时间线上发出流内 `<delegate>` 请求,并以相同方式接收后端异步返回的结果,使外部任务不会阻塞正在进行的对话。(执行这些请求需要 Realtime-Venus-Harness 运行时,详见 [GitHub 仓库](https://github.com/inclusionAI/Realtime-Venus)。)
|
| 41 |
+
- **无需额外训练的长视频记忆:** 保存视觉信息丰富的时刻,检索与问题相关且不重复的证据,并重新组装对应的音视频上下文,无需额外训练。
|
| 42 |
+
- **文本与语音输出:** 结合仓库内附带的 Token2wav 资源与参考音色,在生成回复文本的同时生成原生语音。
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
|
| 44 |
## 3. 📋 模型信息
|
| 45 |
|
|
|
|
| 50 |
| 视觉编码器 | SigLIP2 | 推理时不使用 |
|
| 51 |
| 音频编码器 | Whisper-Medium | Whisper-Medium |
|
| 52 |
| 语言模型骨干 | Qwen3-8B | Qwen3-8B |
|
| 53 |
+
| 语音生成 | 离散 S3 语音 token,配合流式流匹配解码器 | 使用相同解码器,在全双工模式下启用 |
|
| 54 |
| 输入 | 视频/图像、音频和文本 | 音频和文本 |
|
| 55 |
+
| 输出 | 文本,以及可选的语音波形 | 文本和语音波形 |
|
| 56 |
| 上下文长度 | 40,960 tokens | 40,960 tokens |
|
| 57 |
| 权重精度 | BF16 | BF16 |
|
| 58 |
|
| 59 |
## 4. 📊 评测
|
| 60 |
|
| 61 |
+
以下结果均引自 [Realtime-Venus 技术报告](https://arxiv.org/abs/2609.13814)。
|
| 62 |
|
| 63 |
+
<p align="center"><img src="assets/paper-understanding.svg" width="100%" alt="比较 Omni 视频理解能力与 Audio 音频理解能力的雷达图" /><br /><sub>图 1. <a href="https://arxiv.org/html/2609.13814v1#S0.F1">论文</a>中的视频和音频理解评测结果。</sub></p>
|
| 64 |
|
| 65 |
+
<p align="center"><img src="assets/paper-duplex.svg" width="100%" alt="不同语音重叠场景下打断处理与继续发言能力的全双工基准对比" /><br /><sub>图 2. <a href="https://arxiv.org/html/2609.13814v1#S0.F2">论文</a>中的全双工交互评测结果。</sub></p>
|
| 66 |
|
| 67 |
## 5. 🗂️ 仓库结构
|
| 68 |
|
|
|
|
| 71 |
├── Realtime-Venus-Omni/ # 音视频全双工模型
|
| 72 |
│ ├── model-*.safetensors # 分片模型权重
|
| 73 |
│ ├── config.json, *.py # 模型配置与自定义 Transformers 代码
|
| 74 |
+
│ ├── realtime_venus_omni_memory.py # 记忆功能的公开入口
|
| 75 |
+
│ ├── memory_adapter/ # 对话与全双工模式的记忆运行时
|
| 76 |
│ ├── assets/ # 参考音色、Token2wav、演示视频
|
| 77 |
│ └── requirements.txt
|
| 78 |
+
├── Realtime-Venus-Audio/ # 专注于音频的模型
|
| 79 |
│ ├── model-*.safetensors # 分片模型权重
|
| 80 |
│ ├── config.json, *.py # 模型配置与自定义 Transformers 代码
|
| 81 |
│ └── assets/ # 参考音色、Token2wav、演示音频
|
| 82 |
+
├── assets/ # 品牌资源(标志)
|
| 83 |
├── README.md
|
| 84 |
├── README_zh.md
|
| 85 |
└── LICENSE
|
| 86 |
```
|
| 87 |
|
| 88 |
+
下文示例将生成的媒体文件写入 `output/`。重复运行实验时,请使用新的文件名或输出目录。
|
| 89 |
|
| 90 |
## 6. 🛠️ 安装
|
| 91 |
|
| 92 |
+
需要 Python 3.10、CUDA 和 FFmpeg。先下载仓库(两个模型检查点分别位于各自的子目录中),再安装 Python 依赖:
|
|
|
|
| 93 |
|
| 94 |
```bash
|
| 95 |
huggingface-cli download inclusionAI/Realtime-Venus --local-dir .
|
|
|
|
| 98 |
python -m pip install -r Realtime-Venus-Omni/requirements.txt
|
| 99 |
```
|
| 100 |
|
| 101 |
+
下文所有示例路径均相对于下载后的仓库根目录。示例通过 Hugging Face Transformers 加载本地的 `Realtime-Venus-Omni/` 与 `Realtime-Venus-Audio/` 模型。
|
|
|
|
|
|
|
| 102 |
|
| 103 |
## 7. 🎙️ Realtime-Venus-Omni 使用方法
|
| 104 |
|
| 105 |
+
这些示例的独立可运行脚本位于 GitHub 上的 [Omni cookbook](https://github.com/inclusionAI/Realtime-Venus/tree/main/frontend/Realtime-Venus-Omni)。
|
|
|
|
| 106 |
|
| 107 |
### 7.1 🧱 模型初始化
|
| 108 |
|
| 109 |
+
下文示例共用以下模型初始化代码;请在新的 Python 进程中分别运行每个示例。普通对话和全双工对话均会自动加载默认参考音色。
|
| 110 |
|
| 111 |
<details>
|
| 112 |
<summary>点击展开 Omni 模型加载代码。</summary>
|
|
|
|
| 121 |
set_seed(42)
|
| 122 |
print("Loading model ...")
|
| 123 |
model = AutoModel.from_pretrained(
|
| 124 |
+
"./Realtime-Venus-Omni", # or an absolute path to the sub-directory
|
| 125 |
trust_remote_code=True,
|
| 126 |
local_files_only=True,
|
| 127 |
attn_implementation="sdpa",
|
|
|
|
| 135 |
|
| 136 |
### 7.2 🔊 全双工 Omni 模式
|
| 137 |
|
| 138 |
+
调用 `model = model.as_duplex()` 可切换为全双工流式模式:`prepare()` 初始化会话,随后每秒输入由一对 `streaming_prefill()` 和 `streaming_generate()` 调用处理;`as_simplex()` 则切换回离线模式。请在导入 `minicpmo.utils` 之前设置 `MAX_NUM_FRAMES`,否则超过 64 秒的视频会受到默认帧数上限的截断。
|
|
|
|
|
|
|
|
|
|
| 139 |
|
| 140 |
+
字幕字体说明:全双工示例通过 FFmpeg/libass 将回复文本烧录到输出视频中,字体由 fontconfig 查找。渲染中文等非拉丁文字需要系统安装支持中日韩字符的字体,否则相应字符会显示为空白方框。在 Linux 发行版中,可以无需 root 权限安装字体并刷新字体缓存:
|
|
|
|
|
|
|
| 141 |
|
| 142 |
```bash
|
| 143 |
mkdir -p ~/.local/share/fonts
|
|
|
|
| 147 |
fc-cache -f
|
| 148 |
```
|
| 149 |
|
| 150 |
+
也可以通过包管理器安装:Debian/Ubuntu 使用 `apt install -y fonts-noto-cjk`;RHEL/Alibaba Cloud Linux 使用 `yum install -y cjkuni-ukai-fonts cjkuni-uming-fonts`。无需修改代码。
|
|
|
|
|
|
|
| 151 |
|
| 152 |
+
#### 7.2.1 全双工对话
|
| 153 |
|
| 154 |
+
逐秒流式输入演示视频,并在 `question_times` 指定的时间注入与 `questions` 一一对应的文本问题。模型持续聆听,并在回答时生成语音。
|
|
|
|
| 155 |
|
| 156 |
<details>
|
| 157 |
+
<summary>点击展开全双工对话代码。</summary>
|
| 158 |
|
| 159 |
```python
|
| 160 |
import os
|
|
|
|
| 163 |
|
| 164 |
from minicpmo.utils import get_video_frame_audio_segments, generate_duplex_video
|
| 165 |
|
| 166 |
+
model = model.as_duplex() # switch to full-duplex streaming
|
| 167 |
model.prepare()
|
| 168 |
|
| 169 |
video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
|
| 170 |
+
# each question is injected at the corresponding second
|
| 171 |
question_times = [60, 128]
|
| 172 |
questions = [
|
| 173 |
"What do you see in the video so far?",
|
|
|
|
| 208 |
|
| 209 |
</details>
|
| 210 |
|
| 211 |
+
#### 7.2.2 语音输入的全双工对话
|
| 212 |
|
| 213 |
+
流程与上例相同,但问题以语音形式预先混入视频音轨(约第 3 秒,请求在水烧开时发出提醒)。因此,示例不注入文本问题,模型必须从音频中听取问题。
|
|
|
|
| 214 |
|
| 215 |
<details>
|
| 216 |
+
<summary>点击展开语音输入的全双工对话代码。</summary>
|
| 217 |
|
| 218 |
```python
|
| 219 |
import os
|
|
|
|
| 222 |
|
| 223 |
from minicpmo.utils import get_video_frame_audio_segments, generate_duplex_video
|
| 224 |
|
| 225 |
+
model = model.as_duplex() # switch to full-duplex streaming
|
| 226 |
model.prepare()
|
| 227 |
|
| 228 |
video_path = "Realtime-Venus-Omni/assets/speech_in.mp4"
|
|
|
|
| 259 |
|
| 260 |
</details>
|
| 261 |
|
| 262 |
+
#### 7.2.3 带记忆的全双工对话
|
| 263 |
|
| 264 |
+
在进入全双工模式前调用 `model.use_memory(memory_minutes=40)`,即可启用长视频记忆。
|
| 265 |
|
| 266 |
<details>
|
| 267 |
+
<summary>点击展开带记忆的全双工对话代码。</summary>
|
| 268 |
|
| 269 |
```python
|
| 270 |
import os
|
|
|
|
| 273 |
|
| 274 |
from minicpmo.utils import get_video_frame_audio_segments, generate_duplex_video
|
| 275 |
|
| 276 |
+
model.use_memory(memory_minutes=40) # enable long-video memory
|
| 277 |
+
model = model.as_duplex() # switch to full-duplex streaming
|
| 278 |
model.prepare()
|
| 279 |
|
| 280 |
video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
|
|
|
|
| 315 |
|
| 316 |
### 7.3 💬 半双工 Omni 模式
|
| 317 |
|
| 318 |
+
`model.chat(...)` 以完整视频为输入,每次回答一轮问题。`model.init_tts()` 用于启用语音输出。
|
| 319 |
|
| 320 |
+
#### 7.3.1 离线对话
|
| 321 |
|
| 322 |
+
将采样后的视频帧、逐秒音频和问题一起传入一次 `chat()` 调用。128 帧上限(`MAX_NUM_FRAMES`)用于控制视觉输入规模,`max_inp_length=32768` 用于设置输入 token 预算。完整音频仍会保留,因此即使对视频帧进行了采样,超长视频也可能超过该预算。
|
|
|
|
|
|
|
| 323 |
|
| 324 |
<details>
|
| 325 |
+
<summary>点击展开离线对话代码。</summary>
|
| 326 |
|
| 327 |
```python
|
| 328 |
import os
|
|
|
|
| 331 |
|
| 332 |
from minicpmo.utils import get_video_frame_audio_segments
|
| 333 |
|
| 334 |
+
model.init_tts() # enable speech output
|
| 335 |
|
| 336 |
video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
|
| 337 |
question = "What is the color of the cooler labeled PRIME near the team bench?"
|
|
|
|
| 366 |
|
| 367 |
</details>
|
| 368 |
|
| 369 |
+
#### 7.3.2 带记忆的离线对话
|
| 370 |
|
| 371 |
+
在调用对话接口前,通过 `model.use_memory()` 启用记忆功能。检索最多选择 96 个历史帧和 4 个近期帧,并为每个选中帧保留前后各 1 秒的音频。
|
|
|
|
| 372 |
|
| 373 |
<details>
|
| 374 |
+
<summary>点击展开带记忆的离线对话代码。</summary>
|
| 375 |
|
| 376 |
```python
|
| 377 |
import os
|
|
|
|
| 380 |
|
| 381 |
from minicpmo.utils import get_video_frame_audio_segments
|
| 382 |
|
| 383 |
+
model.use_memory() # enable long-video memory
|
| 384 |
+
model.init_tts() # enable speech output
|
| 385 |
|
| 386 |
video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
|
| 387 |
question = "What is the color of the cooler labeled PRIME near the team bench?"
|
|
|
|
| 418 |
|
| 419 |
## 8. 🎧 Realtime-Venus-Audio 使用方法
|
| 420 |
|
| 421 |
+
这些示例的独立可运行脚本位于 GitHub 上的 [Audio cookbook](https://github.com/inclusionAI/Realtime-Venus/tree/main/frontend/Realtime-Venus-Audio)。
|
|
|
|
| 422 |
|
| 423 |
+
Audio 模型支持两种纯音频推理方式:逐轮调用 `model.chat` 生成文本回复,以及使用全双工流式 API 生成语音回复。输入可来自音频或视频文件,统一解码为 16 kHz 单声道音频。
|
|
|
|
|
|
|
| 424 |
|
| 425 |
### 8.1 🧱 模型初始化
|
| 426 |
|
| 427 |
+
下方代码通过 `init_tts=True` 启用语音输出,使同一个 `model` 可用于后文的两个示例。若只需要文本回复,可设置 `init_tts=False` 以加快加载。
|
|
|
|
| 428 |
|
| 429 |
<details>
|
| 430 |
<summary>点击展开 Audio 模型加载代码。</summary>
|
|
|
|
| 448 |
local_files_only=True,
|
| 449 |
attn_implementation="sdpa",
|
| 450 |
torch_dtype=torch.bfloat16,
|
| 451 |
+
init_vision=False, # audio-only usage
|
| 452 |
init_audio=True,
|
| 453 |
+
init_tts=True, # speech output; set False for text-only chat
|
| 454 |
).eval().cuda()
|
| 455 |
print("Model loaded.")
|
| 456 |
```
|
| 457 |
|
| 458 |
</details>
|
| 459 |
|
| 460 |
+
### 8.2 💭 离线对话
|
| 461 |
|
| 462 |
+
对完整音频输入进行一轮确定性推理:将音频(以及可选的文本指令)传入一次 `model.chat()` 调用。
|
|
|
|
| 463 |
|
| 464 |
<details>
|
| 465 |
+
<summary>点击展开离线对话代码。</summary>
|
| 466 |
|
| 467 |
```python
|
| 468 |
import librosa
|
|
|
|
| 491 |
|
| 492 |
</details>
|
| 493 |
|
| 494 |
+
### 8.3 🎙️ 全双工对话
|
| 495 |
|
| 496 |
+
调用 `model.as_duplex(generate_audio=True)` 切换为全双工流式模式:逐秒输入音频,模型持续聆听,并在回答时生成语音。示例在输入末尾追加 10 秒静音,使模型能够在输入结束后说完回复,并将生成的语音写入 `output/audio_full_duplex.wav`。
|
|
|
|
|
|
|
| 497 |
|
| 498 |
<details>
|
| 499 |
+
<summary>点击展开全双工对话代码。</summary>
|
| 500 |
|
| 501 |
```python
|
| 502 |
import librosa
|
| 503 |
import numpy as np
|
| 504 |
import soundfile as sf
|
| 505 |
|
| 506 |
+
duplex = model.as_duplex(generate_audio=True) # full-duplex with speech output
|
| 507 |
duplex.prepare(prompt_wav_path="Realtime-Venus-Audio/assets/HT_ref_audio.wav")
|
| 508 |
|
| 509 |
audio, _ = librosa.load(
|
|
|
|
| 532 |
if result["audio_waveform"] is not None and not result["is_listen"]:
|
| 533 |
timed_audio.append((chunk_index, result["audio_waveform"]))
|
| 534 |
|
| 535 |
+
# stitch the generated speech on its original timeline (24 kHz)
|
| 536 |
sample_rate = 24000
|
| 537 |
total_samples = max(
|
| 538 |
t * sample_rate + len(np.asarray(w, dtype=np.float32).squeeze())
|
|
|
|
| 563 |
|
| 564 |
## 10. 📄 许可证
|
| 565 |
|
| 566 |
+
本仓库包含 [Apache License 2.0](./LICENSE)。另请查阅上游模型、第三方库,以及与本模型配合使用的数据各自的许可证和可接受使用条款。
|
|
|