File size: 25,264 Bytes
04bceb2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
374d7b4
04bceb2
374d7b4
04bceb2
374d7b4
04bceb2
374d7b4
04bceb2
1b66c69
04bceb2
374d7b4
04bceb2
 
 
ad42af5
04bceb2
374d7b4
 
04bceb2
 
ad42af5
04bceb2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ad42af5
04bceb2
ad42af5
04bceb2
ad42af5
04bceb2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5f0742d
 
 
04bceb2
 
 
 
 
 
 
 
 
 
5f0742d
 
 
04bceb2
 
5f0742d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
04bceb2
5f0742d
 
 
 
 
 
 
 
 
04bceb2
 
 
5f0742d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
04bceb2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0ab6c8f
04bceb2
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
---
license: apache-2.0
language:
  - en
  - zh
library_name: transformers
pipeline_tag: any-to-any
tags:
  - multimodal
  - audio
  - video
  - speech
  - streaming
  - full-duplex
  - long-video
  - custom-code
---

<p align="center">
  <img src="./assets/venus-logo-white.gif" alt="Realtime-Venus logo" width="180">
</p>

<h1 align="center" style="text-align: center;">Realtime-Venus</h1>

<p align="center" style="text-align: center;"><strong>A full-duplex interaction system with asynchronous delegation</strong></p>

<p align="center" style="text-align: center;"><strong>English</strong> | <a href="https://huggingface.co/inclusionAI/Realtime-Venus/blob/main/README_zh.md">็ฎ€ไฝ“ไธญๆ–‡</a></p>

<p align="center">
<a href="https://realtime-venus.github.io/"><img src="https://img.shields.io/badge/Project_Page-4c9aff.svg?logo=googlechrome&logoColor=white" alt="Project Page"></a>
<a href="https://huggingface.co/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/Hugging_Face-Realtime--Venus-FFD21E.svg?logo=huggingface&logoColor=000" alt="Realtime-Venus on Hugging Face"></a>
<a href="https://www.modelscope.cn/models/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/ModelScope-Realtime--Venus-624AFF.svg?logo=modelscope&logoColor=white" alt="Realtime-Venus on ModelScope"></a>
<a href="https://arxiv.org/pdf/2609.13814"><img src="https://img.shields.io/badge/arXiv-2609.13814-b31b1b.svg?logo=arxiv&logoColor=white" alt="arXiv"></a>
<a href="https://github.com/inclusionAI/Realtime-Venus"><img src="https://img.shields.io/badge/GitHub-Realtime--Venus-181717.svg?logo=github&logoColor=white" alt="GitHub"></a>
<a href="https://huggingface.co/inclusionAI/Realtime-Venus/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache_2.0-0b7285.svg?logo=apache&logoColor=white" alt="Apache License 2.0"></a>
</p>

<p align="center">
  <a href="https://arxiv.org/pdf/2609.13814">
    <img src="https://arxiv.org/html/2609.13814v1/case.png" alt="Example proactive, delegated, and full-duplex interactions with Realtime-Venus" width="100%">
  </a>
</p>
<p align="center"><em>Realtime-Venus supports proactive audio-visual interaction, asynchronous delegation, and interruption-aware full-duplex dialogue.</em></p>

## 1. ๐Ÿงญ Overview

This repository hosts two checkpoints of the
[Realtime-Venus](https://realtime-venus.github.io/) system:

- **Realtime-Venus-Omni** (`Realtime-Venus-Omni/`): the 9B audio-visual
  interaction model. It continuously watches and listens, decides whether and
  when to respond, and generates text and speech on a shared causal timeline.
  Adapted from MiniCPM-o 4.5, it supports proactive interaction, semantic
  interruption handling, and training-free long-video memory.
- **Realtime-Venus-Audio** (`Realtime-Venus-Audio/`): the audio-focused
  checkpoint on the same streaming backbone, for audio understanding and
  audio-driven conversation with text or speech output.

Both directories contain model weights and custom Hugging Face Transformers
code. The asynchronous Realtime-Venus-Harness and its external tool
integrations live in the
[GitHub repository](https://github.com/inclusionAI/Realtime-Venus).

## 2. โœจ Highlights

- **Native full-duplex conversation:** keeps perceiving while speaking and
  distinguishes backchannels, interruptions, corrections, and redirections.
- **Omni-Proactive interaction:** continuously processes temporally aligned
  video and audio, and initiates a response when an event warrants it โ€” without
  waiting for a user prompt.
- **Delegation:** emits in-stream `<delegate>` requests on the shared causal
  timeline and consumes asynchronous backend results the same way, so external
  tasks never block the ongoing conversation. (Executing requests requires the
  Realtime-Venus-Harness runtime, available in the
  [GitHub repository](https://github.com/inclusionAI/Realtime-Venus).)
- **Training-free long-video Memory:** archives visually informative moments,
  retrieves query-relevant and non-redundant evidence, and reassembles the
  corresponding audio-visual context โ€” no additional training required.
- **Text and speech output:** generates response text together with native
  speech through the bundled Token2wav resources and a reference voice.

## 3. ๐Ÿ“‹ Model Details

| Item | Realtime-Venus-Omni | Realtime-Venus-Audio |
| --- | --- | --- |
| Parameters | 9B | 9B |
| Base architecture | MiniCPM-o 4.5 / Omni-Flow | MiniCPM-o 4.5 / Omni-Flow |
| Visual encoder | SigLIP2 | not used at inference |
| Audio encoder | Whisper-Medium | Whisper-Medium |
| Language backbone | Qwen3-8B | Qwen3-8B |
| Speech generation | Discrete S3 speech tokens with a streaming flow-matching decoder | same decoder, enabled in full-duplex mode |
| Inputs | Video/images, audio, and text | Audio and text |
| Outputs | Text and optional speech waveform | Text and speech waveform |
| Context length | 40,960 tokens | 40,960 tokens |
| Weight dtype | BF16 | BF16 |

## 4. ๐Ÿ“Š Evaluation

All values are reported in the
[Realtime-Venus technical report](https://arxiv.org/pdf/2609.13814).

<p align="center"><img src="assets/paper-understanding.svg" width="100%" alt="Radar charts comparing video understanding for Omni and audio understanding for Audio" /><br /><sub>Figure 1. Video and audio understanding results from the <a href="https://arxiv.org/pdf/2609.13814">paper</a>.</sub></p>

<p align="center"><img src="assets/paper-duplex.svg" width="100%" alt="Full-duplex benchmark comparisons for interruption handling and continuation under different types of overlapping speech" /><br /><sub>Figure 2. Full-duplex interaction results from the <a href="https://arxiv.org/pdf/2609.13814">paper</a>.</sub></p>

## 5. ๐Ÿ—‚๏ธ Repository Layout

```text
.
โ”œโ”€โ”€ Realtime-Venus-Omni/          # Audio-visual full-duplex checkpoint
โ”‚   โ”œโ”€โ”€ model-*.safetensors       # Sharded model weights
โ”‚   โ”œโ”€โ”€ config.json, *.py         # Model config and custom Transformers code
โ”‚   โ”œโ”€โ”€ realtime_venus_omni_memory.py  # Public Memory entry point
โ”‚   โ”œโ”€โ”€ memory_adapter/           # Chat and Duplex Memory runtime
โ”‚   โ”œโ”€โ”€ assets/                   # Reference voice, Token2wav, demo videos
โ”‚   โ””โ”€โ”€ requirements.txt
โ”œโ”€โ”€ Realtime-Venus-Audio/         # Audio-focused checkpoint
โ”‚   โ”œโ”€โ”€ model-*.safetensors       # Sharded model weights
โ”‚   โ”œโ”€โ”€ config.json, *.py         # Model config and custom Transformers code
โ”‚   โ””โ”€โ”€ assets/                   # Reference voice, Token2wav, demo audio
โ”œโ”€โ”€ assets/                      # Brand resources (logo)
โ”œโ”€โ”€ config.yaml                  # Model names and download directory mapping
โ”œโ”€โ”€ download_models.py           # Unified Omni / Audio / all downloader
โ”œโ”€โ”€ README.md
โ”œโ”€โ”€ README_zh.md
โ””โ”€โ”€ LICENSE
```

The examples below write generated media to `output/`. Use a new filename or a
new output directory when repeating an experiment.

## 6. ๐Ÿ› ๏ธ Installation

Running the inference examples requires Python 3.10, CUDA, and FFmpeg. First,
install the download dependencies and fetch the unified downloader from this
Hugging Face repository:

```bash
python -m pip install 'huggingface_hub>=0.34' 'PyYAML>=6.0'
hf download inclusionAI/Realtime-Venus download_models.py --local-dir .
```

Then choose the models to download:

| `--model` | Download |
| --- | --- |
| `omni` | Realtime-Venus-Omni for audio-visual interaction |
| `audio` | Realtime-Venus-Audio for audio understanding and conversation |
| `all` | Both models |

For example, download both models into the current directory:

```bash
python download_models.py --model all --local-dir .
```

Use `--model omni` or `--model audio` to download only the model you need.
The downloader reads this repository's root `config.yaml` and downloads each
selected model's complete directory, including weights, custom code, and
assets. It saves a download manifest and uses the Hugging Face Hub's standard
progress display and cache. Each download uses one repository revision.

Install the inference dependencies after downloading. For Omni or `all`:

```bash
python -m pip install -r Realtime-Venus-Omni/requirements.txt
```

Audio uses the same published dependency list; it does not have a separate
`requirements.txt`. If you downloaded only Audio, fetch that small file first
without downloading the Omni weights:

```bash
hf download inclusionAI/Realtime-Venus Realtime-Venus-Omni/requirements.txt --local-dir .
python -m pip install -r Realtime-Venus-Omni/requirements.txt
```

To download from Python instead, run this once from the directory containing
`download_models.py`. It uses the same downloader as the command above:

```python
from download_models import download_models

paths = download_models(model="omni", local_dir=".")  # "omni", "audio", or "all"
model_dir = paths["omni"]  # pathlib.Path; use paths["audio"] for Audio
```

The inference examples below load the downloaded local model directories. Run
them from the same directory; all asset and output paths are relative to it.

As an independent alternative, the ModelScope CLI (installed separately) can
download the entire mirror repository:

```bash
modelscope download --model inclusionAI/Realtime-Venus --local_dir .
```

This mirror command is separate from the Hugging Face downloader above.

## 7. ๐ŸŽ™๏ธ Realtime-Venus-Omni Usages

Runnable standalone versions of these examples live in the
[Omni cookbook](https://github.com/inclusionAI/Realtime-Venus/tree/main/frontend/Realtime-Venus-Omni)
on GitHub.

### 7.1 ๐Ÿงฑ Model Initialization

The examples below share the following model initialization; run each
example in a fresh Python process.
Chat and Duplex automatically load the default reference voice.

<details>
<summary>Click to show Omni model loading code.</summary>

```python
from pathlib import Path

import torch
from transformers import AutoModel, set_seed

Path("output").mkdir(exist_ok=True)
set_seed(42)
print("Loading model ...")
model = AutoModel.from_pretrained(
    "./Realtime-Venus-Omni",  # or an absolute path to the sub-directory
    trust_remote_code=True,
    local_files_only=True,
    attn_implementation="sdpa",
    torch_dtype=torch.bfloat16,
)
model.eval().cuda()
print("Model loaded.")
```

</details>

### 7.2 ๐Ÿ”Š Duplex Omni Mode

`model = model.as_duplex()` switches the model to full-duplex streaming:
`prepare()` initializes the session, then each second of input is handled by
one `streaming_prefill()` + `streaming_generate()` pair, and `as_simplex()`
switches back to offline mode. Set `MAX_NUM_FRAMES` before importing
`minicpmo.utils`, otherwise videos longer than 64 seconds are truncated to the
default frame cap.

Subtitle font note: Duplex examples burn the response text into the output
video through FFmpeg/libass, which resolves fonts via fontconfig. Rendering
non-Latin responses (e.g. Chinese) requires a CJK-capable font on the system,
otherwise those glyphs show up as empty boxes. On any Linux distribution,
install one without root and refresh the font cache:

```bash
mkdir -p ~/.local/share/fonts
curl --fail --location --retry 3 \
  --output ~/.local/share/fonts/NotoSansCJKsc-Regular.otf \
  https://raw.githubusercontent.com/notofonts/noto-cjk/main/Sans/OTF/SimplifiedChinese/NotoSansCJKsc-Regular.otf
fc-cache -f
```

Package-manager equivalents: `apt install -y fonts-noto-cjk` (Debian/Ubuntu) or
`yum install -y cjkuni-ukai-fonts cjkuni-uming-fonts` (RHEL/Alibaba Cloud Linux).
No code changes are needed.

#### 7.2.1 Duplex Chat

Stream the demo video second by second and inject text questions at the seconds
given by `question_times` (paired with `questions`). The model listens
continuously and speaks when it answers.

<details>
<summary>Click to show the Duplex Chat code.</summary>

```python
import os

os.environ["MAX_NUM_FRAMES"] = "100000"

from minicpmo.utils import get_video_frame_audio_segments, generate_duplex_video

model = model.as_duplex()  # switch to full-duplex streaming
model.prepare()

video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
# each question is injected at the corresponding second
question_times = [60, 128]
questions = [
    "What do you see in the video so far?",
    "What is the color of the cooler labeled PRIME near the team bench?",
]
question_plan = dict(zip(question_times, questions))
print(f"Extracting per-second audio and frames from {video_path} ...")
frames, audios, _ = get_video_frame_audio_segments(
    video_path, stack_frames=1, use_ffmpeg=True, adjust_audio_length=True
)
print(f"Streaming {len(audios)} seconds; questions are injected at {question_times}.")
results, output_audio = [], []
for second, (frame, audio) in enumerate(zip(frames, audios), start=1):
    model.streaming_prefill(
        audio_waveform=audio,
        frame_list=[frame] if frame is not None else None,
        text_list=[question_plan[second]] if second in question_plan else None,
    )
    result = model.streaming_generate()
    print(
        f"[{second}/{len(audios)}]",
        "listen..." if result["is_listen"] else f"speak> {result['text']}",
        flush=True,
    )
    results.append({"chunk_idx": second - 1, **result})
    if result["audio_waveform"] is not None:
        output_audio.append((second - 1, result["audio_waveform"]))

model = model.as_simplex()
print("Muxing the spoken responses into the output video ...")
generate_duplex_video(
    video_path=video_path,
    output_video_path="output/duplex_chat.mp4",
    results_log=results,
    timed_output_audio=output_audio,
)
```

</details>

#### 7.2.2 Speech-In Duplex Chat

Same as above, except the question is spoken and already mixed into the video's
audio track (at ~3 s, asking for an alert when the water boils), so no text is
injected โ€” the model must hear it.

<details>
<summary>Click to show the Speech-In Duplex Chat code.</summary>

```python
import os

os.environ["MAX_NUM_FRAMES"] = "100000"

from minicpmo.utils import get_video_frame_audio_segments, generate_duplex_video

model = model.as_duplex()  # switch to full-duplex streaming
model.prepare()

video_path = "Realtime-Venus-Omni/assets/speech_in.mp4"
print(f"Extracting per-second audio and frames from {video_path} ...")
frames, audios, _ = get_video_frame_audio_segments(
    video_path, stack_frames=1, use_ffmpeg=True, adjust_audio_length=True
)
print(f"Streaming {len(audios)} seconds; the spoken question is already in the audio track.")
results, output_audio = [], []
for second, (frame, audio) in enumerate(zip(frames, audios), start=1):
    model.streaming_prefill(
        audio_waveform=audio,
        frame_list=[frame] if frame is not None else None,
    )
    result = model.streaming_generate()
    print(
        f"[{second}/{len(audios)}]",
        "listen..." if result["is_listen"] else result["text"],
        flush=True,
    )
    results.append({"chunk_idx": second - 1, **result})
    if result["audio_waveform"] is not None:
        output_audio.append((second - 1, result["audio_waveform"]))

model = model.as_simplex()
print("Muxing the spoken responses into the output video ...")
generate_duplex_video(
    video_path=video_path,
    output_video_path="output/duplex_speech_in_chat.mp4",
    results_log=results,
    timed_output_audio=output_audio,
)
```

</details>

#### 7.2.3 Memory Duplex Chat

`model.use_memory(memory_minutes=40)` enables the long-video Memory before
entering duplex mode.

<details>
<summary>Click to show the Memory Duplex Chat code.</summary>

```python
import os

os.environ["MAX_NUM_FRAMES"] = "100000"

from minicpmo.utils import get_video_frame_audio_segments, generate_duplex_video

model.use_memory(memory_minutes=40)  # enable long-video memory
model = model.as_duplex()  # switch to full-duplex streaming
model.prepare()

video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
question = "What is the color of the cooler labeled PRIME near the team bench?"
print(f"Extracting per-second audio and frames from {video_path} ...")
frames, audios, _ = get_video_frame_audio_segments(
    video_path, stack_frames=1, use_ffmpeg=True, adjust_audio_length=True
)
print(f"Streaming {len(audios)} seconds; the text question is injected at second 128.")
results, output_audio = [], []
for second, (frame, audio) in enumerate(zip(frames, audios), start=1):
    model.streaming_prefill(
        audio_waveform=audio,
        frame_list=[frame] if frame is not None else None,
        text_list=[question] if second == 128 else None,
    )
    result = model.streaming_generate()
    print(
        f"[{second}/{len(audios)}]",
        "listen..." if result["is_listen"] else f"speak> {result['text']}",
        flush=True,
    )
    results.append({"chunk_idx": second - 1, **result})
    if result["audio_waveform"] is not None:
        output_audio.append((second - 1, result["audio_waveform"]))

model = model.as_simplex()
print("Muxing the spoken responses into the output video ...")
generate_duplex_video(
    video_path=video_path,
    output_video_path="output/duplex_memory_chat.mp4",
    results_log=results,
    timed_output_audio=output_audio,
)
```

</details>

### 7.3 ๐Ÿ’ฌ Half-Duplex Omni Mode

`model.chat(...)` answers one turn at a time over the whole video.
`model.init_tts()` enables speech output.

#### 7.3.1 Offline Chat

Sampled frames, per-second audio, and the question go into a single `chat()`
call. The 128-frame cap (`MAX_NUM_FRAMES`) limits the visual load, while
`max_inp_length=32768` sets the input-token budget. Full audio is still
retained, so very long videos can exceed that budget even with frame sampling.

<details>
<summary>Click to show the Offline Chat code.</summary>

```python
import os

os.environ.setdefault("MAX_NUM_FRAMES", "128")

from minicpmo.utils import get_video_frame_audio_segments

model.init_tts()  # enable speech output

video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
question = "What is the color of the cooler labeled PRIME near the team bench?"
print(f"Extracting audio and frames from {video_path} ...")
frames, audios, _ = get_video_frame_audio_segments(
    video_path, stack_frames=1
)
content = []
for frame, audio in zip(frames, audios):
    if frame is not None:
        content.append(frame)
    content.append(audio)
content.append(question)

print("Running chat inference ...")
response = model.chat(
    msgs=[{"role": "user", "content": content}],
    max_new_tokens=4096,
    max_inp_length=32768,
    do_sample=True,
    temperature=0.7,
    use_image_id=False,
    max_slice_nums=1,
    use_tts_template=True,
    enable_thinking=False,
    omni_mode=True,
    generate_audio=True,
    output_audio_path="output/offline_chat.wav",
)
print(response)
```

</details>

#### 7.3.2 Memory Offline Chat

`model.use_memory()` enables Memory before the chat call; retrieval selects up
to 96 historical frames plus 4 recent frames, each with ยฑ1 s of audio.

<details>
<summary>Click to show the Memory Offline Chat code.</summary>

```python
import os

os.environ["MAX_NUM_FRAMES"] = "100000"

from minicpmo.utils import get_video_frame_audio_segments

model.use_memory()  # enable long-video memory
model.init_tts()  # enable speech output

video_path = "Realtime-Venus-Omni/assets/sample_1_real.mp4"
question = "What is the color of the cooler labeled PRIME near the team bench?"
print(f"Extracting audio and frames from {video_path} ...")
frames, audios, _ = get_video_frame_audio_segments(
    video_path, stack_frames=1, use_ffmpeg=True, adjust_audio_length=True
)
content = []
for frame, audio in zip(frames, audios):
    if frame is not None:
        content.append(frame)
    content.append(audio)
content.append(question)

print("Running chat inference ...")
response = model.chat(
    msgs=[{"role": "user", "content": content}],
    max_new_tokens=4096,
    max_inp_length=32768,
    do_sample=True,
    temperature=0.7,
    use_image_id=False,
    max_slice_nums=1,
    use_tts_template=True,
    enable_thinking=False,
    omni_mode=True,
    generate_audio=True,
    output_audio_path="output/offline_memory_chat.wav",
)
print(response)
```

</details>

## 8. ๐ŸŽง Realtime-Venus-Audio Usages

Runnable standalone versions of these examples live in the
[Audio cookbook](https://github.com/inclusionAI/Realtime-Venus/tree/main/frontend/Realtime-Venus-Audio)
on GitHub.

The Audio checkpoint runs audio-only inference in two ways: turn-based
`model.chat` (text response) and the full-duplex streaming API (spoken
response). Inputs are decoded as 16 kHz mono audio from any audio or video
file.

### 8.1 ๐Ÿงฑ Model Initialization

Speech output is enabled with `init_tts=True` so the same `model` serves both
examples; use `init_tts=False` for text-only chat to load faster.

<details>
<summary>Click to show Audio model loading code.</summary>

```python
from pathlib import Path

import torch
from transformers import AutoModel, AutoTokenizer, set_seed

Path("output").mkdir(exist_ok=True)
set_seed(42)
print("Loading model ...")
tokenizer = AutoTokenizer.from_pretrained(
    "./Realtime-Venus-Audio", trust_remote_code=True, local_files_only=True,
    fix_mistral_regex=True,
)
model = AutoModel.from_pretrained(
    "./Realtime-Venus-Audio",
    trust_remote_code=True,
    local_files_only=True,
    attn_implementation="sdpa",
    torch_dtype=torch.bfloat16,
    init_vision=False,  # audio-only usage
    init_audio=True,
    init_tts=True,      # speech output; set False for text-only chat
).eval().cuda()
print("Model loaded.")
```

</details>

### 8.2 ๐Ÿ’ญ Offline Chat

One deterministic turn over the full audio input: the audio (plus an optional
text instruction) goes into a single `model.chat()` call.

<details>
<summary>Click to show the Offline Chat code.</summary>

```python
import librosa

print("Loading audio ...")
audio, _ = librosa.load(
    "Realtime-Venus-Audio/assets/case_offline.wav", sr=16000, mono=True
)
msgs = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": [audio, "What is the speaker asking about?"]},
]

print("Running chat inference ...")
answer = model.chat(
    msgs=msgs,
    tokenizer=tokenizer,
    do_sample=False,
    max_new_tokens=2048,
    enable_thinking=False,
    use_tts_template=True,
    generate_audio=False,
)
print(answer)
```

</details>

### 8.3 ๐ŸŽ™๏ธ Duplex Chat

`model.as_duplex(generate_audio=True)` switches to full-duplex streaming:
audio is fed second by second, the model listens continuously and speaks when
it answers. The example appends 10 s of trailing silence so the model can
finish its response after the input ends, and writes the generated speech to
`output/audio_full_duplex.wav`.

<details>
<summary>Click to show the Duplex Chat code.</summary>

```python
import librosa
import numpy as np
import soundfile as sf

duplex = model.as_duplex(generate_audio=True)  # full-duplex with speech output
duplex.prepare(prompt_wav_path="Realtime-Venus-Audio/assets/HT_ref_audio.wav")

audio, _ = librosa.load(
    "Realtime-Venus-Audio/assets/case_duplex.wav", sr=16000, mono=True
)
audio = np.concatenate([audio, np.zeros(10 * 16000, dtype=np.float32)])

chunk_samples = int(duplex.CHUNK_MS * duplex.SAMPLE_RATE / 1000)
total_chunks = max(1, (len(audio) + chunk_samples - 1) // chunk_samples)
timed_audio = []
for chunk_index in range(total_chunks):
    chunk = audio[chunk_index * chunk_samples:(chunk_index + 1) * chunk_samples]
    if len(chunk) < chunk_samples:
        chunk = np.pad(chunk, (0, chunk_samples - len(chunk)))
    duplex.streaming_prefill(audio_waveform=chunk)
    result = duplex.streaming_generate(
        max_new_speak_tokens_per_chunk=20,
        decode_mode="sampling",
        temperature=0.7,
        top_k=20,
        top_p=0.8,
        listen_prob_scale=1.0,
    )
    state = "listen" if result["is_listen"] else f"speak> {result['text']}"
    print(f"[{chunk_index + 1}/{total_chunks}] {state}", flush=True)
    if result["audio_waveform"] is not None and not result["is_listen"]:
        timed_audio.append((chunk_index, result["audio_waveform"]))

# stitch the generated speech on its original timeline (24 kHz)
sample_rate = 24000
total_samples = max(
    t * sample_rate + len(np.asarray(w, dtype=np.float32).squeeze())
    for t, w in timed_audio
)
output = np.zeros(total_samples, dtype=np.float32)
for t, waveform in timed_audio:
    w = np.asarray(waveform, dtype=np.float32).squeeze()
    output[t * sample_rate: t * sample_rate + len(w)] += w
sf.write("output/audio_full_duplex.wav", np.clip(output, -1.0, 1.0), sample_rate)
print("Saved generated speech to output/audio_full_duplex.wav")
```

</details>

## 9. ๐Ÿ“ Citation

If you find Realtime-Venus useful, please cite the technical report:

```bibtex
@article{zhao2026realtime,
  title={{Realtime-Venus}: A full-duplex interaction system with asynchronous delegation},
  author={{Venus Team(Ant Group), Tsinghua University}},
  journal={arXiv preprint arXiv:2609.13814},
  year={2026}
}
```

## 10. ๐Ÿ“„ License

This repository includes an [Apache License 2.0](./LICENSE). Please also review
the licenses and acceptable-use terms of the upstream model, third-party
libraries, and any data used with this checkpoint.