kim2h7903 kumass2020 ddwkim ubuntu commited on
Commit
419adf0
·
0 Parent(s):

Super-squash branch 'main' using huggingface_hub

Browse files

Co-authored-by: kumass2020 <kumass2020@users.noreply.huggingface.co>
Co-authored-by: ddwkim <ddwkim@users.noreply.huggingface.co>
Co-authored-by: ubuntu <ubuntu@users.noreply.huggingface.co>

.gitattributes ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ model-00001-of-00004.safetensors filter=lfs diff=lfs merge=lfs -text
2
+ model-00002-of-00004.safetensors filter=lfs diff=lfs merge=lfs -text
3
+ model-00003-of-00004.safetensors filter=lfs diff=lfs merge=lfs -text
4
+ model-00004-of-00004.safetensors filter=lfs diff=lfs merge=lfs -text
5
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,206 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: transformers
3
+ tags:
4
+ - speech
5
+ - audio
6
+ - multimodal
7
+ license: cc-by-nc-4.0
8
+ language:
9
+ - en
10
+ - ko
11
+ ---
12
+
13
+ # Raon-Speech-9B
14
+
15
+ <div align="center">
16
+ <img class="block dark:hidden" src="assets/Raon-Speech-Gradient-Black.png" alt="Raon-Speech Logo" width="400">
17
+ <img class="hidden dark:block" src="assets/Raon-Speech-Gradient-White.png" alt="Raon-Speech Logo" width="400">
18
+ </div>
19
+
20
+ <p align="center">
21
+ <a href="https://www.krafton.ai/ko/"><img src="https://img.shields.io/badge/Homepage-KRAFTON%20AI-blue?style=flat&logo=google-chrome&logoColor=white" alt="Homepage"></a>
22
+ <a href="https://github.com/krafton-ai/Raon-Speech"><img src="https://img.shields.io/badge/GitHub-Raon-white?style=flat&logo=github&logoColor=black" alt="GitHub"></a>
23
+ <br>
24
+ <a href="https://huggingface.co/KRAFTON"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-KRAFTON-yellow?style=flat" alt="Hugging Face"></a>
25
+ <a href="https://x.com/Krafton_AI"><img src="https://img.shields.io/badge/X-KRAFTON%20AI-white?style=flat&logo=x&logoColor=black" alt="X"></a>
26
+ <br>
27
+ <a href="https://creativecommons.org/licenses/by-nc/4.0/"><img src="https://img.shields.io/badge/License-CC%20BY--NC%204.0-lightgrey?style=flat" alt="License"></a>
28
+ </p>
29
+
30
+ <p align="center">
31
+ Technical Report | Blog (Coming soon)
32
+ <!-- 📄 <a href="https://arxiv.org/abs/YOUR_PAPER_ID">Technical Report</a> | 📝 <a href="https://YOUR_BLOG_URL">Blog</a> -->
33
+ </p>
34
+
35
+ RAON-Speech is a 9B-parameter speech language model that supports state-of-the-art speech understanding, answering and generation in English and Korean.
36
+ This model successfully transforms a pre-trained LLM into a SpeechLM to both understand and generate speech without compromising its original language capabilities.
37
+ It trains on millions of hours of English-Korean speech-text datasets with the following training stages: (1) speech encoder-decoder alignment, (2) end-to-end SpeechLM pre-training, and (3) multi-reward DPO-based post-training.
38
+
39
+ ## Key Features
40
+
41
+ - **End-to-End Speech Language Model**: 9B-parameter multimodal model built on Qwen3 (36 layers, 4096 hidden dim), Qwen3OmniMoeAudioEncoder (24 layers), Mimi codec (32 quantizers), and ECAPA-TDNN speaker encoder.
42
+ - **Bilingual Support**: State-of-the-art speech understanding, answering, and generation in both English and Korean.
43
+ - **Multi-Task Capabilities**: Supports STT (audio → text), TTS (text → audio), TextQA (text + audio → text), and SpeechChat (audio → text) in a single unified model.
44
+ - **Speaker Voice Conditioning**: TTS with optional speaker reference audio for voice cloning via ECAPA-TDNN embeddings.
45
+ - **TTS Continuation**: Generate speech that naturally continues from a reference audio, with prefill-based continuation for seamless prosody.
46
+ - **Multi-Reward DPO Post-Training**: Three-stage training pipeline — (1) speech encoder-decoder alignment, (2) end-to-end SpeechLM pre-training, and (3) multi-reward DPO-based post-training — for high-quality speech generation.
47
+ - **HuggingFace Transformers Integration**: Load and run directly via `AutoModel.from_pretrained` with `trust_remote_code=True` — no custom package installation required.
48
+
49
+ ## Benchmark Results
50
+
51
+ Measured with LibriSpeech test-clean samples on single-GPU setups via streaming TTS. All values are averaged.
52
+
53
+ | Metric | RTX 6000 Pro | L40S |
54
+ |--------|-------------|------|
55
+ | **RTF** | 0.27 (3.7× real-time) | 0.45 (2.2× real-time) |
56
+ | **TTFT** | 617 ms | 887 ms |
57
+ | **TBT** | 135 ms | 233 ms |
58
+
59
+ - **RTF** (Real-Time Factor): Lower is faster. Values below 1.0 mean faster-than-real-time synthesis.
60
+ - **TTFT** (Time to First Token): Latency until the first audio chunk is returned.
61
+ - **TBT** (Time Between Tokens): Average interval between consecutive audio chunks.
62
+
63
+
64
+ ## Requirements
65
+
66
+ ```bash
67
+ pip install transformers>=4.57.1 torch torchaudio soundfile accelerate
68
+
69
+ # Optional
70
+ pip install speechbrain # for TTS with speaker voice conditioning
71
+ pip install gradio # for Gradio demo
72
+ ```
73
+
74
+ ## Quick Start
75
+
76
+ ### Option 1: Load from Hub (recommended)
77
+
78
+ No `pip install raon` needed.
79
+
80
+ ```python
81
+ from transformers import AutoConfig
82
+ from transformers.dynamic_module_utils import get_class_from_dynamic_module
83
+
84
+ MODEL_ID = "KRAFTON/Raon-Speech-9B"
85
+
86
+ _cfg = AutoConfig.from_pretrained(MODEL_ID, trust_remote_code=True)
87
+ RaonPipeline = get_class_from_dynamic_module(
88
+ "modeling_raon.RaonPipeline",
89
+ MODEL_ID,
90
+ revision=getattr(_cfg, "_commit_hash", None),
91
+ )
92
+ del _cfg
93
+
94
+ pipe = RaonPipeline(MODEL_ID, device="cuda", dtype="bfloat16")
95
+ ```
96
+
97
+ ### Option 2: With raon package installed
98
+
99
+ ```bash
100
+ git clone https://github.com/krafton-ai/Raon-Speech.git
101
+ cd Raon-Speech/raon
102
+ pip install -e . # or: uv sync
103
+ ```
104
+
105
+ ```python
106
+ from raon import RaonPipeline
107
+
108
+ # From Hub (local code + Hub weights)
109
+ pipe = RaonPipeline("KRAFTON/Raon-Speech-9B")
110
+
111
+ # From local path
112
+ pipe = RaonPipeline("/path/to/raon-model")
113
+ ```
114
+
115
+ ## Tasks
116
+
117
+ #### STT (Audio → Text)
118
+ ```python
119
+ text = pipe.stt("audio.wav")
120
+ ```
121
+
122
+ #### TTS (Text → Audio)
123
+ ```python
124
+ # Without speaker conditioning
125
+ audio, sr = pipe.tts("Hello, how are you?")
126
+ pipe.save_audio((audio, sr), "output.wav")
127
+
128
+ # With speaker conditioning (requires speechbrain)
129
+ audio, sr = pipe.tts("Hello, how are you?", speaker_audio="speaker_ref.wav")
130
+ ```
131
+
132
+ #### TextQA (Text + Audio → Text)
133
+ ```python
134
+ answer = pipe.textqa("What is the speaker saying?", audio="audio.wav")
135
+ ```
136
+
137
+ #### SpeechChat (Audio → Text)
138
+ ```python
139
+ answer = pipe.speech_chat("question.wav")
140
+ ```
141
+
142
+ #### Chat (Multimodal)
143
+ ```python
144
+ messages = [
145
+ {
146
+ "role": "user",
147
+ "content": [
148
+ {"type": "audio", "audio": "audio.wav"},
149
+ {"type": "text", "text": "Transcribe and summarise this audio."},
150
+ ],
151
+ },
152
+ ]
153
+ response = pipe.chat(messages)
154
+ ```
155
+
156
+ ## Deployment (vLLM-Omni)
157
+
158
+ ####
159
+
160
+ # 1. Clone & Build
161
+ ```bash
162
+ git clone https://github.com/krafton-ai/vllm-omni.git
163
+ cd vllm-omni
164
+ docker build -f docker/Dockerfile.ci -t vllm-omni .
165
+ ```
166
+
167
+ # 2. Serve
168
+ ```bash
169
+ docker run --rm --gpus all \
170
+ --shm-size=16g \
171
+ -p 8000:8000 \
172
+ vllm-omni \
173
+ bash -c "vllm serve KRAFTON/Raon-Speech-9B --omni --port 8000 --trust-remote-code"
174
+ ```
175
+
176
+ # 3. Test — TTS
177
+ ```bash
178
+ curl -X POST http://localhost:8000/v1/audio/speech \
179
+ -H "Content-Type: application/json" \
180
+ -d '{
181
+ "input": "Hello, how are you?",
182
+ "model": "KRAFTON/Raon-Speech-9B",
183
+ "response_format": "wav"
184
+ }' --output output.wav
185
+ ```
186
+
187
+ # 4. Test — TTS with voice cloning
188
+ ```bash
189
+ curl -X POST http://localhost:8000/v1/audio/speech \
190
+ -H "Content-Type: application/json" \
191
+ -d '{
192
+ "input": "Hello, how are you?",
193
+ "model": "KRAFTON/Raon-Speech-9B",
194
+ "ref_audio": "data:audio/wav;base64,'$(base64 -w0 speaker_ref.wav)'",
195
+ "task_type": "Base",
196
+ "response_format": "wav"
197
+ }' --output cloned.wav
198
+ ```
199
+
200
+
201
+ ## License
202
+
203
+ This repository is licensed under the
204
+ [Creative Commons Attribution-NonCommercial 4.0 International License](https://creativecommons.org/licenses/by-nc/4.0/).
205
+
206
+ © 2026 KRAFTON
added_tokens.json ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "<tts_pad>": 151691,
3
+ "<tts_text_bos>": 151680,
4
+ "<tts_text_bos_single>": 151682,
5
+ "<tts_text_eod>": 151681,
6
+ "<|audio_end|>": 151670,
7
+ "<|audio_input_placeholder|>": 151676,
8
+ "<|audio_output_end_pad|>": 151678,
9
+ "<|audio_output_pad|>": 151677,
10
+ "<|audio_output_placeholder|>": 151675,
11
+ "<|audio_pad|>": 151683,
12
+ "<|box_end|>": 151688,
13
+ "<|box_start|>": 151687,
14
+ "<|endoftext|>": 151679,
15
+ "<|object_ref_end|>": 151686,
16
+ "<|object_ref_start|>": 151685,
17
+ "<|quad_end|>": 151690,
18
+ "<|quad_start|>": 151689,
19
+ "<|secondary_audio_pad|>": 151684,
20
+ "<|speaker_embedding_placeholder|>": 151671
21
+ }
assets/Raon-Speech-Gradient-Black.png ADDED
assets/Raon-Speech-Gradient-White.png ADDED
chat_template.jinja ADDED
@@ -0,0 +1,121 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {%- if messages[0].content is string %}
5
+ {{- messages[0].content }}
6
+ {%- else %}
7
+ {%- for content in messages[0].content %}
8
+ {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
9
+ {{- "<|vision_start|><|image_pad|><|vision_end|>" }}
10
+ {%- elif content.type == 'audio' or 'audio' in content or 'audio_url' in content %}
11
+ {{- "<|audio_start|><|audio_pad|><|audio_end|>" }}
12
+ {%- elif content.type == 'video' or 'video' in content %}
13
+ {{- "<|vision_start|><|video_pad|><|vision_end|>" }}
14
+ {%- elif content.type == 'text' %}
15
+ {{- content.text }}
16
+ {%- endif %}
17
+ {%- endfor %}
18
+ {%- endif %}
19
+ {%- endif %}
20
+ {{- '\n\n' }}
21
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
22
+ {%- for tool in tools %}
23
+ {{- "\n" }}
24
+ {{- tool | tojson }}
25
+ {%- endfor %}
26
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
27
+ {%- else %}
28
+ {%- if messages[0].role == 'system' %}
29
+ {%- if messages[0].content is string %}
30
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
31
+ {%- else %}
32
+ {%- for content in messages[0].content %}
33
+ {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
34
+ {{- '<|im_start|>system\n' +"<|vision_start|><|image_pad|><|vision_end|>"+ '<|im_end|>\n' }}
35
+ {%- elif content.type == 'audio' or 'audio' in content or 'audio_url' in content %}
36
+ {{- '<|im_start|>system\n' +"<|audio_start|><|audio_pad|><|audio_end|>"+ '<|im_end|>\n' }}
37
+ {%- elif content.type == 'video' or 'video' in content %}
38
+ {{- '<|im_start|>system\n' +"<|vision_start|><|video_pad|><|vision_end|>"+ '<|im_end|>\n' }}
39
+ {%- elif content.type == 'text' %}
40
+ {{- '<|im_start|>system\n' +content.text+ '<|im_end|>\n' }}
41
+ {%- endif %}
42
+ {%- endfor %}
43
+ {%- endif %}
44
+ {%- endif %}
45
+ {%- endif %}
46
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
47
+ {%- for message in messages[::-1] %}
48
+ {%- set index = (messages|length - 1) - loop.index0 %}
49
+ {%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
50
+ {%- set ns.multi_step_tool = false %}
51
+ {%- set ns.last_query_index = index %}
52
+ {%- endif %}
53
+ {%- endfor %}
54
+ {%- for message in messages %}
55
+ {%- if message.content is string %}
56
+ {%- set content = message.content %}
57
+ {%- else %}
58
+ {%- set content = namespace(text="") %}
59
+ {%- for mcontent in message.content %}
60
+ {%- if mcontent.type == 'image' or 'image' in mcontent or 'image_url' in mcontent %}
61
+ {%- set content.text = content.text~"<|vision_start|><|image_pad|><|vision_end|>" %}
62
+ {%- elif mcontent.type == 'audio' or 'audio' in mcontent or 'audio_url' in mcontent %}
63
+ {%- set content.text = content.text~"<|audio_start|><|audio_pad|><|audio_end|>" %}
64
+ {%- elif mcontent.type == 'video' or 'video' in mcontent %}
65
+ {%- set content.text = content.text~"<|vision_start|><|video_pad|><|vision_end|>" %}
66
+ {%- elif mcontent.type == 'text' %}
67
+ {%- set content.text = content.text~mcontent.text %}
68
+ {%- endif %}
69
+ {%- endfor %}
70
+ {%- set content = content.text %}
71
+ {%- endif %}
72
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
73
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
74
+ {%- elif message.role == "assistant" %}
75
+ {%- set reasoning_content = "" %}
76
+ {%- if message.reasoning_content is string %}
77
+ {%- set reasoning_content = message.reasoning_content %}
78
+ {%- else %}
79
+ {%- if '</think>' in content %}
80
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
81
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
82
+ {%- endif %}
83
+ {%- endif %}
84
+ {%- if loop.index0 > ns.last_query_index %}
85
+ {%- if loop.last or (not loop.last and reasoning_content) %}
86
+ {{- '<|im_start|>' + message.role + '\n' + reasoning_content + content.lstrip('\n') }}
87
+ {%- else %}
88
+ {{- '<|im_start|>' + message.role + '\n' + content }}
89
+ {%- endif %}
90
+ {%- else %}
91
+ {{- '<|im_start|>' + message.role + '\n' + content }}
92
+ {%- endif %}
93
+ {%- if message.tool_calls %}
94
+ {%- for tool_call in message.tool_calls %}
95
+ {%- if (loop.first and content) or (not loop.first) %}{{- '\n' }}{%- endif %}
96
+ {%- if tool_call.function %}
97
+ {%- set tool_call = tool_call.function %}
98
+ {%- endif %}
99
+ {{- '<tool_call>\n{"name": "' }}
100
+ {{- tool_call.name }}
101
+ {{- '", "arguments": ' }}
102
+ {%- if tool_call.arguments is string %}
103
+ {{- tool_call.arguments }}
104
+ {%- else %}
105
+ {{- tool_call.arguments | tojson }}
106
+ {%- endif %}
107
+ {{- '}\n</tool_call>' }}
108
+ {%- endfor %}
109
+ {%- endif %}
110
+ {{- '<|im_end|>\n' }}
111
+ {%- elif message.role == "tool" %}
112
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}{{- '<|im_start|>user' }}{%- endif %}
113
+ {{- '\n<tool_response>\n' }}
114
+ {{- content }}
115
+ {{- '\n</tool_response>' }}
116
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}{{- '<|im_end|>\n' }}{%- endif %}
117
+ {%- endif %}
118
+ {%- endfor %}
119
+ {%- if add_generation_prompt %}
120
+ {{- '<|im_start|>assistant\n' }}
121
+ {%- endif %}
config.json ADDED
@@ -0,0 +1,795 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_name_or_path": "",
3
+ "accept_hidden_layer": -1,
4
+ "acoustic_delay": null,
5
+ "add_cross_attention": false,
6
+ "architectures": [
7
+ "RaonModel"
8
+ ],
9
+ "audio_encoder_config": {
10
+ "_name_or_path": "",
11
+ "activation_dropout": 0,
12
+ "activation_function": "gelu",
13
+ "add_cross_attention": false,
14
+ "architectures": null,
15
+ "attention_dropout": 0,
16
+ "bad_words_ids": null,
17
+ "begin_suppress_tokens": null,
18
+ "bos_token_id": null,
19
+ "chunk_size_feed_forward": 0,
20
+ "conv_chunksize": 500,
21
+ "cross_attention_hidden_size": null,
22
+ "d_model": 1024,
23
+ "decoder_start_token_id": null,
24
+ "diversity_penalty": 0.0,
25
+ "do_sample": false,
26
+ "downsample_hidden_size": 480,
27
+ "dropout": 0,
28
+ "dtype": "bfloat16",
29
+ "early_stopping": false,
30
+ "encoder_attention_heads": 16,
31
+ "encoder_ffn_dim": 4096,
32
+ "encoder_layers": 24,
33
+ "encoder_no_repeat_ngram_size": 0,
34
+ "eos_token_id": null,
35
+ "exponential_decay_length_penalty": null,
36
+ "finetuning_task": null,
37
+ "forced_bos_token_id": null,
38
+ "forced_eos_token_id": null,
39
+ "id2label": {
40
+ "0": "LABEL_0",
41
+ "1": "LABEL_1"
42
+ },
43
+ "initializer_range": 0.02,
44
+ "is_decoder": false,
45
+ "is_encoder_decoder": false,
46
+ "label2id": {
47
+ "LABEL_0": 0,
48
+ "LABEL_1": 1
49
+ },
50
+ "length_penalty": 1.0,
51
+ "max_length": 20,
52
+ "max_source_positions": 1500,
53
+ "min_length": 0,
54
+ "model_type": "qwen3_omni_moe_audio_encoder",
55
+ "n_window": 50,
56
+ "n_window_infer": 800,
57
+ "no_repeat_ngram_size": 0,
58
+ "num_beam_groups": 1,
59
+ "num_beams": 1,
60
+ "num_hidden_layers": 24,
61
+ "num_mel_bins": 128,
62
+ "num_return_sequences": 1,
63
+ "output_attentions": false,
64
+ "output_dim": 2048,
65
+ "output_hidden_states": false,
66
+ "output_scores": false,
67
+ "pad_token_id": 151679,
68
+ "prefix": null,
69
+ "problem_type": null,
70
+ "pruned_heads": {},
71
+ "remove_invalid_values": false,
72
+ "repetition_penalty": 1.0,
73
+ "return_dict": true,
74
+ "return_dict_in_generate": false,
75
+ "sampling_rate": 24000,
76
+ "scale_embedding": false,
77
+ "sep_token_id": null,
78
+ "suppress_tokens": null,
79
+ "task_specific_params": null,
80
+ "temperature": 1.0,
81
+ "tf_legacy_loss": false,
82
+ "tie_encoder_decoder": false,
83
+ "tie_word_embeddings": true,
84
+ "tokenizer_class": null,
85
+ "top_k": 50,
86
+ "top_p": 1.0,
87
+ "torchscript": false,
88
+ "typical_p": 1.0,
89
+ "use_bfloat16": false
90
+ },
91
+ "audio_tokenizer_config": {
92
+ "_frame_rate": 12.5,
93
+ "_name_or_path": "kyutai/mimi",
94
+ "add_cross_attention": false,
95
+ "architectures": [
96
+ "MimiModel"
97
+ ],
98
+ "attention_bias": false,
99
+ "attention_dropout": 0.0,
100
+ "audio_channels": 1,
101
+ "bad_words_ids": null,
102
+ "begin_suppress_tokens": null,
103
+ "bos_token_id": null,
104
+ "chunk_size_feed_forward": 0,
105
+ "codebook_dim": 256,
106
+ "codebook_size": 2048,
107
+ "compress": 2,
108
+ "cross_attention_hidden_size": null,
109
+ "decoder_start_token_id": null,
110
+ "dilation_growth_rate": 2,
111
+ "diversity_penalty": 0.0,
112
+ "do_sample": false,
113
+ "dtype": "bfloat16",
114
+ "early_stopping": false,
115
+ "encoder_no_repeat_ngram_size": 0,
116
+ "eos_token_id": null,
117
+ "exponential_decay_length_penalty": null,
118
+ "finetuning_task": null,
119
+ "forced_bos_token_id": null,
120
+ "forced_eos_token_id": null,
121
+ "head_dim": 64,
122
+ "hidden_act": "gelu",
123
+ "hidden_size": 512,
124
+ "id2label": {
125
+ "0": "LABEL_0",
126
+ "1": "LABEL_1"
127
+ },
128
+ "initializer_range": 0.02,
129
+ "intermediate_size": 2048,
130
+ "is_decoder": false,
131
+ "is_encoder_decoder": false,
132
+ "kernel_size": 7,
133
+ "label2id": {
134
+ "LABEL_0": 0,
135
+ "LABEL_1": 1
136
+ },
137
+ "last_kernel_size": 3,
138
+ "layer_scale_initial_scale": 0.01,
139
+ "length_penalty": 1.0,
140
+ "max_length": 20,
141
+ "max_position_embeddings": 8000,
142
+ "min_length": 0,
143
+ "model_type": "mimi",
144
+ "no_repeat_ngram_size": 0,
145
+ "norm_eps": 1e-05,
146
+ "normalize": false,
147
+ "num_attention_heads": 8,
148
+ "num_beam_groups": 1,
149
+ "num_beams": 1,
150
+ "num_filters": 64,
151
+ "num_hidden_layers": 8,
152
+ "num_key_value_heads": 8,
153
+ "num_quantizers": 32,
154
+ "num_residual_layers": 1,
155
+ "num_return_sequences": 1,
156
+ "num_semantic_quantizers": 1,
157
+ "output_attentions": false,
158
+ "output_hidden_states": false,
159
+ "output_scores": false,
160
+ "pad_mode": "constant",
161
+ "pad_token_id": 151679,
162
+ "prefix": null,
163
+ "problem_type": null,
164
+ "pruned_heads": {},
165
+ "remove_invalid_values": false,
166
+ "repetition_penalty": 1.0,
167
+ "residual_kernel_size": 3,
168
+ "return_dict": true,
169
+ "return_dict_in_generate": false,
170
+ "rope_theta": 10000.0,
171
+ "sampling_rate": 24000,
172
+ "sep_token_id": null,
173
+ "sliding_window": 250,
174
+ "suppress_tokens": null,
175
+ "task_specific_params": null,
176
+ "temperature": 1.0,
177
+ "tf_legacy_loss": false,
178
+ "tie_encoder_decoder": false,
179
+ "tie_word_embeddings": true,
180
+ "tokenizer_class": null,
181
+ "top_k": 50,
182
+ "top_p": 1.0,
183
+ "torchscript": false,
184
+ "trim_right_ratio": 1.0,
185
+ "typical_p": 1.0,
186
+ "upsample_groups": 512,
187
+ "upsampling_ratios": [
188
+ 8,
189
+ 6,
190
+ 5,
191
+ 4
192
+ ],
193
+ "use_bfloat16": true,
194
+ "use_cache": false,
195
+ "use_causal_conv": true,
196
+ "use_conv_shortcut": false,
197
+ "use_streaming": false,
198
+ "vector_quantization_hidden_dimension": 256
199
+ },
200
+ "aut_is_causal": false,
201
+ "bad_words_ids": null,
202
+ "begin_suppress_tokens": null,
203
+ "bos_token_id": null,
204
+ "chunk_size_feed_forward": 0,
205
+ "code_predictor_config": {
206
+ "_name_or_path": "",
207
+ "add_cross_attention": false,
208
+ "architectures": null,
209
+ "attention_bias": false,
210
+ "attention_dropout": 0,
211
+ "bad_words_ids": null,
212
+ "begin_suppress_tokens": null,
213
+ "bos_token_id": null,
214
+ "chunk_size_feed_forward": 0,
215
+ "cross_attention_hidden_size": null,
216
+ "decoder_start_token_id": null,
217
+ "diversity_penalty": 0.0,
218
+ "do_sample": false,
219
+ "dtype": "bfloat16",
220
+ "early_stopping": false,
221
+ "encoder_no_repeat_ngram_size": 0,
222
+ "eos_token_id": null,
223
+ "exponential_decay_length_penalty": null,
224
+ "finetuning_task": null,
225
+ "forced_bos_token_id": null,
226
+ "forced_eos_token_id": null,
227
+ "head_dim": 128,
228
+ "hidden_act": "silu",
229
+ "hidden_size": 1024,
230
+ "id2label": {
231
+ "0": "LABEL_0",
232
+ "1": "LABEL_1"
233
+ },
234
+ "initializer_range": 0.02,
235
+ "intermediate_size": 3072,
236
+ "is_decoder": false,
237
+ "is_encoder_decoder": false,
238
+ "label2id": {
239
+ "LABEL_0": 0,
240
+ "LABEL_1": 1
241
+ },
242
+ "layer_types": [
243
+ "full_attention",
244
+ "full_attention",
245
+ "full_attention",
246
+ "full_attention",
247
+ "full_attention"
248
+ ],
249
+ "length_penalty": 1.0,
250
+ "max_length": 20,
251
+ "max_position_embeddings": 32768,
252
+ "max_window_layers": 28,
253
+ "min_length": 0,
254
+ "model_type": "qwen3_omni_moe_talker_code_predictor",
255
+ "no_repeat_ngram_size": 0,
256
+ "num_attention_heads": 16,
257
+ "num_beam_groups": 1,
258
+ "num_beams": 1,
259
+ "num_code_groups": 16,
260
+ "num_hidden_layers": 5,
261
+ "num_key_value_heads": 8,
262
+ "num_return_sequences": 1,
263
+ "output_attentions": false,
264
+ "output_hidden_states": false,
265
+ "output_scores": false,
266
+ "pad_token_id": 151679,
267
+ "prefix": null,
268
+ "problem_type": null,
269
+ "pruned_heads": {},
270
+ "remove_invalid_values": false,
271
+ "repetition_penalty": 1.0,
272
+ "return_dict": true,
273
+ "return_dict_in_generate": false,
274
+ "rms_norm_eps": 1e-06,
275
+ "rope_scaling": null,
276
+ "rope_theta": 1000000,
277
+ "sep_token_id": null,
278
+ "sliding_window": null,
279
+ "suppress_tokens": null,
280
+ "task_specific_params": null,
281
+ "temperature": 1.0,
282
+ "tf_legacy_loss": false,
283
+ "tie_encoder_decoder": false,
284
+ "tie_word_embeddings": false,
285
+ "tokenizer_class": null,
286
+ "top_k": 50,
287
+ "top_p": 1.0,
288
+ "torchscript": false,
289
+ "typical_p": 1.0,
290
+ "use_bfloat16": true,
291
+ "use_cache": false,
292
+ "use_sliding_window": false,
293
+ "vocab_size": 2048
294
+ },
295
+ "cross_attention_hidden_size": null,
296
+ "decoder_start_token_id": null,
297
+ "depth_loss_weight": 0.25,
298
+ "diversity_penalty": 0.0,
299
+ "do_sample": false,
300
+ "dtype": "bfloat16",
301
+ "early_stopping": false,
302
+ "encoder_no_repeat_ngram_size": 0,
303
+ "eos_token_id": null,
304
+ "exponential_decay_length_penalty": null,
305
+ "finetuning_task": null,
306
+ "forced_bos_token_id": null,
307
+ "forced_eos_token_id": null,
308
+ "hidden_size": null,
309
+ "id2label": {
310
+ "0": "LABEL_0",
311
+ "1": "LABEL_1"
312
+ },
313
+ "input_adaptor_config": {
314
+ "_name_or_path": "",
315
+ "add_cross_attention": false,
316
+ "architectures": null,
317
+ "bad_words_ids": null,
318
+ "begin_suppress_tokens": null,
319
+ "bos_token_id": null,
320
+ "chunk_size_feed_forward": 0,
321
+ "cross_attention_hidden_size": null,
322
+ "decoder_config": null,
323
+ "decoder_start_token_id": null,
324
+ "diversity_penalty": 0.0,
325
+ "do_sample": false,
326
+ "dtype": "bfloat16",
327
+ "early_stopping": false,
328
+ "encoder_no_repeat_ngram_size": 0,
329
+ "eos_token_id": null,
330
+ "exponential_decay_length_penalty": null,
331
+ "finetuning_task": null,
332
+ "forced_bos_token_id": null,
333
+ "forced_eos_token_id": null,
334
+ "hidden_size": null,
335
+ "id2label": {
336
+ "0": "LABEL_0",
337
+ "1": "LABEL_1"
338
+ },
339
+ "input_size": 2048,
340
+ "is_decoder": false,
341
+ "is_encoder_decoder": false,
342
+ "label2id": {
343
+ "LABEL_0": 0,
344
+ "LABEL_1": 1
345
+ },
346
+ "length_penalty": 1.0,
347
+ "max_length": 20,
348
+ "min_length": 0,
349
+ "model_type": "embedding_adaptor",
350
+ "no_repeat_ngram_size": 0,
351
+ "norm_eps": 1e-06,
352
+ "num_beam_groups": 1,
353
+ "num_beams": 1,
354
+ "num_layers": 2,
355
+ "num_return_sequences": 1,
356
+ "output_attentions": false,
357
+ "output_hidden_states": false,
358
+ "output_scores": false,
359
+ "output_size": 4096,
360
+ "output_time_scale": 1,
361
+ "pad_token_id": null,
362
+ "post_norm_init_scale": 0.02,
363
+ "prefix": null,
364
+ "problem_type": null,
365
+ "pruned_heads": {},
366
+ "remove_invalid_values": false,
367
+ "repetition_penalty": 1.0,
368
+ "return_dict": true,
369
+ "return_dict_in_generate": false,
370
+ "sep_token_id": null,
371
+ "suppress_tokens": null,
372
+ "task_specific_params": null,
373
+ "temperature": 1.0,
374
+ "tf_legacy_loss": false,
375
+ "tie_encoder_decoder": false,
376
+ "tie_word_embeddings": true,
377
+ "tokenizer_class": null,
378
+ "top_k": 50,
379
+ "top_p": 1.0,
380
+ "torchscript": false,
381
+ "typical_p": 1.0,
382
+ "use_bfloat16": false,
383
+ "use_post_norm": true
384
+ },
385
+ "input_num_code_groups": null,
386
+ "is_decoder": false,
387
+ "is_encoder_decoder": false,
388
+ "keys_to_ignore_at_inference": [
389
+ "past_key_values"
390
+ ],
391
+ "label2id": {
392
+ "LABEL_0": 0,
393
+ "LABEL_1": 1
394
+ },
395
+ "length_penalty": 1.0,
396
+ "max_length": 20,
397
+ "min_length": 0,
398
+ "model_type": "raon",
399
+ "no_repeat_ngram_size": 0,
400
+ "num_beam_groups": 1,
401
+ "num_beams": 1,
402
+ "num_return_sequences": 1,
403
+ "num_talker_layers": 4,
404
+ "output_adaptor_config": {
405
+ "_name_or_path": "",
406
+ "add_cross_attention": false,
407
+ "architectures": null,
408
+ "bad_words_ids": null,
409
+ "begin_suppress_tokens": null,
410
+ "bos_token_id": null,
411
+ "chunk_size_feed_forward": 0,
412
+ "cross_attention_hidden_size": null,
413
+ "decoder_config": null,
414
+ "decoder_start_token_id": null,
415
+ "diversity_penalty": 0.0,
416
+ "do_sample": false,
417
+ "dtype": "bfloat16",
418
+ "early_stopping": false,
419
+ "encoder_no_repeat_ngram_size": 0,
420
+ "eos_token_id": null,
421
+ "exponential_decay_length_penalty": null,
422
+ "finetuning_task": null,
423
+ "forced_bos_token_id": null,
424
+ "forced_eos_token_id": null,
425
+ "hidden_size": null,
426
+ "id2label": {
427
+ "0": "LABEL_0",
428
+ "1": "LABEL_1"
429
+ },
430
+ "input_size": 512,
431
+ "is_decoder": false,
432
+ "is_encoder_decoder": false,
433
+ "label2id": {
434
+ "LABEL_0": 0,
435
+ "LABEL_1": 1
436
+ },
437
+ "length_penalty": 1.0,
438
+ "max_length": 20,
439
+ "min_length": 0,
440
+ "model_type": "embedding_adaptor",
441
+ "no_repeat_ngram_size": 0,
442
+ "norm_eps": 1e-06,
443
+ "num_beam_groups": 1,
444
+ "num_beams": 1,
445
+ "num_layers": 2,
446
+ "num_return_sequences": 1,
447
+ "output_attentions": false,
448
+ "output_hidden_states": false,
449
+ "output_scores": false,
450
+ "output_size": 4096,
451
+ "output_time_scale": 1,
452
+ "pad_token_id": null,
453
+ "post_norm_init_scale": 0.02,
454
+ "prefix": null,
455
+ "problem_type": null,
456
+ "pruned_heads": {},
457
+ "remove_invalid_values": false,
458
+ "repetition_penalty": 1.0,
459
+ "return_dict": true,
460
+ "return_dict_in_generate": false,
461
+ "sep_token_id": null,
462
+ "suppress_tokens": null,
463
+ "task_specific_params": null,
464
+ "temperature": 1.0,
465
+ "tf_legacy_loss": false,
466
+ "tie_encoder_decoder": false,
467
+ "tie_word_embeddings": true,
468
+ "tokenizer_class": null,
469
+ "top_k": 50,
470
+ "top_p": 1.0,
471
+ "torchscript": false,
472
+ "typical_p": 1.0,
473
+ "use_bfloat16": false,
474
+ "use_post_norm": true
475
+ },
476
+ "output_attentions": false,
477
+ "output_hidden_states": false,
478
+ "output_scores": false,
479
+ "pad_token_id": 151679,
480
+ "prefix": null,
481
+ "problem_type": null,
482
+ "proj_code_bias": true,
483
+ "pruned_heads": {},
484
+ "remove_invalid_values": false,
485
+ "repetition_penalty": 1.0,
486
+ "return_dict": true,
487
+ "return_dict_in_generate": false,
488
+ "sep_token_id": null,
489
+ "speaker_encoder_config": {
490
+ "_name_or_path": "",
491
+ "add_cross_attention": false,
492
+ "architectures": null,
493
+ "bad_words_ids": null,
494
+ "begin_suppress_tokens": null,
495
+ "bos_token_id": null,
496
+ "chunk_size_feed_forward": 0,
497
+ "cross_attention_hidden_size": null,
498
+ "decoder_start_token_id": null,
499
+ "diversity_penalty": 0.0,
500
+ "do_sample": false,
501
+ "dtype": null,
502
+ "early_stopping": false,
503
+ "encoder_no_repeat_ngram_size": 0,
504
+ "encoder_type": "ecapa_tdnn",
505
+ "eos_token_id": null,
506
+ "exponential_decay_length_penalty": null,
507
+ "finetuning_task": null,
508
+ "forced_bos_token_id": null,
509
+ "forced_eos_token_id": null,
510
+ "frame_rate": 12.5,
511
+ "id2label": {
512
+ "0": "LABEL_0",
513
+ "1": "LABEL_1"
514
+ },
515
+ "input_size": 512,
516
+ "is_decoder": false,
517
+ "is_encoder_decoder": false,
518
+ "label2id": {
519
+ "LABEL_0": 0,
520
+ "LABEL_1": 1
521
+ },
522
+ "length_penalty": 1.0,
523
+ "max_length": 20,
524
+ "max_seconds": 10.0,
525
+ "min_length": 0,
526
+ "min_seconds": 2.0,
527
+ "model_type": "speaker_encoder",
528
+ "no_repeat_ngram_size": 0,
529
+ "num_beam_groups": 1,
530
+ "num_beams": 1,
531
+ "num_heads": 8,
532
+ "num_return_sequences": 1,
533
+ "output_attentions": false,
534
+ "output_hidden_states": false,
535
+ "output_scores": false,
536
+ "output_size": 4096,
537
+ "pad_token_id": null,
538
+ "prefix": null,
539
+ "pretrained_dim": 192,
540
+ "pretrained_model_id": "speechbrain/spkrec-ecapa-voxceleb",
541
+ "problem_type": null,
542
+ "pruned_heads": {},
543
+ "remove_invalid_values": false,
544
+ "repetition_penalty": 1.0,
545
+ "return_dict": true,
546
+ "return_dict_in_generate": false,
547
+ "sep_token_id": null,
548
+ "suppress_tokens": null,
549
+ "task_specific_params": null,
550
+ "temperature": 1.0,
551
+ "tf_legacy_loss": false,
552
+ "tie_encoder_decoder": false,
553
+ "tie_word_embeddings": true,
554
+ "tokenizer_class": null,
555
+ "top_k": 50,
556
+ "top_p": 1.0,
557
+ "torchscript": false,
558
+ "typical_p": 1.0,
559
+ "use_bfloat16": false
560
+ },
561
+ "supports_audio_input": true,
562
+ "supports_audio_output": true,
563
+ "suppress_tokens": null,
564
+ "talker_config": {
565
+ "_name_or_path": "",
566
+ "add_cross_attention": false,
567
+ "architectures": null,
568
+ "attention_bias": false,
569
+ "attention_dropout": 0.0,
570
+ "bad_words_ids": null,
571
+ "begin_suppress_tokens": null,
572
+ "bos_token_id": null,
573
+ "chunk_size_feed_forward": 0,
574
+ "cross_attention_hidden_size": null,
575
+ "decoder_start_token_id": null,
576
+ "diversity_penalty": 0.0,
577
+ "do_sample": false,
578
+ "dtype": null,
579
+ "early_stopping": false,
580
+ "encoder_no_repeat_ngram_size": 0,
581
+ "eos_token_id": null,
582
+ "exponential_decay_length_penalty": null,
583
+ "finetuning_task": null,
584
+ "forced_bos_token_id": null,
585
+ "forced_eos_token_id": null,
586
+ "head_dim": 128,
587
+ "hidden_act": "silu",
588
+ "hidden_size": 2048,
589
+ "id2label": {
590
+ "0": "LABEL_0",
591
+ "1": "LABEL_1"
592
+ },
593
+ "initializer_range": 0.02,
594
+ "intermediate_size": 6144,
595
+ "is_decoder": false,
596
+ "is_encoder_decoder": false,
597
+ "label2id": {
598
+ "LABEL_0": 0,
599
+ "LABEL_1": 1
600
+ },
601
+ "layer_types": [
602
+ "full_attention",
603
+ "full_attention",
604
+ "full_attention",
605
+ "full_attention"
606
+ ],
607
+ "length_penalty": 1.0,
608
+ "max_length": 20,
609
+ "max_position_embeddings": 32768,
610
+ "max_window_layers": 28,
611
+ "min_length": 0,
612
+ "model_type": "qwen3",
613
+ "no_repeat_ngram_size": 0,
614
+ "num_attention_heads": 16,
615
+ "num_beam_groups": 1,
616
+ "num_beams": 1,
617
+ "num_hidden_layers": 4,
618
+ "num_key_value_heads": 8,
619
+ "num_return_sequences": 1,
620
+ "output_attentions": false,
621
+ "output_hidden_states": false,
622
+ "output_scores": false,
623
+ "pad_token_id": 151679,
624
+ "prefix": null,
625
+ "problem_type": null,
626
+ "pruned_heads": {},
627
+ "remove_invalid_values": false,
628
+ "repetition_penalty": 1.0,
629
+ "return_dict": true,
630
+ "return_dict_in_generate": false,
631
+ "rms_norm_eps": 1e-06,
632
+ "rope_scaling": null,
633
+ "rope_theta": 10000.0,
634
+ "sep_token_id": null,
635
+ "sliding_window": null,
636
+ "suppress_tokens": null,
637
+ "task_specific_params": null,
638
+ "temperature": 1.0,
639
+ "tf_legacy_loss": false,
640
+ "tie_encoder_decoder": false,
641
+ "tie_word_embeddings": false,
642
+ "tokenizer_class": null,
643
+ "top_k": 50,
644
+ "top_p": 1.0,
645
+ "torchscript": false,
646
+ "typical_p": 1.0,
647
+ "use_bfloat16": false,
648
+ "use_cache": true,
649
+ "use_sliding_window": false,
650
+ "vocab_size": 153723
651
+ },
652
+ "task_specific_params": null,
653
+ "temperature": 1.0,
654
+ "text_model_config": {
655
+ "_name_or_path": "",
656
+ "add_cross_attention": false,
657
+ "architectures": [
658
+ "Qwen3ForCausalLM"
659
+ ],
660
+ "attention_bias": false,
661
+ "attention_dropout": 0.0,
662
+ "bad_words_ids": null,
663
+ "begin_suppress_tokens": null,
664
+ "bos_token_id": null,
665
+ "chunk_size_feed_forward": 0,
666
+ "cross_attention_hidden_size": null,
667
+ "decoder_start_token_id": null,
668
+ "diversity_penalty": 0.0,
669
+ "do_sample": false,
670
+ "dtype": "bfloat16",
671
+ "early_stopping": false,
672
+ "encoder_no_repeat_ngram_size": 0,
673
+ "eos_token_id": null,
674
+ "exponential_decay_length_penalty": null,
675
+ "finetuning_task": null,
676
+ "forced_bos_token_id": null,
677
+ "forced_eos_token_id": null,
678
+ "head_dim": 128,
679
+ "hidden_act": "silu",
680
+ "hidden_size": 4096,
681
+ "id2label": {
682
+ "0": "LABEL_0",
683
+ "1": "LABEL_1"
684
+ },
685
+ "initializer_range": 0.02,
686
+ "intermediate_size": 12288,
687
+ "is_decoder": false,
688
+ "is_encoder_decoder": false,
689
+ "label2id": {
690
+ "LABEL_0": 0,
691
+ "LABEL_1": 1
692
+ },
693
+ "layer_types": [
694
+ "full_attention",
695
+ "full_attention",
696
+ "full_attention",
697
+ "full_attention",
698
+ "full_attention",
699
+ "full_attention",
700
+ "full_attention",
701
+ "full_attention",
702
+ "full_attention",
703
+ "full_attention",
704
+ "full_attention",
705
+ "full_attention",
706
+ "full_attention",
707
+ "full_attention",
708
+ "full_attention",
709
+ "full_attention",
710
+ "full_attention",
711
+ "full_attention",
712
+ "full_attention",
713
+ "full_attention",
714
+ "full_attention",
715
+ "full_attention",
716
+ "full_attention",
717
+ "full_attention",
718
+ "full_attention",
719
+ "full_attention",
720
+ "full_attention",
721
+ "full_attention",
722
+ "full_attention",
723
+ "full_attention",
724
+ "full_attention",
725
+ "full_attention",
726
+ "full_attention",
727
+ "full_attention",
728
+ "full_attention",
729
+ "full_attention"
730
+ ],
731
+ "length_penalty": 1.0,
732
+ "max_length": 20,
733
+ "max_position_embeddings": 262144,
734
+ "max_window_layers": 36,
735
+ "min_length": 0,
736
+ "model_type": "qwen3",
737
+ "no_repeat_ngram_size": 0,
738
+ "num_attention_heads": 32,
739
+ "num_beam_groups": 1,
740
+ "num_beams": 1,
741
+ "num_hidden_layers": 36,
742
+ "num_key_value_heads": 8,
743
+ "num_return_sequences": 1,
744
+ "output_attentions": false,
745
+ "output_hidden_states": false,
746
+ "output_scores": false,
747
+ "pad_token_id": null,
748
+ "prefix": null,
749
+ "problem_type": null,
750
+ "pruned_heads": {},
751
+ "remove_invalid_values": false,
752
+ "repetition_penalty": 1.0,
753
+ "return_dict": true,
754
+ "return_dict_in_generate": false,
755
+ "rms_norm_eps": 1e-06,
756
+ "rope_scaling": null,
757
+ "rope_theta": 5000000,
758
+ "sep_token_id": null,
759
+ "sliding_window": null,
760
+ "suppress_tokens": null,
761
+ "task_specific_params": null,
762
+ "temperature": 1.0,
763
+ "tf_legacy_loss": false,
764
+ "tie_encoder_decoder": false,
765
+ "tie_word_embeddings": false,
766
+ "tokenizer_class": null,
767
+ "top_k": 50,
768
+ "top_p": 1.0,
769
+ "torchscript": false,
770
+ "typical_p": 1.0,
771
+ "use_bfloat16": true,
772
+ "use_cache": true,
773
+ "use_sliding_window": false,
774
+ "vocab_size": 153723
775
+ },
776
+ "tf_legacy_loss": false,
777
+ "thinker_to_talker_intermediate_size": 6144,
778
+ "thinker_to_talker_pre_norm": false,
779
+ "thinker_to_talker_projection_mode": "mlp",
780
+ "tie_encoder_decoder": false,
781
+ "tie_word_embeddings": true,
782
+ "tokenizer_class": null,
783
+ "top_k": 50,
784
+ "top_p": 1.0,
785
+ "torchscript": false,
786
+ "transformers_version": "4.57.3",
787
+ "typical_p": 1.0,
788
+ "use_bfloat16": false,
789
+ "use_duplex_end_pad": false,
790
+ "use_inline_text_prediction": false,
791
+ "auto_map": {
792
+ "AutoConfig": "configuration_raon.RaonConfig",
793
+ "AutoModel": "modeling_raon.RaonModel"
794
+ }
795
+ }
configuration_raon.py ADDED
@@ -0,0 +1,510 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # AUTO-GENERATED — do not edit manually. Run build_hub_files.py to regenerate.
2
+ from __future__ import annotations
3
+
4
+ from copy import deepcopy
5
+ from typing import Any
6
+
7
+ from transformers import PretrainedConfig, MimiConfig, Qwen3Config
8
+ from transformers.models.qwen3_omni_moe.configuration_qwen3_omni_moe import (
9
+ Qwen3OmniMoeAudioEncoderConfig,
10
+ Qwen3OmniMoeTalkerCodePredictorConfig,
11
+ Qwen3OmniMoeTextConfig,
12
+ )
13
+
14
+ # ── from modules/embedding.py ──
15
+
16
+ class EmbeddingAdaptorConfig(PretrainedConfig):
17
+ """Configuration for EmbeddingAdaptor.
18
+
19
+ Controls the projection from audio encoder embeddings to LM embedding space,
20
+ including the time-scale ratio, MLP depth, optional transformer decoder, and
21
+ optional post-projection RMSNorm.
22
+
23
+ Args:
24
+ input_size: Feature dimension of the encoder output (e.g. 512 for Mimi).
25
+ output_size: Feature dimension expected by the LM (e.g. 4096 for Qwen3-7B).
26
+ output_time_scale: Ratio of output frames to input frames. Values >= 1
27
+ upsample (expand time); values < 1 downsample (compress time).
28
+ Must be a reciprocal integer in either direction.
29
+ num_layers: Number of MLP layers (1 or 2). Ignored in transformer mode.
30
+ hidden_size: Hidden dimension for the 2-layer MLP. Defaults to output_size.
31
+ decoder_config: If provided, uses a lightweight Qwen3 transformer instead
32
+ of an MLP for the adaptor projection.
33
+ use_post_norm: If True, apply RMSNorm to the output embeddings.
34
+ norm_eps: Epsilon for RMSNorm.
35
+ post_norm_init_scale: If set, initialize RMSNorm weight to this value
36
+ (useful for residual scaling at initialisation).
37
+ """
38
+
39
+ model_type = "embedding_adaptor"
40
+
41
+ def __init__(
42
+ self,
43
+ input_size: int = 512,
44
+ output_size: int = 4096,
45
+ output_time_scale: float = 1.0,
46
+ num_layers: int = 1,
47
+ hidden_size: int | None = None,
48
+ decoder_config: dict[str, Any] | Qwen3Config | None = None,
49
+ use_post_norm: bool = False,
50
+ norm_eps: float = 1e-6,
51
+ post_norm_init_scale: float | None = None,
52
+ **kwargs: Any,
53
+ ) -> None:
54
+ super().__init__(**kwargs)
55
+ self.input_size = input_size
56
+ self.output_size = output_size
57
+ self.output_time_scale = output_time_scale
58
+ self.num_layers = num_layers
59
+ self.hidden_size = hidden_size
60
+ self.use_post_norm = use_post_norm
61
+ self.norm_eps = norm_eps
62
+ self.post_norm_init_scale = post_norm_init_scale
63
+
64
+ # Parse decoder_config for transformer adaptor mode
65
+ if isinstance(decoder_config, dict):
66
+ decoder_config = Qwen3Config(**decoder_config)
67
+ self.decoder_config = decoder_config
68
+
69
+
70
+ # ── from modules/speaker_encoder.py ──
71
+
72
+ class SpeakerEncoderConfig(PretrainedConfig):
73
+ """Configuration for SpeakerEncoder: input/output sizes, attention heads, and frame window."""
74
+
75
+ model_type = "speaker_encoder"
76
+
77
+ def __init__(
78
+ self,
79
+ input_size: int = 512,
80
+ output_size: int = 4096,
81
+ num_heads: int = 8,
82
+ min_seconds: float = 2.0,
83
+ max_seconds: float = 10.0,
84
+ frame_rate: float = 12.5,
85
+ encoder_type: str = "from_scratch",
86
+ pretrained_model_id: str | None = None,
87
+ pretrained_dim: int | None = None,
88
+ **kwargs: Any,
89
+ ) -> None:
90
+ super().__init__(**kwargs)
91
+ self.input_size = input_size
92
+ self.output_size = output_size
93
+ self.num_heads = num_heads
94
+ self.min_seconds = min_seconds
95
+ self.max_seconds = max_seconds
96
+ self.frame_rate = frame_rate
97
+ self.encoder_type = encoder_type
98
+ self.pretrained_model_id = pretrained_model_id
99
+ self.pretrained_dim = pretrained_dim
100
+
101
+
102
+ # ── from modules/voxtral_encoder.py ──
103
+
104
+ class VoxtralRealtimeEncoderConfig(PretrainedConfig):
105
+ """Configuration for the Voxtral Realtime audio encoder.
106
+
107
+ Stores both the encoder architecture parameters and the projector/downsample
108
+ settings needed to reconstruct the full audio pipeline.
109
+ """
110
+
111
+ model_type = "voxtral_realtime_encoder"
112
+
113
+ def __init__(
114
+ self,
115
+ hidden_size: int = 1280,
116
+ intermediate_size: int = 5120,
117
+ num_hidden_layers: int = 32,
118
+ num_attention_heads: int = 32,
119
+ num_key_value_heads: int | None = None,
120
+ activation_function: str = "gelu",
121
+ num_mel_bins: int = 128,
122
+ initializer_range: float = 0.02,
123
+ attention_dropout: float = 0.0,
124
+ hidden_act: str = "silu",
125
+ max_position_embeddings: int = 1500,
126
+ rms_norm_eps: float = 1e-5,
127
+ rope_theta: float = 10000.0,
128
+ sliding_window: int = 750,
129
+ head_dim: int = 64,
130
+ downsample_factor: int = 4,
131
+ projector_hidden_act: str = "gelu",
132
+ projector_output_size: int | None = None,
133
+ output_embedding_scale: float = 1.0,
134
+ skip_projector: bool = False,
135
+ attn_implementation: str = "eager",
136
+ **kwargs: Any,
137
+ ) -> None:
138
+ super().__init__(**kwargs)
139
+ self.hidden_size = hidden_size
140
+ self.intermediate_size = intermediate_size
141
+ self.num_hidden_layers = num_hidden_layers
142
+ self.num_attention_heads = num_attention_heads
143
+ self.num_key_value_heads = num_key_value_heads if num_key_value_heads is not None else num_attention_heads
144
+ self.activation_function = activation_function
145
+ self.num_mel_bins = num_mel_bins
146
+ self.initializer_range = initializer_range
147
+ self.attention_dropout = attention_dropout
148
+ self.hidden_act = hidden_act
149
+ self.max_position_embeddings = max_position_embeddings
150
+ self.rms_norm_eps = rms_norm_eps
151
+ self.rope_theta = rope_theta
152
+ self.sliding_window = sliding_window
153
+ self.head_dim = head_dim if head_dim is not None else hidden_size // num_attention_heads
154
+ self.downsample_factor = downsample_factor
155
+ self.projector_hidden_act = projector_hidden_act
156
+ self.projector_output_size = projector_output_size
157
+ self.output_embedding_scale = output_embedding_scale
158
+ self.skip_projector = skip_projector
159
+ self._attn_implementation = attn_implementation
160
+
161
+ # Aliases expected by the encoder layers.
162
+ self.encoder_layers = num_hidden_layers
163
+ self.encoder_attention_heads = num_attention_heads
164
+
165
+ @classmethod
166
+ def from_pretrained(
167
+ cls,
168
+ pretrained_model_name_or_path: str,
169
+ **kwargs: Any,
170
+ ) -> "VoxtralRealtimeEncoderConfig":
171
+ """Load config from a Voxtral Realtime checkpoint.
172
+
173
+ Reads ``config.json`` and extracts the ``audio_config`` sub-dict along
174
+ with top-level ``downsample_factor``, ``projector_hidden_act``, and
175
+ ``text_config.hidden_size`` (used as ``projector_output_size``).
176
+
177
+ Works with both the full ``voxtral_realtime`` model config and a
178
+ standalone ``voxtral_realtime_encoder`` config.
179
+
180
+ Args:
181
+ pretrained_model_name_or_path: HuggingFace model ID or local path.
182
+
183
+ Returns:
184
+ Populated ``VoxtralRealtimeEncoderConfig``.
185
+ """
186
+ import json
187
+ import os
188
+
189
+ from huggingface_hub import hf_hub_download
190
+
191
+ is_local = os.path.isdir(pretrained_model_name_or_path)
192
+ if is_local:
193
+ config_path = os.path.join(pretrained_model_name_or_path, "config.json")
194
+ else:
195
+ config_path = hf_hub_download(
196
+ repo_id=pretrained_model_name_or_path,
197
+ filename="config.json",
198
+ )
199
+
200
+ with open(config_path) as f:
201
+ full_config = json.load(f)
202
+
203
+ # If this is the full model config, extract the audio sub-config.
204
+ if "audio_config" in full_config:
205
+ audio_cfg = full_config["audio_config"]
206
+ downsample_factor = full_config.get("downsample_factor", 4)
207
+ projector_hidden_act = full_config.get("projector_hidden_act", "gelu")
208
+ text_hidden_size = full_config.get("text_config", {}).get("hidden_size")
209
+ else:
210
+ # Standalone encoder config (e.g. saved by us).
211
+ audio_cfg = full_config
212
+ downsample_factor = audio_cfg.get("downsample_factor", 4)
213
+ projector_hidden_act = audio_cfg.get("projector_hidden_act", "gelu")
214
+ text_hidden_size = audio_cfg.get("projector_output_size")
215
+
216
+ # The upstream rope_theta lives inside rope_parameters.
217
+ rope_params = audio_cfg.get("rope_parameters") or {}
218
+ rope_theta = rope_params.get("rope_theta", audio_cfg.get("rope_theta", 10000.0))
219
+
220
+ return cls(
221
+ hidden_size=audio_cfg.get("hidden_size", 1280),
222
+ intermediate_size=audio_cfg.get("intermediate_size", 5120),
223
+ num_hidden_layers=audio_cfg.get("num_hidden_layers", 32),
224
+ num_attention_heads=audio_cfg.get("num_attention_heads", 32),
225
+ num_key_value_heads=audio_cfg.get("num_key_value_heads"),
226
+ activation_function=audio_cfg.get("activation_function", "gelu"),
227
+ num_mel_bins=audio_cfg.get("num_mel_bins", 128),
228
+ initializer_range=audio_cfg.get("initializer_range", 0.02),
229
+ attention_dropout=audio_cfg.get("attention_dropout", 0.0),
230
+ hidden_act=audio_cfg.get("hidden_act", "silu"),
231
+ max_position_embeddings=audio_cfg.get("max_position_embeddings", 1500),
232
+ rms_norm_eps=audio_cfg.get("rms_norm_eps", 1e-5),
233
+ rope_theta=rope_theta,
234
+ sliding_window=audio_cfg.get("sliding_window", 750),
235
+ head_dim=audio_cfg.get("head_dim", 64),
236
+ downsample_factor=downsample_factor,
237
+ projector_hidden_act=projector_hidden_act,
238
+ projector_output_size=text_hidden_size,
239
+ **kwargs,
240
+ )
241
+
242
+
243
+ # ---------------------------------------------------------------------------
244
+ # Conv1d padding cache (for streaming)
245
+ # ---------------------------------------------------------------------------
246
+
247
+
248
+ # ── from models/raon.py ──
249
+
250
+ TEXT_MODEL_CONFIGS: dict[str, type[PretrainedConfig]] = {
251
+ Qwen3Config.model_type: Qwen3Config,
252
+ }
253
+
254
+
255
+ class RaonConfig(PretrainedConfig):
256
+ """Configuration class for RaonModel."""
257
+
258
+ model_type = "raon"
259
+ has_no_defaults_at_init = True
260
+ text_model_config: PretrainedConfig = None
261
+ audio_encoder_config: Qwen3OmniMoeAudioEncoderConfig | VoxtralRealtimeEncoderConfig = None
262
+ audio_tokenizer_config: MimiConfig = None
263
+ input_adaptor_config: EmbeddingAdaptorConfig = None
264
+ output_adaptor_config: EmbeddingAdaptorConfig = None
265
+ code_predictor_config: Qwen3OmniMoeTalkerCodePredictorConfig = None
266
+ speaker_encoder_config: SpeakerEncoderConfig | None = None
267
+ # Note: speaker_encoder_config is intentionally excluded from sub_configs.
268
+ # It is optional (can be None), and transformers' _get_dtype unconditionally
269
+ # calls sub_config.dtype on every entry, which crashes on None.
270
+ # Deserialization from dict is handled in __init__ instead.
271
+ sub_configs = {
272
+ "text_model_config": PretrainedConfig,
273
+ "audio_encoder_config": PretrainedConfig,
274
+ "audio_tokenizer_config": PretrainedConfig,
275
+ "input_adaptor_config": EmbeddingAdaptorConfig,
276
+ "output_adaptor_config": EmbeddingAdaptorConfig,
277
+ "code_predictor_config": Qwen3OmniMoeTalkerCodePredictorConfig,
278
+ }
279
+
280
+ def __init__(
281
+ self,
282
+ *,
283
+ text_model_config: dict[str, Any] | PretrainedConfig | None = None,
284
+ audio_encoder_config: dict[str, Any] | Qwen3OmniMoeAudioEncoderConfig | VoxtralRealtimeEncoderConfig | None = None,
285
+ audio_tokenizer_config: dict[str, Any] | MimiConfig | None = None,
286
+ input_adaptor_config: dict[str, Any] | EmbeddingAdaptorConfig | None = None,
287
+ output_adaptor_config: dict[str, Any] | EmbeddingAdaptorConfig | None = None,
288
+ code_predictor_config: dict[str, Any] | Qwen3OmniMoeTalkerCodePredictorConfig | None = None,
289
+ speaker_encoder_config: dict[str, Any] | SpeakerEncoderConfig | None = None,
290
+ num_talker_layers: int = 0,
291
+ supports_audio_input: bool = True,
292
+ supports_audio_output: bool = True,
293
+ aut_is_causal: bool = False,
294
+ proj_code_bias: bool = False,
295
+ accept_hidden_layer: int = -1,
296
+ talker_config: dict[str, Any] | PretrainedConfig | None = None,
297
+ thinker_to_talker_pre_norm: bool = False,
298
+ sequence_mode: str | None = None,
299
+ use_sil_token: bool = False,
300
+ no_audio_in_sil: bool = False,
301
+ text_lookahead: int = 0,
302
+ use_duplex_end_pad: bool = False,
303
+ speaker_embedding_to_code_predictor: bool = True,
304
+ duplex_pad_token_id: int | None = None,
305
+ duplex_end_pad_token_id: int | None = None,
306
+ duplex_sil_token_id: int | None = None,
307
+ duplex_bc_token_id: int | None = None,
308
+ use_backchannel_token: bool = False,
309
+ bc_loss_weight: float = 1.0,
310
+ speaker_token_id: int | None = None,
311
+ audio_input_token_id: int | None = None,
312
+ audio_output_token_id: int | None = None,
313
+ audio_start_token_id: int | None = None,
314
+ im_start_token_id: int | None = None,
315
+ text_loss_weight: float = 1.0,
316
+ sil_loss_weight: float = 1.0,
317
+ epad_loss_weight: float = 0.0,
318
+ semantic_loss_weight: float = 1.0,
319
+ acoustic_loss_weights: list[float] | None = None,
320
+ audio_lm_head_enabled: bool = True,
321
+ delays: list[int] | None = None,
322
+ **kwargs: Any,
323
+ ) -> None:
324
+ super().__init__(**kwargs)
325
+
326
+ # Ensure auto_map is always serialized for trust_remote_code Hub loading.
327
+ if not hasattr(self, "auto_map") or not self.auto_map:
328
+ self.auto_map = {
329
+ "AutoConfig": "configuration_raon.RaonConfig",
330
+ "AutoModel": "modeling_raon.RaonModel",
331
+ }
332
+
333
+ assert text_model_config is not None, "RaonConfig: `text_model_config` is required."
334
+ assert audio_encoder_config is not None, "RaonConfig: `audio_encoder_config` is required."
335
+ assert audio_tokenizer_config is not None, "RaonConfig: `audio_tokenizer_config` is required."
336
+ assert input_adaptor_config is not None, "RaonConfig: `input_adaptor_config` is required."
337
+ assert output_adaptor_config is not None, "RaonConfig: `output_adaptor_config` is required."
338
+ assert code_predictor_config is not None, "RaonConfig: `code_predictor_config` is required."
339
+
340
+ if isinstance(text_model_config, dict):
341
+ model_type = text_model_config.get("model_type", Qwen3Config.model_type)
342
+ text_model_config = TEXT_MODEL_CONFIGS[model_type](**text_model_config)
343
+
344
+ # Convert sub-configs from dict or generic PretrainedConfig to specific types.
345
+ # The generic PretrainedConfig case occurs when transformers' sub_configs mechanism
346
+ # auto-deserializes before __init__ runs (e.g. with trust_remote_code Hub loading).
347
+ def _to_dict(cfg: Any) -> dict[str, Any]:
348
+ """Convert a config to dict, handling both dict and PretrainedConfig."""
349
+ if isinstance(cfg, dict):
350
+ return cfg
351
+ return cfg.to_dict()
352
+
353
+ if isinstance(audio_encoder_config, dict) or (
354
+ isinstance(audio_encoder_config, PretrainedConfig)
355
+ and not isinstance(audio_encoder_config, (Qwen3OmniMoeAudioEncoderConfig, VoxtralRealtimeEncoderConfig))
356
+ ):
357
+ d = _to_dict(audio_encoder_config)
358
+ model_type = d.get("model_type", Qwen3OmniMoeAudioEncoderConfig.model_type)
359
+ if model_type == Qwen3OmniMoeAudioEncoderConfig.model_type:
360
+ audio_encoder_config = Qwen3OmniMoeAudioEncoderConfig(**d)
361
+ elif model_type == "voxtral_realtime_encoder":
362
+ audio_encoder_config = VoxtralRealtimeEncoderConfig(**d)
363
+ else:
364
+ raise ValueError(
365
+ f"Unsupported audio_encoder model_type: {model_type!r}. "
366
+ "Expected 'qwen3_omni_moe_audio_encoder' or 'voxtral_realtime_encoder'."
367
+ )
368
+
369
+ if isinstance(audio_tokenizer_config, dict) or (
370
+ isinstance(audio_tokenizer_config, PretrainedConfig) and not isinstance(audio_tokenizer_config, MimiConfig)
371
+ ):
372
+ audio_tokenizer_config = MimiConfig(**_to_dict(audio_tokenizer_config))
373
+
374
+ if isinstance(input_adaptor_config, dict) or (
375
+ isinstance(input_adaptor_config, PretrainedConfig)
376
+ and not isinstance(input_adaptor_config, EmbeddingAdaptorConfig)
377
+ ):
378
+ input_adaptor_config = EmbeddingAdaptorConfig(**_to_dict(input_adaptor_config))
379
+
380
+ if isinstance(output_adaptor_config, dict) or (
381
+ isinstance(output_adaptor_config, PretrainedConfig)
382
+ and not isinstance(output_adaptor_config, EmbeddingAdaptorConfig)
383
+ ):
384
+ output_adaptor_config = EmbeddingAdaptorConfig(**_to_dict(output_adaptor_config))
385
+
386
+ if isinstance(code_predictor_config, dict) or (
387
+ isinstance(code_predictor_config, PretrainedConfig)
388
+ and not isinstance(code_predictor_config, Qwen3OmniMoeTalkerCodePredictorConfig)
389
+ ):
390
+ code_predictor_config = Qwen3OmniMoeTalkerCodePredictorConfig(**_to_dict(code_predictor_config))
391
+
392
+ if isinstance(speaker_encoder_config, dict) or (
393
+ isinstance(speaker_encoder_config, PretrainedConfig)
394
+ and not isinstance(speaker_encoder_config, SpeakerEncoderConfig)
395
+ ):
396
+ speaker_encoder_config = SpeakerEncoderConfig(**_to_dict(speaker_encoder_config))
397
+
398
+ if isinstance(talker_config, dict) or (
399
+ isinstance(talker_config, PretrainedConfig) and type(talker_config) is PretrainedConfig
400
+ ):
401
+ d = _to_dict(talker_config) if talker_config is not None else {}
402
+ talker_model_type = d.get("model_type", Qwen3Config.model_type)
403
+ talker_config = TEXT_MODEL_CONFIGS[talker_model_type](**d)
404
+
405
+ assert isinstance(
406
+ audio_encoder_config, (Qwen3OmniMoeAudioEncoderConfig, VoxtralRealtimeEncoderConfig, MimiConfig)
407
+ ), "audio_encoder_config must be Qwen3OmniMoeAudioEncoderConfig, VoxtralRealtimeEncoderConfig, or MimiConfig."
408
+ assert isinstance(audio_tokenizer_config, MimiConfig), "audio_tokenizer_config must be MimiConfig."
409
+ assert isinstance(input_adaptor_config, EmbeddingAdaptorConfig), (
410
+ "input_adaptor_config must be EmbeddingAdaptorConfig."
411
+ )
412
+ assert isinstance(output_adaptor_config, EmbeddingAdaptorConfig), (
413
+ "output_adaptor_config must be EmbeddingAdaptorConfig."
414
+ )
415
+ assert isinstance(code_predictor_config, Qwen3OmniMoeTalkerCodePredictorConfig), (
416
+ "code_predictor_config must be Qwen3OmniMoeTalkerCodePredictorConfig."
417
+ )
418
+ assert isinstance(text_model_config, PretrainedConfig), "text_model_config must be PretrainedConfig."
419
+ assert speaker_encoder_config is None or isinstance(speaker_encoder_config, SpeakerEncoderConfig), (
420
+ "speaker_encoder_config must be None or SpeakerEncoderConfig."
421
+ )
422
+
423
+ self.text_model_config = text_model_config
424
+ self.audio_encoder_config = audio_encoder_config
425
+ self.audio_tokenizer_config = audio_tokenizer_config
426
+ self.input_adaptor_config = input_adaptor_config
427
+ self.output_adaptor_config = output_adaptor_config
428
+ self.code_predictor_config = code_predictor_config
429
+ self.speaker_encoder_config = speaker_encoder_config
430
+ self.num_talker_layers = num_talker_layers
431
+ self.supports_audio_input = supports_audio_input
432
+ self.supports_audio_output = supports_audio_output
433
+ self.aut_is_causal = aut_is_causal
434
+ self.proj_code_bias = proj_code_bias
435
+ self.accept_hidden_layer = accept_hidden_layer
436
+ self.talker_config = talker_config
437
+ self.thinker_to_talker_pre_norm = thinker_to_talker_pre_norm
438
+ self.sequence_mode = sequence_mode
439
+ self.use_sil_token = use_sil_token
440
+ self.no_audio_in_sil = no_audio_in_sil
441
+ self.text_lookahead = int(text_lookahead)
442
+ self.use_duplex_end_pad = use_duplex_end_pad
443
+ self.speaker_embedding_to_code_predictor = speaker_embedding_to_code_predictor
444
+ self.duplex_pad_token_id = duplex_pad_token_id
445
+ self.duplex_end_pad_token_id = duplex_end_pad_token_id
446
+ self.duplex_sil_token_id = duplex_sil_token_id
447
+ self.duplex_bc_token_id = duplex_bc_token_id
448
+ self.use_backchannel_token = use_backchannel_token
449
+ self.bc_loss_weight = bc_loss_weight
450
+ self.speaker_token_id = speaker_token_id
451
+ self.audio_input_token_id = audio_input_token_id
452
+ self.audio_output_token_id = audio_output_token_id
453
+ self.audio_start_token_id = audio_start_token_id
454
+ self.im_start_token_id = im_start_token_id
455
+ self.text_loss_weight = text_loss_weight
456
+ self.sil_loss_weight = sil_loss_weight
457
+ self.epad_loss_weight = epad_loss_weight
458
+ self.semantic_loss_weight = semantic_loss_weight
459
+ self.acoustic_loss_weights = acoustic_loss_weights
460
+ self.audio_lm_head_enabled = audio_lm_head_enabled
461
+ self.delays = delays
462
+
463
+ if supports_audio_output and audio_lm_head_enabled:
464
+ assert talker_config is not None, "RaonConfig: `talker_config` is required when audio output is enabled."
465
+ assert num_talker_layers > 0, "RaonConfig: `num_talker_layers` must be positive when audio output is enabled."
466
+
467
+ def _get_non_default_generation_parameters(self) -> dict[str, Any]:
468
+ return {}
469
+
470
+ def to_diff_dict(self) -> dict[str, Any]:
471
+ """Return config as a dict suitable for diffing."""
472
+ return self.to_dict()
473
+
474
+
475
+ class RaonDuplexConfig(RaonConfig):
476
+ """Configuration alias for full-duplex checkpoints (model_type='raon_duplex')."""
477
+
478
+ model_type = "raon_duplex"
479
+
480
+ def __init__(self, **kwargs: Any) -> None:
481
+ # Duplex-specific defaults; overridden by values in config.json when present.
482
+ kwargs.setdefault("sequence_mode", "uta")
483
+ kwargs.setdefault("use_sil_token", True)
484
+ kwargs.setdefault("no_audio_in_sil", False)
485
+ kwargs.setdefault("text_lookahead", 0)
486
+ kwargs.setdefault("use_duplex_end_pad", True)
487
+ kwargs.setdefault("duplex_pad_token_id", 151677)
488
+ kwargs.setdefault("duplex_end_pad_token_id", 151678)
489
+ kwargs.setdefault("duplex_sil_token_id", 151672)
490
+ kwargs.setdefault("duplex_bc_token_id", 151673)
491
+ kwargs.setdefault("speaker_token_id", 151671)
492
+ kwargs.setdefault("audio_input_token_id", 151676)
493
+ kwargs.setdefault("audio_output_token_id", 151675)
494
+ kwargs.setdefault("audio_start_token_id", 151669)
495
+ kwargs.setdefault("im_start_token_id", 151644)
496
+ # Loss weight defaults (overridable at training time via duplex_train args)
497
+ kwargs.setdefault("text_loss_weight", 1.0)
498
+ kwargs.setdefault("sil_loss_weight", 1.0)
499
+ kwargs.setdefault("epad_loss_weight", 0.0)
500
+ kwargs.setdefault("semantic_loss_weight", 1.0)
501
+ kwargs.setdefault("acoustic_loss_weights", None)
502
+ super().__init__(**kwargs)
503
+ self.auto_map = {
504
+ "AutoConfig": "configuration_raon.RaonDuplexConfig",
505
+ "AutoModel": "modeling_raon.RaonDuplexModel",
506
+ }
507
+
508
+
509
+ # Duplex model — same architecture, different model_type for HF registry
510
+
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model-00001-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dff45b01197c041324d528c65d5551670c1292aa40874beea4c571bfaae2d7c5
3
+ size 4916897336
model-00002-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2e24f7df5232cf22e66b276bf3f545c78d649fd53fd66e9282a65da5fda3277d
3
+ size 4915961072
model-00003-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:20c5e7c016455faa9a854af8e6b2f9b391e5a59facd603b70ea427db522a3bb2
3
+ size 4983069200
model-00004-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7776e8b1c4bb0a1ea47c1adb98d689cd488164bc067dfea1e8bfc5aeb4a964ab
3
+ size 3290040026
model.safetensors.index.json ADDED
The diff for this file is too large to render. See raw diff
 
modeling_raon.py ADDED
The diff for this file is too large to render. See raw diff
 
special_tokens_map.json ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "additional_special_tokens": [
3
+ "<|audio_input_placeholder|>"
4
+ ],
5
+ "audio_bos_token": "<|audio_start|>",
6
+ "audio_eos_token": "<|audio_end|>",
7
+ "audio_token": "<|audio_pad|>",
8
+ "eos_token": {
9
+ "content": "<|im_end|>",
10
+ "lstrip": false,
11
+ "normalized": false,
12
+ "rstrip": false,
13
+ "single_word": false
14
+ },
15
+ "image_token": "<|image_pad|>",
16
+ "pad_token": {
17
+ "content": "<|endoftext|>",
18
+ "lstrip": false,
19
+ "normalized": false,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ },
23
+ "video_token": "<|video_pad|>",
24
+ "vision_bos_token": "<|vision_start|>",
25
+ "vision_eos_token": "<|vision_end|>"
26
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c865b4da59eb914ef2e574352533b8b09c8f8ae7c5ef64b0a3ddb35d884e3d00
3
+ size 11426101
tokenizer_config.json ADDED
@@ -0,0 +1,346 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": false,
3
+ "add_prefix_space": false,
4
+ "added_tokens_decoder": {
5
+ "151644": {
6
+ "content": "<|im_start|>",
7
+ "lstrip": false,
8
+ "normalized": false,
9
+ "rstrip": false,
10
+ "single_word": false,
11
+ "special": true
12
+ },
13
+ "151645": {
14
+ "content": "<|im_end|>",
15
+ "lstrip": false,
16
+ "normalized": false,
17
+ "rstrip": false,
18
+ "single_word": false,
19
+ "special": true
20
+ },
21
+ "151652": {
22
+ "content": "<|vision_start|>",
23
+ "lstrip": false,
24
+ "normalized": false,
25
+ "rstrip": false,
26
+ "single_word": false,
27
+ "special": true
28
+ },
29
+ "151653": {
30
+ "content": "<|vision_end|>",
31
+ "lstrip": false,
32
+ "normalized": false,
33
+ "rstrip": false,
34
+ "single_word": false,
35
+ "special": true
36
+ },
37
+ "151654": {
38
+ "content": "<|vision_pad|>",
39
+ "lstrip": false,
40
+ "normalized": false,
41
+ "rstrip": false,
42
+ "single_word": false,
43
+ "special": true
44
+ },
45
+ "151655": {
46
+ "content": "<|image_pad|>",
47
+ "lstrip": false,
48
+ "normalized": false,
49
+ "rstrip": false,
50
+ "single_word": false,
51
+ "special": true
52
+ },
53
+ "151656": {
54
+ "content": "<|video_pad|>",
55
+ "lstrip": false,
56
+ "normalized": false,
57
+ "rstrip": false,
58
+ "single_word": false,
59
+ "special": true
60
+ },
61
+ "151657": {
62
+ "content": "<tool_call>",
63
+ "lstrip": false,
64
+ "normalized": false,
65
+ "rstrip": false,
66
+ "single_word": false,
67
+ "special": false
68
+ },
69
+ "151658": {
70
+ "content": "</tool_call>",
71
+ "lstrip": false,
72
+ "normalized": false,
73
+ "rstrip": false,
74
+ "single_word": false,
75
+ "special": false
76
+ },
77
+ "151659": {
78
+ "content": "<|fim_prefix|>",
79
+ "lstrip": false,
80
+ "normalized": false,
81
+ "rstrip": false,
82
+ "single_word": false,
83
+ "special": false
84
+ },
85
+ "151660": {
86
+ "content": "<|fim_middle|>",
87
+ "lstrip": false,
88
+ "normalized": false,
89
+ "rstrip": false,
90
+ "single_word": false,
91
+ "special": false
92
+ },
93
+ "151661": {
94
+ "content": "<|fim_suffix|>",
95
+ "lstrip": false,
96
+ "normalized": false,
97
+ "rstrip": false,
98
+ "single_word": false,
99
+ "special": false
100
+ },
101
+ "151662": {
102
+ "content": "<|fim_pad|>",
103
+ "lstrip": false,
104
+ "normalized": false,
105
+ "rstrip": false,
106
+ "single_word": false,
107
+ "special": false
108
+ },
109
+ "151663": {
110
+ "content": "<|repo_name|>",
111
+ "lstrip": false,
112
+ "normalized": false,
113
+ "rstrip": false,
114
+ "single_word": false,
115
+ "special": false
116
+ },
117
+ "151664": {
118
+ "content": "<|file_sep|>",
119
+ "lstrip": false,
120
+ "normalized": false,
121
+ "rstrip": false,
122
+ "single_word": false,
123
+ "special": false
124
+ },
125
+ "151665": {
126
+ "content": "<tool_response>",
127
+ "lstrip": false,
128
+ "normalized": false,
129
+ "rstrip": false,
130
+ "single_word": false,
131
+ "special": false
132
+ },
133
+ "151666": {
134
+ "content": "</tool_response>",
135
+ "lstrip": false,
136
+ "normalized": false,
137
+ "rstrip": false,
138
+ "single_word": false,
139
+ "special": false
140
+ },
141
+ "151667": {
142
+ "content": "<think>",
143
+ "lstrip": false,
144
+ "normalized": false,
145
+ "rstrip": false,
146
+ "single_word": false,
147
+ "special": false
148
+ },
149
+ "151668": {
150
+ "content": "</think>",
151
+ "lstrip": false,
152
+ "normalized": false,
153
+ "rstrip": false,
154
+ "single_word": false,
155
+ "special": false
156
+ },
157
+ "151669": {
158
+ "content": "<|audio_start|>",
159
+ "lstrip": false,
160
+ "normalized": false,
161
+ "rstrip": false,
162
+ "single_word": false,
163
+ "special": true
164
+ },
165
+ "151670": {
166
+ "content": "<|audio_end|>",
167
+ "lstrip": false,
168
+ "normalized": false,
169
+ "rstrip": false,
170
+ "single_word": false,
171
+ "special": true
172
+ },
173
+ "151671": {
174
+ "content": "<|speaker_embedding_placeholder|>",
175
+ "lstrip": false,
176
+ "normalized": false,
177
+ "rstrip": false,
178
+ "single_word": false,
179
+ "special": true
180
+ },
181
+ "151675": {
182
+ "content": "<|audio_output_placeholder|>",
183
+ "lstrip": false,
184
+ "normalized": false,
185
+ "rstrip": false,
186
+ "single_word": false,
187
+ "special": true
188
+ },
189
+ "151676": {
190
+ "content": "<|audio_input_placeholder|>",
191
+ "lstrip": false,
192
+ "normalized": false,
193
+ "rstrip": false,
194
+ "single_word": false,
195
+ "special": true
196
+ },
197
+ "151677": {
198
+ "content": "<|audio_output_pad|>",
199
+ "lstrip": false,
200
+ "normalized": false,
201
+ "rstrip": false,
202
+ "single_word": false,
203
+ "special": true
204
+ },
205
+ "151678": {
206
+ "content": "<|audio_output_end_pad|>",
207
+ "lstrip": false,
208
+ "normalized": false,
209
+ "rstrip": false,
210
+ "single_word": false,
211
+ "special": true
212
+ },
213
+ "151679": {
214
+ "content": "<|endoftext|>",
215
+ "lstrip": false,
216
+ "normalized": false,
217
+ "rstrip": false,
218
+ "single_word": false,
219
+ "special": true
220
+ },
221
+ "151680": {
222
+ "content": "<tts_text_bos>",
223
+ "lstrip": false,
224
+ "normalized": false,
225
+ "rstrip": false,
226
+ "single_word": false,
227
+ "special": true
228
+ },
229
+ "151681": {
230
+ "content": "<tts_text_eod>",
231
+ "lstrip": false,
232
+ "normalized": false,
233
+ "rstrip": false,
234
+ "single_word": false,
235
+ "special": true
236
+ },
237
+ "151682": {
238
+ "content": "<tts_text_bos_single>",
239
+ "lstrip": false,
240
+ "normalized": false,
241
+ "rstrip": false,
242
+ "single_word": false,
243
+ "special": true
244
+ },
245
+ "151683": {
246
+ "content": "<|audio_pad|>",
247
+ "lstrip": false,
248
+ "normalized": false,
249
+ "rstrip": false,
250
+ "single_word": false,
251
+ "special": true
252
+ },
253
+ "151684": {
254
+ "content": "<|secondary_audio_pad|>",
255
+ "lstrip": false,
256
+ "normalized": false,
257
+ "rstrip": false,
258
+ "single_word": false,
259
+ "special": true
260
+ },
261
+ "151685": {
262
+ "content": "<|object_ref_start|>",
263
+ "lstrip": false,
264
+ "normalized": false,
265
+ "rstrip": false,
266
+ "single_word": false,
267
+ "special": true
268
+ },
269
+ "151686": {
270
+ "content": "<|object_ref_end|>",
271
+ "lstrip": false,
272
+ "normalized": false,
273
+ "rstrip": false,
274
+ "single_word": false,
275
+ "special": true
276
+ },
277
+ "151687": {
278
+ "content": "<|box_start|>",
279
+ "lstrip": false,
280
+ "normalized": false,
281
+ "rstrip": false,
282
+ "single_word": false,
283
+ "special": true
284
+ },
285
+ "151688": {
286
+ "content": "<|box_end|>",
287
+ "lstrip": false,
288
+ "normalized": false,
289
+ "rstrip": false,
290
+ "single_word": false,
291
+ "special": true
292
+ },
293
+ "151689": {
294
+ "content": "<|quad_start|>",
295
+ "lstrip": false,
296
+ "normalized": false,
297
+ "rstrip": false,
298
+ "single_word": false,
299
+ "special": true
300
+ },
301
+ "151690": {
302
+ "content": "<|quad_end|>",
303
+ "lstrip": false,
304
+ "normalized": false,
305
+ "rstrip": false,
306
+ "single_word": false,
307
+ "special": true
308
+ },
309
+ "151691": {
310
+ "content": "<tts_pad>",
311
+ "lstrip": false,
312
+ "normalized": false,
313
+ "rstrip": false,
314
+ "single_word": false,
315
+ "special": true
316
+ }
317
+ },
318
+ "additional_special_tokens": [
319
+ "<|audio_input_placeholder|>"
320
+ ],
321
+ "audio_bos_token": "<|audio_start|>",
322
+ "audio_eos_token": "<|audio_end|>",
323
+ "audio_token": "<|audio_pad|>",
324
+ "bos_token": null,
325
+ "clean_up_tokenization_spaces": false,
326
+ "eos_token": "<|im_end|>",
327
+ "errors": "replace",
328
+ "extra_special_tokens": {
329
+ "audio_bos_token": "<|audio_start|>",
330
+ "audio_eos_token": "<|audio_end|>",
331
+ "audio_token": "<|audio_pad|>",
332
+ "image_token": "<|image_pad|>",
333
+ "video_token": "<|video_pad|>",
334
+ "vision_bos_token": "<|vision_start|>",
335
+ "vision_eos_token": "<|vision_end|>"
336
+ },
337
+ "image_token": "<|image_pad|>",
338
+ "model_max_length": 131072,
339
+ "pad_token": "<|endoftext|>",
340
+ "split_special_tokens": false,
341
+ "tokenizer_class": "Qwen2Tokenizer",
342
+ "unk_token": null,
343
+ "video_token": "<|video_pad|>",
344
+ "vision_bos_token": "<|vision_start|>",
345
+ "vision_eos_token": "<|vision_end|>"
346
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff