File size: 24,306 Bytes
54edb09
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
122fc80
 
54edb09
122fc80
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
54edb09
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b0f00f6
54edb09
 
 
b0f00f6
 
 
 
 
54edb09
84bdf64
 
 
 
 
54edb09
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
# VOICEVOX setup (VOIC-01)

Everything in this document was **executed**, not inferred. `01-RESEARCH.md`'s code sample was
marked `[A]`; every name below was confirmed either against the release's own materials or by
running it against an installed `voicevox_core` 0.17.0 on Python 3.12.12.

Reconnaissance date: **2026-08-27**.

---

## Confirmed API surface — voicevox_core 0.17.0

Import root: `from voicevox_core.blocking import Onnxruntime, OpenJtalk, Synthesizer, VoiceModelFile`
(`voicevox_core.asyncio` mirrors it; we use `blocking` because synthesis is CPU-bound and
`Synthesizer` is internally mutex-guarded anyway).

| Call | Confirmed signature | Source |
|---|---|---|
| `Onnxruntime.load_once(filename=...)` | `load_once(*, filename: str = LIB_RECOMMENDED_VERSIONED_FILENAME) -> Onnxruntime` | [`docs/guide/user/usage.md` @ tag 0.17.0](https://github.com/VOICEVOX/voicevox_core/blob/0.17.0/docs/guide/user/usage.md) + `dir(Onnxruntime)` on the installed wheel |
| `OpenJtalk(dict_dir)` | `OpenJtalk(open_jtalk_dict_dir: str \| PathLike)` | same `usage.md`; executed |
| `Synthesizer(onnxruntime, open_jtalk)` | `Synthesizer(Onnxruntime, OpenJtalk, *, acceleration_mode=..., cpu_num_threads=...)` | same `usage.md`; executed |
| `VoiceModelFile.open(path)` | classmethod, returns a context manager; `.id`, `.metas` | same `usage.md`; executed |
| `synthesizer.load_voice_model(model)` | `load_voice_model(model, *, on_existing=...)` — `on_existing` is new in 0.17.0 | [release 0.17.0 body](https://github.com/VOICEVOX/voicevox_core/releases/tag/0.17.0); executed |
| `synthesizer.create_audio_query(text, style_id)` | `-> voicevox_core.AudioQuery` | same `usage.md`; executed |
| `synthesizer.synthesis(audio_query, style_id)` | `-> bytes` (RIFF WAV) | same `usage.md`; executed |
| `Onnxruntime.LIB_RECOMMENDED_VERSION` | `"1.23.2"` | executed against the installed wheel |
| `Onnxruntime.LIB_RECOMMENDED_NAME` | `"voicevox_onnxruntime"` | executed |
| `Onnxruntime.LIB_RECOMMENDED_VERSIONED_FILENAME` | `"voicevox_onnxruntime.dll"` on Windows; platform-dependent | executed |

`AudioQuery`, `AccentPhrase` and `Mora` are **real Python `dataclasses`** (`dataclasses.is_dataclass`
returns `True` for all three), which is what makes the plain-dict conversion in `tts.py` exact
rather than hand-transcribed.

| Type | `dataclasses.fields()` | Source |
|---|---|---|
| `AudioQuery` | `accent_phrases, speed_scale, pitch_scale, intonation_scale, volume_scale, pre_phoneme_length, post_phoneme_length, output_sampling_rate, output_stereo, kana` | executed |
| `AccentPhrase` | `moras, accent, pause_mora, is_interrogative` | executed |
| `Mora` | `text, vowel, vowel_length, pitch, consonant, consonant_length` | executed |

---

## Corrections to 01-RESEARCH.md

Five differences. All five matter to a later plan.

**1. `voicevox_vvm` is on 0.17.0, not 0.16.4 — and the version-track mismatch does not exist.**
RESEARCH warned that "core is on 0.17.0 but the vvm repo's latest release is 0.16.4 (2026-04-30)"
and told this plan to confirm the pairing. There is in fact a
[`voicevox_vvm` 0.17.0 release](https://github.com/VOICEVOX/voicevox_vvm/releases/tag/0.17.0)
published **2026-08-12**, one day *before* core 0.17.0. The core 0.17.0 release notes link to it
directly. We use the matched pair. See "Version pairing" below for why this is not optional.

**2. `AudioQuery` has no `.json()` method.** The plan proposed `json.loads(query.json())` to reach
a plain dict. That attribute does not exist in 0.17.0 (`hasattr(q, "json")` is `False`; so is
`to_json`). The supported conversion is `dataclasses.asdict(query)`, which returns snake_case keys.
`tts.py` therefore converts to the VOICEVOX **ENGINE** JSON schema explicitly — see "AudioQuery
serialisation" below.

**3. `AudioQuery` has no `pauseLength` / `pauseLengthScale`.** RESEARCH's `<interfaces>` schema and
its timing-pipeline step 3 both reference them. They are **not** fields of `voicevox_core`
0.17.0's `AudioQuery` (they exist in the separate VOICEVOX *ENGINE* HTTP product, not in CORE).
**Consequence for plan 01-06:** pipeline step 3 (`pauseLength` override, then `pauseLengthScale`
multiply) is a **no-op on this stack**. Do not implement it against a key that will never be
present; if it is implemented defensively, it must tolerate the key being absent.

**4. `SpeakerMeta` is `CharacterMeta`.** `usage.md`'s pasted `pprint` output still shows
`SpeakerMeta(...)`, but 0.17.0 exports `CharacterMeta`. The `.speaker_uuid` *attribute* is
unchanged. Cosmetic, but it will bite anyone who imports the name.

**5. The Open JTalk dictionary's copyright holder is Nara Institute of Science and Technology,
not Nagoya Institute of Technology.** See "Open JTalk dictionary licence" below — the two
institutions are genuinely both involved and RESEARCH conflated them.

Everything else RESEARCH inferred was correct, including the whole `Onnxruntime` → `OpenJtalk` →
`Synthesizer` → `VoiceModelFile` construction order and the exact Linux wheel filename.

**6. `speedScale` is applied to the frame count, not to the phoneme length.** This is the most
consequential correction in this document and it has its own section below.

---

## Frame quantisation — read this before writing `visemes.py`

**Plan 01-06 depends on this section.** 01-RESEARCH.md's `build_timeline` gets the order wrong,
and the error is invisible at normal speed, which is exactly what makes it dangerous.

RESEARCH's pipeline divides each phoneme length by `speedScale` and *then* quantises:

```python
frames = round(length / speed_scale * 93.75)          # WRONG
```

VOICEVOX CORE 0.17.0 actually quantises **first**, at speed 1.0, and then divides the resulting
**frame count** and rounds again:

```python
frames = round(round(length * 93.75) / speed_scale)   # CORRECT
```

At `speedScale == 1.0` the two are identical, so RESEARCH's version looks verified. They diverge
everywhere else.

Measured: 4 sentences × 6 speed values (1.0, 0.9, 0.75, 0.5, 1.25, 1.5), predicted total frames
compared against the true frame count of the synthesised WAV (`getnframes() // 256`):

| Formula | Correct |
|---|---|
| `round(length / speed * 93.75)` (01-RESEARCH.md) | **8 / 24** |
| `round(round(length * 93.75) / speed)` | **24 / 24** |

Worst observed error from the wrong formula: **5 frames ≈ 53 ms** on the long sentence at
`speedScale = 1.25`. That is well past the ±1 frame (10.667 ms) tolerance `test_no_drift_long_utterance`
is specified with, and it is exactly the symptom `01-RESEARCH.md` Pitfall 5 describes — "normal
speed syncs, slow speed drifts" — reached from the other direction.

Everything else about the pipeline is confirmed and should be implemented as RESEARCH describes:

- Flatten each accent phrase's `moras`, then its `pause_mora` if present — pause **after**.
- Wrap the whole sequence in `prePhonemeLength` / `postPhonemeLength` silence moras, and scale
  those by `speedScale` like any other phoneme. They are not exempt.
- Emit the consonant (mouth closed) before the vowel within a mora.
- Accumulate the quantised frame counts, never the raw floats. Per-phoneme rounding error is up to
  ±0.5 frame ≈ ±5.3 ms, and the long fixture has 61 phonemes.
- Use banker's rounding. Python's built-in `round()` already is; JavaScript's `Math.round()` is
  round-half-up and would disagree, which is an independent reason the timeline is built in Python
  and the browser only plays it.
- Steps involving `pauseLength` / `pauseLengthScale` are no-ops — those fields do not exist here.

The committed fixtures are the proof: `tests/fixtures/make_synth_fixtures.py` asserts the correct
formula reproduces each fixture's true frame count before it will write anything, so the fixture
set cannot silently drift away from the engine.

| Fixture | speedScale | Frames | Duration (s) |
|---|---|---|---|
| `short` | 1.0 | 99 | 1.056 |
| `long` | 1.0 | 516 | 5.504 |
| `slow` | 0.75 | 692 | 7.381333333333333 |

Note that `slow / long = 1.341085`, **not** exactly `1/0.75 = 1.333333`. Re-quantisation after
scaling means the realised ratio lands near the requested one, not on it. A test that asserts
"exactly 1/0.75× longer" against *durations* will fail; assert the timeline matches the WAV
instead, which is the property that actually matters for lip-sync.

---

## AudioQuery serialisation

`SynthResult.audio_query` is a plain dict in the **VOICEVOX ENGINE schema**, which is
deliberately mixed-case: nested structures keep snake_case, top-level scalar parameters are
camelCase. This is not a style choice — it is the wire format every VOICEVOX consumer expects,
and it is what the committed fixtures record.

| `voicevox_core` dataclass field | ENGINE JSON key |
|---|---|
| `accent_phrases` | `accent_phrases` (unchanged) |
| `speed_scale` | `speedScale` |
| `pitch_scale` | `pitchScale` |
| `intonation_scale` | `intonationScale` |
| `volume_scale` | `volumeScale` |
| `pre_phoneme_length` | `prePhonemeLength` |
| `post_phoneme_length` | `postPhonemeLength` |
| `output_sampling_rate` | `outputSamplingRate` |
| `output_stereo` | `outputStereo` |
| `kana` | `kana` (unchanged) |

Mora keys (`text`, `consonant`, `consonant_length`, `vowel`, `vowel_length`, `pitch`) and accent
phrase keys (`moras`, `accent`, `pause_mora`, `is_interrogative`) pass through unchanged.
`consonant` / `consonant_length` are `None` for vowel-only moras and for `pau`.

---

## Runtime assets

| Asset | Path in repo | Size | Acquired from | Committed via LFS? |
|---|---|---|---|---|
| Open JTalk dict | `voicevox/open_jtalk_dic_utf_8-1.11/` | 107,304,813 B total (`sys.dic` alone 103,073,776 B) | `https://downloads.sourceforge.net/project/open-jtalk/Dictionary/open_jtalk_dic-1.11/open_jtalk_dic_utf_8-1.11.tar.gz` | **yes** — via `voicevox/.gitattributes` (`*.dic`, `*.bin`) |
| Voice model | `voicevox/model/zundamon.vvm` | 59,308,488 B | `voicevox_vvm` 0.17.0 asset, original filename **`0.vvm`** | **yes** — root `.gitattributes` `*.vvm` |
| VOICEVOX ONNX Runtime | *not in the repo* — fetched to `voicevox_runtime/` (gitignored) | 8.2 MB (linux-x64 CPU `.tgz`) | `github.com/VOICEVOX/onnxruntime-builder` release `voicevox_onnxruntime-1.23.2` | **NO — never vendored** (option **c**, below) |

```text
SHA256 zundamon.vvm: ecd35374d4182cd883cba5040376f7f888cc6ba248b1c2f4cea07cdb34bb1318
speaker_uuid:        388f246b-8c41-4ac1-8e2d-5d79f3ff56d9
style_id (ノーマル):  3
```

`style_id` **3** is the integer constant `SPEAKER_STYLE_ID` in `src/japanese_avatar/voice/tts.py`.

`0.vvm` also carries 四国めたん (2/0/6/4), 春日部つむぎ (8) and 雨晴はう (10). Only ずんだもん is
credited and used; the others are present because the model file is packaged that way upstream.
Confirmed against the `metas.json` inside the downloaded file and against the character↔style
table in the release's own `README.txt`.

### Why only one `.vvm`

The `voicevox_vvm` 0.17.0 release ships 29 assets — 25 talk models `0.vvm`..`24.vvm`, `n0.vvm`
(Nemo), `s0.vvm` (song), plus `README.txt` and `TERMS.txt`. Each talk model is 57-67 MB. The
character↔file index lives in `README.txt` (section 音声モデル(.vvm)ファイルと声…の対応表), which
resolves ずんだもん ノーマル to **`0.vvm`, style ID 3**. Only that one file was downloaded.

---

## ONNX Runtime acquisition — **option (c)**, lazy first-use download

Chosen: **(c) a lazy first-use download, guarded so it never happens at import time.**

Reasons, in the plan's stated order of preference:

- **(a) a pip-installable companion distribution — does not exist.** Verified: `pypi.org/pypi/voicevox-onnxruntime/json`,
  `voicevox_onnxruntime` and `voicevox-core` all return **HTTP 404**. The runtime is published
  only as `.tgz` / `.zip` archives on
  `github.com/VOICEVOX/onnxruntime-builder/releases/tag/voicevox_onnxruntime-1.23.2`.
  There is nothing that could be added to `requirements.txt`.
- **(b) the official downloader binary run at build time — not available on this platform.** A
  Hugging Face **Gradio-SDK** Space has no arbitrary build hook; its build phase runs pip against
  `requirements.txt` (optionally `pre-requirements.txt`) and apt against `packages.txt`. Neither can
  execute `./download`. Option (b) remains the correct choice for local development and for any
  future Docker-SDK variant, and it is what a contributor may use instead of the automatic path.
- **(c) chosen.** `get_synthesizer()` resolves the runtime in this order: `VOICEVOX_ORT_PATH` env
  var → an existing copy under `voicevox_runtime/` → download the official archive for the current
  platform → fall back to `Onnxruntime.load_once()`'s own library search path. The linux-x64 CPU
  archive is **8.2 MB**, so a cold wake pays roughly a second of network, once, and only inside
  the lazily-constructed synthesizer — never at module import. Plan 01-08's `warmup()` call pays it
  during app startup so the first visitor does not.

This keeps the repo compliant with the VOICEVOX ソフトウェア利用規約 禁止事項 (unauthorised
redistribution of the software is forbidden): the runtime and the wheel are always *referenced*
from their official release URLs, never copied into this repository. The asymmetry with the voice
model is deliberate and is spelled out under "Licence compliance" below.

---

## Version pairing

**voicevox_core 0.17.0 + voicevox_vvm 0.17.0 — verified compatible, executed.**

The pairing question RESEARCH raised is real but resolves the other way round from what it
assumed. `voicevox_vvm` 0.17.0's `0.vvm` reports `vvm_format_version: 2` in its internal
`manifest.json`, and every style in it has `type: "streaming_talk"`. Both the format version and
the `StyleType::StreamingTalk` variant were **introduced in core 0.17.0**. So:

- core 0.17.0 + vvm **0.17.0** → verified working (audio synthesised, see below).
- core 0.17.0 + vvm 0.16.4 → the old `vvm_format_version` is still readable by 0.17.0, but there is
  no reason to take it; 0.16.4 predates the format the current core is built around.
- core **0.16.x** + vvm 0.17.0 → would **not** work: an older core cannot read format 2.

Evidence, executed: loading `voicevox/model/zundamon.vvm` into a 0.17.0 `Synthesizer` succeeded and
`create_audio_query("こんにちは", 3)` → `synthesis(...)` returned a 50,732-byte RIFF WAV at
24000 Hz / mono / 16-bit, duration 1.056 s.

Pin both to **0.17.0**. If either is bumped, regenerate the fixtures with
`tests/fixtures/make_synth_fixtures.py` — `tests/test_tts_contract.py::test_fixtures_match_current_engine`
exists to make that failure loud rather than mysterious.

---

## Cold-start posture

Import-time work is **zero**: `import japanese_avatar.voice.tts` constructs nothing, touches no
file and opens no socket, so `app.py` stays importable on a machine with no wheel installed. All
loading happens inside `get_synthesizer()`, which is `functools.lru_cache(maxsize=1)`-wrapped, and
`warmup()` exists purely so plan 01-08 can pay that cost during startup.

Measured on Windows / Python 3.12.12, warm disk, runtime already present:

| Stage | Time |
|---|---|
| `import japanese_avatar.voice.tts` | 0.110 s (pure Python import; no file or socket touched) |
| `Onnxruntime.load_once(...)` | 0.093 s |
| `OpenJtalk(dict_dir)` | 0.001 s |
| `VoiceModelFile.open` + `load_voice_model` | 1.742 s |
| **Total import-to-ready (`warmup()`)** | **1.73 s** |
| Second `warmup()` call (cache hit) | 0.000001 s |

Per-turn, once warm: `create_audio_query` **1.2 ms**, `synthesis` **1215 ms** for 「こんにちは」.
Both are recorded through `TurnTimings`, so plan 01-08's latency harness gets them for free.

**The lazy runtime download is exercised, not assumed.** Pointing `VOICEVOX_ORT_DIR` at an empty
directory and calling `warmup()` fetched, unpacked and loaded the runtime and then synthesised
successfully: **2.66 s** total, versus 1.73 s warm — so the download costs about **0.9 s** on a
cold container. That is the entire cold-start penalty this design carries.

The ~1.7 s is model deserialisation, not I/O, so it is paid on every process start regardless of
caching. The one thing that *would* have dominated cold start — pulling ~166 MB of dictionary and
voice model on each 48 h wake onto ephemeral Space disk — is eliminated by committing both through
Git LFS, so they arrive with the clone. Only the 8.2 MB runtime is fetched, and only once per
container.

---

## Licence compliance (voice assets)

The VRM's equivalent lives in `docs/ASSETS.md`, owned by plan 01-01. Plan 01-09 merges both into
`LICENSES.md`.

- **Software** — VOICEVOX ソフトウェア利用規約, <https://voicevox.hiroshiba.jp/term/>.
  Commercial and non-commercial use permitted. 禁止事項 forbids redistributing the software in
  whole or part without authorisation, so `voicevox_core` is installed from the official release
  URL pinned in `requirements.txt` and **neither the wheel nor the ONNX Runtime binary is
  vendored into this repository**.
- **Voice model** — VOICEVOX 音声モデル 利用規約, published in the
  [`voicevox_vvm` README](https://github.com/VOICEVOX/voicevox_vvm/blob/main/README.md) and shipped
  as `README.txt` / `TERMS.txt` alongside the release assets. Clause 2 reads
  「アプリケーションに組み込んで再配布することができます」 — **embedded redistribution is explicitly
  permitted**, which is what licenses committing `zundamon.vvm` to this repo.
- **Character** — ずんだもん, SSS LLC, <https://zunko.jp/con_ongen_kiyaku.html>. One document covers
  nine characters. Commercial and non-commercial use permitted *with credit*; without credit,
  per-character licensing is ¥400,000 (+ tax).

### The credit string

```text
VOICEVOX:ずんだもん
```

Exact, including the ASCII colon and no spaces. Later plans render it; this document is where the
string is defined.

**Placement requirement** (character terms clause 2):
「アプリなどでの利用の場合は、アプリの紹介画面などに記載をお願いします。(少し探せばわかる場所に)」
— an app introduction / about screen, in a place findable with a little looking. Phase 1 satisfies
this with a persistently visible footer line next to the avatar **plus** an About/Credits surface.

**Flow-down obligation** (software clause 3 / voice-model clause 4): when audio generated here is
made available to others, those others must be bound to the same terms. Anywhere audio is
user-obtainable (playback, replay, download), display a terms notice to the effect of *"Synthesised
audio is provided under the VOICEVOX and VOICEVOX:ずんだもん terms of use; by using it you agree to
comply with them"*, linking both URLs above.

**The credit obligation survives an engine swap.** The official Q&A answers
「音声の中間表現(AudioQuery/FrameAudioQuery)を VOICEVOX 以外の音声合成に利用した場合はクレジット記載が
必要ですか?」with 「必要です。」 Using the `AudioQuery` — which is precisely what the viseme
timeline consumes — triggers the credit requirement on its own.

### There is no "ask the rights holder" step, ever

SSS LLC 免責条項 2: 「本ガイドラインに該当するかどうかのご質問については、原則お答えしておりません。
特に無償利用の範囲についての問い合わせについては回答いたしかねます。…できるかぎりガイドラインを自身で
読み込んだ上での判断をお願いいたします。」 They decline "does my use qualify?" questions as a matter
of policy. The compliance posture is: read the guideline (done, above), comply visibly, document
here. Recorded so it is never re-litigated in a later phase.

### Pre-cleared swap-in: VOICEVOX Nemo

If the character terms ever become inconvenient, **VOICEVOX Nemo** is the drop-in replacement and
the swap is a two-line change (`SPEAKER_STYLE_ID` plus the model path):

- Model file: `n0.vvm` from the same `voicevox_vvm` 0.17.0 release; nine voices, style IDs
  10000-10008 (男声1-3, 女声1-6), all ノーマル.
- Credit string is just `VOICEVOX Nemo` — no character name, no third-party rights holder.
- Terms: <https://voicevox.hiroshiba.jp/nemo/term/>, restated in the `voicevox_vvm` README.
- No ¥400,000 credit-omission clause, because there is no separate character licensor.

Cost of the swap: the loss of a recognisable character identity. Recommendation remains ずんだもん.

### Open JTalk dictionary licence

BSD-3-Clause. The shipped `voicevox/open_jtalk_dic_utf_8-1.11/COPYING` reads verbatim:

> Copyright (c) 2009, **Nara Institute of Science and Technology**, Japan.
> All rights reserved.
>
> Redistribution and use in source and binary forms, with or without modification, are permitted
> provided that the following conditions are met: Redistributions of source code must retain the
> above copyright notice, this list of conditions and the following disclaimer. Redistributions in
> binary form must reproduce the above copyright notice, this list of conditions and the following
> disclaimer in the documentation and/or other materials provided with the distribution. Neither
> the name of the Nara Institute of Science and Technology (NAIST) nor the names of its
> contributors may be used to endorse or promote products derived from this software without
> specific prior written permission.

**Correction to 01-RESEARCH.md:** RESEARCH attributes this dictionary to "Nagoya Institute of
Technology / HTS Working Group". Both institutions are genuinely in the picture and RESEARCH
conflated them:

- The **dictionary data** (`open_jtalk_dic_utf_8-1.11`, derived from the NAIST Japanese
  Dictionary / IPAdic lineage) is © 2009 **Nara** Institute of Science and Technology, BSD-3-Clause
  — this is the notice that must be reproduced for the files committed here.
- **Open JTalk itself** — the engine whose dictionary format this is, and whose code is linked
  inside `voicevox_core` — is from the **Nagoya Institute of Technology** and the HTS Working
  Group, also under a modified BSD licence.

`LICENSES.md` (plan 01-09) must carry **both** attributions; reproducing only Nagoya would fail to
satisfy the BSD notice on the files actually redistributed in this repo.

---

## Local development

The Space consumes `requirements.txt`, which pins the **Linux** wheel. Local development on
Windows needs the `win_amd64` asset from the same release. It is wired into `pyproject.toml` as
the `voice` extra with marker-differentiated `[tool.uv.sources]`, so one lockfile serves both:

```console
uv sync --extra dev --extra voice
```

Never put a Windows wheel in `requirements.txt` — that file is consumed by a Linux builder.

To obtain the ONNX Runtime explicitly rather than letting `get_synthesizer()` fetch it, use the
official downloader (option **b**) and point `VOICEVOX_ORT_PATH` at the result:

```console
curl -sSfL https://github.com/VOICEVOX/voicevox_core/releases/download/0.17.0/download-linux-x64 -o download
chmod +x download
./download --only onnxruntime -o ./voicevox_runtime
```

`voicevox_runtime/` is gitignored.

### Environment variables

| Variable | Default | Purpose |
|---|---|---|
| `VOICEVOX_DICT_DIR` | `voicevox/open_jtalk_dic_utf_8-1.11` | Open JTalk dictionary directory |
| `VOICEVOX_VVM_PATH` | `voicevox/model/zundamon.vvm` | Voice model file |
| `VOICEVOX_ORT_PATH` | *(auto-resolved)* | Explicit path to the VOICEVOX ONNX Runtime shared library |
| `VOICEVOX_ORT_DIR` | `voicevox_runtime` | Where the runtime is cached / downloaded to |
| `VOICEVOX_ORT_NO_DOWNLOAD` | unset | Set to `1` to forbid the lazy download entirely |

---

## Zero GPU, permanently

Nothing in `src/japanese_avatar/voice/` imports `spaces` or `torch`, and no `@spaces.GPU`
decorator exists anywhere on the synthesis path. Synthesis is CPU-only by design — that is the
whole reason this engine was chosen, since it means speech output consumes **none** of a visitor's
daily ZeroGPU quota. `tests/test_tts_contract.py::test_no_gpu_imports_on_synthesis_path` walks the
package with `ast` and fails if that ever changes; plan 01-08 adds the wider turn-path scan.