# VOICEVOX setup (VOIC-01) Everything in this document was **executed**, not inferred. `01-RESEARCH.md`'s code sample was marked `[A]`; every name below was confirmed either against the release's own materials or by running it against an installed `voicevox_core` 0.17.0 on Python 3.12.12. Reconnaissance date: **2026-08-27**. --- ## Confirmed API surface — voicevox_core 0.17.0 Import root: `from voicevox_core.blocking import Onnxruntime, OpenJtalk, Synthesizer, VoiceModelFile` (`voicevox_core.asyncio` mirrors it; we use `blocking` because synthesis is CPU-bound and `Synthesizer` is internally mutex-guarded anyway). | Call | Confirmed signature | Source | |---|---|---| | `Onnxruntime.load_once(filename=...)` | `load_once(*, filename: str = LIB_RECOMMENDED_VERSIONED_FILENAME) -> Onnxruntime` | [`docs/guide/user/usage.md` @ tag 0.17.0](https://github.com/VOICEVOX/voicevox_core/blob/0.17.0/docs/guide/user/usage.md) + `dir(Onnxruntime)` on the installed wheel | | `OpenJtalk(dict_dir)` | `OpenJtalk(open_jtalk_dict_dir: str \| PathLike)` | same `usage.md`; executed | | `Synthesizer(onnxruntime, open_jtalk)` | `Synthesizer(Onnxruntime, OpenJtalk, *, acceleration_mode=..., cpu_num_threads=...)` | same `usage.md`; executed | | `VoiceModelFile.open(path)` | classmethod, returns a context manager; `.id`, `.metas` | same `usage.md`; executed | | `synthesizer.load_voice_model(model)` | `load_voice_model(model, *, on_existing=...)` — `on_existing` is new in 0.17.0 | [release 0.17.0 body](https://github.com/VOICEVOX/voicevox_core/releases/tag/0.17.0); executed | | `synthesizer.create_audio_query(text, style_id)` | `-> voicevox_core.AudioQuery` | same `usage.md`; executed | | `synthesizer.synthesis(audio_query, style_id)` | `-> bytes` (RIFF WAV) | same `usage.md`; executed | | `Onnxruntime.LIB_RECOMMENDED_VERSION` | `"1.23.2"` | executed against the installed wheel | | `Onnxruntime.LIB_RECOMMENDED_NAME` | `"voicevox_onnxruntime"` | executed | | `Onnxruntime.LIB_RECOMMENDED_VERSIONED_FILENAME` | `"voicevox_onnxruntime.dll"` on Windows; platform-dependent | executed | `AudioQuery`, `AccentPhrase` and `Mora` are **real Python `dataclasses`** (`dataclasses.is_dataclass` returns `True` for all three), which is what makes the plain-dict conversion in `tts.py` exact rather than hand-transcribed. | Type | `dataclasses.fields()` | Source | |---|---|---| | `AudioQuery` | `accent_phrases, speed_scale, pitch_scale, intonation_scale, volume_scale, pre_phoneme_length, post_phoneme_length, output_sampling_rate, output_stereo, kana` | executed | | `AccentPhrase` | `moras, accent, pause_mora, is_interrogative` | executed | | `Mora` | `text, vowel, vowel_length, pitch, consonant, consonant_length` | executed | --- ## Corrections to 01-RESEARCH.md Five differences. All five matter to a later plan. **1. `voicevox_vvm` is on 0.17.0, not 0.16.4 — and the version-track mismatch does not exist.** RESEARCH warned that "core is on 0.17.0 but the vvm repo's latest release is 0.16.4 (2026-04-30)" and told this plan to confirm the pairing. There is in fact a [`voicevox_vvm` 0.17.0 release](https://github.com/VOICEVOX/voicevox_vvm/releases/tag/0.17.0) published **2026-08-12**, one day *before* core 0.17.0. The core 0.17.0 release notes link to it directly. We use the matched pair. See "Version pairing" below for why this is not optional. **2. `AudioQuery` has no `.json()` method.** The plan proposed `json.loads(query.json())` to reach a plain dict. That attribute does not exist in 0.17.0 (`hasattr(q, "json")` is `False`; so is `to_json`). The supported conversion is `dataclasses.asdict(query)`, which returns snake_case keys. `tts.py` therefore converts to the VOICEVOX **ENGINE** JSON schema explicitly — see "AudioQuery serialisation" below. **3. `AudioQuery` has no `pauseLength` / `pauseLengthScale`.** RESEARCH's `` schema and its timing-pipeline step 3 both reference them. They are **not** fields of `voicevox_core` 0.17.0's `AudioQuery` (they exist in the separate VOICEVOX *ENGINE* HTTP product, not in CORE). **Consequence for plan 01-06:** pipeline step 3 (`pauseLength` override, then `pauseLengthScale` multiply) is a **no-op on this stack**. Do not implement it against a key that will never be present; if it is implemented defensively, it must tolerate the key being absent. **4. `SpeakerMeta` is `CharacterMeta`.** `usage.md`'s pasted `pprint` output still shows `SpeakerMeta(...)`, but 0.17.0 exports `CharacterMeta`. The `.speaker_uuid` *attribute* is unchanged. Cosmetic, but it will bite anyone who imports the name. **5. The Open JTalk dictionary's copyright holder is Nara Institute of Science and Technology, not Nagoya Institute of Technology.** See "Open JTalk dictionary licence" below — the two institutions are genuinely both involved and RESEARCH conflated them. Everything else RESEARCH inferred was correct, including the whole `Onnxruntime` → `OpenJtalk` → `Synthesizer` → `VoiceModelFile` construction order and the exact Linux wheel filename. ### Bonus: the frame-quantisation algorithm is now empirically verified RESEARCH's `build_timeline` pipeline (flatten → wrap in pre/post silence → divide by `speedScale` → quantise each phoneme with `round(sec * 93.75)` using banker's rounding → accumulate) was replayed against a real 36-mora synthesis. The accumulated timeline total and the true WAV duration agree to **2.7e-15 s** — i.e. exactly, to float representation. Plan 01-06 can implement that algorithm as written and expect bit-exact agreement. --- ## AudioQuery serialisation `SynthResult.audio_query` is a plain dict in the **VOICEVOX ENGINE schema**, which is deliberately mixed-case: nested structures keep snake_case, top-level scalar parameters are camelCase. This is not a style choice — it is the wire format every VOICEVOX consumer expects, and it is what the committed fixtures record. | `voicevox_core` dataclass field | ENGINE JSON key | |---|---| | `accent_phrases` | `accent_phrases` (unchanged) | | `speed_scale` | `speedScale` | | `pitch_scale` | `pitchScale` | | `intonation_scale` | `intonationScale` | | `volume_scale` | `volumeScale` | | `pre_phoneme_length` | `prePhonemeLength` | | `post_phoneme_length` | `postPhonemeLength` | | `output_sampling_rate` | `outputSamplingRate` | | `output_stereo` | `outputStereo` | | `kana` | `kana` (unchanged) | Mora keys (`text`, `consonant`, `consonant_length`, `vowel`, `vowel_length`, `pitch`) and accent phrase keys (`moras`, `accent`, `pause_mora`, `is_interrogative`) pass through unchanged. `consonant` / `consonant_length` are `None` for vowel-only moras and for `pau`. --- ## Runtime assets | Asset | Path in repo | Size | Acquired from | Committed via LFS? | |---|---|---|---|---| | Open JTalk dict | `voicevox/open_jtalk_dic_utf_8-1.11/` | 107,304,813 B total (`sys.dic` alone 103,073,776 B) | `https://downloads.sourceforge.net/project/open-jtalk/Dictionary/open_jtalk_dic-1.11/open_jtalk_dic_utf_8-1.11.tar.gz` | **yes** — via `voicevox/.gitattributes` (`*.dic`, `*.bin`) | | Voice model | `voicevox/model/zundamon.vvm` | 59,308,488 B | `voicevox_vvm` 0.17.0 asset, original filename **`0.vvm`** | **yes** — root `.gitattributes` `*.vvm` | | VOICEVOX ONNX Runtime | *not in the repo* — fetched to `voicevox_runtime/` (gitignored) | 8.2 MB (linux-x64 CPU `.tgz`) | `github.com/VOICEVOX/onnxruntime-builder` release `voicevox_onnxruntime-1.23.2` | **NO — never vendored** (option **c**, below) | ```text SHA256 zundamon.vvm: ecd35374d4182cd883cba5040376f7f888cc6ba248b1c2f4cea07cdb34bb1318 speaker_uuid: 388f246b-8c41-4ac1-8e2d-5d79f3ff56d9 style_id (ノーマル): 3 ``` `style_id` **3** is the integer constant `SPEAKER_STYLE_ID` in `src/japanese_avatar/voice/tts.py`. `0.vvm` also carries 四国めたん (2/0/6/4), 春日部つむぎ (8) and 雨晴はう (10). Only ずんだもん is credited and used; the others are present because the model file is packaged that way upstream. Confirmed against the `metas.json` inside the downloaded file and against the character↔style table in the release's own `README.txt`. ### Why only one `.vvm` The `voicevox_vvm` 0.17.0 release ships 29 assets — 25 talk models `0.vvm`..`24.vvm`, `n0.vvm` (Nemo), `s0.vvm` (song), plus `README.txt` and `TERMS.txt`. Each talk model is 57-67 MB. The character↔file index lives in `README.txt` (section 音声モデル(.vvm)ファイルと声…の対応表), which resolves ずんだもん ノーマル to **`0.vvm`, style ID 3**. Only that one file was downloaded. --- ## ONNX Runtime acquisition — **option (c)**, lazy first-use download Chosen: **(c) a lazy first-use download, guarded so it never happens at import time.** Reasons, in the plan's stated order of preference: - **(a) a pip-installable companion distribution — does not exist.** Verified: `pypi.org/pypi/voicevox-onnxruntime/json`, `voicevox_onnxruntime` and `voicevox-core` all return **HTTP 404**. The runtime is published only as `.tgz` / `.zip` archives on `github.com/VOICEVOX/onnxruntime-builder/releases/tag/voicevox_onnxruntime-1.23.2`. There is nothing that could be added to `requirements.txt`. - **(b) the official downloader binary run at build time — not available on this platform.** A Hugging Face **Gradio-SDK** Space has no arbitrary build hook; its build phase runs pip against `requirements.txt` (optionally `pre-requirements.txt`) and apt against `packages.txt`. Neither can execute `./download`. Option (b) remains the correct choice for local development and for any future Docker-SDK variant, and it is what a contributor may use instead of the automatic path. - **(c) chosen.** `get_synthesizer()` resolves the runtime in this order: `VOICEVOX_ORT_PATH` env var → an existing copy under `voicevox_runtime/` → download the official archive for the current platform → fall back to `Onnxruntime.load_once()`'s own library search path. The linux-x64 CPU archive is **8.2 MB**, so a cold wake pays roughly a second of network, once, and only inside the lazily-constructed synthesizer — never at module import. Plan 01-08's `warmup()` call pays it during app startup so the first visitor does not. This keeps the repo compliant with the VOICEVOX ソフトウェア利用規約 禁止事項 (unauthorised redistribution of the software is forbidden): the runtime and the wheel are always *referenced* from their official release URLs, never copied into this repository. The asymmetry with the voice model is deliberate and is spelled out under "Licence compliance" below. --- ## Version pairing **voicevox_core 0.17.0 + voicevox_vvm 0.17.0 — verified compatible, executed.** The pairing question RESEARCH raised is real but resolves the other way round from what it assumed. `voicevox_vvm` 0.17.0's `0.vvm` reports `vvm_format_version: 2` in its internal `manifest.json`, and every style in it has `type: "streaming_talk"`. Both the format version and the `StyleType::StreamingTalk` variant were **introduced in core 0.17.0**. So: - core 0.17.0 + vvm **0.17.0** → verified working (audio synthesised, see below). - core 0.17.0 + vvm 0.16.4 → the old `vvm_format_version` is still readable by 0.17.0, but there is no reason to take it; 0.16.4 predates the format the current core is built around. - core **0.16.x** + vvm 0.17.0 → would **not** work: an older core cannot read format 2. Evidence, executed: loading `voicevox/model/zundamon.vvm` into a 0.17.0 `Synthesizer` succeeded and `create_audio_query("こんにちは", 3)` → `synthesis(...)` returned a 50,732-byte RIFF WAV at 24000 Hz / mono / 16-bit, duration 1.056 s. Pin both to **0.17.0**. If either is bumped, regenerate the fixtures with `tests/fixtures/make_synth_fixtures.py` — `tests/test_tts_contract.py::test_fixtures_match_current_engine` exists to make that failure loud rather than mysterious. --- ## Cold-start posture Import-time work is **zero**: `import japanese_avatar.voice.tts` constructs nothing, touches no file and opens no socket, so `app.py` stays importable on a machine with no wheel installed. All loading happens inside `get_synthesizer()`, which is `functools.lru_cache(maxsize=1)`-wrapped, and `warmup()` exists purely so plan 01-08 can pay that cost during startup. Measured on Windows / Python 3.12.12, warm disk, runtime already present: | Stage | Time | |---|---| | `Onnxruntime.load_once(...)` | 0.093 s | | `OpenJtalk(dict_dir)` | 0.001 s | | `VoiceModelFile.open` + `load_voice_model` | 1.742 s | | **Total import-to-ready (`warmup()`)** | **1.84 s** | The ~1.7 s is model deserialisation, not I/O, so it is paid on every process start regardless of caching. The one thing that *would* have dominated cold start — pulling ~166 MB of dictionary and voice model on each 48 h wake onto ephemeral Space disk — is eliminated by committing both through Git LFS, so they arrive with the clone. Only the 8.2 MB runtime is fetched, and only once per container. --- ## Licence compliance (voice assets) The VRM's equivalent lives in `docs/ASSETS.md`, owned by plan 01-01. Plan 01-09 merges both into `LICENSES.md`. - **Software** — VOICEVOX ソフトウェア利用規約, . Commercial and non-commercial use permitted. 禁止事項 forbids redistributing the software in whole or part without authorisation, so `voicevox_core` is installed from the official release URL pinned in `requirements.txt` and **neither the wheel nor the ONNX Runtime binary is vendored into this repository**. - **Voice model** — VOICEVOX 音声モデル 利用規約, published in the [`voicevox_vvm` README](https://github.com/VOICEVOX/voicevox_vvm/blob/main/README.md) and shipped as `README.txt` / `TERMS.txt` alongside the release assets. Clause 2 reads 「アプリケーションに組み込んで再配布することができます」 — **embedded redistribution is explicitly permitted**, which is what licenses committing `zundamon.vvm` to this repo. - **Character** — ずんだもん, SSS LLC, . One document covers nine characters. Commercial and non-commercial use permitted *with credit*; without credit, per-character licensing is ¥400,000 (+ tax). ### The credit string ```text VOICEVOX:ずんだもん ``` Exact, including the ASCII colon and no spaces. Later plans render it; this document is where the string is defined. **Placement requirement** (character terms clause 2): 「アプリなどでの利用の場合は、アプリの紹介画面などに記載をお願いします。(少し探せばわかる場所に)」 — an app introduction / about screen, in a place findable with a little looking. Phase 1 satisfies this with a persistently visible footer line next to the avatar **plus** an About/Credits surface. **Flow-down obligation** (software clause 3 / voice-model clause 4): when audio generated here is made available to others, those others must be bound to the same terms. Anywhere audio is user-obtainable (playback, replay, download), display a terms notice to the effect of *"Synthesised audio is provided under the VOICEVOX and VOICEVOX:ずんだもん terms of use; by using it you agree to comply with them"*, linking both URLs above. **The credit obligation survives an engine swap.** The official Q&A answers 「音声の中間表現(AudioQuery/FrameAudioQuery)を VOICEVOX 以外の音声合成に利用した場合はクレジット記載が 必要ですか?」with 「必要です。」 Using the `AudioQuery` — which is precisely what the viseme timeline consumes — triggers the credit requirement on its own. ### There is no "ask the rights holder" step, ever SSS LLC 免責条項 2: 「本ガイドラインに該当するかどうかのご質問については、原則お答えしておりません。 特に無償利用の範囲についての問い合わせについては回答いたしかねます。…できるかぎりガイドラインを自身で 読み込んだ上での判断をお願いいたします。」 They decline "does my use qualify?" questions as a matter of policy. The compliance posture is: read the guideline (done, above), comply visibly, document here. Recorded so it is never re-litigated in a later phase. ### Pre-cleared swap-in: VOICEVOX Nemo If the character terms ever become inconvenient, **VOICEVOX Nemo** is the drop-in replacement and the swap is a two-line change (`SPEAKER_STYLE_ID` plus the model path): - Model file: `n0.vvm` from the same `voicevox_vvm` 0.17.0 release; nine voices, style IDs 10000-10008 (男声1-3, 女声1-6), all ノーマル. - Credit string is just `VOICEVOX Nemo` — no character name, no third-party rights holder. - Terms: , restated in the `voicevox_vvm` README. - No ¥400,000 credit-omission clause, because there is no separate character licensor. Cost of the swap: the loss of a recognisable character identity. Recommendation remains ずんだもん. ### Open JTalk dictionary licence BSD-3-Clause. The shipped `voicevox/open_jtalk_dic_utf_8-1.11/COPYING` reads verbatim: > Copyright (c) 2009, **Nara Institute of Science and Technology**, Japan. > All rights reserved. > > Redistribution and use in source and binary forms, with or without modification, are permitted > provided that the following conditions are met: Redistributions of source code must retain the > above copyright notice, this list of conditions and the following disclaimer. Redistributions in > binary form must reproduce the above copyright notice, this list of conditions and the following > disclaimer in the documentation and/or other materials provided with the distribution. Neither > the name of the Nara Institute of Science and Technology (NAIST) nor the names of its > contributors may be used to endorse or promote products derived from this software without > specific prior written permission. **Correction to 01-RESEARCH.md:** RESEARCH attributes this dictionary to "Nagoya Institute of Technology / HTS Working Group". Both institutions are genuinely in the picture and RESEARCH conflated them: - The **dictionary data** (`open_jtalk_dic_utf_8-1.11`, derived from the NAIST Japanese Dictionary / IPAdic lineage) is © 2009 **Nara** Institute of Science and Technology, BSD-3-Clause — this is the notice that must be reproduced for the files committed here. - **Open JTalk itself** — the engine whose dictionary format this is, and whose code is linked inside `voicevox_core` — is from the **Nagoya Institute of Technology** and the HTS Working Group, also under a modified BSD licence. `LICENSES.md` (plan 01-09) must carry **both** attributions; reproducing only Nagoya would fail to satisfy the BSD notice on the files actually redistributed in this repo. --- ## Local development The Space consumes `requirements.txt`, which pins the **Linux** wheel. Local development on Windows needs the `win_amd64` asset from the same release. It is wired into `pyproject.toml` as the `voice` extra with marker-differentiated `[tool.uv.sources]`, so one lockfile serves both: ```console uv sync --extra dev --extra voice ``` Never put a Windows wheel in `requirements.txt` — that file is consumed by a Linux builder. To obtain the ONNX Runtime explicitly rather than letting `get_synthesizer()` fetch it, use the official downloader (option **b**) and point `VOICEVOX_ORT_PATH` at the result: ```console curl -sSfL https://github.com/VOICEVOX/voicevox_core/releases/download/0.17.0/download-linux-x64 -o download chmod +x download ./download --only onnxruntime -o ./voicevox_runtime ``` `voicevox_runtime/` is gitignored. ### Environment variables | Variable | Default | Purpose | |---|---|---| | `VOICEVOX_DICT_DIR` | `voicevox/open_jtalk_dic_utf_8-1.11` | Open JTalk dictionary directory | | `VOICEVOX_VVM_PATH` | `voicevox/model/zundamon.vvm` | Voice model file | | `VOICEVOX_ORT_PATH` | *(auto-resolved)* | Explicit path to the VOICEVOX ONNX Runtime shared library | | `VOICEVOX_ORT_DIR` | `voicevox_runtime` | Where the runtime is cached / downloaded to | | `VOICEVOX_ORT_NO_DOWNLOAD` | unset | Set to `1` to forbid the lazy download entirely | --- ## Zero GPU, permanently Nothing in `src/japanese_avatar/voice/` imports `spaces` or `torch`, and no `@spaces.GPU` decorator exists anywhere on the synthesis path. Synthesis is CPU-only by design — that is the whole reason this engine was chosen, since it means speech output consumes **none** of a visitor's daily ZeroGPU quota. `tests/test_tts_contract.py::test_no_gpu_imports_on_synthesis_path` walks the package with `ast` and fails if that ever changes; plan 01-08 adds the wider turn-path scan.