Spaces:
Running on Zero
VOICEVOX setup (VOIC-01)
Everything in this document was executed, not inferred. 01-RESEARCH.md's code sample was
marked [A]; every name below was confirmed either against the release's own materials or by
running it against an installed voicevox_core 0.17.0 on Python 3.12.12.
Reconnaissance date: 2026-08-27.
Confirmed API surface — voicevox_core 0.17.0
Import root: from voicevox_core.blocking import Onnxruntime, OpenJtalk, Synthesizer, VoiceModelFile
(voicevox_core.asyncio mirrors it; we use blocking because synthesis is CPU-bound and
Synthesizer is internally mutex-guarded anyway).
| Call | Confirmed signature | Source |
|---|---|---|
Onnxruntime.load_once(filename=...) |
load_once(*, filename: str = LIB_RECOMMENDED_VERSIONED_FILENAME) -> Onnxruntime |
docs/guide/user/usage.md @ tag 0.17.0 + dir(Onnxruntime) on the installed wheel |
OpenJtalk(dict_dir) |
OpenJtalk(open_jtalk_dict_dir: str | PathLike) |
same usage.md; executed |
Synthesizer(onnxruntime, open_jtalk) |
Synthesizer(Onnxruntime, OpenJtalk, *, acceleration_mode=..., cpu_num_threads=...) |
same usage.md; executed |
VoiceModelFile.open(path) |
classmethod, returns a context manager; .id, .metas |
same usage.md; executed |
synthesizer.load_voice_model(model) |
load_voice_model(model, *, on_existing=...) — on_existing is new in 0.17.0 |
release 0.17.0 body; executed |
synthesizer.create_audio_query(text, style_id) |
-> voicevox_core.AudioQuery |
same usage.md; executed |
synthesizer.synthesis(audio_query, style_id) |
-> bytes (RIFF WAV) |
same usage.md; executed |
Onnxruntime.LIB_RECOMMENDED_VERSION |
"1.23.2" |
executed against the installed wheel |
Onnxruntime.LIB_RECOMMENDED_NAME |
"voicevox_onnxruntime" |
executed |
Onnxruntime.LIB_RECOMMENDED_VERSIONED_FILENAME |
"voicevox_onnxruntime.dll" on Windows; platform-dependent |
executed |
AudioQuery, AccentPhrase and Mora are real Python dataclasses (dataclasses.is_dataclass
returns True for all three), which is what makes the plain-dict conversion in tts.py exact
rather than hand-transcribed.
| Type | dataclasses.fields() |
Source |
|---|---|---|
AudioQuery |
accent_phrases, speed_scale, pitch_scale, intonation_scale, volume_scale, pre_phoneme_length, post_phoneme_length, output_sampling_rate, output_stereo, kana |
executed |
AccentPhrase |
moras, accent, pause_mora, is_interrogative |
executed |
Mora |
text, vowel, vowel_length, pitch, consonant, consonant_length |
executed |
Corrections to 01-RESEARCH.md
Five differences. All five matter to a later plan.
1. voicevox_vvm is on 0.17.0, not 0.16.4 — and the version-track mismatch does not exist.
RESEARCH warned that "core is on 0.17.0 but the vvm repo's latest release is 0.16.4 (2026-04-30)"
and told this plan to confirm the pairing. There is in fact a
voicevox_vvm 0.17.0 release
published 2026-08-12, one day before core 0.17.0. The core 0.17.0 release notes link to it
directly. We use the matched pair. See "Version pairing" below for why this is not optional.
2. AudioQuery has no .json() method. The plan proposed json.loads(query.json()) to reach
a plain dict. That attribute does not exist in 0.17.0 (hasattr(q, "json") is False; so is
to_json). The supported conversion is dataclasses.asdict(query), which returns snake_case keys.
tts.py therefore converts to the VOICEVOX ENGINE JSON schema explicitly — see "AudioQuery
serialisation" below.
3. AudioQuery has no pauseLength / pauseLengthScale. RESEARCH's <interfaces> schema and
its timing-pipeline step 3 both reference them. They are not fields of voicevox_core
0.17.0's AudioQuery (they exist in the separate VOICEVOX ENGINE HTTP product, not in CORE).
Consequence for plan 01-06: pipeline step 3 (pauseLength override, then pauseLengthScale
multiply) is a no-op on this stack. Do not implement it against a key that will never be
present; if it is implemented defensively, it must tolerate the key being absent.
4. SpeakerMeta is CharacterMeta. usage.md's pasted pprint output still shows
SpeakerMeta(...), but 0.17.0 exports CharacterMeta. The .speaker_uuid attribute is
unchanged. Cosmetic, but it will bite anyone who imports the name.
5. The Open JTalk dictionary's copyright holder is Nara Institute of Science and Technology, not Nagoya Institute of Technology. See "Open JTalk dictionary licence" below — the two institutions are genuinely both involved and RESEARCH conflated them.
Everything else RESEARCH inferred was correct, including the whole Onnxruntime → OpenJtalk →
Synthesizer → VoiceModelFile construction order and the exact Linux wheel filename.
6. speedScale is applied to the frame count, not to the phoneme length. This is the most
consequential correction in this document and it has its own section below.
Frame quantisation — read this before writing visemes.py
Plan 01-06 depends on this section. 01-RESEARCH.md's build_timeline gets the order wrong,
and the error is invisible at normal speed, which is exactly what makes it dangerous.
RESEARCH's pipeline divides each phoneme length by speedScale and then quantises:
frames = round(length / speed_scale * 93.75) # WRONG
VOICEVOX CORE 0.17.0 actually quantises first, at speed 1.0, and then divides the resulting frame count and rounds again:
frames = round(round(length * 93.75) / speed_scale) # CORRECT
At speedScale == 1.0 the two are identical, so RESEARCH's version looks verified. They diverge
everywhere else.
Measured: 4 sentences × 6 speed values (1.0, 0.9, 0.75, 0.5, 1.25, 1.5), predicted total frames
compared against the true frame count of the synthesised WAV (getnframes() // 256):
| Formula | Correct |
|---|---|
round(length / speed * 93.75) (01-RESEARCH.md) |
8 / 24 |
round(round(length * 93.75) / speed) |
24 / 24 |
Worst observed error from the wrong formula: 5 frames ≈ 53 ms on the long sentence at
speedScale = 1.25. That is well past the ±1 frame (10.667 ms) tolerance test_no_drift_long_utterance
is specified with, and it is exactly the symptom 01-RESEARCH.md Pitfall 5 describes — "normal
speed syncs, slow speed drifts" — reached from the other direction.
Everything else about the pipeline is confirmed and should be implemented as RESEARCH describes:
- Flatten each accent phrase's
moras, then itspause_moraif present — pause after. - Wrap the whole sequence in
prePhonemeLength/postPhonemeLengthsilence moras, and scale those byspeedScalelike any other phoneme. They are not exempt. - Emit the consonant (mouth closed) before the vowel within a mora.
- Accumulate the quantised frame counts, never the raw floats. Per-phoneme rounding error is up to ±0.5 frame ≈ ±5.3 ms, and the long fixture has 61 phonemes.
- Use banker's rounding. Python's built-in
round()already is; JavaScript'sMath.round()is round-half-up and would disagree, which is an independent reason the timeline is built in Python and the browser only plays it. - Steps involving
pauseLength/pauseLengthScaleare no-ops — those fields do not exist here.
The committed fixtures are the proof: tests/fixtures/make_synth_fixtures.py asserts the correct
formula reproduces each fixture's true frame count before it will write anything, so the fixture
set cannot silently drift away from the engine.
| Fixture | speedScale | Frames | Duration (s) |
|---|---|---|---|
short |
1.0 | 99 | 1.056 |
long |
1.0 | 516 | 5.504 |
slow |
0.75 | 692 | 7.381333333333333 |
Note that slow / long = 1.341085, not exactly 1/0.75 = 1.333333. Re-quantisation after
scaling means the realised ratio lands near the requested one, not on it. A test that asserts
"exactly 1/0.75× longer" against durations will fail; assert the timeline matches the WAV
instead, which is the property that actually matters for lip-sync.
AudioQuery serialisation
SynthResult.audio_query is a plain dict in the VOICEVOX ENGINE schema, which is
deliberately mixed-case: nested structures keep snake_case, top-level scalar parameters are
camelCase. This is not a style choice — it is the wire format every VOICEVOX consumer expects,
and it is what the committed fixtures record.
voicevox_core dataclass field |
ENGINE JSON key |
|---|---|
accent_phrases |
accent_phrases (unchanged) |
speed_scale |
speedScale |
pitch_scale |
pitchScale |
intonation_scale |
intonationScale |
volume_scale |
volumeScale |
pre_phoneme_length |
prePhonemeLength |
post_phoneme_length |
postPhonemeLength |
output_sampling_rate |
outputSamplingRate |
output_stereo |
outputStereo |
kana |
kana (unchanged) |
Mora keys (text, consonant, consonant_length, vowel, vowel_length, pitch) and accent
phrase keys (moras, accent, pause_mora, is_interrogative) pass through unchanged.
consonant / consonant_length are None for vowel-only moras and for pau.
Runtime assets
| Asset | Path in repo | Size | Acquired from | Committed via LFS? |
|---|---|---|---|---|
| Open JTalk dict | voicevox/open_jtalk_dic_utf_8-1.11/ |
107,304,813 B total (sys.dic alone 103,073,776 B) |
https://downloads.sourceforge.net/project/open-jtalk/Dictionary/open_jtalk_dic-1.11/open_jtalk_dic_utf_8-1.11.tar.gz |
yes — via voicevox/.gitattributes (*.dic, *.bin) |
| Voice model | voicevox/model/zundamon.vvm |
59,308,488 B | voicevox_vvm 0.17.0 asset, original filename 0.vvm |
yes — root .gitattributes *.vvm |
| VOICEVOX ONNX Runtime | not in the repo — fetched to voicevox_runtime/ (gitignored) |
8.2 MB (linux-x64 CPU .tgz) |
github.com/VOICEVOX/onnxruntime-builder release voicevox_onnxruntime-1.23.2 |
NO — never vendored (option c, below) |
SHA256 zundamon.vvm: ecd35374d4182cd883cba5040376f7f888cc6ba248b1c2f4cea07cdb34bb1318
speaker_uuid: 388f246b-8c41-4ac1-8e2d-5d79f3ff56d9
style_id (ノーマル): 3
style_id 3 is the integer constant SPEAKER_STYLE_ID in src/japanese_avatar/voice/tts.py.
0.vvm also carries 四国めたん (2/0/6/4), 春日部つむぎ (8) and 雨晴はう (10). Only ずんだもん is
credited and used; the others are present because the model file is packaged that way upstream.
Confirmed against the metas.json inside the downloaded file and against the character↔style
table in the release's own README.txt.
Why only one .vvm
The voicevox_vvm 0.17.0 release ships 29 assets — 25 talk models 0.vvm..24.vvm, n0.vvm
(Nemo), s0.vvm (song), plus README.txt and TERMS.txt. Each talk model is 57-67 MB. The
character↔file index lives in README.txt (section 音声モデル(.vvm)ファイルと声…の対応表), which
resolves ずんだもん ノーマル to 0.vvm, style ID 3. Only that one file was downloaded.
ONNX Runtime acquisition — option (c), lazy first-use download
Chosen: (c) a lazy first-use download, guarded so it never happens at import time.
Reasons, in the plan's stated order of preference:
- (a) a pip-installable companion distribution — does not exist. Verified:
pypi.org/pypi/voicevox-onnxruntime/json,voicevox_onnxruntimeandvoicevox-coreall return HTTP 404. The runtime is published only as.tgz/.ziparchives ongithub.com/VOICEVOX/onnxruntime-builder/releases/tag/voicevox_onnxruntime-1.23.2. There is nothing that could be added torequirements.txt. - (b) the official downloader binary run at build time — not available on this platform. A
Hugging Face Gradio-SDK Space has no arbitrary build hook; its build phase runs pip against
requirements.txt(optionallypre-requirements.txt) and apt againstpackages.txt. Neither can execute./download. Option (b) remains the correct choice for local development and for any future Docker-SDK variant, and it is what a contributor may use instead of the automatic path. - (c) chosen.
get_synthesizer()resolves the runtime in this order:VOICEVOX_ORT_PATHenv var → an existing copy undervoicevox_runtime/→ download the official archive for the current platform → fall back toOnnxruntime.load_once()'s own library search path. The linux-x64 CPU archive is 8.2 MB, so a cold wake pays roughly a second of network, once, and only inside the lazily-constructed synthesizer — never at module import. Plan 01-08'swarmup()call pays it during app startup so the first visitor does not.
This keeps the repo compliant with the VOICEVOX ソフトウェア利用規約 禁止事項 (unauthorised redistribution of the software is forbidden): the runtime and the wheel are always referenced from their official release URLs, never copied into this repository. The asymmetry with the voice model is deliberate and is spelled out under "Licence compliance" below.
Version pairing
voicevox_core 0.17.0 + voicevox_vvm 0.17.0 — verified compatible, executed.
The pairing question RESEARCH raised is real but resolves the other way round from what it
assumed. voicevox_vvm 0.17.0's 0.vvm reports vvm_format_version: 2 in its internal
manifest.json, and every style in it has type: "streaming_talk". Both the format version and
the StyleType::StreamingTalk variant were introduced in core 0.17.0. So:
- core 0.17.0 + vvm 0.17.0 → verified working (audio synthesised, see below).
- core 0.17.0 + vvm 0.16.4 → the old
vvm_format_versionis still readable by 0.17.0, but there is no reason to take it; 0.16.4 predates the format the current core is built around. - core 0.16.x + vvm 0.17.0 → would not work: an older core cannot read format 2.
Evidence, executed: loading voicevox/model/zundamon.vvm into a 0.17.0 Synthesizer succeeded and
create_audio_query("こんにちは", 3) → synthesis(...) returned a 50,732-byte RIFF WAV at
24000 Hz / mono / 16-bit, duration 1.056 s.
Pin both to 0.17.0. If either is bumped, regenerate the fixtures with
tests/fixtures/make_synth_fixtures.py — tests/test_tts_contract.py::test_fixtures_match_current_engine
exists to make that failure loud rather than mysterious.
Cold-start posture
Import-time work is zero: import japanese_avatar.voice.tts constructs nothing, touches no
file and opens no socket, so app.py stays importable on a machine with no wheel installed. All
loading happens inside get_synthesizer(), which is functools.lru_cache(maxsize=1)-wrapped, and
warmup() exists purely so plan 01-08 can pay that cost during startup.
Measured on Windows / Python 3.12.12, warm disk, runtime already present:
| Stage | Time |
|---|---|
import japanese_avatar.voice.tts |
0.110 s (pure Python import; no file or socket touched) |
Onnxruntime.load_once(...) |
0.093 s |
OpenJtalk(dict_dir) |
0.001 s |
VoiceModelFile.open + load_voice_model |
1.742 s |
Total import-to-ready (warmup()) |
1.73 s |
Second warmup() call (cache hit) |
0.000001 s |
Per-turn, once warm: create_audio_query 1.2 ms, synthesis 1215 ms for 「こんにちは」.
Both are recorded through TurnTimings, so plan 01-08's latency harness gets them for free.
The lazy runtime download is exercised, not assumed. Pointing VOICEVOX_ORT_DIR at an empty
directory and calling warmup() fetched, unpacked and loaded the runtime and then synthesised
successfully: 2.66 s total, versus 1.73 s warm — so the download costs about 0.9 s on a
cold container. That is the entire cold-start penalty this design carries.
The ~1.7 s is model deserialisation, not I/O, so it is paid on every process start regardless of caching. The one thing that would have dominated cold start — pulling ~166 MB of dictionary and voice model on each 48 h wake onto ephemeral Space disk — is eliminated by committing both through Git LFS, so they arrive with the clone. Only the 8.2 MB runtime is fetched, and only once per container.
Licence compliance (voice assets)
The VRM's equivalent lives in docs/ASSETS.md, owned by plan 01-01. Plan 01-09 merges both into
LICENSES.md.
- Software — VOICEVOX ソフトウェア利用規約, https://voicevox.hiroshiba.jp/term/.
Commercial and non-commercial use permitted. 禁止事項 forbids redistributing the software in
whole or part without authorisation, so
voicevox_coreis installed from the official release URL pinned inrequirements.txtand neither the wheel nor the ONNX Runtime binary is vendored into this repository. - Voice model — VOICEVOX 音声モデル 利用規約, published in the
voicevox_vvmREADME and shipped asREADME.txt/TERMS.txtalongside the release assets. Clause 2 reads 「アプリケーションに組み込んで再配布することができます」 — embedded redistribution is explicitly permitted, which is what licenses committingzundamon.vvmto this repo. - Character — ずんだもん, SSS LLC, https://zunko.jp/con_ongen_kiyaku.html. One document covers nine characters. Commercial and non-commercial use permitted with credit; without credit, per-character licensing is ¥400,000 (+ tax).
The credit string
VOICEVOX:ずんだもん
Exact, including the ASCII colon and no spaces. Later plans render it; this document is where the string is defined.
Placement requirement (character terms clause 2): 「アプリなどでの利用の場合は、アプリの紹介画面などに記載をお願いします。(少し探せばわかる場所に)」 — an app introduction / about screen, in a place findable with a little looking. Phase 1 satisfies this with a persistently visible footer line next to the avatar plus an About/Credits surface.
Flow-down obligation (software clause 3 / voice-model clause 4): when audio generated here is made available to others, those others must be bound to the same terms. Anywhere audio is user-obtainable (playback, replay, download), display a terms notice to the effect of "Synthesised audio is provided under the VOICEVOX and VOICEVOX:ずんだもん terms of use; by using it you agree to comply with them", linking both URLs above.
The credit obligation survives an engine swap. The official Q&A answers
「音声の中間表現(AudioQuery/FrameAudioQuery)を VOICEVOX 以外の音声合成に利用した場合はクレジット記載が
必要ですか?」with 「必要です。」 Using the AudioQuery — which is precisely what the viseme
timeline consumes — triggers the credit requirement on its own.
There is no "ask the rights holder" step, ever
SSS LLC 免責条項 2: 「本ガイドラインに該当するかどうかのご質問については、原則お答えしておりません。 特に無償利用の範囲についての問い合わせについては回答いたしかねます。…できるかぎりガイドラインを自身で 読み込んだ上での判断をお願いいたします。」 They decline "does my use qualify?" questions as a matter of policy. The compliance posture is: read the guideline (done, above), comply visibly, document here. Recorded so it is never re-litigated in a later phase.
Pre-cleared swap-in: VOICEVOX Nemo
If the character terms ever become inconvenient, VOICEVOX Nemo is the drop-in replacement and
the swap is a two-line change (SPEAKER_STYLE_ID plus the model path):
- Model file:
n0.vvmfrom the samevoicevox_vvm0.17.0 release; nine voices, style IDs 10000-10008 (男声1-3, 女声1-6), all ノーマル. - Credit string is just
VOICEVOX Nemo— no character name, no third-party rights holder. - Terms: https://voicevox.hiroshiba.jp/nemo/term/, restated in the
voicevox_vvmREADME. - No ¥400,000 credit-omission clause, because there is no separate character licensor.
Cost of the swap: the loss of a recognisable character identity. Recommendation remains ずんだもん.
Open JTalk dictionary licence
BSD-3-Clause. The shipped voicevox/open_jtalk_dic_utf_8-1.11/COPYING reads verbatim:
Copyright (c) 2009, Nara Institute of Science and Technology, Japan. All rights reserved.
Redistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met: Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer. Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following disclaimer in the documentation and/or other materials provided with the distribution. Neither the name of the Nara Institute of Science and Technology (NAIST) nor the names of its contributors may be used to endorse or promote products derived from this software without specific prior written permission.
Correction to 01-RESEARCH.md: RESEARCH attributes this dictionary to "Nagoya Institute of Technology / HTS Working Group". Both institutions are genuinely in the picture and RESEARCH conflated them:
- The dictionary data (
open_jtalk_dic_utf_8-1.11, derived from the NAIST Japanese Dictionary / IPAdic lineage) is © 2009 Nara Institute of Science and Technology, BSD-3-Clause — this is the notice that must be reproduced for the files committed here. - Open JTalk itself — the engine whose dictionary format this is, and whose code is linked
inside
voicevox_core— is from the Nagoya Institute of Technology and the HTS Working Group, also under a modified BSD licence.
LICENSES.md (plan 01-09) must carry both attributions; reproducing only Nagoya would fail to
satisfy the BSD notice on the files actually redistributed in this repo.
Local development
The Space consumes requirements.txt, which pins the Linux wheel. Local development on
Windows needs the win_amd64 asset from the same release. It is wired into pyproject.toml as
the voice extra with marker-differentiated [tool.uv.sources], so one lockfile serves both:
uv sync --extra dev --extra voice
Never put a Windows wheel in requirements.txt — that file is consumed by a Linux builder.
To obtain the ONNX Runtime explicitly rather than letting get_synthesizer() fetch it, use the
official downloader (option b) and point VOICEVOX_ORT_PATH at the result:
curl -sSfL https://github.com/VOICEVOX/voicevox_core/releases/download/0.17.0/download-linux-x64 -o download
chmod +x download
./download --only onnxruntime -o ./voicevox_runtime
voicevox_runtime/ is gitignored.
Environment variables
| Variable | Default | Purpose |
|---|---|---|
VOICEVOX_DICT_DIR |
voicevox/open_jtalk_dic_utf_8-1.11 |
Open JTalk dictionary directory |
VOICEVOX_VVM_PATH |
voicevox/model/zundamon.vvm |
Voice model file |
VOICEVOX_ORT_PATH |
(auto-resolved) | Explicit path to the VOICEVOX ONNX Runtime shared library |
VOICEVOX_ORT_DIR |
voicevox_runtime |
Where the runtime is cached / downloaded to |
VOICEVOX_ORT_NO_DOWNLOAD |
unset | Set to 1 to forbid the lazy download entirely |
Zero GPU, permanently
Nothing in src/japanese_avatar/voice/ imports spaces or torch, and no @spaces.GPU
decorator exists anywhere on the synthesis path. Synthesis is CPU-only by design — that is the
whole reason this engine was chosen, since it means speech output consumes none of a visitor's
daily ZeroGPU quota. tests/test_tts_contract.py::test_no_gpu_imports_on_synthesis_path walks the
package with ast and fails if that ever changes; plan 01-08 adds the wider turn-path scan.