WolfDavid commited on
Commit
c3f6d17
·
1 Parent(s): 3c77193

feat(01-09): add LICENSES.md, prove the credits and the VRM meta on the Space, record the SC-4 run

Browse files

LICENSES.md (DPLY-04) merges docs/ASSETS.md and docs/VOICEVOX-SETUP.md into one table
row per shipped or runtime-loaded asset - the VRM, VOICEVOX CORE, the VOICEVOX ONNX
Runtime, the voice model, the zundamon character, the Open JTalk dictionary, three.js,
@pixiv /three-vrm, @huggingface/transformers and the Whisper weights - with every cell
filled, the permission each one is used under, the exact credit strings, the model-terms
clause the committed .vvm relies on quoted verbatim, the character terms' fee and
prohibitions, the Nemo swap-in, the NAIST notice reproduced, and a "not used" section
stating that AvatarSample_A/B/C are not CC0. Two facts corrected against primary sources
while writing it: voicevox_core 0.17.0 is MIT (the LGPL/dual licence applied to prebuilt
cores below 0.16), and the whisper-base WEIGHTS are Apache-2.0 per the model card while
the openai/whisper code is MIT.

Two machine-readable blocks make the document and the shipped assets unable to disagree:
test_vrm_meta_matches_licenses reads the deployed page's vrmMeta and compares it with the
vrm_meta block through one normalising helper (VRM 0.0 and 1.0 key names) - 14 fields,
0 mismatches; test_credits_visible reads the required credit strings from the credits
block and asserts them in the footer, the flow-down notice (72 px under Replay, both
terms linked), the About panel and the server-delivered GET /config, with no interaction.
Gradio 6 serves an SSR shell, so the root HTML carries 0 hits by design (docs/HOSTING.md
finding 5); the server-delivered assertion is on /config.

SC-4: the entire tests/e2e/test_avatar_loop.py ran against Space revision c911d74 with
DISABLE_GPU=1 read back from the Space variables in the same script - 15 passed, 0
failed. Recorded under docs/HOSTING.md "SC-4 run record" with the latency numbers the
run produced for plan 01-10.

Files changed (3) hide show
  1. LICENSES.md +232 -0
  2. docs/HOSTING.md +28 -0
  3. tests/e2e/test_avatar_loop.py +153 -0
LICENSES.md ADDED
@@ -0,0 +1,232 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Third-party licences (DPLY-04)
2
+
3
+ Every third-party asset this project ships, fetches at runtime, or loads into a visitor's
4
+ browser, with the permission each one is used under and how that permission is honoured.
5
+
6
+ This document **merges** `docs/ASSETS.md` (the VRM, plan 01-01) and `docs/VOICEVOX-SETUP.md`
7
+ (the three VOICEVOX licence layers and the dictionary, plan 01-04); it does not re-research
8
+ either. Every runtime-library licence below was re-verified on 2026-09-06 against the package
9
+ registry and the upstream repository, and the VRM's row is verified automatically:
10
+ `tests/e2e/test_avatar_loop.py::test_vrm_meta_matches_licenses` reads the shipped file's own
11
+ embedded `VRMC_vrm.meta` on the deployed Space and compares it with the `vrm_meta` block in
12
+ this file, so the document and the binary cannot silently disagree.
13
+
14
+ The repository's own code is MIT (see the Space `README.md` front-matter). Nothing below is
15
+ covered by that licence.
16
+
17
+ ## Summary
18
+
19
+ | Asset | Version | Author / rights holder | Source URL | Licence / terms URL | Permission relied on | Required credit string | How we comply |
20
+ |---|---|---|---|---|---|---|---|
21
+ | **VRM character** `avatar/assets/tutor.vrm` | `VRM1_Constraint_Twist_Sample` v1.0.1, VRM 1.0 | pixiv Inc. — `(c) 2022 pixiv Inc.` | https://github.com/vrm-c/vrm-specification/tree/master/samples/VRM1_Constraint_Twist_Sample | https://vrm.dev/licenses/1.0/ (VRM Public License 1.0, with per-file flags embedded in the file) | Embedded `VRMC_vrm.meta`: `allowRedistribution: true`, `avatarPermission: everyone`, `modification: allowModificationRedistribution`, `commercialUsage: corporation` | **none** (`creditNotation: unnecessary`); credited voluntarily as `VRM1_Constraint_Twist_Sample (c) 2022 pixiv Inc. - VRM Public License 1.0` | Committed through Git LFS unmodified; credit rendered in the footer, the About panel and the README; the one restriction (`allowAntisocialOrHateUsage: false`) is stated below; sourcing gate re-raised in plan 01-10 |
22
+ | **VOICEVOX CORE** (software) | `voicevox_core` 0.17.0 (Python wheel, abi3) | Hiroshiba Kazuyuki (ヒホ) / the VOICEVOX project | https://github.com/VOICEVOX/voicevox_core/releases/tag/0.17.0 | https://voicevox.hiroshiba.jp/term/ (VOICEVOX ソフトウェア利用規約); code and build artefacts MIT — https://github.com/VOICEVOX/voicevox_core/blob/0.17.0/LICENSE | Software terms: commercial and non-commercial use permitted; generated audio may be provided onward **subject to the flow-down obligation of clause 3**. MIT for the linked code | `VOICEVOX:<character>` — satisfied by the character row below; plus the terms notice wherever audio is obtainable | Installed from the official release URL pinned in `requirements.txt`; the wheel is **never vendored** (禁止事項 forbids unauthorised redistribution of the software); flow-down notice rendered next to the Replay control |
23
+ | **VOICEVOX ONNX Runtime** | `voicevox_onnxruntime` 1.23.2 | VOICEVOX project build of Microsoft ONNX Runtime | https://github.com/VOICEVOX/onnxruntime-builder/releases/tag/voicevox_onnxruntime-1.23.2 | MIT (both `VOICEVOX/onnxruntime-builder` and `microsoft/onnxruntime`) | MIT | none | Fetched lazily at runtime into the gitignored `voicevox_runtime/` from the official release; never committed or redistributed (`docs/VOICEVOX-SETUP.md` § ONNX Runtime acquisition) |
24
+ | **VOICEVOX voice model** `voicevox/model/zundamon.vvm` | `voicevox_vvm` 0.17.0, file `0.vvm` (SHA-256 `ecd35374d4182cd883cba5040376f7f888cc6ba248b1c2f4cea07cdb34bb1318`) | the VOICEVOX project (model file); voices inside it belong to their character licensors | https://github.com/VOICEVOX/voicevox_vvm/releases/tag/0.17.0 | VOICEVOX 音声モデル 利用規約 — https://github.com/VOICEVOX/voicevox_vvm/blob/main/README.md (shipped as `README.txt` / `TERMS.txt` with the release) | Clause 2: 「アプリケーションに組み込んで再配布することができます」 — embedded redistribution is explicitly permitted; clause 4 flow-down; a credit that shows VOICEVOX was used is required | `VOICEVOX:ずんだもん` (only ずんだもん, style 3, is used; the other voices packaged in `0.vvm` are neither used nor credited) | Committed through Git LFS under clause 2; credit and flow-down notice rendered; the model is used only through VOICEVOX CORE |
25
+ | **ずんだもん** (character voice) | — (voice library shipped inside `0.vvm`) | SSS LLC (合同会社SSS) | https://zunko.jp/ | https://zunko.jp/con_ongen_kiyaku.html (音源利用規約, one document covering nine characters) | Commercial and non-commercial use of generated audio permitted **with credit**; app use requires the credit on an introduction/about screen | **`VOICEVOX:ずんだもん`** — exact, ASCII colon, no spaces | Persistent footer credit + About panel on first paint (server-delivered, not injected); flow-down notice linking both terms URLs; no prohibited-use path exists (see below) |
26
+ | **Open JTalk dictionary** `voicevox/open_jtalk_dic_utf_8-1.11/` | `open_jtalk_dic_utf_8-1.11` | © 2009 **Nara Institute of Science and Technology** (NAIST), Japan | https://downloads.sourceforge.net/project/open-jtalk/Dictionary/open_jtalk_dic-1.11/open_jtalk_dic_utf_8-1.11.tar.gz | BSD-3-Clause — `voicevox/open_jtalk_dic_utf_8-1.11/COPYING` (reproduced below) | BSD-3-Clause redistribution in binary form, with the notice reproduced | none (attribution notice) | `COPYING` is committed alongside the dictionary files; the notice is reproduced in this file; the About panel names NAIST. Open JTalk itself (Nagoya Institute of Technology / HTS Working Group, modified BSD) is linked inside VOICEVOX CORE and is attributed too |
27
+ | **three.js** | 0.185.1 | three.js authors (mrdoob et al.) | https://github.com/mrdoob/three.js | MIT — https://github.com/mrdoob/three.js/blob/master/LICENSE | MIT | none | Loaded in the visitor's browser from `https://esm.sh/three@0.185.1` today; plan 01-10 vendors it under `avatar/vendor/` with its LICENSE file |
28
+ | **@pixiv/three-vrm** | 3.5.5 | pixiv Inc. | https://github.com/pixiv/three-vrm | MIT — https://github.com/pixiv/three-vrm/blob/dev/LICENSE | MIT | none | Loaded from `https://esm.sh/@pixiv/three-vrm@3.5.5?deps=three@0.185.1` today; vendored in plan 01-10 with its LICENSE file |
29
+ | **@huggingface/transformers** (transformers.js) | 4.2.0 | Hugging Face | https://github.com/huggingface/transformers.js | Apache-2.0 — https://github.com/huggingface/transformers.js/blob/main/LICENSE | Apache-2.0 | none (NOTICE reproduction if vendored) | Loaded from `https://esm.sh/@huggingface/transformers@4.2.0` today; vendored in plan 01-10 with LICENSE (and NOTICE, if the package ships one) |
30
+ | **Whisper ASR model** `onnx-community/whisper-base` (q4 ONNX) | `onnx-community/whisper-base`, converted from `openai/whisper-base` | OpenAI (model); onnx-community (ONNX conversion) | https://huggingface.co/onnx-community/whisper-base | Weights: Apache-2.0 per the `openai/whisper-base` model card (`license: apache-2.0`; the conversion card declares `base_model: openai/whisper-base` and no licence of its own). Source code of `openai/whisper`: MIT | Apache-2.0 use of the weights in the visitor's browser | none (attribution) | **Not in this repository**: the visitor's browser downloads it from huggingface.co on the first accepted push (135.8 MB); attributed here and in the About panel ("Whisper"). The "better accuracy" model `whisper-large-v3-turbo` is not offered yet |
31
+
32
+ Nine rows for eight asset classes: the ONNX Runtime is listed separately from VOICEVOX CORE
33
+ because it is fetched by a different mechanism and under a different licence.
34
+
35
+ ## Machine-readable blocks
36
+
37
+ The deployed suite parses these. Keep the keys and the values exactly as the shipped assets
38
+ and the rendered page have them.
39
+
40
+ The VRM's embedded `VRMC_vrm.meta`, as read out of `avatar/assets/tutor.vrm` itself (VRM
41
+ 1.0 key names; `thumbnailImage` is an image index and is omitted):
42
+
43
+ ```vrm_meta
44
+ name: VRM1_Constraint_Twist_Sample
45
+ version: v1.0.1
46
+ authors: pixiv Inc.
47
+ copyrightInformation: (c) 2022 pixiv Inc.
48
+ licenseUrl: https://vrm.dev/licenses/1.0/
49
+ avatarPermission: everyone
50
+ commercialUsage: corporation
51
+ creditNotation: unnecessary
52
+ allowRedistribution: true
53
+ modification: allowModificationRedistribution
54
+ allowAntisocialOrHateUsage: false
55
+ allowExcessivelySexualUsage: true
56
+ allowExcessivelyViolentUsage: true
57
+ allowPoliticalOrReligiousUsage: true
58
+ ```
59
+
60
+ The credit strings the page must render (`tests/e2e/test_avatar_loop.py::test_credits_visible`
61
+ checks the footer, the About panel and the server-delivered config for them):
62
+
63
+ ```credits
64
+ voice: VOICEVOX:ずんだもん
65
+ avatar: VRM1_Constraint_Twist_Sample (c) 2022 pixiv Inc. - VRM Public License 1.0
66
+ ```
67
+
68
+ ## VRM character — `avatar/assets/tutor.vrm`
69
+
70
+ > **TEMPORARY DEV ASSET.** Legally clean as recorded here, but it ships a spec sample's name.
71
+ > Plan 01-10 must re-raise the sourcing checkpoint (keep it, or author a bespoke VRoid Studio
72
+ > character with a deliberately set `vrm.meta`) before the phase-exit licensing audit. See
73
+ > `docs/ASSETS.md` § Open item for plan 01-10.
74
+
75
+ VRM Public License 1.0 is a **template, not a fixed grant**: redistribution, avatar use,
76
+ modification and commercial use are per-file flags set by the licensor inside the model's own
77
+ `VRMC_vrm.meta`. Two files citing the same licence URL can grant materially different rights.
78
+ The facts above were therefore parsed out of this file's glTF JSON chunk (10,776,032 bytes,
79
+ SHA-256 `12c2b97e95e700783a6a550dc0eee2d7880aeedccef9ae67bc4c5a2f0f2631a2`), never off a README.
80
+
81
+ Permissions relied on, and what each one covers here:
82
+
83
+ | Flag | Value | What it licenses in this project |
84
+ |---|---|---|
85
+ | `allowRedistribution` | `true` | Committing the file to this public repository and serving it from the Space |
86
+ | `avatarPermission` | `everyone` | Any visitor uses it as the tutor avatar |
87
+ | `modification` | `allowModificationRedistribution` | The rest pose applied at mount, and any future retargeting, still redistributable |
88
+ | `commercialUsage` | `corporation` | The broadest tier; a public portfolio Space is covered unambiguously |
89
+ | `creditNotation` | `unnecessary` | No credit is required. One is rendered anyway |
90
+ | `allowAntisocialOrHateUsage` | **`false`** | The one restriction. A Japanese-language tutor has no such use path; stated so it is not silently dropped |
91
+
92
+ ## VOICEVOX — three licence layers, read separately
93
+
94
+ The voice output involves three rights holders, and their terms differ in one place that looks
95
+ like a contradiction until both are read.
96
+
97
+ **1. Software** — VOICEVOX ソフトウェア利用規約, https://voicevox.hiroshiba.jp/term/. Commercial
98
+ and non-commercial use permitted. Clause 3 attaches a **flow-down obligation**: when audio
99
+ generated here is made available to others, those others must be bound to the same terms.
100
+ 禁止事項 forbids redistributing the software in whole or in part without authorisation, so the
101
+ `voicevox_core` wheel and the ONNX Runtime binary are always **referenced** from their official
102
+ release URLs and never copied into this repository. Separately from the terms of use, the
103
+ VOICEVOX CORE **code and build artefacts** are MIT-licensed as of 0.17.0 (`LICENSE`, © 2021
104
+ Hiroshiba Kazuyuki; README § ライセンス). *Correction to earlier project notes:* prebuilt cores
105
+ below version 0.16 were under a different (LGPL v3 / commercial dual) licence, and that is the
106
+ version some of this project's planning text still describes; the shipped 0.17.0 is MIT.
107
+
108
+ **2. Voice model** — VOICEVOX 音声モデル 利用規約, in the `voicevox_vvm` README and shipped as
109
+ `README.txt` / `TERMS.txt` with the release. Clause 2 reads
110
+ 「アプリケーションに組み込んで再配布することができます」 — **embedded redistribution is explicitly
111
+ permitted**, which is why `voicevox/model/zundamon.vvm` is committed (through Git LFS) while the
112
+ software is not. That asymmetry is deliberate on VOICEVOX's side and on ours: the software terms
113
+ forbid redistributing the *software*; the model terms permit redistributing the *model* inside an
114
+ application. Clause 4 carries the same flow-down obligation, and a credit showing that VOICEVOX
115
+ was used is required.
116
+
117
+ **3. Character** — ずんだもん, SSS LLC, https://zunko.jp/con_ongen_kiyaku.html. Audio generated
118
+ with the ずんだもん voice library may be used commercially and non-commercially **provided the
119
+ credit `VOICEVOX:ずんだもん` is displayed**. Without the credit, per-character licensing is
120
+ **¥400,000 (+ tax) per character** — which is why the credit is not optional and why it is
121
+ rendered on first paint rather than behind a click. Placement (clause 2): for app use, on an
122
+ introduction/about screen, somewhere findable with a little looking
123
+ (「アプリの紹介画面などに記載をお願いします。(少し探せばわかる場所に)」); Phase 1 renders it in a
124
+ persistent footer beside the avatar **and** in the About panel, open on first load.
125
+
126
+ Prohibited uses under the character terms, recorded so later phases can check against them:
127
+ 公序良俗に反する利用; 政治的・宗教的活動; 情報商材; **意図的な**虚偽情報の作成・拡散; 風俗営業;
128
+ 反社会的勢力による利用. The falsehood clause is scoped to *intentional* creation or spread —
129
+ relevant once an LLM produces the avatar's lines (Phase 3): a tutor that is sometimes wrong is not
130
+ in scope of it; a product designed to mislead would be.
131
+
132
+ There is no escalation path and none should be sought: SSS LLC's 免責条項 2 states that they do
133
+ not, as a rule, answer "does my use qualify" questions
134
+ (「本ガイドラインに該当するかどうかのご質問については、原則お答えしておりません。」). The posture is:
135
+ read the guideline, comply visibly, document here.
136
+
137
+ The credit obligation survives an engine swap: the official VOICEVOX Q&A answers that using the
138
+ intermediate `AudioQuery` with another synthesiser still requires the credit
139
+ (「必要です。」). The viseme timeline in this project is built from the `AudioQuery`, so the credit
140
+ is owed on that basis alone.
141
+
142
+ **Pre-cleared swap-in: VOICEVOX Nemo.** If the character terms ever become inconvenient,
143
+ `n0.vvm` from the same `voicevox_vvm` 0.17.0 release (nine voices, style IDs 10000–10008) is a
144
+ two-line change. Credit is just `VOICEVOX Nemo`; the terms live entirely in the `voicevox_vvm`
145
+ README (https://voicevox.hiroshiba.jp/nemo/term/); there is no third-party rights holder and no
146
+ ¥400,000 credit-omission clause. The cost is the loss of a recognisable character.
147
+
148
+ ## Open JTalk dictionary — `voicevox/open_jtalk_dic_utf_8-1.11/`
149
+
150
+ BSD-3-Clause. The shipped `COPYING` reads, verbatim:
151
+
152
+ > Copyright (c) 2009, Nara Institute of Science and Technology, Japan.
153
+ > All rights reserved.
154
+ >
155
+ > Redistribution and use in source and binary forms, with or without modification, are permitted
156
+ > provided that the following conditions are met: Redistributions of source code must retain the
157
+ > above copyright notice, this list of conditions and the following disclaimer. Redistributions in
158
+ > binary form must reproduce the above copyright notice, this list of conditions and the following
159
+ > disclaimer in the documentation and/or other materials provided with the distribution. Neither
160
+ > the name of the Nara Institute of Science and Technology (NAIST) nor the names of its
161
+ > contributors may be used to endorse or promote products derived from this software without
162
+ > specific prior written permission.
163
+ >
164
+ > THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS" AND ANY EXPRESS OR
165
+ > IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND
166
+ > FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR
167
+ > CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR
168
+ > CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
169
+ > SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY
170
+ > THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR
171
+ > OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE
172
+ > POSSIBILITY OF SUCH DAMAGE.
173
+
174
+ Two institutions are genuinely involved, and earlier research conflated them:
175
+
176
+ - The **dictionary data** committed here (`open_jtalk_dic_utf_8-1.11`, NAIST Japanese Dictionary /
177
+ IPAdic lineage) is © 2009 **Nara Institute of Science and Technology** — the notice above is the
178
+ one that must be reproduced for the files actually redistributed in this repository.
179
+ - **Open JTalk itself** — the engine whose dictionary format this is, linked inside
180
+ `voicevox_core` — is by the **Nagoya Institute of Technology** and the **HTS Working Group**,
181
+ under a modified BSD licence. It is not redistributed by this repository; it is attributed
182
+ because the About panel names the engine.
183
+
184
+ ## Runtime libraries loaded into the visitor's browser
185
+
186
+ | Library | Version | Licence | Copyright | Loaded from (today) |
187
+ |---|---|---|---|---|
188
+ | three.js | 0.185.1 | MIT | © 2010-2026 three.js authors | `https://esm.sh/three@0.185.1` (+ `examples/jsm/loaders/GLTFLoader.js`) |
189
+ | @pixiv/three-vrm | 3.5.5 | MIT | © 2019-2026 pixiv Inc. | `https://esm.sh/@pixiv/three-vrm@3.5.5?deps=three@0.185.1` |
190
+ | @huggingface/transformers | 4.2.0 | Apache-2.0 | © Hugging Face | `https://esm.sh/@huggingface/transformers@4.2.0` |
191
+
192
+ All three are fetched by the visitor's browser from esm.sh at runtime in this revision; **plan
193
+ 01-10 vendors them** under `avatar/vendor/` with each package's LICENSE (and NOTICE where one
194
+ exists) alongside, at which point this table's last column changes to the vendored path. MIT and
195
+ Apache-2.0 both permit that redistribution with the notice preserved.
196
+
197
+ The ASR model (`onnx-community/whisper-base`, q4) is downloaded by the browser from
198
+ huggingface.co on the first accepted push-to-talk utterance and cached by the browser; it is
199
+ never served from this repository or the Space. `docs/ASR-TIERS.md` records why this model and
200
+ this quantisation.
201
+
202
+ ## Not used, and why
203
+
204
+ - **`AvatarSample_A`, `AvatarSample_B` and `AvatarSample_C` are NOT CC0.** Copyright is retained
205
+ and VRoid Hub's conditions apply; VRoid's own FAQ prohibits redistributing them under a CC0
206
+ designation. The claim is repeated confidently across tutorials and blog posts, and it is
207
+ false. Recorded here so it is not reintroduced into this repository.
208
+ - CC0 **cannot be set** on VRoid Hub at all, so any "CC0 on VRoid Hub" claim is false by
209
+ construction; and enabling "Allow third-party usage" on VRoid Hub **overwrites** a model's
210
+ embedded licence with the Hub's own conditions, so a Hub-sourced file's embedded flags may not
211
+ say what its author intended. Both are reasons the VRM row above reads the flags out of the
212
+ shipped file.
213
+ - **Hosted ASR on accelerated hardware (tier D)** is deliberately not built: it would burn the
214
+ visitor's ZeroGPU quota, contradicting SC-4. `tests/test_transport_seam.py` fails if one appears.
215
+ - **Style-Bert-VITS2** (AGPL-3.0), **pykakasi** (GPL-3.0) and **edge-tts** (undocumented
216
+ endpoint) were rejected at stack selection for licence reasons and do not appear anywhere in
217
+ the tree. Python dependencies that are not assets (`gradio`, `spaces`, `huggingface_hub`,
218
+ `numpy`) are ordinary Apache-2.0 / BSD packages installed from PyPI and are not redistributed.
219
+
220
+ ## Verification record
221
+
222
+ | Check | Result (2026-09-06) |
223
+ |---|---|
224
+ | VRM `VRMC_vrm.meta` parsed from the shipped file | matches the `vrm_meta` block above, key for key |
225
+ | `voicevox_core` 0.17.0 licence | README § ライセンス at tag 0.17.0: MIT; `LICENSE` © 2021 Hiroshiba Kazuyuki; GitHub API `MIT` |
226
+ | `voicevox_vvm` clause 2 | present verbatim in the `main` README (line 21) |
227
+ | `VOICEVOX/onnxruntime-builder`, `microsoft/onnxruntime` | GitHub API: MIT, MIT |
228
+ | `three@0.185.1`, `@pixiv/three-vrm@3.5.5`, `@huggingface/transformers@4.2.0` | npm registry: MIT, MIT, Apache-2.0 (all three are also the current `latest`) |
229
+ | `openai/whisper-base` model card | HF API `cardData.license = apache-2.0`; `openai/whisper` repo MIT |
230
+ | `onnx-community/whisper-base` card | no licence field; `base_model: openai/whisper-base` |
231
+ | Open JTalk `COPYING` | first line: `Copyright (c) 2009, Nara Institute of Science and Technology, Japan.` |
232
+ | Terms URLs answer | voicevox.hiroshiba.jp/term/ 200 · zunko.jp/con_ongen_kiyaku.html 200 · voicevox.hiroshiba.jp/nemo/term/ 200 · vrm.dev/licenses/1.0/ 200 |
docs/HOSTING.md CHANGED
@@ -149,6 +149,34 @@ python -c "from huggingface_hub import HfApi; print(HfApi().get_space_variables(
149
  `AVATAR_TRANSPORT` is deliberately **not** set, so the Space uses the component's default,
150
  `inline` — the transport the spike is testing. Setting it to `iframe` is the entire fallback.
151
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
152
  ## Why not the alternatives
153
 
154
  - **`cpu-basic` (free)** — no longer available. Attempting it returned **HTTP 402 Payment Required**
 
149
  `AVATAR_TRANSPORT` is deliberately **not** set, so the Space uses the component's default,
150
  `inline` — the transport the spike is testing. Setting it to `iframe` is the entire fallback.
151
 
152
+ ### SC-4 run record — the whole loop with the GPU disabled (plan 01-09)
153
+
154
+ The platform forces one `@spaces.GPU` function to exist (`app.py::zerogpu_probe`, finding 4
155
+ below), so "no GPU function exists" is not an available proof. The honest proof is that the
156
+ whole turn loop completes without ever reaching one — and reaching it would raise, because it
157
+ checks `DISABLE_GPU` on every call.
158
+
159
+ | | |
160
+ |---|---|
161
+ | Date | 2026-09-06, 03:27–03:32 UTC |
162
+ | Space revision | `c911d74b378981e6ba0d0111eb570ed3b20927db` (plan 01-08 code, deployed by the owner) |
163
+ | Runtime at run time | `RUNNING`, hardware `zero-a10g`, read from `space_info()` in the same script |
164
+ | `DISABLE_GPU` | **`1`**, read back from `get_space_variables()` in the same script (set 2026-09-05 by plan 01-05) |
165
+ | Command | `pytest tests/e2e/test_avatar_loop.py -q --space-url https://wolfdavid-japanese-learning-avatar.hf.space` |
166
+ | Result | **15 passed, 0 failed** (14 test functions; `test_silence_rejected` runs twice) — exit 0 |
167
+ | Turns that completed | typed turn ×5 (incl. 20 wave-5 interactions in `test_no_remount`), push-to-talk ×2 (WASM tier), replay, slower re-synthesis; silence and cafe noise produced 0 turns |
168
+
169
+ Every row in `01-VALIDATION.md`'s Per-Task Verification Map that names a
170
+ `tests/e2e/test_avatar_loop.py` node was green in this run. The two ASR rows ran on the WASM
171
+ tier because headless Chromium has no WebGPU adapter; neither tier touches the server.
172
+
173
+ Numbers this run recorded for plan 01-10's latency harness (all on the Space's CPU, one
174
+ visitor): `synthesis_ms` 3088–6724 for こんにちは across four typed turns in the session,
175
+ 10194–10292 for the 36-mora long sentence (normal / speedScale 0.75); dispatch→speech-start
176
+ 4170–7252 ms for こんにちは; replay 0 ms and 0 requests; page→ready 4.0–16.8 s with
177
+ `tutor.vrm` 3.0–7.4 s of that. Wake was 0.2 s every time — the Space was already running, so
178
+ true cold-from-sleep remains 01-10's manual row.
179
+
180
  ## Why not the alternatives
181
 
182
  - **`cpu-basic` (free)** — no longer available. Attempting it returned **HTTP 402 Payment Required**
tests/e2e/test_avatar_loop.py CHANGED
@@ -18,6 +18,8 @@ or a decoded ``AudioBuffer.duration``, never a timeline inference.
18
 
19
  from __future__ import annotations
20
 
 
 
21
  import time
22
  from pathlib import Path
23
 
@@ -872,3 +874,154 @@ def test_asr_wasm_fallback(
872
  assert page_errors == [], f"an uncaught exception reached the page: {page_errors}"
873
  assert errors == [], f"the facade emitted error events on the fallback path: {errors}"
874
  speech_events.wait_for(page, "speech-end", timeout_ms=SPEECH_END_TIMEOUT_MS)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
18
 
19
  from __future__ import annotations
20
 
21
+ import json
22
+ import re
23
  import time
24
  from pathlib import Path
25
 
 
874
  assert page_errors == [], f"an uncaught exception reached the page: {page_errors}"
875
  assert errors == [], f"the facade emitted error events on the fallback path: {errors}"
876
  speech_events.wait_for(page, "speech-end", timeout_ms=SPEECH_END_TIMEOUT_MS)
877
+
878
+
879
+ # ================================================================ credits and licences
880
+
881
+ # The exact flow-down sentence rendered next to the Replay control (blocks.py).
882
+ FLOW_DOWN_SENTENCE = "by using it you agree to comply with them"
883
+ VOICEVOX_TERMS_HOST = "voicevox.hiroshiba.jp/term"
884
+ ZUNDAMON_TERMS_HOST = "zunko.jp/con_ongen_kiyaku.html"
885
+ # The notice sits directly under the Replay / Slower row; this is a generous bound on the
886
+ # vertical gap between the bottom of the Replay button and the top of the notice.
887
+ NOTICE_MAX_GAP_PX = 160
888
+
889
+ # VRM 0.0 meta uses different key names for the same facts. Normalised in ONE place so the
890
+ # assertions below never need an `or` fallback.
891
+ VRM0_TO_VRM1_KEYS = {
892
+ "title": "name",
893
+ "author": "authors",
894
+ "otherLicenseUrl": "licenseUrl",
895
+ "reference": "references",
896
+ }
897
+
898
+
899
+ def licenses_block(name: str) -> dict[str, str]:
900
+ """Parse a ```<name> fenced block of ``key: value`` lines out of LICENSES.md."""
901
+ text = LICENSES_MD.read_text(encoding="utf-8")
902
+ match = re.search(rf"```{re.escape(name)}\n(.*?)\n```", text, re.S)
903
+ assert match, f"LICENSES.md has no ```{name} block"
904
+ block: dict[str, str] = {}
905
+ for line in match.group(1).splitlines():
906
+ if not line.strip() or line.lstrip().startswith("#"):
907
+ continue
908
+ key, sep, value = line.partition(":")
909
+ assert sep, f"LICENSES.md ```{name} line is not `key: value`: {line!r}"
910
+ block[key.strip()] = " ".join(value.split())
911
+ return block
912
+
913
+
914
+ def normalise_vrm_meta(meta: dict) -> dict[str, str]:
915
+ """Flatten a VRM meta block - VRM 0.0 or 1.0 key names - to VRM 1.0 names with
916
+ string values: lists joined with ' | ', booleans as 'true'/'false', whitespace
917
+ collapsed, nulls dropped. The documented helper the assertions compare through."""
918
+ out: dict[str, str] = {}
919
+ for key, value in meta.items():
920
+ if value is None:
921
+ continue
922
+ name = VRM0_TO_VRM1_KEYS.get(key, key)
923
+ if isinstance(value, list):
924
+ text = " | ".join(str(v) for v in value)
925
+ elif isinstance(value, bool):
926
+ text = "true" if value else "false"
927
+ else:
928
+ text = str(value)
929
+ out[name] = " ".join(text.split())
930
+ return out
931
+
932
+
933
+ @pytest.mark.deployed
934
+ def test_credits_visible(page, space_url, warm_space):
935
+ """DPLY-04. On first load, with NO interaction: the VOICEVOX:ずんだもん credit is visible,
936
+ the flow-down notice sits by the Replay control with both terms linked, the About panel
937
+ carries both terms URLs and the VRM credit, and the credit is delivered by the server
938
+ rather than injected by this project's JavaScript.
939
+
940
+ The strings asserted are the ones LICENSES.md declares as required, parsed from its
941
+ ```credits block, so the document and the page cannot disagree.
942
+ """
943
+ credits = licenses_block("credits")
944
+ voice, avatar = credits["voice"], credits["avatar"]
945
+ assert voice == "VOICEVOX:ずんだもん", f"LICENSES.md declares the voice credit as {voice!r}"
946
+
947
+ page.goto(space_url, timeout=STAGE_ATTACHED_TIMEOUT_MS + 30_000)
948
+ page.wait_for_selector("#credits", state="visible", timeout=STAGE_ATTACHED_TIMEOUT_MS)
949
+
950
+ footer = page.locator("#credits")
951
+ assert footer.is_visible()
952
+ footer_text = footer.inner_text()
953
+ assert voice in footer_text, f"#credits reads {footer_text!r}"
954
+ assert avatar in footer_text, f"#credits reads {footer_text!r}"
955
+
956
+ notice = page.locator("#terms-notice")
957
+ assert notice.is_visible()
958
+ notice_text = notice.inner_text()
959
+ assert FLOW_DOWN_SENTENCE in notice_text, f"#terms-notice reads {notice_text!r}"
960
+ assert voice in notice_text
961
+ notice_links = notice.locator("a").evaluate_all("els => els.map((e) => e.href)")
962
+ assert any(VOICEVOX_TERMS_HOST in h for h in notice_links), notice_links
963
+ assert any(ZUNDAMON_TERMS_HOST in h for h in notice_links), notice_links
964
+
965
+ replay = page.locator("#replay-button")
966
+ assert replay.is_visible()
967
+ replay_box = replay.bounding_box()
968
+ notice_box = notice.bounding_box()
969
+ assert replay_box and notice_box
970
+ gap = notice_box["y"] - (replay_box["y"] + replay_box["height"])
971
+ print(
972
+ f"[deployed] credits: footer {footer_text!r}; notice {gap:.0f}px below Replay; "
973
+ f"notice links {notice_links}"
974
+ )
975
+ assert -1 <= gap <= NOTICE_MAX_GAP_PX, (
976
+ f"#terms-notice is {gap:.0f}px from the Replay button; it must sit with the controls"
977
+ )
978
+
979
+ about = page.locator("#about-panel")
980
+ assert about.count() == 1
981
+ about_text = about.inner_text()
982
+ about_links = about.locator("a").evaluate_all("els => els.map((e) => e.href)")
983
+ assert any(VOICEVOX_TERMS_HOST in h for h in about_links), about_links
984
+ assert any(ZUNDAMON_TERMS_HOST in h for h in about_links), about_links
985
+ assert voice in about_text
986
+ assert avatar in about_text, f"the About panel does not carry the VRM credit: {about_text!r}"
987
+
988
+ # Server-delivered, not injected by avatar/host.js: Gradio 6 serves an SSR shell and
989
+ # ships the component tree - including every server-rendered gr.HTML value - from
990
+ # GET /config (docs/HOSTING.md, first-deploy finding 5). The root HTML is reported for
991
+ # the record; the assertion is on what the server sends before any project script runs.
992
+ config = requests.get(f"{space_url}/config", timeout=30)
993
+ config.raise_for_status()
994
+ served = json.dumps(config.json(), ensure_ascii=False)
995
+ root_hits = requests.get(space_url, timeout=30).text.count(voice)
996
+ print(f"[deployed] credit in GET /config: {served.count(voice)}; in the SSR shell: {root_hits}")
997
+ assert voice in served, "the VOICEVOX credit is not in the server-delivered config"
998
+ assert avatar in served
999
+
1000
+
1001
+ @pytest.mark.deployed
1002
+ def test_vrm_meta_matches_licenses(page, space_url, warm_space, wait_for_avatar_ready):
1003
+ """DPLY-04. The shipped VRM's own embedded licence metadata agrees with LICENSES.md.
1004
+
1005
+ Read from the deployed page's ``vrmMeta`` (the stage copies the scalar fields of
1006
+ ``vrm.meta`` out of the loaded file) and compared with the ```vrm_meta block, key for
1007
+ key, through one normalising helper. If either side changes, this fails.
1008
+ """
1009
+ expected = licenses_block("vrm_meta")
1010
+ for key in ("name", "authors", "licenseUrl"):
1011
+ assert key in expected, f"LICENSES.md vrm_meta block lacks {key}"
1012
+
1013
+ wait_for_avatar_ready(page, space_url)
1014
+ debug = read_debug(page)
1015
+ meta = debug.get("vrmMeta") if debug else None
1016
+ assert meta, "vrmMeta is missing from the deployed debug surface; the VRM loaded without it"
1017
+ actual = normalise_vrm_meta(meta)
1018
+ mismatches = {k: (v, actual.get(k)) for k, v in expected.items() if actual.get(k) != v}
1019
+ print(
1020
+ f"[deployed] vrmMeta ({len(actual)} fields, metaVersion {actual.get('metaVersion')!r}) "
1021
+ f"vs LICENSES.md ({len(expected)} fields): {len(mismatches)} mismatch(es) {mismatches}"
1022
+ )
1023
+ assert not mismatches, (
1024
+ "the shipped VRM's embedded meta disagrees with LICENSES.md "
1025
+ f"(documented, actual): {mismatches}"
1026
+ )
1027
+ assert debug["vrmMetaTitle"] == expected["name"]