# Third-party licences (DPLY-04) Every third-party asset this project ships, fetches at runtime, or loads into a visitor's browser, with the permission each one is used under and how that permission is honoured. This document **merges** `docs/ASSETS.md` (the VRM, plan 01-01) and `docs/VOICEVOX-SETUP.md` (the three VOICEVOX licence layers and the dictionary, plan 01-04); it does not re-research either. Every runtime-library licence below was re-verified on 2026-09-06 against the package registry and the upstream repository, and the VRM's row is verified automatically: `tests/e2e/test_avatar_loop.py::test_vrm_meta_matches_licenses` reads the shipped file's own embedded `VRMC_vrm.meta` on the deployed Space and compares it with the `vrm_meta` block in this file, so the document and the binary cannot silently disagree. The repository's own code is MIT (see the Space `README.md` front-matter). Nothing below is covered by that licence. ## Summary | Asset | Version | Author / rights holder | Source URL | Licence / terms URL | Permission relied on | Required credit string | How we comply | |---|---|---|---|---|---|---|---| | **VRM character** `avatar/assets/tutor.vrm` | `VRM1_Constraint_Twist_Sample` v1.0.1, VRM 1.0 | pixiv Inc. — `(c) 2022 pixiv Inc.` | https://github.com/vrm-c/vrm-specification/tree/master/samples/VRM1_Constraint_Twist_Sample | https://vrm.dev/licenses/1.0/ (VRM Public License 1.0, with per-file flags embedded in the file) | Embedded `VRMC_vrm.meta`: `allowRedistribution: true`, `avatarPermission: everyone`, `modification: allowModificationRedistribution`, `commercialUsage: corporation` | **none** (`creditNotation: unnecessary`); credited voluntarily as `VRM1_Constraint_Twist_Sample (c) 2022 pixiv Inc. - VRM Public License 1.0` | Committed through Git LFS unmodified; credit rendered in the footer, the About panel and the README; the one restriction (`allowAntisocialOrHateUsage: false`) is stated below; sourcing gate re-raised and closed in plan 01-10 (kept; `docs/ASSETS.md`) | | **VOICEVOX CORE** (software) | `voicevox_core` 0.17.0 (Python wheel, abi3) | Hiroshiba Kazuyuki (ヒホ) / the VOICEVOX project | https://github.com/VOICEVOX/voicevox_core/releases/tag/0.17.0 | https://voicevox.hiroshiba.jp/term/ (VOICEVOX ソフトウェア利用規約); code and build artefacts MIT — https://github.com/VOICEVOX/voicevox_core/blob/0.17.0/LICENSE | Software terms: commercial and non-commercial use permitted; generated audio may be provided onward **subject to the flow-down obligation of clause 3**. MIT for the linked code | `VOICEVOX:` — satisfied by the character row below; plus the terms notice wherever audio is obtainable | Installed from the official release URL pinned in `requirements.txt`; the wheel is **never vendored** (禁止事項 forbids unauthorised redistribution of the software); flow-down notice rendered next to the Replay control | | **VOICEVOX ONNX Runtime** | `voicevox_onnxruntime` 1.23.2 | VOICEVOX project build of Microsoft ONNX Runtime | https://github.com/VOICEVOX/onnxruntime-builder/releases/tag/voicevox_onnxruntime-1.23.2 | MIT (both `VOICEVOX/onnxruntime-builder` and `microsoft/onnxruntime`) | MIT | none | Fetched lazily at runtime into the gitignored `voicevox_runtime/` from the official release; never committed or redistributed (`docs/VOICEVOX-SETUP.md` § ONNX Runtime acquisition) | | **VOICEVOX voice model** `voicevox/model/zundamon.vvm` | `voicevox_vvm` 0.17.0, file `0.vvm` (SHA-256 `ecd35374d4182cd883cba5040376f7f888cc6ba248b1c2f4cea07cdb34bb1318`) | the VOICEVOX project (model file); voices inside it belong to their character licensors | https://github.com/VOICEVOX/voicevox_vvm/releases/tag/0.17.0 | VOICEVOX 音声モデル 利用規約 — https://github.com/VOICEVOX/voicevox_vvm/blob/main/README.md (shipped as `README.txt` / `TERMS.txt` with the release) | Clause 2: 「アプリケーションに組み込んで再配布することができます」 — embedded redistribution is explicitly permitted; clause 4 flow-down; a credit that shows VOICEVOX was used is required | `VOICEVOX:雨晴はう` (only 雨晴はう, style 10, is used; the other voices packaged in `0.vvm` — 四国めたん, ずんだもん, 春日部つむぎ — are neither used nor credited) | Committed through Git LFS under clause 2; credit and flow-down notice rendered; the model is used only through VOICEVOX CORE | | **雨晴はう** (character voice) | — (voice library shipped inside `0.vvm`) | Amehare Project (©2021-2026 Amehare Project, あめはれくりにっく) | https://amehau.com/ | https://amehau.com/?page_id=225 (雨晴はう 利用規約; returns HTTP 403 to non-browser user agents — read with a Chrome UA, HTTP 200, 13,954 bytes, 2026-09-12); restated by VOICEVOX at https://github.com/VOICEVOX/voicevox_resource/blob/main/character_info/雨晴はう_3474ee95-c274-47f9-aa1a-8322163d96f1/policy.md | 「個人様/企業様ともに商用・非商用問わず利用可能です。ただし音声のR18指定利用はご遠慮ください。」 and 「VOICEVOXのクレジット記載は必須となります。」; the character-name credit is optional: 「キャラクター名のクレジット表記は必須ではありませんが、記載があれば嬉しいです。(例:VOICEVOX:雨晴はう)」 | **`VOICEVOX:雨晴はう`** — the terms' own example, exact, ASCII colon, no spaces; satisfies the mandatory VOICEVOX credit and the welcomed character credit in one string | Persistent footer credit + About panel on first paint (server-delivered, not injected); flow-down notice linking both terms URLs; the audio is generated directly by VOICEVOX CORE, so the voice-changer prohibition does not apply; no prohibited-use path exists (see below) | | **Open JTalk dictionary** `voicevox/open_jtalk_dic_utf_8-1.11/` | `open_jtalk_dic_utf_8-1.11` | © 2009 **Nara Institute of Science and Technology** (NAIST), Japan | https://downloads.sourceforge.net/project/open-jtalk/Dictionary/open_jtalk_dic-1.11/open_jtalk_dic_utf_8-1.11.tar.gz | BSD-3-Clause — `voicevox/open_jtalk_dic_utf_8-1.11/COPYING` (reproduced below) | BSD-3-Clause redistribution in binary form, with the notice reproduced | none (attribution notice) | `COPYING` is committed alongside the dictionary files; the notice is reproduced in this file; the About panel names NAIST. Open JTalk itself (Nagoya Institute of Technology / HTS Working Group, modified BSD) is linked inside VOICEVOX CORE and is attributed too | | **three.js** (incl. its `GLTFLoader` addon) | 0.185.1 | three.js authors (mrdoob et al.) | https://github.com/mrdoob/three.js | MIT — https://github.com/mrdoob/three.js/blob/master/LICENSE | MIT: redistribution permitted with the copyright and permission notice kept | none | **Vendored into this repository** (plan 01-10) as `avatar/vendor/three.mjs` and `avatar/vendor/GLTFLoader.mjs`, the package's `LICENSE` copied alongside as `avatar/vendor/LICENSE-three.txt`; served by the Space itself, no CDN in the render path. Provenance, hashes and the regeneration command: `avatar/vendor/README.md` | | **@pixiv/three-vrm** | 3.5.5 | pixiv Inc. | https://github.com/pixiv/three-vrm | MIT — https://github.com/pixiv/three-vrm/blob/dev/LICENSE | MIT: redistribution permitted with the copyright and permission notice kept | none | **Vendored into this repository** (plan 01-10) as `avatar/vendor/three-vrm.mjs` (the build against three@0.185.1), the package's `LICENSE` copied alongside as `avatar/vendor/LICENSE-three-vrm.txt`; served by the Space itself | | **@huggingface/transformers** (transformers.js) | 4.2.0 | Hugging Face | https://github.com/huggingface/transformers.js | Apache-2.0 — https://github.com/huggingface/transformers.js/blob/main/LICENSE | Apache-2.0 use in the visitor's browser | none | **Not vendored, deliberately** (plan 01-10): still loaded by the visitor's browser from `https://esm.sh/@huggingface/transformers@4.2.0` on the first accepted push-to-talk, because the library fetches its own ONNX Runtime WASM assets and the Whisper model from the Hugging Face CDN regardless, so vendoring the shim alone would remove no dependency (`avatar/vendor/README.md`). Nothing of it is redistributed by this repository, so no NOTICE reproduction is triggered | | **Whisper ASR model** `onnx-community/whisper-base` (q4 ONNX) | `onnx-community/whisper-base`, converted from `openai/whisper-base` | OpenAI (model); onnx-community (ONNX conversion) | https://huggingface.co/onnx-community/whisper-base | Weights: Apache-2.0 per the `openai/whisper-base` model card (`license: apache-2.0`; the conversion card declares `base_model: openai/whisper-base` and no licence of its own). Source code of `openai/whisper`: MIT | Apache-2.0 use of the weights in the visitor's browser | none (attribution) | **Not in this repository**: the visitor's browser downloads it from huggingface.co on the first accepted push (135.8 MB); attributed here and in the About panel ("Whisper"). The "better accuracy" model `whisper-large-v3-turbo` is not offered yet | | **JMdict** — the compact projection `data/jmdict/jmdict-compact.json.gz` | jmdict-simplified release `3.6.2+20260831182826` (JSON `version` `3.6.2`, `dictDate` `2026-08-31`), 218,672 entries | **James William Breen and The Electronic Dictionary Research and Development Group** — copyright assigned to the Group in March 2000 and held by both, per the licence statement itself | https://www.edrdg.org/wiki/index.php/JMdict-EDICT_Dictionary_Project | https://www.edrdg.org/edrdg/licence.html — Creative Commons Attribution-ShareAlike Licence (V4.0) | Share and Remix. §3 permits building on the files provided the result carries the same or a compatible licence and the acknowledgement conditions below are met; the statement puts **no restriction on commercial use** once they are | `Dictionary: JMdict (EDRDG, CC BY-SA 4.0)` — plus the full acknowledgement sentence quoted in the prose section below | The compact file is committed through Git LFS as a **CC BY-SA 4.0 derivative** (`data/jmdict/README.md` says so). The credit is in the persistent footer (the licence requires the acknowledgement "on each screen display" for a WWW server displaying words from the files) **and** on the About panel (the licence's separate requirement for app-style "About" screens), each with links to the project page and the licence page | | **jmdict-simplified** (the JSON packaging JMdict is taken from) | release `3.6.2+20260831182826`, published 2026-08-31 | **Dmitry Shpika** (`scriptin`) | https://github.com/scriptin/jmdict-simplified | https://github.com/scriptin/jmdict-simplified/blob/master/LICENSE.txt — CC BY-SA 4.0 (GitHub API `license.spdx_id` = `CC-BY-SA-4.0`). Its README is explicit that the JMdict-derived files stay under the EDRDG licence: "All derived files are distributed under the same license, as the original license requires it" | CC BY-SA 4.0 for the project's own files; the EDRDG licence for the dictionary data inside them | none of its own | **Build-time input only** — `scripts/build_jmdict.py` downloads the pinned release tarball (SHA-256 pinned in `data/jmdict/README.md`) and projects it; nothing of the packaging is redistributed. Credited by name in `data/jmdict/README.md` and in the About panel | | **JLPT vocabulary lists** `data/jlpt/n1.csv` … `n5.csv` | stephenmk/yomitan-jlpt-vocab release tag `2025.08.01.0`, published 2025-08-01; 8,293 rows, 7,748 unique JMdict ids | **Stephen Kraus** (`stephenmk`, the repository and the JMdict-id mapping) over **Jonathan Waller**'s JLPT Resources data | https://github.com/stephenmk/yomitan-jlpt-vocab and https://www.tanos.co.uk/jlpt/ | Repository: CC BY-SA 4.0 (`LICENSE.txt`, GitHub API `CC-BY-SA-4.0`). Waller's data: https://www.tanos.co.uk/jlpt/sharing/ — Creative Commons "BY", "use anything here however you like (commercial or non-commercial), but credit my site" | Redistribution of the CSVs byte-for-byte under CC BY-SA 4.0, with credit to Waller as the CC BY upstream | `JLPT levels: Jonathan Waller's JLPT Resources via stephenmk/yomitan-jlpt-vocab` | Both licence texts sit beside the data: `data/jlpt/LICENSE-yomitan-jlpt-vocab.txt` is the repository's own `LICENSE.txt` at the tag. Credit in the footer and the About panel, with links to `tanos.co.uk/jlpt/` and the GitHub repository. Any redistributed derivative of the CSVs stays CC BY-SA | | **Kanji JLPT levels** `data/jlpt/kanji_levels.json` | projection of the `jlpt_new` field of `kanji.json` at davidluzgouveia/kanji-data commit `00fd7079c3890f430759536f91aa5e854ec0ca4f`; 2,211 kanji | **David Gouveia** (`davidluzgouveia`) — MIT — over the same Jonathan Waller lists (the repo's README names "Jonathan Waller's JLPT Resources page" as the source of its JLPT field) | https://github.com/davidluzgouveia/kanji-data | MIT — the repository's `LICENSE`, "Copyright (c) 2019 David Gouveia" (GitHub API `MIT`) | MIT redistribution of a derived file, with the notice kept | none (attribution) | `data/jlpt/LICENSE-kanji-data.txt` is the repository's `LICENSE` at that commit, committed beside the extracted file. **Only `jlpt_new` is extracted** — the Waller-derived field — so no KANJIDIC (EDRDG) field of `kanji.json` is redistributed here. Waller is credited by the same footer and About line as the vocabulary lists | | **OPUS-MT ja→en** `data/mt/opus-mt-ja-en-ct2-int8/` | `Helsinki-NLP/opus-mt-ja-en`, Hub revision `0770961a39ba6bd66305b149c3f4110bcafca2e6` (OPUS-MT release `opus-2019-12-18`); converted to CTranslate2 int8, `model.bin` 77,339,435 B | **Helsinki-NLP Research Group, University of Helsinki** (the HF organisation's own `fullname`; the committed `NOTICE` names it "Language Technology Research Group at the University of Helsinki") | https://huggingface.co/Helsinki-NLP/opus-mt-ja-en | Apache-2.0 — the model card at that exact revision carries `license: apache-2.0` in its front-matter; the Hub API reports `cardData.license = apache-2.0` and the tag `license:apache-2.0` | Apache-2.0 §4: redistribution of a **modified** work (the int8 conversion) with the licence text, the retained notices and a statement of the changes | `Translation: OPUS-MT (Helsinki-NLP, Apache-2.0)` | Weights and both `.spm` files committed through Git LFS; `data/mt/LICENSE-apache-2.0.txt` is the full Apache-2.0 text and `data/mt/NOTICE` names the source model, the revision and the conversion (the §4(b) change statement). Credit in the footer and the About panel | | **CTranslate2** | `ctranslate2==4.8.2` | OpenNMT | https://github.com/OpenNMT/CTranslate2 | MIT (GitHub API `license.spdx_id` = `MIT`; PyPI `info.license` = `MIT`) | MIT use of an installed library | none | **Installed from PyPI, never redistributed** — pinned in `pyproject.toml` / `requirements.txt`. The About panel names it as the translation runtime | | **SentencePiece** | `sentencepiece==0.2.2` | Google | https://github.com/google/sentencepiece | Apache-2.0 (GitHub API `license.spdx_id` = `Apache-2.0`; PyPI carries it as the PEP 639 `license_expression` `Apache-2.0`, with the older `info.license` field empty) | Apache-2.0 use of an installed library | none | **Installed from PyPI, never redistributed.** The two `.spm` model files under `data/mt/` are OPUS-MT's, not SentencePiece's, and are covered by the OPUS-MT row. The About panel names it | | **SudachiPy + SudachiDict-core** | `SudachiPy 0.6.11`, `SudachiDict-core 20260723` (`system.dic`, 217,466,039 B, memory-mapped from the installed wheel) | **Works Applications Co., Ltd.** — `SudachiDict` README: "SudachiDict by Works Applications Co., Ltd. is licensed under the Apache License, Version 2.0", Copyright (c) 2017-2023 | https://github.com/WorksApplications/sudachi.rs and https://github.com/WorksApplications/SudachiDict | Apache-2.0 for both. sudachi.rs: GitHub API `Apache-2.0`. SudachiDict: GitHub API reports **no** licence because the file is named `LICENSE-2.0.txt` rather than `LICENSE`; the file itself is the Apache-2.0 text and the repo's `LEGAL` states "All the files in this distribution are covered under the Apache License version 2.0" | Apache-2.0 use of installed wheels | none | **Installed from PyPI, never redistributed** — the 217 MB `system.dic` arrives with the wheel on the Space build and is not committed. Both PyPI records declare `Apache-2.0`. The dictionary bundles UniDic (© 2011-2013 The UniDic Consortium, BSD-3-Clause) and part of NEologd, as its `LEGAL` records; neither is redistributed by this repository. The About panel names SudachiPy and SudachiDict-core | | **jaconv** | `jaconv==0.5.0` | ikegami-yukino | https://github.com/ikegami-yukino/jaconv | MIT (GitHub API `license.spdx_id` = `MIT`; PyPI `info.license` = `MIT License`) | MIT use of an installed library | none | **Installed from PyPI, never redistributed.** Used for kana normalisation in `nlp/units.py`. The About panel names it | Ten rows for the eight asset classes 01-VALIDATION.md names: the ONNX Runtime is listed separately from VOICEVOX CORE because it is fetched by a different mechanism and under a different licence, and the Whisper weights are listed because the visitor's browser downloads them even though this repository never ships them. The nine rows after that are Phase 2's language assets (plan 02-11), in the order the prose section below discusses them: the dictionary and its packaging, the two JLPT level sources, the translation model, and the four PyPI runtimes that are installed but never redistributed. Every one of them was verified against its own primary source on 2026-09-07 — see § Verification record for what was read and what it said. ## Machine-readable blocks The deployed suite parses these. Keep the keys and the values exactly as the shipped assets and the rendered page have them. The VRM's embedded `VRMC_vrm.meta`, as read out of `avatar/assets/tutor.vrm` itself (VRM 1.0 key names; `thumbnailImage` is an image index and is omitted): ```vrm_meta name: VRM1_Constraint_Twist_Sample version: v1.0.1 authors: pixiv Inc. copyrightInformation: (c) 2022 pixiv Inc. licenseUrl: https://vrm.dev/licenses/1.0/ avatarPermission: everyone commercialUsage: corporation creditNotation: unnecessary allowRedistribution: true modification: allowModificationRedistribution allowAntisocialOrHateUsage: false allowExcessivelySexualUsage: true allowExcessivelyViolentUsage: true allowPoliticalOrReligiousUsage: true ``` The credit strings the page must render (`tests/e2e/test_avatar_loop.py::test_credits_visible` checks the footer, the About panel and the server-delivered config for them): ```credits voice: VOICEVOX:雨晴はう avatar: VRM1_Constraint_Twist_Sample (c) 2022 pixiv Inc. - VRM Public License 1.0 dictionary: Dictionary: JMdict (EDRDG, CC BY-SA 4.0) jlpt: JLPT levels: Jonathan Waller's JLPT Resources via stephenmk/yomitan-jlpt-vocab translation: Translation: OPUS-MT (Helsinki-NLP, Apache-2.0) ``` The three Phase 2 keys are added by plan 02-11; the first two are Phase 1's, byte for byte. `test_credits_visible` now loops over **every** key of this block and asserts each value in the footer, in the About panel and in the server-delivered `GET /config`, so adding a key here without adding it to `src/japanese_avatar/ui/blocks.py` fails the deployed suite. ## VRM character — `avatar/assets/tutor.vrm` > **FINAL.** The sourcing gate armed in plan 01-01 was re-raised at plan 01-10's phase-exit > checkpoint (2026-09-06) and the owner kept this file as the shipped character: the embedded > meta permits redistribution, modification and avatar use with no credit owed, and the page > credits pixiv Inc. regardless. Decision and rationale: `docs/ASSETS.md` § Sourcing decision. VRM Public License 1.0 is a **template, not a fixed grant**: redistribution, avatar use, modification and commercial use are per-file flags set by the licensor inside the model's own `VRMC_vrm.meta`. Two files citing the same licence URL can grant materially different rights. The facts above were therefore parsed out of this file's glTF JSON chunk (10,776,032 bytes, SHA-256 `12c2b97e95e700783a6a550dc0eee2d7880aeedccef9ae67bc4c5a2f0f2631a2`), never off a README. Permissions relied on, and what each one covers here: | Flag | Value | What it licenses in this project | |---|---|---| | `allowRedistribution` | `true` | Committing the file to this public repository and serving it from the Space | | `avatarPermission` | `everyone` | Any visitor uses it as the tutor avatar | | `modification` | `allowModificationRedistribution` | The rest pose applied at mount, and any future retargeting, still redistributable | | `commercialUsage` | `corporation` | The broadest tier; a public portfolio Space is covered unambiguously | | `creditNotation` | `unnecessary` | No credit is required. One is rendered anyway | | `allowAntisocialOrHateUsage` | **`false`** | The one restriction. A Japanese-language tutor has no such use path; stated so it is not silently dropped | ## VOICEVOX — three licence layers, read separately The voice output involves three rights holders, and their terms differ in one place that looks like a contradiction until both are read. **1. Software** — VOICEVOX ソフトウェア利用規約, https://voicevox.hiroshiba.jp/term/. Commercial and non-commercial use permitted. Clause 3 attaches a **flow-down obligation**: when audio generated here is made available to others, those others must be bound to the same terms. 禁止事項 forbids redistributing the software in whole or in part without authorisation, so the `voicevox_core` wheel and the ONNX Runtime binary are always **referenced** from their official release URLs and never copied into this repository. Separately from the terms of use, the VOICEVOX CORE **code and build artefacts** are MIT-licensed as of 0.17.0 (`LICENSE`, © 2021 Hiroshiba Kazuyuki; README § ライセンス). *Correction to earlier project notes:* prebuilt cores below version 0.16 were under a different (LGPL v3 / commercial dual) licence, and that is the version some of this project's planning text still describes; the shipped 0.17.0 is MIT. **2. Voice model** — VOICEVOX 音声モデル 利用規約, in the `voicevox_vvm` README and shipped as `README.txt` / `TERMS.txt` with the release. Clause 2 (quoted verbatim in the table above) states that the model may be embedded in an application and redistributed — **embedded redistribution is explicitly permitted**, which is why `voicevox/model/zundamon.vvm` is committed (through Git LFS) while the software is not. That asymmetry is deliberate on VOICEVOX's side and on ours: the software terms forbid redistributing the *software*; the model terms permit redistributing the *model* inside an application. Clause 4 carries the same flow-down obligation, and a credit showing that VOICEVOX was used is required. **3. Character** — 雨晴はう, Amehare Project (©2021-2026 Amehare Project, あめはれくりにっく), https://amehau.com/?page_id=225 (「雨晴はう 利用規約」). The page returns **HTTP 403 to non-browser user agents**; it was read with a Chrome UA (HTTP 200, 13,954 bytes) on 2026-09-12 — a verifier whose script gets a 403 has hit that wall, not a takedown. VOICEVOX's own restatement, shipped in `voicevox_resource` as `character_info/雨晴はう_3474ee95-c274-47f9-aa1a-8322163d96f1/policy.md`, reads verbatim: 「雨晴はうの音声ライブラリを用いて生成した音声は、「VOICEVOX:雨晴はう」とクレジットを記載すれば、商用・非商用で利用可能です。」 The character's own terms, section 音声(VOICEVOX_雨晴はう)の利用, verbatim and in order: - 「VOICEVOXの利用規約を守ってご使用下さい。」 - 「個人様/企業様ともに商用・非商用問わず利用可能です。ただし音声のR18指定利用はご遠慮ください。」 - 「キャラクター名のクレジット表記は必須ではありませんが、記載があれば嬉しいです。(例:VOICEVOX:雨晴はう)」 - 「VOICEVOXのクレジット記載は必須となります。仕様上で記載が出来ない場合はお問合せください。」 - 「VOICEVOXで直接生成していない音声及びVOICEVOX出力音声を利用して作成したボイスチェンジャーモデルの利用を禁止。」 The terms list approved exceptions to this line (WEB版VOICEVOX / 文章(コメント)読み上げ / ゆかりねっと等). This project's audio is generated directly by VOICEVOX CORE from the `.vvm`, and no voice-changer model is trained or used, so the prohibition does not apply here. So the VOICEVOX credit is **mandatory** and the character-name credit is optional-but-welcomed; the string `VOICEVOX:雨晴はう` — the terms' own example — satisfies both at once, and Phase 1's placement standard (a persistent footer beside the avatar **and** the About panel, open on first load) is kept unchanged. 禁止事項 (general), verbatim, recorded so later phases can check against them: 法令に違反する行為又は犯罪行為に関連する行為 / 第三者に不利益、損害、不快感を与える行為 / 第三者に対する詐欺または脅迫行為・公序良俗に著しく反する行為 / 特定の思想運動の勧誘・社会問題の特定の主張に当たる行為団体 / 許可無しに公式イラストを利用した二次配布、商品の販売行為 / キャラクターの著作情報を偽る行為、自作発言 / 故意に医療知識の誤った主張をする及び誤解を生む行為. お願い: 倫理的配慮が欠如している作品の公開はお控えください。 The clause Phase 3 must be checked against is 「故意に医療知識の誤った主張をする及び誤解を生む行為」: the character is a nurse, and once an LLM produces the avatar's lines, a tutor that is sometimes wrong is not in scope of an *intentional* false medical claim; a product designed to mislead would be — the same reading the ずんだもん "intentional falsehood" clause was given. **There is no ¥400,000 credit-omission clause for this character.** The ずんだもん terms this section described until 2026-09-12 carried one (per-character licensing at ¥400,000 + tax without the credit), which is why that explanation stood here. 雨晴はう's terms contain no such fee. The credit is still rendered on first paint rather than behind a click, because the VOICEVOX credit is mandatory (「VOICEVOXのクレジット記載は必須となります。仕様上で記載が出来ない場合はお問合せください。」) and because consuming the `AudioQuery` requires it on its own (next paragraph). There is no escalation path and none should be sought: the 雨晴はう terms carry a disclaimer of liability, state that they may change at any time, and state that questions about other parties' terms are not answered. The posture is unchanged: read the terms (with a browser UA — see the 403 caveat above), comply visibly, document here. The credit obligation survives an engine swap: the official VOICEVOX Q&A answers that using the intermediate `AudioQuery` with another synthesiser still requires the credit (「必要です。」). The viseme timeline in this project is built from the `AudioQuery`, so the credit is owed on that basis alone. **Pre-cleared swap-in: VOICEVOX Nemo.** If the character terms ever become inconvenient, `n0.vvm` from the same `voicevox_vvm` 0.17.0 release (nine voices, style IDs 10000–10008) is a two-line change. Credit is just `VOICEVOX Nemo`; the terms live entirely in the `voicevox_vvm` README (https://voicevox.hiroshiba.jp/nemo/term/); there is no third-party rights holder. 雨晴はう has no ¥400,000 credit-omission clause either, so what Nemo removes is the third-party rights holder, not a fee. The cost is the loss of a recognisable character. ## Open JTalk dictionary — `voicevox/open_jtalk_dic_utf_8-1.11/` BSD-3-Clause. The shipped `COPYING` reads, verbatim: > Copyright (c) 2009, Nara Institute of Science and Technology, Japan. > All rights reserved. > > Redistribution and use in source and binary forms, with or without modification, are permitted > provided that the following conditions are met: Redistributions of source code must retain the > above copyright notice, this list of conditions and the following disclaimer. Redistributions in > binary form must reproduce the above copyright notice, this list of conditions and the following > disclaimer in the documentation and/or other materials provided with the distribution. Neither > the name of the Nara Institute of Science and Technology (NAIST) nor the names of its > contributors may be used to endorse or promote products derived from this software without > specific prior written permission. > > THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS" AND ANY EXPRESS OR > IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND > FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR > CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR > CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR > SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY > THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR > OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE > POSSIBILITY OF SUCH DAMAGE. Two institutions are genuinely involved, and earlier research conflated them: - The **dictionary data** committed here (`open_jtalk_dic_utf_8-1.11`, NAIST Japanese Dictionary / IPAdic lineage) is © 2009 **Nara Institute of Science and Technology** — the notice above is the one that must be reproduced for the files actually redistributed in this repository. - **Open JTalk itself** — the engine whose dictionary format this is, linked inside `voicevox_core` — is by the **Nagoya Institute of Technology** and the **HTS Working Group**, under a modified BSD licence. It is not redistributed by this repository; it is attributed because the About panel names the engine. ## Phase 2 language assets Three of Phase 2's assets are **redistributed** by this repository and therefore carry real obligations; the rest are ordinary PyPI installs. Read in that order. ### JMdict — the acknowledgement is not optional, and it is owed twice The EDRDG licence statement (read in full at https://www.edrdg.org/edrdg/licence.html on 2026-09-07) makes JMdict available under a **Creative Commons Attribution-ShareAlike Licence (V4.0)** and then spells out what attribution means for something shaped like this Space. Its conditions, in its own terms: - Acknowledge **the usage and the source of the files** in the documentation, publicity material and WWW site of the package or server. `LICENSES.md`, `data/jmdict/README.md`, `README.md` and the About panel each do this. - Provide **links** to the licence and to the documentation at the EDRDG site. Both links ride on the footer credit and in the About paragraph. - «If a WWW server is providing a dictionary function or an on-screen display of words from the files, the acknowledgement must be made on **each screen display**, e.g. in the form of a message at the foot of the screen or page.» The word-lookup popover (plan 02-08) is exactly a dictionary function, so the credit is a **persistent footer line**, not a link behind a click. - For app-style surfaces, the acknowledgement must additionally be reachable from a menu on a screen such as "About" — and it is **not sufficient** to put it only on a launch page. Hence the About panel carries it as well as the footer. The two placements are cumulative, not alternatives, which is why both are asserted by `test_credits_visible`. - **Copyright is held by James William BREEN and The Electronic Dictionary Research and Development Group**, and users of the material must **not** claim copyright over it. This repository claims copyright over its own code (MIT) and over nothing in `data/jmdict/`. - Adding material to the files does not diminish the Group's copyright, and — provided the above is met — the statement places **no restriction on commercial use**. The sentence carried verbatim in the About panel, in `README.md` and in `data/jmdict/README.md`: > Dictionary data: **JMdict** — Copyright (c) James William Breen and The Electronic Dictionary > Research and Development Group, used under the Creative Commons Attribution-ShareAlike Licence > (V4.0). https://www.edrdg.org/wiki/index.php/JMdict-EDICT_Dictionary_Project · > https://www.edrdg.org/edrdg/licence.html `data/jmdict/jmdict-compact.json.gz` is a projection of jmdict-simplified's `jmdict-eng` JSON — per entry, up to three kanji forms, three kana forms and three senses of up to three English glosses each. That is a **derivative of JMdict**, so ShareAlike applies to it and it is itself CC BY-SA 4.0; `data/jmdict/README.md` states this. jmdict-simplified is the packaging (its own files are CC BY-SA 4.0, and its README says its JMdict-derived files stay under the EDRDG licence); it is a build-time input, credited by name and never redistributed. ### The JLPT levels are estimates, and the document says so **There is no official JLPT vocabulary list.** The upstream repository says it plainly — «Official vocabulary lists do not exist for the JLPT, so these lists are essentially an educated guess» — and its author records having met an N1-listed word on the N2 exam. Every level this project shows a learner is therefore a reconstruction of Jonathan Waller's, not a fact from the JLPT organisers, and the UI says `N1+ / beyond lists` rather than inventing a level for a word that is on no list (D-10). This is a property of the data, not a defect in the analyser. Two licences stack on that data. Waller's own terms (https://www.tanos.co.uk/jlpt/sharing/) are **Creative Commons "BY"**: "use anything here however you like (commercial or non-commercial), but credit my site. (A link would be nice.)" — so the credit line names him and the About panel links `tanos.co.uk/jlpt/`. stephenmk/yomitan-jlpt-vocab, which added the JMdict entry id to every row and is what makes the level join an id join rather than a string match, is **CC BY-SA 4.0**; its `LICENSE.txt` at the pinned tag is committed as `data/jlpt/LICENSE-yomitan-jlpt-vocab.txt` and any redistributed derivative of the CSVs stays CC BY-SA. The CSVs are shipped byte-for-byte. The kanji axis comes from davidluzgouveia/kanji-data (**MIT**, © 2019 David Gouveia), whose `LICENSE` is committed as `data/jlpt/LICENSE-kanji-data.txt`. Only its `jlpt_new` field is extracted, and that field's own upstream is the same Waller lists — so the kanji levels are estimates for the same reason, and no KANJIDIC (EDRDG) field of that repository is redistributed here. ### The translation model is a modified Apache-2.0 work `Helsinki-NLP/opus-mt-ja-en` is Apache-2.0 (the model card at the pinned revision declares `license: apache-2.0`). What is committed here is **not** that model: it is a CTranslate2 int8 conversion produced by `scripts/convert_mt.py`. Apache-2.0 §4 governs redistribution of a modified work, so `data/mt/` carries the full licence text (`LICENSE-apache-2.0.txt`) and a `NOTICE` that names the source model, the exact Hub revision and the fact of the conversion — the §4(b) "prominent notices stating that You changed the files". The two SentencePiece model files are the upstream repository's own, copied unmodified, and are covered by the same licence. CTranslate2 (MIT) and SentencePiece (Apache-2.0) are the runtime; both are installed from PyPI and neither is redistributed, so no NOTICE reproduction is triggered for them. The same is true of SudachiPy, SudachiDict-core and jaconv: the 217 MB `system.dic` is downloaded by pip on the Space build and is not in this repository. ## Runtime libraries loaded into the visitor's browser | Library | Version | Licence | Copyright | Loaded from | |---|---|---|---|---| | three.js | 0.185.1 | MIT | © 2010-2026 three.js authors | `avatar/vendor/three.mjs` + `avatar/vendor/GLTFLoader.mjs`, served by the Space (`/gradio_api/file=avatar/vendor/…`); licence text `avatar/vendor/LICENSE-three.txt` | | @pixiv/three-vrm | 3.5.5 | MIT | © 2019-2026 pixiv Inc. | `avatar/vendor/three-vrm.mjs`, served by the Space; licence text `avatar/vendor/LICENSE-three-vrm.txt` | | @huggingface/transformers | 4.2.0 | Apache-2.0 | © Hugging Face | `https://esm.sh/@huggingface/transformers@4.2.0` — the one remaining CDN load, on the push-to-talk path only, deliberately (see the table above and `avatar/vendor/README.md`) | The two rendering libraries were **vendored by plan 01-10**: `scripts/vendor_modules.py` downloads the exact esm.sh builds the browser executed before (three@0.185.1's es2022 build, its GLTFLoader, and @pixiv/three-vrm@3.5.5 built against that three), rewrites their imports so all three share one `./three.mjs`, asserts no absolute or root-relative import survives, and copies each package's `LICENSE` beside the modules — MIT permits exactly that redistribution provided the notice is kept. `avatar/vendor/README.md` records every file's SHA256, byte size and source URL; `tests/test_vendor.py` fails if a module is hand-edited or a CDN URL returns to the render path. transformers.js stays on the CDN because its own runtime assets and the Whisper model come from the Hugging Face CDN whatever the shim's origin; the render path is what must survive a CDN outage, and it now does. The ASR model (`onnx-community/whisper-base`, q4) is downloaded by the browser from huggingface.co on the first accepted push-to-talk utterance and cached by the browser; it is never served from this repository or the Space. `docs/ASR-TIERS.md` records why this model and this quantisation. ## Not used, and why - **`AvatarSample_A`, `AvatarSample_B` and `AvatarSample_C` are NOT CC0.** Copyright is retained and VRoid Hub's conditions apply; VRoid's own FAQ prohibits redistributing them under a CC0 designation. The claim is repeated confidently across tutorials and blog posts, and it is false. Recorded here so it is not reintroduced into this repository. - CC0 **cannot be set** on VRoid Hub at all, so any "CC0 on VRoid Hub" claim is false by construction; and enabling "Allow third-party usage" on VRoid Hub **overwrites** a model's embedded licence with the Hub's own conditions, so a Hub-sourced file's embedded flags may not say what its author intended. Both are reasons the VRM row above reads the flags out of the shipped file. - **Hosted ASR on accelerated hardware (tier D)** is deliberately not built: it would burn the visitor's ZeroGPU quota, contradicting SC-4. `tests/test_transport_seam.py` fails if one appears. - **Style-Bert-VITS2** (AGPL-3.0), **pykakasi** (GPL-3.0) and **edge-tts** (undocumented endpoint) were rejected at stack selection for licence reasons and do not appear anywhere in the tree. Python dependencies that are not assets (`gradio`, `spaces`, `huggingface_hub`, `numpy`) are ordinary Apache-2.0 / BSD packages installed from PyPI and are not redistributed. ## Completeness review (plan 01-10 Task 3, 2026-09-06) The eight-asset checklist from `01-VALIDATION.md`, re-read against the sources in this tree by the execution agent on the owner's instruction (the owner asked for the manual rows to be done on their behalf; the legal judgement is therefore the agent's reading, recorded so the owner can overrule it). **Result: no gaps.** | # | Asset | Row present | Checked against | |---|---|---|---| | 1 | VRM character | yes | `docs/ASSETS.md` (same source, licence URL and embedded flags); `tests/e2e/test_avatar_loop.py::test_vrm_meta_matches_licenses` green on the deployed file. Sourcing gate closed: the file is final | | 2 | VOICEVOX software | yes | `docs/VOICEVOX-SETUP.md` (terms URL `voicevox.hiroshiba.jp/term/`, MIT at tag 0.17.0); `requirements.txt` installs exactly the 0.17.0 release wheel from the official URL | | 3 | VOICEVOX voice model (`.vvm`) | yes | `voicevox_vvm` 0.17.0 release in `docs/VOICEVOX-SETUP.md`; `sha256sum voicevox/model/zundamon.vvm` = `ecd35374…1318`, the value in the row; the file is an LFS object | | 4 | 雨晴はう character terms | yes | `docs/VOICEVOX-SETUP.md` (`amehau.com/?page_id=225`, browser UA required — 403 otherwise); the credit string in the `credits` block is what `test_credits_visible` asserts; NOT yet re-asserted on the Space (still on 7eb3e94 with ずんだもん until the owner pushes) | | 5 | Open JTalk dictionary | yes | `voicevox/open_jtalk_dic_utf_8-1.11/COPYING` first line `Copyright (c) 2009, Nara Institute of Science and Technology, Japan.`; the Nagoya Institute of Technology / HTS Working Group attribution for Open JTalk itself, as `docs/VOICEVOX-SETUP.md` records; the About panel names both institutions | | 6 | three.js | yes | `avatar/vendor/README.md` (three@0.185.1 es2022 build, SHA256 recorded, `tests/test_vendor.py` re-checks); `avatar/vendor/LICENSE-three.txt` is the MIT notice © 2010-2026 three.js authors | | 7 | @pixiv/three-vrm | yes | `avatar/vendor/README.md` (3.5.5 built against three@0.185.1); `avatar/vendor/LICENSE-three-vrm.txt` is the MIT notice © 2019-2026 pixiv Inc. | | 8 | @huggingface/transformers | yes | `avatar/asr.js` loads exactly `https://esm.sh/@huggingface/transformers@4.2.0`, the version in the row; not redistributed, so no NOTICE reproduction is owed; the Whisper weights it downloads have their own row | Every cell of the summary table is filled (no `TODO`, `TBD` or empty cell); the two machine-readable blocks are the ones the deployed suite parsed green. Nothing in this document is a claim the agent is unwilling to publish. ### Phase 2 assets (plan 02-11, 2026-09-07) The nine assets plan 02-11 is responsible for, each checked against **this tree** rather than against the plan that named it. **Result: no gaps.** | # | Asset | Row present | Checked against the tree | |---|---|---|---| | 1 | JMdict compact projection | yes | `data/jmdict/jmdict-compact.json.gz` present, 7,662,785 B, `git check-attr filter` = `lfs`, first two bytes `1f 8b` (gzip, not an LFS pointer). `data/jmdict/README.md` records the release, the `dictDate`, the input tarball's SHA-256 and states the projection is itself CC BY-SA 4.0. `tests/test_data_assets.py` pins its SHA-256 and entry count | | 2 | jmdict-simplified (packaging) | yes | Build-time only: `scripts/build_jmdict.py` holds the pinned release URL; no packaging file is in the tree. Credited in `data/jmdict/README.md` and the About panel | | 3 | JLPT vocabulary lists | yes | `data/jlpt/n1.csv`…`n5.csv` present (199,103 / 95,356 / 88,623 / 24,772 / 24,010 B), byte-for-byte upstream; `data/jlpt/LICENSE-yomitan-jlpt-vocab.txt` present, 20,131 B, first line `Attribution-ShareAlike 4.0 International` | | 4 | Kanji JLPT levels | yes | `data/jlpt/kanji_levels.json` present, 28,746 B; `data/jlpt/LICENSE-kanji-data.txt` present, 1,070 B, reads `MIT License` / `Copyright (c) 2019 David Gouveia` | | 5 | OPUS-MT ja→en (CT2 int8) | yes | `data/mt/opus-mt-ja-en-ct2-int8/model.bin` present, 77,339,435 B, LFS; `data/mt/LICENSE-apache-2.0.txt` present, 11,358 B (the Apache-2.0 text); `data/mt/NOTICE` present, 446 B, naming the model, the revision `0770961a…` and the conversion | | 6 | CTranslate2 | yes | `ctranslate2==4.8.2` pinned in `pyproject.toml` and `requirements.txt`; nothing of it in the tree | | 7 | SentencePiece | yes | `sentencepiece==0.2.2` pinned in both; nothing of it in the tree (the two `.spm` files are OPUS-MT's) | | 8 | SudachiPy + SudachiDict-core | yes | `SudachiPy==0.6.11`, `SudachiDict-core==20260723` pinned in both; `system.dic` (217,466,039 B) lives in the installed wheel under `.venv/`, not in the repository | | 9 | jaconv | yes | `jaconv==0.5.0` pinned in both; nothing of it in the tree | The three redistributing rows (1, 3–5) each have their licence text committed **beside the data**, which is the obligation that a `LICENSES.md` row alone would not discharge. ## Verification record | Check | Result (2026-09-06) | |---|---| | VRM `VRMC_vrm.meta` parsed from the shipped file | matches the `vrm_meta` block above, key for key | | `voicevox_core` 0.17.0 licence | README § ライセンス at tag 0.17.0: MIT; `LICENSE` © 2021 Hiroshiba Kazuyuki; GitHub API `MIT` | | `voicevox_vvm` clause 2 | present verbatim in the `main` README (line 21) | | `VOICEVOX/onnxruntime-builder`, `microsoft/onnxruntime` | GitHub API: MIT, MIT | | `three@0.185.1`, `@pixiv/three-vrm@3.5.5`, `@huggingface/transformers@4.2.0` | npm registry: MIT, MIT, Apache-2.0 (all three are also the current `latest`) | | Vendored `avatar/vendor/*.mjs` and the two `LICENSE-*.txt` (plan 01-10) | SHA256 and byte size of each recorded in `avatar/vendor/README.md` from the fetch; `tests/test_vendor.py::test_vendored_hashes_match_readme` re-checks them on every quick loop; the licence texts are the packages' own `LICENSE` files from the npm tarballs (unpkg), both MIT | | `openai/whisper-base` model card | HF API `cardData.license = apache-2.0`; `openai/whisper` repo MIT | | `onnx-community/whisper-base` card | no licence field; `base_model: openai/whisper-base` | | Open JTalk `COPYING` | first line: `Copyright (c) 2009, Nara Institute of Science and Technology, Japan.` | | Terms URLs answer | voicevox.hiroshiba.jp/term/ 200 · zunko.jp/con_ongen_kiyaku.html 200 (historical — the ずんだもん row, not re-fetched since the 2026-09-12 voice switch) · voicevox.hiroshiba.jp/nemo/term/ 200 · vrm.dev/licenses/1.0/ 200 · amehau.com/?page_id=225 403 (non-browser UA) / 200 with a Chrome UA, 13,954 bytes, 2026-09-12 | ### Phase 2 assets (plan 02-11) Nine assets, each read from **its own** primary source on 2026-09-07 — the upstream repository's licence file, the model card at the pinned revision, the dictionary project's own licence page, the package registry's metadata — never from `02-RESEARCH.md`, from an earlier SUMMARY, or from the plan's draft table. Where a primary source disagreed with the plan, the source won and the disagreement is written out below. | Asset | Primary source read | What it said (2026-09-07) | |---|---|---| | JMdict | https://www.edrdg.org/edrdg/licence.html, fetched and read in full (HTTP 200) | §3: "made available under a Creative Commons Attribution-ShareAlike Licence (V4.0)". Copyright "held by James William BREEN and The Electronic Dictionary Research and Development Group". Attribution conditions read verbatim: acknowledge usage and source in documentation/publicity/WWW site; provide links to the licence and documentation; **"the acknowledgement must be made on each screen display"** for a WWW server displaying words from the files; for app surfaces, a separate "About"-style screen, and a launch page alone is "not sufficient"; users must NOT claim copyright; no restriction on commercial use. The project page https://www.edrdg.org/wiki/index.php/JMdict-EDICT_Dictionary_Project answered 200 | | jmdict-simplified | GitHub API `repos/scriptin/jmdict-simplified`; `LICENSE.txt` and `README.md` on `master`; release tag `3.6.2+20260831182826` | API `license.spdx_id` = **`CC-BY-SA-4.0`**. `LICENSE.txt` first line: `Attribution-ShareAlike 4.0 International`. README § License: the JMdict XML files "are the property of the Electronic Dictionary Research and Development Group … All derived files are distributed under the same license". Release tag exists, `published_at` `2026-08-31T18:28:35Z` — consistent with the `dictDate` `2026-08-31` in `data/jmdict/README.md`. Owner `scriptin` resolves to **Dmitry Shpika** | | JLPT vocabulary lists | GitHub API `repos/stephenmk/yomitan-jlpt-vocab`; the release tag `2025.08.01.0`; the README at that tag; the committed `LICENSE-yomitan-jlpt-vocab.txt` | API `license.spdx_id` = **`CC-BY-SA-4.0`**. Tag exists, `published_at` `2025-08-01T15:46:28Z`. README § Attribution: "JLPT data is sourced from Jonathan Waller's JLPT Resources page under the terms of the Creative Commons BY license", and § Limitations: **"Official vocabulary lists do not exist for the JLPT, so these lists are essentially an educated guess."** Owner `stephenmk` resolves to **Stephen Kraus**. The committed licence text's first line is `Attribution-ShareAlike 4.0 International` | | Jonathan Waller's JLPT Resources | https://www.tanos.co.uk/jlpt/sharing/ (HTTP 200), read this session | "Everything on this site (that I'm not selling), is licenced under Creative Commons 'BY'. Basically this means… use anything here however you like (commercial or non-commercial), but credit my site. (A link would be nice.)" Signed "© Jonathan Waller". So: **CC BY, credit + link** | | Kanji levels (kanji-data) | GitHub API `repos/davidluzgouveia/kanji-data`; the README at commit `00fd7079c3890f430759536f91aa5e854ec0ca4f`; the committed `LICENSE-kanji-data.txt` | API `license.spdx_id` = **`MIT`**; the committed file reads `MIT License` / `Copyright (c) 2019 David Gouveia`. README § References names KANJIDIC for the kanji data and **"Jonathan Waller's JLPT Resources page"** for the JLPT field — confirming the `jlpt_new` lineage this project relies on. Owner resolves to **David Gouveia** | | OPUS-MT ja→en | Hub API `api/models/Helsinki-NLP/opus-mt-ja-en`; the model card `README.md` fetched **at revision `0770961a39ba6bd66305b149c3f4110bcafca2e6`** | API: `sha` = `0770961a39ba6bd66305b149c3f4110bcafca2e6` — identical to the revision `data/mt/README.md` pins; `cardData.license` = **`apache-2.0`**; tag `license:apache-2.0`. The card at that revision carries `license: apache-2.0` in its front-matter and names the release `opus-2019-12-18`. HF organisation `Helsinki-NLP` has `fullname` "Helsinki-NLP Research Group", details "At the University of Helsinki" | | CTranslate2 | GitHub API `repos/OpenNMT/CTranslate2`; PyPI `pypi/ctranslate2/json`; installed wheel metadata | API `MIT`. PyPI `info.license` = `MIT`, current version `4.8.2`. Installed distribution reports `4.8.2`, `License: MIT` — the pin and the installed build agree | | SentencePiece | GitHub API `repos/google/sentencepiece`; PyPI `pypi/sentencepiece/json`; installed wheel metadata | API **`Apache-2.0`**. PyPI: the legacy `info.license` field is **empty**; the PEP 639 `info.license_expression` is **`Apache-2.0`**. The installed wheel's `License` field is likewise empty for the same reason. Recorded as read rather than smoothed over: the authoritative registry statement is the `license_expression`, and it agrees with the repository | | SudachiPy | GitHub API `repos/WorksApplications/sudachi.rs`; PyPI `pypi/SudachiPy/json`; installed wheel metadata | API **`Apache-2.0`**. PyPI `info.license` = `Apache-2.0`, current version `0.6.11`, home page the `python/` subdirectory of `sudachi.rs`. Installed distribution reports `0.6.11`, `License: Apache-2.0` | | SudachiDict-core | GitHub API `repos/WorksApplications/SudachiDict`; the repo's `LICENSE-2.0.txt`, `README.md` and `LEGAL`; PyPI `pypi/SudachiDict-core/json`; installed wheel metadata | **GitHub API reports `license: null`** — because the file is named `LICENSE-2.0.txt`, which GitHub's detector does not recognise. The file itself is the Apache License 2.0 text; `README.md` § Licenses states "SudachiDict by Works Applications Co., Ltd. is licensed under the Apache License, Version 2.0", Copyright (c) 2017-2023; `LEGAL` states "All the files in this distribution are covered under the Apache License version 2.0". PyPI `info.license` = `Apache-2.0`, current version `20260723`; installed distribution reports `20260723`, `License: Apache-2.0`. `LEGAL` also records that the dictionary contains part of **UniDic** (© 2011-2013 The UniDic Consortium, BSD-3-Clause) and part of **NEologd** — neither is redistributed here | | jaconv | GitHub API `repos/ikegami-yukino/jaconv`; PyPI `pypi/jaconv/json`; installed wheel metadata | API `MIT`. PyPI `info.license` = `MIT License`, current version `0.5.0`. Installed distribution reports `0.5.0`, `License: MIT License` | | Pinned versions vs installed | `importlib.metadata` in `.venv` | SudachiPy `0.6.11`, SudachiDict-core `20260723`, jaconv `0.5.0`, ctranslate2 `4.8.2`, sentencepiece `0.2.2` — every one identical to the pin in `pyproject.toml` / `requirements.txt` and to the version this document's rows name | | Phase 2 source URLs answer | HTTP `HEAD`/`GET` on the fourteen URLs the nine rows cite | All **200**: edrdg.org project page and licence page, creativecommons.org/licenses/by-sa/4.0/, github.com/scriptin/jmdict-simplified, github.com/stephenmk/yomitan-jlpt-vocab, tanos.co.uk/jlpt/ and /jlpt/sharing/, github.com/davidluzgouveia/kanji-data, huggingface.co/Helsinki-NLP/opus-mt-ja-en, github.com/OpenNMT/CTranslate2, github.com/google/sentencepiece, github.com/WorksApplications/sudachi.rs, github.com/WorksApplications/SudachiDict, github.com/ikegami-yukino/jaconv | **Two places where a primary source corrected the plan**, both recorded in the rows above rather than quietly reconciled: 1. The plan's row 8 asserted "Apache-2.0 (both)" for SudachiPy and SudachiDict on the strength of the GitHub API. The API returns **no licence at all** for `WorksApplications/SudachiDict`. The Apache-2.0 conclusion still holds, but it rests on the repository's `LICENSE-2.0.txt`, `LEGAL` and `README.md` — not on the API — and it comes with a bundled-data note (UniDic, NEologd) that the plan did not mention. 2. The plan's row 5 named the OPUS-MT rights holder "Language Technology Research Group, University of Helsinki". The Hub organisation's own `fullname` is "**Helsinki-NLP Research Group**" at the University of Helsinki, and the model card names no individual author. The row uses the Hub's wording and notes that the already-committed `data/mt/NOTICE` uses the longer historical name; the `NOTICE` is left untouched because it is the artefact the conversion shipped with.