--- title: Japanese Learning Avatar emoji: 🗾 colorFrom: pink colorTo: indigo sdk: gradio sdk_version: 6.22.0 python_version: 3.12.12 app_file: app.py pinned: false license: mit short_description: Talk to a lip-synced VRM avatar in Japanese --- # Japanese Learning Avatar An animated 3D avatar tutor that teaches Japanese through spoken conversation. A VRM anime-style character listens to you speak Japanese and answers aloud, with lip-sync driven by real per-mora phoneme timings rather than by microphone-style mouth flapping. The avatar is rendered **in your browser** with three.js and `@pixiv/three-vrm`. Nothing about the character is generated server-side, so it paints and starts breathing while the Python backend is still waking up. ## Scope of this phase **Phase 1 has no AI tutoring in it yet.** This phase exists to prove the transport — that a VRM avatar can live inside a Gradio `gr.HTML` component on a real Hugging Face Space, that speech can be synthesised with per-mora timings, and that the microphone round trip works without spending any GPU quota. The tutoring brain, level gating and progress tracking come in later phases. What is live here: - A VRM avatar with idle life — blinking, breathing, a slight sway — and a visible thinking pose while a reply is being synthesised. - Japanese speech synthesis with mora-accurate visemes (VOICEVOX, CPU only). Type Japanese and press Enter, or click "Say hello", and the avatar says it back with lip-sync. - Push-to-talk capture with in-browser speech recognition (Whisper via transformers.js), which costs **zero** ZeroGPU quota and therefore does not eat into a visitor's daily allowance. Hold the button, speak, and the transcript is echoed back aloud. - Replay (instant, no server round trip) and Slower (re-synthesised at 0.75x with a rebuilt lip-sync timeline), with per-stage latency shown under the controls. **The avatar repeats what you say.** There is no tutor yet; that is Phase 3. ## Run it locally ```bash uv sync --extra dev uv run python app.py ``` Requires Python 3.12.12 (see `.python-version`). The Space itself is pinned to Gradio 6.22.0 and Python 3.12.12 — ZeroGPU only provides 3.10.13 and 3.12.12, so those pins are a platform constraint, not a preference. Tests: ```bash uv run pytest tests/ -q --ignore=tests/e2e # quick loop uv run pytest tests/e2e/ -q # browser suite (run it whole, never per-file) ``` ## Credits The in-app footer and About panel are the binding credit surfaces; this section mirrors them so the Space's own page carries the same strings. - **Voice: VOICEVOX:ずんだもん** — speech is synthesised with [VOICEVOX CORE](https://voicevox.hiroshiba.jp/term/) (software terms) using the ずんだもん voice by SSS LLC ([character terms](https://zunko.jp/con_ongen_kiyaku.html)). *Synthesised audio is provided under the VOICEVOX and VOICEVOX:ずんだもん terms of use; by using it you agree to comply with them.* - **Avatar: VRM1_Constraint_Twist_Sample (c) 2022 pixiv Inc. — VRM Public License 1.0** ([terms](https://vrm.dev/licenses/1.0/)), from the official VRM specification samples. The file's embedded `VRMC_vrm.meta` grants redistribution, avatar use by everyone, modification and commercial use, and requires no credit; it is credited anyway. - **Open JTalk dictionary** `open_jtalk_dic_utf_8-1.11` — BSD-3-Clause, (c) 2009 Nara Institute of Science and Technology. Open JTalk itself is by the Nagoya Institute of Technology and the HTS Working Group, also under a modified BSD licence. - **Runtime libraries** — [three.js](https://threejs.org/) (MIT), [@pixiv/three-vrm](https://github.com/pixiv/three-vrm) (MIT), [@huggingface/transformers](https://github.com/huggingface/transformers.js) (Apache-2.0). Provenance and the full licence facts are recorded in `docs/ASSETS.md` and `docs/VOICEVOX-SETUP.md`, and are consolidated into `LICENSES.md` before the phase ships. ## Licence MIT for this repository's own code. Third-party assets and models carry their own terms — see above.