File size: 4,098 Bytes
28febab
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
eeb4e4d
 
 
 
28febab
 
eeb4e4d
 
 
 
 
28febab
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
eeb4e4d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
28febab
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
---
title: Japanese Learning Avatar
emoji: πŸ—Ύ
colorFrom: pink
colorTo: indigo
sdk: gradio
sdk_version: 6.22.0
python_version: 3.12.12
app_file: app.py
pinned: false
license: mit
short_description: Talk to a lip-synced VRM avatar in Japanese
---

# Japanese Learning Avatar

An animated 3D avatar tutor that teaches Japanese through spoken conversation. A VRM
anime-style character listens to you speak Japanese and answers aloud, with lip-sync
driven by real per-mora phoneme timings rather than by microphone-style mouth flapping.

The avatar is rendered **in your browser** with three.js and `@pixiv/three-vrm`. Nothing
about the character is generated server-side, so it paints and starts breathing while the
Python backend is still waking up.

## Scope of this phase

**Phase 1 has no AI tutoring in it yet.** This phase exists to prove the transport β€” that a
VRM avatar can live inside a Gradio `gr.HTML` component on a real Hugging Face Space, that
speech can be synthesised with per-mora timings, and that the microphone round trip works
without spending any GPU quota. The tutoring brain, level gating and progress tracking come
in later phases.

What is live here:

- A VRM avatar with idle life β€” blinking, breathing, a slight sway β€” and a visible
  thinking pose while a reply is being synthesised.
- Japanese speech synthesis with mora-accurate visemes (VOICEVOX, CPU only). Type Japanese
  and press Enter, or click "Say hello", and the avatar says it back with lip-sync.
- Push-to-talk capture with in-browser speech recognition (Whisper via transformers.js),
  which costs **zero** ZeroGPU quota and therefore does not eat into a visitor's daily
  allowance. Hold the button, speak, and the transcript is echoed back aloud.
- Replay (instant, no server round trip) and Slower (re-synthesised at 0.75x with a rebuilt
  lip-sync timeline), with per-stage latency shown under the controls.

**The avatar repeats what you say.** There is no tutor yet; that is Phase 3.

## Run it locally

```bash
uv sync --extra dev
uv run python app.py
```

Requires Python 3.12.12 (see `.python-version`). The Space itself is pinned to Gradio
6.22.0 and Python 3.12.12 β€” ZeroGPU only provides 3.10.13 and 3.12.12, so those pins are a
platform constraint, not a preference.

Tests:

```bash
uv run pytest tests/ -q --ignore=tests/e2e   # quick loop
uv run pytest tests/e2e/ -q                  # browser suite (run it whole, never per-file)
```

## Credits

The in-app footer and About panel are the binding credit surfaces; this section mirrors
them so the Space's own page carries the same strings.

- **Voice: VOICEVOX:γšγ‚“γ γ‚‚γ‚“** β€” speech is synthesised with
  [VOICEVOX CORE](https://voicevox.hiroshiba.jp/term/) (software terms) using the γšγ‚“γ γ‚‚γ‚“
  voice by SSS LLC ([character terms](https://zunko.jp/con_ongen_kiyaku.html)).
  *Synthesised audio is provided under the VOICEVOX and VOICEVOX:γšγ‚“γ γ‚‚γ‚“ terms of use; by
  using it you agree to comply with them.*
- **Avatar: VRM1_Constraint_Twist_Sample (c) 2022 pixiv Inc. β€” VRM Public License 1.0**
  ([terms](https://vrm.dev/licenses/1.0/)), from the official VRM specification samples. The
  file's embedded `VRMC_vrm.meta` grants redistribution, avatar use by everyone, modification
  and commercial use, and requires no credit; it is credited anyway.
- **Open JTalk dictionary** `open_jtalk_dic_utf_8-1.11` β€” BSD-3-Clause, (c) 2009 Nara
  Institute of Science and Technology. Open JTalk itself is by the Nagoya Institute of
  Technology and the HTS Working Group, also under a modified BSD licence.
- **Runtime libraries** β€” [three.js](https://threejs.org/) (MIT),
  [@pixiv/three-vrm](https://github.com/pixiv/three-vrm) (MIT),
  [@huggingface/transformers](https://github.com/huggingface/transformers.js) (Apache-2.0).

Provenance and the full licence facts are recorded in `docs/ASSETS.md` and
`docs/VOICEVOX-SETUP.md`, and are consolidated into `LICENSES.md` before the phase ships.

## Licence

MIT for this repository's own code. Third-party assets and models carry their own terms β€”
see above.