Spaces:
Running on Zero
Running on Zero
File size: 2,491 Bytes
28febab | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 | ---
title: Japanese Learning Avatar
emoji: 🗾
colorFrom: pink
colorTo: indigo
sdk: gradio
sdk_version: 6.22.0
python_version: 3.12.12
app_file: app.py
pinned: false
license: mit
short_description: Talk to a lip-synced VRM avatar in Japanese
---
# Japanese Learning Avatar
An animated 3D avatar tutor that teaches Japanese through spoken conversation. A VRM
anime-style character listens to you speak Japanese and answers aloud, with lip-sync
driven by real per-mora phoneme timings rather than by microphone-style mouth flapping.
The avatar is rendered **in your browser** with three.js and `@pixiv/three-vrm`. Nothing
about the character is generated server-side, so it paints and starts breathing while the
Python backend is still waking up.
## Scope of this phase
**Phase 1 has no AI tutoring in it yet.** This phase exists to prove the transport — that a
VRM avatar can live inside a Gradio `gr.HTML` component on a real Hugging Face Space, that
speech can be synthesised with per-mora timings, and that the microphone round trip works
without spending any GPU quota. The tutoring brain, level gating and progress tracking come
in later phases.
What is live here:
- A VRM avatar with idle life — blinking, breathing, a slight sway.
- Japanese speech synthesis with mora-accurate visemes (VOICEVOX, CPU only).
- Push-to-talk capture with in-browser speech recognition (Whisper via transformers.js),
which costs **zero** ZeroGPU quota and therefore does not eat into a visitor's daily
allowance.
## Run it locally
```bash
uv sync --extra dev
uv run python app.py
```
Requires Python 3.12.12 (see `.python-version`). The Space itself is pinned to Gradio
6.22.0 and Python 3.12.12 — ZeroGPU only provides 3.10.13 and 3.12.12, so those pins are a
platform constraint, not a preference.
Tests:
```bash
uv run pytest tests/ -q --ignore=tests/e2e # quick loop
uv run pytest tests/e2e/ -q # browser suite (run it whole, never per-file)
```
## Credits
<!-- Plan 01-08 replaces this block with the rendered VOICEVOX credit strings and the
flow-down terms notice. The in-app footer (elem_id="credits") is the binding
surface for DPLY-04; this section mirrors it. -->
Credits and third-party licences are recorded in `docs/ASSETS.md` and
`docs/VOICEVOX-SETUP.md`, and are consolidated into `LICENSES.md` before the phase ships.
## Licence
MIT for this repository's own code. Third-party assets and models carry their own terms —
see above.
|