Spaces:
Running on Zero
Running on Zero
docs: project research
Browse files- .planning/research/ARCHITECTURE.md +572 -0
- .planning/research/FEATURES.md +415 -0
- .planning/research/PITFALLS.md +592 -0
- .planning/research/STACK.md +369 -0
- .planning/research/SUMMARY.md +186 -0
.planning/research/ARCHITECTURE.md
ADDED
|
@@ -0,0 +1,572 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Architecture Research
|
| 2 |
+
|
| 3 |
+
**Domain:** Real-time 3D-avatar conversational language tutor on Hugging Face Spaces (Gradio backend + in-browser three.js/VRM frontend + full-duplex voice + hybrid LLM + persistent learner state)
|
| 4 |
+
**Researched:** 2026-08-08
|
| 5 |
+
**Confidence:** MEDIUM-HIGH (platform APIs verified against official Gradio/HF docs; pedagogical-layer patterns verified against 2025-2026 papers; latency numbers are estimates flagged LOW)
|
| 6 |
+
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
## Standard Architecture
|
| 10 |
+
|
| 11 |
+
Systems of this shape converge on **five layers with one hard boundary in the middle**: the browser owns rendering and animation timing; the server owns cognition and state; a single narrow message contract crosses between them. Almost every failure mode in this domain comes from letting those two sides leak into each other.
|
| 12 |
+
|
| 13 |
+
The closest verified prior art is [`victor/gemma-avatar`](https://huggingface.co/spaces/victor/gemma-avatar) — an HF-staff-built Space with pipeline `speak → silero-VAD → parakeet (STT) → gemma-4 (LLM) → Qwen3-TTS → avatar`, rendered with three.js + TalkingHead, lip-sync via an `AudioWorklet` doing MFCC analysis (~50 ms), transported over WebSocket with PCM16 @ 16 kHz, `sdk: docker`. It validates the layer split below — and its choices (Docker SDK, cloud-hosted frontier LLM, ARKit-viseme RPM avatars) are the exact places this project should deliberately diverge (Gradio SDK, free CPU baseline, VRM visemes).
|
| 14 |
+
|
| 15 |
+
### System Overview
|
| 16 |
+
|
| 17 |
+
```
|
| 18 |
+
┌──────────────────────────────────────────────────────────────────────────┐
|
| 19 |
+
│ BROWSER (owns: rendering, animation clock, mic capture, playback) │
|
| 20 |
+
├──────────────────────────────────────────────────────────────────────────┤
|
| 21 |
+
│ ┌────────────┐ ┌────────────┐ ┌────────────┐ ┌──────────────────┐ │
|
| 22 |
+
│ │ VRM Stage │ │ Lip-Sync │ │ Audio │ │ Mic Capture │ │
|
| 23 |
+
│ │ three.js + │◄─│ Driver │◄─│ Queue │ │ MediaRecorder / │ │
|
| 24 |
+
│ │ three-vrm │ │ (timeline │ │ (WebAudio) │ │ AudioWorklet+VAD │ │
|
| 25 |
+
│ │ idle/blink │ │ + envelope│ │ │ │ │ │
|
| 26 |
+
│ └────────────┘ │ fallback) │ └────────────┘ └────────┬─────────┘ │
|
| 27 |
+
│ ▲ └────────────┘ ▲ │ │
|
| 28 |
+
│ └──────────── avatar.speak() ───┘ │ │
|
| 29 |
+
├═══════════════════════ TRANSPORT BOUNDARY ═══════════════════════════════┤
|
| 30 |
+
│ Python→JS: AvatarDirective {audio, visemes[], expression, subtitle} │
|
| 31 |
+
│ JS→Python: UserUtterance {audio|text, ts} · UiIntent {mode, answer} │
|
| 32 |
+
├══════════════════════════════════════════════════════════════════════════┤
|
| 33 |
+
│ GRADIO APP SHELL (Blocks; gr.HTML custom component hosts the stage) │
|
| 34 |
+
├──────────────────────────────────────────────────────────────────────────┤
|
| 35 |
+
│ VOICE PIPELINE │ TUTOR CORE │
|
| 36 |
+
│ ┌──────┐ ┌──────┐ │ ┌────────────────────────────────────────┐ │
|
| 37 |
+
│ │ VAD │ │ ASR │ │ │ SessionOrchestrator (turn lifecycle) │ │
|
| 38 |
+
│ └──┬───┘ └──┬───┘ │ ├──────────┬───────────┬─────────────────┤ │
|
| 39 |
+
│ └────────┘ │ │ Free │ Drill │ RolePlay │ │
|
| 40 |
+
│ ┌──────┐ ┌───────────┐ │ │ Convo │ Session │ Scene │ │
|
| 41 |
+
│ │ TTS │→│ Viseme │ │ └──────────┴───────────┴─────────────────┘ │
|
| 42 |
+
│ └──────┘ │ Generator │ │ ┌────────────┐ ┌─────────────┐ │
|
| 43 |
+
│ └───────────┘ │ │LevelPolicy │ │ OutputGuard │ │
|
| 44 |
+
├──────────────────────────┴──┴────────────┴─┴──────┬──────┴───────────────┤
|
| 45 |
+
│ SHARED SERVICES │ │
|
| 46 |
+
│ ┌───────────────┐ ┌──────────────────┐ ┌──────────▼────────────┐ │
|
| 47 |
+
│ │ LLMGateway │ │ JapaneseAnalyzer │ │ AssessmentEngine │ │
|
| 48 |
+
│ │ local│HF│BYOK │ │ fugashi+UniDic │ │ FSRS + level estimate │ │
|
| 49 |
+
│ └───────────────┘ │ JLPT tags, kana │ └───────────────────────┘ │
|
| 50 |
+
│ └──────────────────┘ │
|
| 51 |
+
├──────────────────────────────────────────────────────────────────────────┤
|
| 52 |
+
│ PERSISTENCE │
|
| 53 |
+
│ ┌──────────────┐ ┌──────────────────┐ ┌───────────────────────────┐ │
|
| 54 |
+
│ │ HF OAuth │ │ Postgres (Neon) │ │ ContentStore (repo files) │ │
|
| 55 |
+
│ │ identity │ │ learner state │ │ lessons/scenarios/vocab │ │
|
| 56 |
+
│ └──────────────┘ └──────────────────┘ └───────────────────────────┘ │
|
| 57 |
+
└──────────────────────────────────────────────────────────────────────────┘
|
| 58 |
+
▲ external: Anthropic/OpenAI (BYOK) · HF Inference Providers
|
| 59 |
+
```
|
| 60 |
+
|
| 61 |
+
### Component Responsibilities
|
| 62 |
+
|
| 63 |
+
| Component | Responsibility (owns) | Typical Implementation |
|
| 64 |
+
|-----------|----------------------|------------------------|
|
| 65 |
+
| **VRM Stage** | Scene, camera, lights, VRM model, idle/blink/breath loop, expression blending. Runs its own `requestAnimationFrame` clock. | `three` + `@pixiv/three-vrm` (v3.x) loaded via ESM CDN or bundled asset |
|
| 66 |
+
| **Lip-Sync Driver** | Turns a viseme timeline into per-frame `expressionManager.setValue('aa', w)` calls; falls back to RMS envelope when no timeline | Plain JS module, ~150 LOC; `AudioContext.currentTime` as master clock |
|
| 67 |
+
| **Audio Queue** | Sequences TTS clips, exposes precise playback time, emits `speech-start` / `speech-end` | WebAudio `AudioBufferSourceNode` chain |
|
| 68 |
+
| **Mic Capture** | getUserMedia, encoding, push-to-talk or VAD-gated turn detection | `MediaRecorder` (v1) → `AudioWorklet` + PCM16 (v2) |
|
| 69 |
+
| **Gradio App Shell** | Layout, session `gr.State`, event wiring, auth button. **No business logic.** | `gr.Blocks` in `app.py` |
|
| 70 |
+
| **VAD / ASR** | Endpointing + Japanese transcription with confidence + (optionally) word timings | Silero VAD; kotoba-whisper / faster-whisper int8 server-side, or browser ASR |
|
| 71 |
+
| **TTS + Viseme Generator** | Japanese audio synthesis **and** the mora-accurate viseme timeline | VOICEVOX/AivisSpeech-compatible engine (`audio_query` → mora durations), or TTS + `pyopenjtalk` g2p estimation |
|
| 72 |
+
| **SessionOrchestrator** | Turn lifecycle: receive utterance → route to active mode → assemble directive → persist evidence | Plain Python class; the only thing `app.py` calls |
|
| 73 |
+
| **Modes** (Free / Drill / RolePlay) | Mode-specific state machine; produces a `TurnRequest`, consumes a `TurnResult` | Three classes behind one ABC — *not* three code paths |
|
| 74 |
+
| **LevelPolicy** | Resolves learner level → allowed kanji/grammar/vocab sets, target sentence length, speech rate | Data-driven from JLPT lists; injected into every mode |
|
| 75 |
+
| **OutputGuard** | Deterministic post-generation validation of tutor Japanese against LevelPolicy; repair/regenerate loop | `JapaneseAnalyzer` + set membership. No LLM. |
|
| 76 |
+
| **LLMGateway** | One `chat(messages, schema, stream)` interface over local / HF-routed / BYOK providers, with capability flags | llama-cpp-python (GBNF grammar), `huggingface_hub.InferenceClient`, `anthropic`/`openai` SDKs |
|
| 77 |
+
| **JapaneseAnalyzer** | Tokenization, readings/furigana, lemma, POS, JLPT level per token — used for *both* grading input and guarding output | `fugashi` + `unidic-lite`, JMdict via `jamdict`, JLPT tag tables |
|
| 78 |
+
| **AssessmentEngine** | Turn evidence → learner-state deltas: level estimate, per-item scheduling | FSRS (`py-fsrs`) for vocab/grammar items; rolling level estimator |
|
| 79 |
+
| **ContentStore** | Lessons, drills, scenario scripts, vocab lists as versioned data | YAML/JSONL in the repo, loaded at boot |
|
| 80 |
+
| **Persistence Repository** | The *only* module that speaks SQL | SQLAlchemy over serverless Postgres |
|
| 81 |
+
|
| 82 |
+
---
|
| 83 |
+
|
| 84 |
+
## The Transport Boundary (most important decision)
|
| 85 |
+
|
| 86 |
+
### Recommended: Gradio 6 `gr.HTML` custom component
|
| 87 |
+
|
| 88 |
+
Gradio 6's `gr.HTML` gained a genuine custom-component API that removes almost all the reason to build a packaged custom component or a separate SPA. Verified parameters ([official docs](https://www.gradio.app/docs/gradio/html), [guide](https://www.gradio.app/guides/custom-HTML-components)):
|
| 89 |
+
|
| 90 |
+
| Parameter | What it gives you here |
|
| 91 |
+
|-----------|------------------------|
|
| 92 |
+
| `head` | Raw HTML injected into `<head>` — loads three.js + `@pixiv/three-vrm` from CDN. **Scripts are deduplicated by `src`.** |
|
| 93 |
+
| `html_template` | The `<canvas>` + overlay markup, with `${}` / Handlebars interpolation |
|
| 94 |
+
| `css_template` | Component-scoped styling |
|
| 95 |
+
| `js_on_load` | Boot code with access to `element`, `props`, `trigger(eventName, data)`, `upload(file)`, and `server` |
|
| 96 |
+
| `server_functions` | Python functions callable from JS as `await server.my_func(...)` |
|
| 97 |
+
| `props` | Extra template variables; JS sets `props.value = x` to re-render; Python-driven prop updates fire JS `watch()` callbacks |
|
| 98 |
+
| `container` / `padding` | Both default `False` — no Gradio chrome around the canvas |
|
| 99 |
+
|
| 100 |
+
**Why this over the alternatives:**
|
| 101 |
+
|
| 102 |
+
| Option | Verdict |
|
| 103 |
+
|--------|---------|
|
| 104 |
+
| `gr.HTML` custom component | **Recommended for v1.** No npm/build toolchain, no publish loop, stays on `sdk: gradio`, bidirectional Python↔JS is first-class. |
|
| 105 |
+
| Packaged custom component (`gradio cc create`, Svelte + Vite) | Defer. Gives TypeScript + real bundling + a publishable PyPI artifact (nice portfolio bonus), but adds a build/version-pin loop to every avatar tweak. Migrate later if the JS exceeds ~1500 LOC. |
|
| 106 |
+
| `gr.HTML` + `<iframe>` + `postMessage` | Avoid. Adds an origin boundary, breaks HF's iframe/cookie handling, and `gr.HTML` custom components make it unnecessary. |
|
| 107 |
+
| FastAPI shell + `gr.mount_gradio_app` + static SPA | Only if you need your own WebSocket/WebRTC signalling. Confirmed to work (`gr.mount_gradio_app(app, demo, path=..., auth_dependency=...)`), but pushes the Space to `sdk: docker` and costs the "Gradio demo" framing. |
|
| 108 |
+
|
| 109 |
+
**Non-negotiable discipline:** keep the actual avatar code in `avatar/avatar.js` (loaded via `head`), and let `js_on_load` be ~20 lines of glue. Expose a narrow, transport-agnostic API:
|
| 110 |
+
|
| 111 |
+
```js
|
| 112 |
+
// avatar/avatar.js — the ONLY surface Python knows about
|
| 113 |
+
window.Avatar = {
|
| 114 |
+
mount(canvasEl, vrmUrl),
|
| 115 |
+
speak({ audioUrl, visemes, expression, subtitle }), // returns Promise, resolves at speech-end
|
| 116 |
+
setMood(name, weight),
|
| 117 |
+
setListening(bool),
|
| 118 |
+
setThinking(bool),
|
| 119 |
+
on(event, cb) // 'speech-start' | 'speech-end' | 'ready' | 'error'
|
| 120 |
+
};
|
| 121 |
+
```
|
| 122 |
+
|
| 123 |
+
If that contract holds, swapping Gradio → FastAPI → WebRTC later is a ~50-line change instead of a rewrite.
|
| 124 |
+
|
| 125 |
+
---
|
| 126 |
+
|
| 127 |
+
## Recommended Project Structure
|
| 128 |
+
|
| 129 |
+
```
|
| 130 |
+
app.py # Gradio Blocks assembly ONLY — wiring, no logic
|
| 131 |
+
README.md # HF Space YAML frontmatter (sdk, sdk_version, hf_oauth)
|
| 132 |
+
requirements.txt
|
| 133 |
+
|
| 134 |
+
avatar/ # served to the browser (gr.set_static_paths)
|
| 135 |
+
├── avatar.js # window.Avatar facade — the transport contract
|
| 136 |
+
├── vrm-stage.js # three.js scene, VRM load, idle/blink/breathe loop
|
| 137 |
+
├── lipsync.js # viseme timeline player + RMS-envelope fallback
|
| 138 |
+
├── audio-queue.js # WebAudio sequencing, clock, speech-start/end events
|
| 139 |
+
├── mic.js # capture, encoding, push-to-talk / VAD gating
|
| 140 |
+
└── assets/
|
| 141 |
+
├── tutor.vrm # Git LFS (5-30 MB)
|
| 142 |
+
└── expressions.json
|
| 143 |
+
|
| 144 |
+
src/japanese_avatar/
|
| 145 |
+
├── ui/
|
| 146 |
+
│ ├── blocks.py # layout composition
|
| 147 |
+
│ ├── avatar_component.py # the gr.HTML wrapper (head/template/js_on_load/server_functions)
|
| 148 |
+
│ └── panels/ # subtitle+furigana, corrections, progress, settings
|
| 149 |
+
├── voice/
|
| 150 |
+
│ ├── asr.py # Transcript{text, confidence, words?}
|
| 151 |
+
│ ├── tts.py # SynthResult{wav_bytes, duration, moras?}
|
| 152 |
+
│ ├── visemes.py # -> VisemeTimeline (the contract type)
|
| 153 |
+
│ └── vad.py
|
| 154 |
+
├── tutor/
|
| 155 |
+
│ ├── orchestrator.py # SessionOrchestrator
|
| 156 |
+
│ ├── modes/ # base.py, free_conversation.py, drill.py, roleplay.py
|
| 157 |
+
│ ├── level_policy.py
|
| 158 |
+
│ ├── output_guard.py
|
| 159 |
+
│ └── prompts/ # versioned .md/.jinja templates, NOT inline strings
|
| 160 |
+
├── llm/
|
| 161 |
+
│ ├── gateway.py # provider resolution + capability flags
|
| 162 |
+
│ ├── providers/ # local_llamacpp.py, hf_inference.py, anthropic.py, openai.py
|
| 163 |
+
│ └── schemas.py # TurnRequest / TurnResult pydantic models
|
| 164 |
+
├── jp/
|
| 165 |
+
│ ├── analyzer.py # fugashi + UniDic
|
| 166 |
+
│ ├── furigana.py
|
| 167 |
+
│ └── jlpt.py # token -> N5..N1 level tables
|
| 168 |
+
├── learner/
|
| 169 |
+
│ ├── assessment.py
|
| 170 |
+
│ ├── scheduler.py # FSRS wrapper
|
| 171 |
+
│ ├── repository.py # the ONLY module with SQL
|
| 172 |
+
│ └── models.py
|
| 173 |
+
├── content/
|
| 174 |
+
│ ├── lessons/*.yaml
|
| 175 |
+
│ ├── scenarios/*.yaml
|
| 176 |
+
│ └── vocab/*.jsonl
|
| 177 |
+
└── db/
|
| 178 |
+
├── engine.py
|
| 179 |
+
└── migrations/
|
| 180 |
+
```
|
| 181 |
+
|
| 182 |
+
### Structure Rationale
|
| 183 |
+
|
| 184 |
+
- **`avatar/` is a peer of `src/`, not a child.** It is a separate runtime with its own lifecycle; nesting it under Python packaging invites someone to generate JS from Python strings.
|
| 185 |
+
- **`tutor/modes/` behind one ABC.** Free conversation, drills, and role-play are the three headline features and they *will* be tempted into three parallel implementations. One interface producing `TurnRequest` is what makes "adaptive N5→N2" a shared parameter rather than three separate difficulty systems.
|
| 186 |
+
- **`jp/` is deliberately LLM-free.** It is used by both the input grader and the output guard; keeping it deterministic is what makes level enforcement provable.
|
| 187 |
+
- **`llm/` isolates the hybrid backend** so that modes never learn which provider answered.
|
| 188 |
+
- **`content/` is data, not code.** Lessons authored as files can be reviewed, diffed, and expanded without touching the engine — and pedagogical quality is an explicit project goal.
|
| 189 |
+
|
| 190 |
+
---
|
| 191 |
+
|
| 192 |
+
## Architectural Patterns
|
| 193 |
+
|
| 194 |
+
### Pattern 1: Viseme Timeline as the Lip-Sync Contract
|
| 195 |
+
|
| 196 |
+
**What:** The server ships the browser a precomputed, absolute-time viseme schedule alongside the audio. The browser never guesses phonemes.
|
| 197 |
+
|
| 198 |
+
```python
|
| 199 |
+
# voice/visemes.py
|
| 200 |
+
VisemeEvent = TypedDict("VisemeEvent", {
|
| 201 |
+
"t": float, # seconds from clip start
|
| 202 |
+
"dur": float,
|
| 203 |
+
"viseme": str, # 'aa' | 'ih' | 'ou' | 'ee' | 'oh' | 'sil'
|
| 204 |
+
"weight": float, # 0..1
|
| 205 |
+
})
|
| 206 |
+
```
|
| 207 |
+
|
| 208 |
+
**Why this is unusually good for Japanese:** Japanese is mora-timed and every mora resolves to one of exactly five vowels (a/i/u/e/o), which map **1:1** onto the VRM 1.0 preset expressions `aa`, `ih`, `ou`, `ee`, `oh` ([three-vrm expression presets](https://github.com/pixiv/three-vrm)). Unlike English, no phoneme→viseme collapsing table is needed. Better still, VOICEVOX-compatible engines return per-mora `consonant_length` and `vowel_length` from `audio_query` *before* synthesis — an exact timeline for free, no audio analysis ([VOICEVOX AudioQuery model](https://deepwiki.com/VOICEVOX/voicevox_engine/5.2-audio-query-and-synthesis-models)). This is the single largest quality-per-effort win available in this project.
|
| 209 |
+
|
| 210 |
+
**Three-tier degradation (build all three behind the same contract):**
|
| 211 |
+
|
| 212 |
+
| Tier | Source | Quality | When |
|
| 213 |
+
|------|--------|---------|------|
|
| 214 |
+
| A | TTS engine mora durations (VOICEVOX / AivisSpeech `audio_query`) | Exact | Preferred path |
|
| 215 |
+
| B | `pyopenjtalk` g2p → mora sequence, durations distributed over measured audio length | Good | Any TTS without timing output |
|
| 216 |
+
| C | WebAudio `AnalyserNode` RMS/MFCC envelope → mouth openness | Flappy but always works | Fallback; also the only option for browser `speechSynthesis` |
|
| 217 |
+
|
| 218 |
+
**Trade-offs:** Tier A couples you to a VOICEVOX-API-compatible engine (a sidecar process, likely `sdk: docker`, ~1 GB image). Tier C is what `gemma-avatar` ships and is genuinely acceptable — build C first (it unblocks the whole pipeline in an afternoon), then upgrade.
|
| 219 |
+
|
| 220 |
+
### Pattern 2: Structured Turn Result (never free prose)
|
| 221 |
+
|
| 222 |
+
**What:** Every tutor generation returns one JSON object serving all UI surfaces from a single call.
|
| 223 |
+
|
| 224 |
+
```python
|
| 225 |
+
class TurnResult(BaseModel):
|
| 226 |
+
reply_ja: str
|
| 227 |
+
reply_kana: str # for furigana rendering
|
| 228 |
+
translation_en: str
|
| 229 |
+
corrections: list[Correction] # span, fixed, why_en, grammar_point_id
|
| 230 |
+
vocab_introduced: list[str]
|
| 231 |
+
expression: Literal["neutral","happy","surprised","thinking","sad"]
|
| 232 |
+
scene_action: Literal["continue","advance","end"] | None
|
| 233 |
+
difficulty_signal: Literal["too_easy","ok","too_hard"] | None
|
| 234 |
+
```
|
| 235 |
+
|
| 236 |
+
**Why:** the avatar needs `expression`, the subtitle panel needs `reply_ja` + `reply_kana`, the correction panel needs `corrections`, the assessment engine needs `vocab_introduced` + `difficulty_signal`, and the TTS needs `reply_ja` alone. Extracting these from prose with a second LLM call doubles latency and cost on a free-CPU budget.
|
| 237 |
+
|
| 238 |
+
**Trade-off / gotcha:** sub-2B local models are unreliable JSON emitters. Handle it at the provider layer, not by prompting harder — **GBNF grammar-constrained decoding** in llama.cpp for the local path, JSON-schema / tool-use for frontier providers, and a `supports_json_schema` capability flag so modes can request a simplified schema when a provider can't guarantee it.
|
| 239 |
+
|
| 240 |
+
### Pattern 3: Deterministic Level Guard (counters "alignment drift")
|
| 241 |
+
|
| 242 |
+
**What:** After generation, tokenize `reply_ja` with `JapaneseAnalyzer`, check every token's JLPT tag and every kanji against `LevelPolicy`'s allowed sets, and reject/repair violations before the text ever reaches TTS.
|
| 243 |
+
|
| 244 |
+
```python
|
| 245 |
+
# tutor/output_guard.py
|
| 246 |
+
def guard(result: TurnResult, policy: LevelPolicy) -> GuardVerdict:
|
| 247 |
+
toks = analyzer.tokenize(result.reply_ja)
|
| 248 |
+
over = [t for t in toks if t.jlpt_level < policy.min_level # harder than target
|
| 249 |
+
and t.lemma not in policy.explicitly_taught]
|
| 250 |
+
bad_kanji = policy.kanji_violations(result.reply_ja)
|
| 251 |
+
long = len(toks) > policy.max_tokens_per_sentence
|
| 252 |
+
return GuardVerdict(ok=not (over or bad_kanji or long),
|
| 253 |
+
violations=over + bad_kanji, retry_hint=...)
|
| 254 |
+
```
|
| 255 |
+
|
| 256 |
+
**Why it matters:** peer-reviewed 2025 work on CEFR-prompted LLM tutors documents *alignment drift* — models hold a requested proficiency level for a few turns and then regress toward their default register, and the effect is **sharply worse on open 7B-12B models** ([Alignment Drift in CEFR-prompted LLMs, BEA 2025](https://aclanthology.org/2025.bea-1.6/)). Since this project's free baseline is a *sub-2B* model, prompt-only level control will not hold. A deterministic guard is the fix, and it is cheap.
|
| 257 |
+
|
| 258 |
+
**Bonus:** it is also the most legible engineering story in the project — "deterministic guardrails wrapping a stochastic model, with a live violation-rate metric" is exactly the kind of thing an AI/ML hiring manager reads as senior judgment. Surface the counter in the UI.
|
| 259 |
+
|
| 260 |
+
**Trade-offs:** a retry loop costs latency; cap at 1-2 retries then fall back to a simplification pass or a templated safe response. Guard the *tutor's* output, not the learner's.
|
| 261 |
+
|
| 262 |
+
### Pattern 4: Two-Tier Voice Transport (turn-based now, streaming later)
|
| 263 |
+
|
| 264 |
+
**What:** Ship a turn-based pipeline, and design the `Avatar.speak()` contract so a streaming transport can be dropped in without touching the tutor core.
|
| 265 |
+
|
| 266 |
+
```
|
| 267 |
+
TIER 1 (v1, ship this):
|
| 268 |
+
mic → MediaRecorder blob → server.transcribe() → orchestrator
|
| 269 |
+
→ TTS + viseme gen → AvatarDirective → Avatar.speak()
|
| 270 |
+
|
| 271 |
+
TIER 2 (later):
|
| 272 |
+
mic → AudioWorklet PCM16 → WebSocket/WebRTC → server VAD (silero)
|
| 273 |
+
→ streaming ASR → streaming LLM tokens → sentence-chunked TTS
|
| 274 |
+
→ chunked {audio, visemes} → Avatar.enqueue() ← same contract, plural
|
| 275 |
+
```
|
| 276 |
+
|
| 277 |
+
**Tier-2 option — FastRTC:** [`gradio-app/fastrtc`](https://github.com/gradio-app/fastrtc) turns a Python generator into a WebRTC/WebSocket stream with built-in `ReplyOnPause` VAD turn-taking, mounts onto FastAPI (`/webrtc/offer`), supports `AdditionalOutputs` for pushing non-audio data (i.e. viseme timelines) to the client, and enforces a `concurrency_limit`. **On HF Spaces it requires a TURN server** — HF/Cloudflare provide 10 GB/month free via `get_cloudflare_turn_credentials_async(hf_token=...)` ([FastRTC deployment docs](https://fastrtc.org/deployment/)).
|
| 278 |
+
⚠️ **Risk (MEDIUM confidence):** FastRTC's last release was v0.0.34 (Nov 2025), before Gradio 6. Verify Gradio 6 compatibility empirically before committing a phase to it; the WebSocket path is the lower-risk half of the library.
|
| 279 |
+
|
| 280 |
+
**Why not stream on day one:** VAD tuning, barge-in/interruption, echo cancellation, TURN, and partial-transcript UX are collectively a project of their own, and none of them are the *tutoring* value. Turn-based with a visible "listening / thinking / speaking" avatar state reads as intentional, not slow.
|
| 281 |
+
|
| 282 |
+
### Pattern 5: Provider Resolution Chain (hybrid backend)
|
| 283 |
+
|
| 284 |
+
**What:** `LLMGateway` resolves a provider per request by a fixed order, and every mode is written against capability flags rather than provider names.
|
| 285 |
+
|
| 286 |
+
```
|
| 287 |
+
explicit user selection
|
| 288 |
+
→ user-supplied BYOK key present in session state (Anthropic / OpenAI)
|
| 289 |
+
→ user's HF OAuth token with `inference-api` scope (spends THEIR credits)
|
| 290 |
+
→ local llama.cpp model in-process [always-available floor, $0]
|
| 291 |
+
```
|
| 292 |
+
|
| 293 |
+
**Capability flags** (`supports_json_schema`, `supports_streaming`, `max_context`, `latency_class`) let a mode degrade: e.g. RolePlay requests a richer schema when available and a two-field schema on the local floor.
|
| 294 |
+
|
| 295 |
+
**Verified economics:** HF's free-tier Inference Provider allowance is **$0.10/month** and PRO is **$2/month** ([HF pricing 2026](https://huggingface.co/docs/api-inference/en/pricing)) — so routing *your* token for all visitors is not viable. Routing the *visitor's* token via the `inference-api` OAuth scope is architecturally elegant but still tiny. **The local CPU model must be a genuinely usable floor, not a token gesture.**
|
| 296 |
+
|
| 297 |
+
---
|
| 298 |
+
|
| 299 |
+
## Data Flow
|
| 300 |
+
|
| 301 |
+
### Turn Flow (v1, turn-based)
|
| 302 |
+
|
| 303 |
+
```
|
| 304 |
+
[learner presses talk / VAD fires]
|
| 305 |
+
│
|
| 306 |
+
▼
|
| 307 |
+
mic.js ── audio blob ──► server.submit_utterance(blob, session) (gr.HTML server_functions)
|
| 308 |
+
│
|
| 309 |
+
▼
|
| 310 |
+
voice/asr.py ──► Transcript{text, confidence}
|
| 311 |
+
│
|
| 312 |
+
▼
|
| 313 |
+
tutor/orchestrator.py
|
| 314 |
+
├─► jp/analyzer.py ........ tokenize learner text, JLPT-tag, detect errors
|
| 315 |
+
├─► learner/repository.py . load LearnerState (level, due items, recent mistakes)
|
| 316 |
+
├���► tutor/modes/<active> .. build TurnRequest(system policy + context + constraints)
|
| 317 |
+
├─► llm/gateway.py ........ chat(TurnRequest, schema=TurnResult) ──► TurnResult
|
| 318 |
+
├─► tutor/output_guard.py . validate reply_ja vs LevelPolicy ──[fail]──┐
|
| 319 |
+
│ ▲ │
|
| 320 |
+
│ └──────────────── repair / regenerate (max 2) ◄───────────────┘
|
| 321 |
+
├─► voice/tts.py .......... SynthResult{wav, duration, moras?}
|
| 322 |
+
├─► voice/visemes.py ...... VisemeTimeline (tier A/B/C)
|
| 323 |
+
├─► learner/assessment.py . evidence ──► FSRS updates, level estimate delta
|
| 324 |
+
└─► learner/repository.py . persist turn + state (async, non-blocking)
|
| 325 |
+
│
|
| 326 |
+
▼
|
| 327 |
+
AvatarDirective { audio_url, visemes[], expression, subtitle{ja,kana,en}, corrections[] }
|
| 328 |
+
│
|
| 329 |
+
▼ (props update → JS watch())
|
| 330 |
+
avatar.js → audio-queue.js (schedules buffer) + lipsync.js (schedules visemes on same clock)
|
| 331 |
+
│
|
| 332 |
+
▼
|
| 333 |
+
[avatar speaks; 'speech-end' → trigger('turn_complete') → UI re-enables mic]
|
| 334 |
+
```
|
| 335 |
+
|
| 336 |
+
### State Management
|
| 337 |
+
|
| 338 |
+
Three stores, deliberately separated — conflating them is the #1 structural mistake here:
|
| 339 |
+
|
| 340 |
+
```
|
| 341 |
+
BROWSER (ephemeral, per-frame) : VRM pose, expression weights, audio clock, mic state
|
| 342 |
+
→ NEVER round-trips to Python
|
| 343 |
+
|
| 344 |
+
SESSION (gr.State / gr.BrowserState) : active mode, scene position, dialogue window,
|
| 345 |
+
BYOK key (memory only, never persisted), UI prefs
|
| 346 |
+
|
| 347 |
+
DURABLE (Postgres, keyed on HF user) : level estimate, per-item FSRS state, vocab exposure,
|
| 348 |
+
mistake history, session summaries
|
| 349 |
+
```
|
| 350 |
+
|
| 351 |
+
Rule: the browser holds *presentation* state, `gr.State` holds *conversation* state, the DB holds *learning* state. A learner should be able to hard-refresh mid-conversation and lose only the dialogue window.
|
| 352 |
+
|
| 353 |
+
### Key Data Flows
|
| 354 |
+
|
| 355 |
+
1. **Level adaptation loop (the product's core):** turn evidence → `AssessmentEngine` → durable level estimate → `LevelPolicy` → injected into the next `TurnRequest` *and* enforced by `OutputGuard`. Note it closes through **two** paths — generation-time hint and validation-time enforcement — because the generation hint alone provably drifts.
|
| 356 |
+
2. **Spaced repetition loop:** `user_vocab_state` due items → `DrillSession` selects → learner responds → grade → FSRS updates stability/difficulty/next-due. FSRS v5 (`py-fsrs`) needs ~20-30% fewer reviews than SM-2 at equal retention and is the current state of the art; there is no reason to hand-roll SM-2.
|
| 357 |
+
3. **Content flow:** `content/*.yaml` loaded at boot into an in-memory index → modes select by (level, topic, due-ness) → item IDs referenced in the DB. **Content lives in the repo; only learner state lives in the DB.** This keeps the DB tiny and content diffable.
|
| 358 |
+
4. **Auth flow:** `gr.LoginButton` → HF OAuth → `gr.OAuthProfile` parameter injected into handlers → `profile.username` (or OIDC `sub`) is the DB key. When `profile is None`, the *same* code path runs against an ephemeral in-session learner — mandatory, since free-tier visitors must get a working experience.
|
| 359 |
+
|
| 360 |
+
---
|
| 361 |
+
|
| 362 |
+
## Persistence: the platform constraint that shapes everything
|
| 363 |
+
|
| 364 |
+
**Verified, and it invalidates the obvious design:** HF Spaces disk is ephemeral — *"Every Space comes with a small amount of disk storage. This disk space is ephemeral, meaning its content will be lost if your Space restarts or is stopped"* ([Disk usage on Spaces](https://huggingface.co/docs/hub/en/spaces-storage)). The current recommended persistence mechanism is **Storage Buckets**, which are explicitly **S3-like, non-versioned object storage** on the Xet backend ([Storage Buckets docs](https://huggingface.co/docs/hub/en/storage-buckets)).
|
| 365 |
+
|
| 366 |
+
Object storage mounted as a volume is **not** a safe home for a live SQLite file (no POSIX locking semantics, no transactional durability, corruption under concurrent writers). Neither is a `CommitScheduler`-to-Dataset pattern, which is append-oriented and eventually consistent — fine for logs, wrong for mutable learner state.
|
| 367 |
+
|
| 368 |
+
**Recommendation: external serverless Postgres — Neon free tier.**
|
| 369 |
+
|
| 370 |
+
| Option | Verdict |
|
| 371 |
+
|--------|---------|
|
| 372 |
+
| **Neon** | ✅ Scale-to-zero after ~5 min idle, ~500 ms resume, project stays reachable indefinitely, 0.5 GB free (learner state is kilobytes/user). |
|
| 373 |
+
| Supabase | ⚠️ Free projects **pause after 7 days of inactivity** and need manual revival — fatal for a portfolio Space a recruiter opens once a month. |
|
| 374 |
+
| Turso/libSQL | ✅ Viable (SQLite semantics, generous free tier, kept warm), but a non-Postgres story is a weaker portfolio signal. |
|
| 375 |
+
| Space disk / Storage Bucket | ❌ Ephemeral / object storage. Use buckets only for model weights and exported analytics. |
|
| 376 |
+
|
| 377 |
+
Connection string goes in **Space Secrets**; use Neon's pooled endpoint (a 2-vCPU Space with several concurrent users will otherwise exhaust connections).
|
| 378 |
+
|
| 379 |
+
### Schema sketch
|
| 380 |
+
|
| 381 |
+
```
|
| 382 |
+
users (id, hf_sub, hf_username, created_at, settings_json)
|
| 383 |
+
sessions (id, user_id, mode, started_at, ended_at)
|
| 384 |
+
turns (id, session_id, role, text_ja, asr_confidence,
|
| 385 |
+
corrections_json, latency_ms, provider, created_at) -- append-only
|
| 386 |
+
level_estimates (id, user_id, jlpt_estimate, confidence, evidence_json, created_at)
|
| 387 |
+
vocab_items (id, lemma, reading, jlpt_level, gloss_en) -- content mirror
|
| 388 |
+
user_vocab_state (user_id, vocab_item_id, stability, difficulty,
|
| 389 |
+
due_at, reps, lapses, last_review) -- FSRS
|
| 390 |
+
grammar_points (id, code, jlpt_level, title)
|
| 391 |
+
user_grammar_state (user_id, grammar_point_id, ...FSRS...)
|
| 392 |
+
```
|
| 393 |
+
|
| 394 |
+
`turns` append-only is what makes the assessment engine improvable later without data loss — you can recompute level estimates from history when the algorithm changes.
|
| 395 |
+
|
| 396 |
+
**Never persist BYOK API keys.** Session memory only, `type="password"`, scrubbed from logs and from any turn record.
|
| 397 |
+
|
| 398 |
+
---
|
| 399 |
+
|
| 400 |
+
## Suggested Build Order
|
| 401 |
+
|
| 402 |
+
Ordered by **risk retirement and hard dependency**, not by feature glamour.
|
| 403 |
+
|
| 404 |
+
| # | Component | Depends on | Why here |
|
| 405 |
+
|---|-----------|-----------|----------|
|
| 406 |
+
| **0** | Space skeleton: `sdk: gradio` + pinned `sdk_version` (6.x), README frontmatter, `hf_oauth: true`, a deploy smoke test | — | The project's documented historical gotcha is Python/Gradio version pinning. Prove deployability on day one, with an empty app. |
|
| 407 |
+
| **1** | **Avatar stage**: `gr.HTML` component, three.js + three-vrm via `head`, VRM loads, idle/blink/breathe, `Avatar.speak()` driven by a **hardcoded** viseme timeline + a canned WAV | 0 | Highest novelty × highest risk, and it needs *zero* AI. If VRM-in-Gradio doesn't work, everything downstream is void. Freeze the `AvatarDirective` contract here. |
|
| 408 |
+
| **2** | **Voice out**: TTS + viseme generator (tier C envelope first, then B, then A) | 1 | Avatar speaks arbitrary Japanese typed into a box. First genuinely demoable artifact. |
|
| 409 |
+
| **3** | **Voice in**: ASR + mic capture, wired as an echo bot (avatar repeats what you said) | 2 | Closes the loop with no LLM. Measure real end-to-end latency **now** — it constrains every later decision. |
|
| 410 |
+
| **4** | **LLM gateway + structured turn + free conversation** (one provider only) | 3 | It becomes a talking Japanese partner. Establish `TurnResult` schema and grammar-constrained local decoding here. |
|
| 411 |
+
| **5** | **`jp/` analyzer + LevelPolicy + OutputGuard + corrections UI** | 4 | The step that converts a chatbot into a *tutor*. Also where N5→N2 becomes a real parameter. |
|
| 412 |
+
| **6** | **Persistence**: HF OAuth + Postgres + repository + FSRS + progress UI | 5 | Deliberately *after* the tutor loop: the schema depends on what `AssessmentEngine` actually emits. Designing the DB first guarantees a migration. |
|
| 413 |
+
| **7** | **Content system**: drills + role-play scenarios on the shared mode interface | 5, 6 | Needs the mode ABC (5) and item-level scheduling (6) to exist first. |
|
| 414 |
+
| **8** | **Hybrid routing + BYOK UI + polish** (expressions, gestures, mobile layout, loading UX) | 4, 7 | Additive; the gateway abstraction from 4 means this is configuration, not surgery. |
|
| 415 |
+
| **9** | *(Optional)* **Streaming transport upgrade** (FastRTC / WebSocket PCM) | 8 | Pure latency work behind an unchanged contract. Worth a dedicated research pass before committing. |
|
| 416 |
+
|
| 417 |
+
**Ordering rationale:** phases 1-3 retire the browser/audio/rendering risk — the part with the least prior art in a Gradio context and the part that cannot be worked around if it fails. Phases 4-7 are well-trodden LLM-application work that iterates fast *once the transport contract is frozen*. Persistence sits after the tutor loop because its schema is downstream of assessment design. Streaming is last because it changes no interfaces.
|
| 418 |
+
|
| 419 |
+
**Phases most likely to need dedicated deeper research:** 1 (Gradio 6 `gr.HTML` custom-component specifics + VRM asset licensing), 2 (Japanese TTS engine choice and its licensing/hosting footprint), 9 (FastRTC × Gradio 6 compatibility, TURN).
|
| 420 |
+
|
| 421 |
+
---
|
| 422 |
+
|
| 423 |
+
## Latency Budget
|
| 424 |
+
|
| 425 |
+
Design target: **first avatar audio within 2.5 s of the learner finishing speaking.** Anything beyond ~4 s reads as broken.
|
| 426 |
+
|
| 427 |
+
| Stage | Free CPU + local model | BYOK frontier + cloud TTS |
|
| 428 |
+
|-------|------------------------|---------------------------|
|
| 429 |
+
| Endpointing (VAD silence) | 300-500 ms | 300-500 ms |
|
| 430 |
+
| Upload + decode | 100-300 ms | 100-300 ms |
|
| 431 |
+
| ASR (short utterance) | 0.5-1.5 s | 0.3-0.8 s |
|
| 432 |
+
| LLM (short structured reply) | 1.5-4 s | 0.6-1.5 s |
|
| 433 |
+
| OutputGuard (+ retry) | 20 ms (+1 LLM round on fail) | 20 ms |
|
| 434 |
+
| TTS + viseme gen | 0.5-2 s | 0.3-0.8 s |
|
| 435 |
+
| **Total** | **≈ 3-8 s** | **≈ 1.6-4 s** |
|
| 436 |
+
|
| 437 |
+
⚠️ **Confidence: LOW** — these are estimates from general small-model CPU benchmarks (roughly 15-50 tok/s for well-quantized sub-2B models on modern CPUs) applied to a 2 vCPU Space, not measured on this stack. **Instrument every stage from phase 3 and let real numbers drive the phase-9 decision.**
|
| 438 |
+
|
| 439 |
+
Perceptual mitigations that are cheaper than optimization: switch the avatar to a "thinking" pose the instant ASR returns; play a short natural filler (「えーと」/ a nod) while generating; stream the first sentence's audio before the rest is synthesized; render the subtitle before the audio starts.
|
| 440 |
+
|
| 441 |
+
---
|
| 442 |
+
|
| 443 |
+
## Scaling Considerations
|
| 444 |
+
|
| 445 |
+
| Scale | Architecture adjustments |
|
| 446 |
+
|-------|--------------------------|
|
| 447 |
+
| 0-1k visitors/mo (realistic for a portfolio Space) | Single free CPU Space (2 vCPU / 16 GB), in-process everything, Neon free tier. No changes needed. |
|
| 448 |
+
| 1k-20k | Move ASR to the browser (Web Speech API `lang="ja-JP"`, or transformers.js Whisper on WebGPU) to free server CPU; add `concurrency_limit` on heavy events; Neon pooled connections; cache TTS for fixed lesson prompts. |
|
| 449 |
+
| 20k+ | Not a realistic target for this project. If reached: split voice workers from the Gradio process, push LLM to Inference Providers, front audio assets with a CDN. |
|
| 450 |
+
|
| 451 |
+
### Scaling Priorities (what breaks first)
|
| 452 |
+
|
| 453 |
+
1. **CPU contention between ASR, TTS, and LLM in one 2-vCPU process.** Two concurrent speakers is enough. *Fix:* a bounded worker queue with explicit `concurrency_limit`, plus offloading ASR to the browser — the highest-leverage single change available (it also improves privacy and cuts latency).
|
| 454 |
+
2. **Blocking the event loop.** One synchronous inference call stalls *every* connected user. *Fix:* async handlers with `run_in_executor`, and never do model work inline in a Gradio event.
|
| 455 |
+
3. **Cold start.** Free Spaces sleep after inactivity (MEDIUM confidence: ~48 h), and cold boot must re-download models to ephemeral disk. *Fix:* small models, lazy-load ASR/TTS on first use, and a friendly first-load UI — the avatar can render and idle while models warm.
|
| 456 |
+
4. **DB connections** — trivially fixed with a pooled endpoint, but easy to hit from a multi-worker container.
|
| 457 |
+
5. **Asset weight.** VRM (5-30 MB) + three.js (~600 KB) on first paint. *Fix:* Git LFS or CDN, progressive loading screen, and start the scene before the VRM finishes.
|
| 458 |
+
|
| 459 |
+
---
|
| 460 |
+
|
| 461 |
+
## Anti-Patterns
|
| 462 |
+
|
| 463 |
+
### Anti-Pattern 1: The Mega-Prompt Tutor
|
| 464 |
+
**What people do:** one giant system prompt ("You are a friendly Japanese teacher, adapt to the student's level…") and parse prose replies with regex.
|
| 465 |
+
**Why it's wrong:** level adherence provably decays over turns (documented alignment drift, and worse on small open models — which is exactly this project's free baseline); prose gives you no furigana, no structured corrections, no expression cue, and no assessment signal.
|
| 466 |
+
**Instead:** structured `TurnResult` + `LevelPolicy` injection + deterministic `OutputGuard`.
|
| 467 |
+
|
| 468 |
+
### Anti-Pattern 2: Streaming/WebRTC First
|
| 469 |
+
**What people do:** start with FastRTC/WebRTC full-duplex because "full-duplex voice in v1" is on the requirements list.
|
| 470 |
+
**Why it's wrong:** TURN configuration, VAD tuning, barge-in, and echo cancellation consume the whole schedule and none of them teach anyone Japanese. FastRTC's Gradio-6 compatibility is also currently unverified.
|
| 471 |
+
**Instead:** turn-based with explicit avatar listening/thinking/speaking states, `AvatarDirective` frozen early, streaming as a later transport swap.
|
| 472 |
+
|
| 473 |
+
### Anti-Pattern 3: Amplitude-Only Lip-Sync as the Design
|
| 474 |
+
**What people do:** drive one "mouth open" blendshape from audio RMS forever.
|
| 475 |
+
**Why it's wrong:** the jaw flaps; there is no vowel identity; for Japanese this discards a nearly free 1:1 mora→viseme mapping and forfeits the most visible quality differentiator.
|
| 476 |
+
**Instead:** the timeline contract, with the envelope as tier-C fallback only.
|
| 477 |
+
|
| 478 |
+
### Anti-Pattern 4: State on the Space Filesystem
|
| 479 |
+
**What people do:** SQLite at `./data/app.db`, or a SQLite file inside a mounted Storage Bucket.
|
| 480 |
+
**Why it's wrong:** Space disk is explicitly ephemeral; Storage Buckets are S3-like object storage with no POSIX locking. Silent, total data loss on restart — and "measurably improve over time" is the stated core value.
|
| 481 |
+
**Instead:** external serverless Postgres; buckets only for weights and exported analytics.
|
| 482 |
+
|
| 483 |
+
### Anti-Pattern 5: Driving Animation from Python
|
| 484 |
+
**What people do:** send per-frame expression weights, or ask Python "what should the mouth look like now?"
|
| 485 |
+
**Why it's wrong:** every Gradio round-trip is tens of milliseconds; animation needs sub-16 ms. Playback and rendering must share one clock, and that clock is `AudioContext.currentTime` in the browser.
|
| 486 |
+
**Instead:** Python sends *intents* (a directive per utterance); the browser owns the clock and the interpolation.
|
| 487 |
+
|
| 488 |
+
### Anti-Pattern 6: Three Feature Silos
|
| 489 |
+
**What people do:** `conversation.py`, `lessons.py`, `roleplay.py`, each with its own prompt, its own level handling, and its own progress writes.
|
| 490 |
+
**Why it's wrong:** adaptive difficulty then has to be implemented three times and will be inconsistent three ways; requirements explicitly want level-parameterization "from day one."
|
| 491 |
+
**Instead:** one mode ABC producing `TurnRequest`; `LevelPolicy`, `OutputGuard`, `LLMGateway`, and `AssessmentEngine` are shared by construction.
|
| 492 |
+
|
| 493 |
+
### Anti-Pattern 7: Server-Side BYOK Key Storage
|
| 494 |
+
**What people do:** save the visitor's Anthropic/OpenAI key to the DB "for convenience."
|
| 495 |
+
**Why it's wrong:** it turns a portfolio demo into a credential-breach liability, and it will be noticed by exactly the audience you are trying to impress.
|
| 496 |
+
**Instead:** session memory only, password-type input, explicit "not stored" copy in the UI, scrubbed from logs and turn records.
|
| 497 |
+
|
| 498 |
+
### Anti-Pattern 8: Vendor-Coupled Viseme Generation
|
| 499 |
+
**What people do:** call VOICEVOX's `audio_query` directly from the tutor orchestrator.
|
| 500 |
+
**Why it's wrong:** the TTS engine is the most likely component to change (licensing, hosting cost, voice quality), and coupling drags the whole pipeline with it.
|
| 501 |
+
**Instead:** `voice/tts.py` returns `SynthResult`, `voice/visemes.py` returns `VisemeTimeline`; the orchestrator knows neither engine.
|
| 502 |
+
|
| 503 |
+
---
|
| 504 |
+
|
| 505 |
+
## Integration Points
|
| 506 |
+
|
| 507 |
+
### External Services
|
| 508 |
+
|
| 509 |
+
| Service | Integration Pattern | Notes / gotchas |
|
| 510 |
+
|---------|---------------------|-----------------|
|
| 511 |
+
| **HF OAuth** | `hf_oauth: true` in README frontmatter → `gr.LoginButton` + `gr.OAuthProfile` / `gr.OAuthToken` params. Space gets `OAUTH_CLIENT_ID`, `OAUTH_CLIENT_SECRET`, `OAUTH_SCOPES`, `OPENID_PROVIDER_URL`. Default expiry 480 min (max 43200). | `openid profile` always included; add `inference-api` only if routing through user credits. **`gr.LoginButton` does not restrict access** — profile is simply `None` for anonymous users, so the anonymous path must work. Use `target=_blank` when the Space runs in an iframe or cookies break on some browsers. Test locally with `HF_TOKEN`. |
|
| 512 |
+
| **Postgres (Neon)** | Connection string in Space Secrets; SQLAlchemy; pooled endpoint | ~500 ms cold resume after 5 min idle — hide behind the avatar's idle animation. Never in `requirements`-visible config. |
|
| 513 |
+
| **Anthropic / OpenAI (BYOK)** | Session-scoped client constructed per request from `gr.State` | Never persisted. Rate-limit and surface provider errors as in-character avatar responses, not stack traces. |
|
| 514 |
+
| **HF Inference Providers** | `InferenceClient(token=oauth_token.token)` with `inference-api` scope | Free allowance is $0.10/mo (PRO $2/mo) — supplementary, never the baseline. |
|
| 515 |
+
| **TTS engine (VOICEVOX / AivisSpeech)** | HTTP sidecar: `POST /audio_query` → mora timings → `POST /synthesis` | Requires `sdk: docker` and a ~1 GB image; character-voice **licensing requires attribution** — verify terms per voice before shipping. This is the main force that could push the Space off `sdk: gradio`. |
|
| 516 |
+
| **Cloudflare TURN** (phase 9 only) | `get_cloudflare_turn_credentials_async(hf_token=...)` | 10 GB/mo free via the HF partnership; required for WebRTC on Spaces. |
|
| 517 |
+
| **CDN (esm.sh / unpkg / jsDelivr)** | `head=` script/importmap for three.js + `@pixiv/three-vrm` | Scripts dedupe by `src`. Pin exact versions — an unpinned three.js major will break three-vrm. Consider vendoring the files into the repo for reproducibility. |
|
| 518 |
+
|
| 519 |
+
### Internal Boundaries
|
| 520 |
+
|
| 521 |
+
| Boundary | Communication | Notes |
|
| 522 |
+
|----------|---------------|-------|
|
| 523 |
+
| Browser ↔ Gradio | `server_functions` (JS→Py) + prop updates/`watch()` (Py→JS) + `trigger()` custom events | The one contract to keep stable: `AvatarDirective` / `UserUtterance`. Everything else can churn. |
|
| 524 |
+
| `app.py` ↔ tutor core | Direct calls into `SessionOrchestrator` only | `app.py` must contain no branching business logic — it is wiring. |
|
| 525 |
+
| Modes ↔ LLM | `TurnRequest` / `TurnResult` pydantic models via `LLMGateway` | Modes never see provider identity; only capability flags. |
|
| 526 |
+
| Tutor ↔ JP analysis | Synchronous, deterministic, in-process | Fast (< 20 ms); safe to call multiple times per turn. |
|
| 527 |
+
| Tutor ↔ persistence | `learner/repository.py` only | Single SQL surface makes swapping Neon → anything a one-file change. |
|
| 528 |
+
| Voice ↔ everything | `SynthResult` / `Transcript` / `VisemeTimeline` dataclasses | Insulates the most swap-prone components (ASR/TTS engines). |
|
| 529 |
+
| Content ↔ modes | Read-only in-memory index loaded at boot | Content is repo data; hot-reload in dev only. |
|
| 530 |
+
|
| 531 |
+
---
|
| 532 |
+
|
| 533 |
+
## Sources
|
| 534 |
+
|
| 535 |
+
**Official documentation (HIGH confidence)**
|
| 536 |
+
- [gr.HTML component reference](https://www.gradio.app/docs/gradio/html) — `head`, `html_template`, `css_template`, `js_on_load`, `server_functions`, `props`, container/padding defaults, script dedup by `src`
|
| 537 |
+
- [Custom HTML Components guide](https://www.gradio.app/guides/custom-HTML-components) — `trigger()`, `server.*`, `props.value` re-render, `watch()`
|
| 538 |
+
- [Sharing Your App](https://gradio.app/guides/sharing-your-app) — `gr.LoginButton`, `gr.OAuthProfile`, `gr.OAuthToken`, `mount_gradio_app(auth_dependency=...)`
|
| 539 |
+
- [Adding a Sign-In with HF button to your Space](https://huggingface.co/docs/hub/en/spaces-oauth) — frontmatter keys, scopes, env vars, redirect URIs, iframe/cookie caveat
|
| 540 |
+
- [Disk usage on Spaces](https://huggingface.co/docs/hub/en/spaces-storage) — **disk is ephemeral**
|
| 541 |
+
- [Storage Buckets](https://huggingface.co/docs/hub/en/storage-buckets) — S3-like, non-versioned, mutable object storage
|
| 542 |
+
- [Spaces ZeroGPU](https://huggingface.co/docs/hub/en/spaces-zerogpu) — quota tiers
|
| 543 |
+
- [HF Inference pricing](https://huggingface.co/docs/api-inference/en/pricing) — free $0.10/mo, PRO $2/mo
|
| 544 |
+
- [FastRTC](https://github.com/gradio-app/fastrtc) and [deployment guide](https://fastrtc.org/deployment/) — `ReplyOnPause`, `/webrtc/offer`, `AdditionalOutputs`, `concurrency_limit`, Cloudflare TURN
|
| 545 |
+
- [@pixiv/three-vrm](https://github.com/pixiv/three-vrm) + [migration guide 1.0](https://pixiv.github.io/three-vrm/docs/documents/migration-guide-1.0.html) — `expressionManager.setValue('aa', w)`, presets incl. aa/ih/ou/ee/oh
|
| 546 |
+
- [fugashi](https://github.com/polm/fugashi) / [How to Tokenize Japanese in Python](https://www.dampfkraft.com/nlp/how-to-tokenize-japanese.html) — MeCab + UniDic
|
| 547 |
+
- [awesome-japanese-nlp-resources](https://github.com/taishi-i/awesome-japanese-nlp-resources) — jamdict/JMdict, JLPT tagging resources
|
| 548 |
+
|
| 549 |
+
**Reference implementations & prior art (MEDIUM-HIGH)**
|
| 550 |
+
- [`victor/gemma-avatar` Space](https://huggingface.co/spaces/victor/gemma-avatar) — verified README frontmatter and pipeline description (docker SDK, VAD→STT→LLM→TTS, TalkingHead + three.js, AudioWorklet MFCC lip-sync ~50 ms, WebSocket PCM16 16 kHz)
|
| 551 |
+
- [met4citizen/TalkingHead](https://github.com/met4citizen/TalkingHead) — lip-sync language modules; no Japanese module built in; external TTS with word timings or viseme IDs can substitute
|
| 552 |
+
- [VOICEVOX AudioQuery / Mora model](https://deepwiki.com/VOICEVOX/voicevox_engine/5.2-audio-query-and-synthesis-models) — `consonant_length`, `vowel_length` per mora
|
| 553 |
+
- [Non-audio-based lip-sync from VOICEVOX AudioQuery (JP)](https://zenn.dev/mochineko/articles/d14278b45240da) — the timeline technique in practice
|
| 554 |
+
|
| 555 |
+
**Research (HIGH for the claim cited)**
|
| 556 |
+
- [Alignment Drift in CEFR-prompted LLMs for Interactive Spanish Tutoring, BEA 2025](https://aclanthology.org/2025.bea-1.6/) ([preprint](https://arxiv.org/pdf/2505.08351)) — prompt-only level control degrades over turns; markedly worse on 7-12B open models
|
| 557 |
+
|
| 558 |
+
**Ecosystem comparisons (MEDIUM — multiple sources agree, no single authority)**
|
| 559 |
+
- Neon vs Supabase vs Turso free-tier behavior (Supabase 7-day pause; Neon scale-to-zero ~500 ms resume) — [buildmvpfast](https://www.buildmvpfast.com/blog/neon-vs-supabase-vs-turso-serverless-postgres-mvp-2026), [agentdeals](https://agentdeals.dev/database-free-tier-comparison-2026)
|
| 560 |
+
- [py-fsrs / FSRS vs SM-2](https://deckstudy.com/blog/fsrs-vs-sm2-modern-spaced-repetition) — ~20-30% fewer reviews at equal retention
|
| 561 |
+
- Japanese ASR options — [2026 Japanese ASR benchmark](https://neosophie.com/en/blog/20260226-japanese-asr-benchmark), [kotoba-whisper-v2.0](https://huggingface.co/kotoba-tech/kotoba-whisper-v2.0)
|
| 562 |
+
- Browser ASR — [Web Speech API support](https://developer.mozilla.org/en-US/docs/Web/API/SpeechRecognition) (Chrome/Edge/Safari; **not Firefox**; server-backed, `ja-JP` supported), [Transformers.js](https://huggingface.co/docs/transformers.js/index) WebGPU
|
| 563 |
+
|
| 564 |
+
**Confidence caveats**
|
| 565 |
+
- ⚠️ FastRTC × Gradio 6 compatibility — **unverified**; last FastRTC release predates Gradio 6. Test before planning a phase around it.
|
| 566 |
+
- ⚠️ Latency budget table — **LOW**; extrapolated from general small-model CPU benchmarks, not measured on 2-vCPU Spaces hardware. Instrument from phase 3.
|
| 567 |
+
- ⚠️ Free-Space sleep interval (~48 h) — **MEDIUM**; widely reported, not re-verified against current docs.
|
| 568 |
+
- ⚠️ VOICEVOX character-voice license terms — **MEDIUM**; engine is open source and free for commercial/non-commercial use, but per-character attribution requirements must be checked individually before shipping.
|
| 569 |
+
|
| 570 |
+
---
|
| 571 |
+
*Architecture research for: real-time 3D-avatar Japanese language tutor on Hugging Face Spaces*
|
| 572 |
+
*Researched: 2026-08-08*
|
.planning/research/FEATURES.md
ADDED
|
@@ -0,0 +1,415 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Feature Research
|
| 2 |
+
|
| 3 |
+
**Domain:** Conversational AI language tutor with animated 3D avatar (Japanese, JLPT N5→N2), delivered as a public Hugging Face Space
|
| 4 |
+
**Researched:** 2026-08-08
|
| 5 |
+
**Confidence:** MEDIUM-HIGH (vendor feature pages and official JLPT spec = HIGH; review-aggregator claims and engagement research = MEDIUM; "no direct competitor exists" claims = LOW-MEDIUM)
|
| 6 |
+
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
## Competitive Landscape (context for the categories below)
|
| 10 |
+
|
| 11 |
+
Four distinct product families overlap on this project, and each contributes a different set of user expectations:
|
| 12 |
+
|
| 13 |
+
| Family | Examples | What learners expect from it |
|
| 14 |
+
|--------|----------|------------------------------|
|
| 15 |
+
| **AI speaking tutors** | Langua/LanguaTalk, TalkPal, Speak, Talkio, ISSEN | Natural voice conversation, in-line corrections, post-session feedback report, roleplay + debate topics, saved-vocab SRS |
|
| 16 |
+
| **Avatar/character conversation** | Duolingo Video Call ("Lily"), Novatar AI, AITOMO | A face that reacts, remembers you, and animates expressively while talking |
|
| 17 |
+
| **Japanese SRS/reference apps** | WaniKani, Bunpro, Renshuu, Anki | JLPT-tagged content, spaced repetition, furigana, correct readings, grammar explanations that are *right* |
|
| 18 |
+
| **Structured courseware** | Busuu, Speak, Duolingo core | Placement test, level progression, lesson units, measurable progress |
|
| 19 |
+
|
| 20 |
+
**Key competitive finding:** the avatar family and the pedagogy family barely intersect. Duolingo's Lily is expressive and memorable but is a Max-tier add-on grafted onto a gamified course, with Japanese support only recently added. The anime-avatar apps (Novatar AI, AITOMO — both VRM-capable) are companion/roleplay products where language practice is a marketing angle, not a curriculum. Langua has the deepest pedagogy (multiple correction styles, feedback reports, mistake-driven drills, SRS) but **no avatar at all**. That gap — *serious JLPT-grounded pedagogy behind an expressive VRM face* — is where this project's differentiation lives. (Confidence: MEDIUM — based on published feature pages and reviews, not exhaustive market scan.)
|
| 21 |
+
|
| 22 |
+
---
|
| 23 |
+
|
| 24 |
+
## Feature Landscape
|
| 25 |
+
|
| 26 |
+
### Table Stakes (Users Expect These)
|
| 27 |
+
|
| 28 |
+
Missing any of these makes the product feel broken, not minimal. Split into three groups because this product inherits expectations from three categories at once.
|
| 29 |
+
|
| 30 |
+
#### A. Voice conversation loop
|
| 31 |
+
|
| 32 |
+
| Feature | Why Expected | Complexity | Notes |
|
| 33 |
+
|---------|--------------|------------|-------|
|
| 34 |
+
| Mic capture + Japanese ASR with visible transcript | Every AI speaking app shows you what it heard; without it a mis-transcription looks like the tutor being stupid | MEDIUM | Whisper-class model. Show the transcript *before* the reply so mis-hears are attributable. Confidence score should be exposed internally to drive avatar "confused" expression |
|
| 35 |
+
| Avatar speaks Japanese aloud (TTS) | Core stated value; a silent tutor is a chatbot | MEDIUM | Japanese TTS quality varies wildly; pitch-accent-aware voices matter (see differentiators) |
|
| 36 |
+
| Lip-sync synchronized to TTS audio | An animated face whose mouth doesn't match is *worse* than a static image — uncanny and immediately reads as broken | MEDIUM | Japanese is mora-timed with 5 vowel visemes (a/i/u/e/o) + closed — genuinely simpler than English. `@pixiv/three-vrm` exposes VRM blendshape presets `aa/ih/ou/ee/oh` that map 1:1 |
|
| 37 |
+
| Perceived response latency under ~1.5s | Human turn-handoff is ~200ms; >500ms is noticeable, >1500ms "conversations feel broken" | **HIGH** | The single biggest make-or-break on HF free CPU. Mitigate with streaming TTS, first-sentence-first playback, and avatar "thinking" animation to cover the gap |
|
| 38 |
+
| Replay last utterance / "say that again" | Learners miss things constantly; without replay they abandon | LOW | Cache last audio blob |
|
| 39 |
+
| "Speak slower" control | Standard in Langua, expected by N5-N4 learners; JLPT N5/N4 listening is explicitly defined as "spoken slowly" | LOW | TTS rate parameter; persist as a user preference, not a per-turn button press |
|
| 40 |
+
| Text input fallback | Mic permission denial, noisy environments, shared/office machines, HF Space visitors on locked-down browsers | LOW | Also the graceful degradation path when ASR is cold-starting |
|
| 41 |
+
| Tap-to-interrupt / stop speaking | If the avatar starts a 20-second monologue with no stop button, users close the tab | LOW | Note: this is *not* full barge-in (see anti-features) |
|
| 42 |
+
|
| 43 |
+
#### B. Japanese-language-specific table stakes
|
| 44 |
+
|
| 45 |
+
These have no analogue in a Spanish or French tutor. Getting them wrong is what marks a Japanese app as built by someone who doesn't study Japanese.
|
| 46 |
+
|
| 47 |
+
| Feature | Why Expected | Complexity | Notes |
|
| 48 |
+
|---------|--------------|------------|-------|
|
| 49 |
+
| Furigana over kanji, toggleable | Universal in credible Japanese apps; without it N5-N3 learners literally cannot read the transcript | MEDIUM | Needs a morphological tokenizer + reading generator (MeCab/fugashi/SudachiPy class). Toggle states should be: always / above-my-level-only / never — the middle option is what serious learners want |
|
| 50 |
+
| Click-any-word → reading, meaning, JLPT level | Direct expectation set by Langua ("full transcripts with click-to-translate any word") and every Japanese reader tool | MEDIUM | Requires the same tokenizer + a JMdict-class dictionary. Do NOT ask the LLM for definitions — deterministic dictionary lookup is faster, cheaper, and correct |
|
| 51 |
+
| Full-sentence translation on demand (hidden by default) | Comprehensible-input principle: translation must be available but not the default reading path | LOW | One click per line; never auto-show |
|
| 52 |
+
| Correct politeness register for the situation | Using casual speech in a shop roleplay teaches an actively harmful habit; Japanese learners notice immediately | MEDIUM | System-prompt + scenario metadata must pin register. Also needs to be *visible* to the learner (see differentiators) |
|
| 53 |
+
| JLPT-level-appropriate output | The reason this exists over ChatGPT. An N5 learner drowned in N2 vocabulary quits | HIGH | Hardest table-stakes item. Prompting alone is unreliable — needs post-generation vocabulary/kanji gating against JLPT lists |
|
| 54 |
+
|
| 55 |
+
#### C. Learning-system table stakes
|
| 56 |
+
|
| 57 |
+
| Feature | Why Expected | Complexity | Notes |
|
| 58 |
+
|---------|--------------|------------|-------|
|
| 59 |
+
| Corrections of learner's Japanese with a short explanation | The defining feature of the category; TalkPal, Langua, Speak, Busuu all do it | MEDIUM-HIGH | Correction *quality* is the risk, not the mechanism — see anti-features on LLM grammar hallucination |
|
| 60 |
+
| Choice of correction timing/style | Langua ships "written or verbal, subtle or explicit"; interrupting every error kills conversational flow, never correcting kills learning | MEDIUM | Recommended default: non-blocking inline marker during conversation + full report after. Research consensus is this hybrid |
|
| 61 |
+
| End-of-session feedback report | Universal across Langua, Busuu, Speak. Lets the learner focus on speaking during the call and reflect after | MEDIUM | Errors + rule + corrected form + new vocab encountered |
|
| 62 |
+
| Level selection at onboarding (N5/N4/N3/N2) | Users will not tolerate being spoken to at the wrong level for even one session | LOW | Self-declaration is table stakes; *assessment* is a differentiator |
|
| 63 |
+
| Saved vocabulary list + spaced-repetition review | Langua saves words from conversation into an SRS queue; Bunpro/WaniKani/Renshuu users assume SRS exists | MEDIUM | Use FSRS, do not invent a scheduler (see anti-features) |
|
| 64 |
+
| Structured drills separate from conversation (kana/vocab/grammar) | Explicit project requirement; also how Japanese learners actually study — the standard stack is "grammar course + SRS + speaking tool" | MEDIUM | This project's pitch is collapsing that three-app stack into one |
|
| 65 |
+
| Roleplay scenario library with a stated goal | Ordering food / directions / etc. is the canonical demo. Scenario design research: scenarios need explicit goals and exit conditions, not open-ended chat | MEDIUM | See "Role-play scenario design" section below |
|
| 66 |
+
| Persistent progress across sessions and devices | "Measurably improve over time because the avatar remembers them" is the stated Core Value; browser-local storage fails the moment the HF Space restarts or the user switches machines | MEDIUM | HF OAuth + hosted Postgres |
|
| 67 |
+
| Zero-setup working demo (no API key, no signup, no payment) | Portfolio constraint: a hiring manager or casual visitor who hits a key-entry wall bounces in 5 seconds | MEDIUM | Anonymous session must reach a talking avatar within seconds; account only needed to *persist* |
|
| 68 |
+
| Mobile-responsive layout | Half of HF Space traffic from social links is mobile; a broken canvas layout reads as unfinished | MEDIUM | 3D canvas sizing + touch push-to-talk are the risk areas |
|
| 69 |
+
|
| 70 |
+
---
|
| 71 |
+
|
| 72 |
+
### Differentiators (Competitive Advantage)
|
| 73 |
+
|
| 74 |
+
Ordered roughly by differentiation-per-unit-effort. The top five are where this product wins.
|
| 75 |
+
|
| 76 |
+
| Feature | Value Proposition | Complexity | Notes |
|
| 77 |
+
|---------|-------------------|------------|-------|
|
| 78 |
+
| **Pedagogically-driven avatar expression** — expression is a function of *tutor state*, not sentiment of text | This is the actual novel idea. Avatar tilts head/looks puzzled when ASR confidence is low, brightens on a correct N-level-appropriate answer, adopts a listening pose during your turn, thinks visibly during latency. Converts unavoidable latency into character | MEDIUM | VRM expression presets (`happy/angry/sad/relaxed/surprised`) + blink/saccade/idle sway. The mapping table (tutor state → expression + intensity + duration) is the design artifact. Cheap to build, high perceived value |
|
| 79 |
+
| **JLPT-grounded content spine with hard level gating** | Makes the "adaptive N5→N2" claim real instead of prompt-flavored. Content tagged by level: N5 ≈ 800 words/80 kanji, N4 ≈ 1500/170, N3 ≈ 3750/370, N2 ≈ 6000/380 | HIGH | Deterministic post-check: tokenize the model's draft reply, flag any token above the learner's level, regenerate or auto-gloss. This is the feature that makes a small open model *usable* — a constrained N5 reply is easy; an unconstrained one is where small models fail |
|
| 80 |
+
| **Pitch-accent visualization and feedback** | Almost nobody does this in a conversation product. Prior art is OJAD/Suzuki-kun (U. Tokyo) — pitch contour over kana with accent-nucleus marking. Instantly signals "built by someone who takes Japanese seriously," and is a genuinely strong ML-portfolio artifact | HIGH | Two halves: (a) show target contour on the avatar's line — solvable with an accent dictionary; (b) score the learner's contour via F0 extraction + DTW against the target. Ship (a) first, (b) as a later phase |
|
| 81 |
+
| **Mistake-driven drill generation (the conversation↔drill loop)** | Errors made speaking become tomorrow's SRS cards and targeted grammar drills. Langua ("personalized grammar drills based on your frequent mistakes") and Speak (regenerates sentence patterns targeting your errors) both prove the pattern; nobody does it with an avatar | MEDIUM | This is the connective tissue that makes conversation + drills one product rather than two tabs. Highest-value non-avatar differentiator |
|
| 82 |
+
| **Conversational level assessment (no quiz)** | Busuu-style adaptive placement takes ~5 min of multiple choice. Doing it as a 3-minute *chat* where the avatar probes upward and stops when you falter is both better UX and a better demo | MEDIUM | Report an honest band ("solid N4, reaching N3") rather than a hard label — matches how credible JLPT estimators behave. Re-estimate continuously from live conversation performance |
|
| 83 |
+
| **Explicit politeness-register dial (タメ口 / 丁寧語 / 敬語)** | Japanese-specific, pedagogically important, and near-absent from AI tutors. Learner picks register, or the scenario dictates it (convenience store = 敬語 from staff, 丁寧 from you), and register errors are corrected as errors | MEDIUM | Also a great avatar hook: the same character speaking casually vs. formally, with matching posture/bow animation |
|
| 84 |
+
| **Avatar memory of the learner** | Duolingo ships "she remembers what you've said in past calls"; Langua ships "smart memory (can be disabled)". Directly serves the Core Value | MEDIUM | Structured memory (interests, name, weak grammar points, last topics) — not raw transcript stuffing. Must be user-visible and user-deletable |
|
| 85 |
+
| **In-character immersion mode** | Avatar stays in the scenario role and won't drop to English/explanation unless you press a dedicated "break character" button. Real immersion that most AI tutors break the moment you struggle | LOW | Cheap: prompt discipline + a UI affordance. Big perceived difference |
|
| 86 |
+
| **Hybrid free / BYOK frontier backend with visible delta** | Everyone gets a working tutor; power users unlock frontier quality with their own key. As a portfolio piece it demonstrates model-routing and cost-aware design | MEDIUM | Surface the difference honestly ("running on the free local model — add a key for richer corrections") rather than hiding it |
|
| 87 |
+
| **Shadowing mode** | Listen → repeat → scored on timing/rhythm/pitch. The single highest-ROI Japanese speaking exercise, rarely productized | MEDIUM | Reuses the pronunciation-scoring pipeline; strong standalone demo |
|
| 88 |
+
| **Engineering transparency panel** | Portfolio-specific: an optional panel showing ASR model + latency breakdown (ASR / LLM / TTS / render), token counts, level-gating decisions. Consumer apps never do this; a Google AI/ML hiring manager will read it | LOW | Directly serves the "would a hiring manager be impressed" quality bar at very low cost |
|
| 89 |
+
| **Anki/CSV export of saved vocabulary** | Serious Japanese learners already have an Anki deck; export removes lock-in anxiety and signals respect for their workflow | LOW | Trivial to build, disproportionate trust value |
|
| 90 |
+
| **Post-session audio summary** | Langua ships audio recaps for hands-free review. Fits the avatar naturally — the character recaps your session in level-appropriate Japanese | LOW | Reuses TTS; nice closing beat to a session |
|
| 91 |
+
|
| 92 |
+
---
|
| 93 |
+
|
| 94 |
+
### Anti-Features (Commonly Requested, Often Problematic)
|
| 95 |
+
|
| 96 |
+
| Feature | Why Requested | Why Problematic | Alternative |
|
| 97 |
+
|---------|---------------|-----------------|-------------|
|
| 98 |
+
| **Streaks, XP, leaderboards, daily-goal gamification** | Duolingo made it the genre default; looks like retention | Already Out of Scope in PROJECT.md, and the evidence supports that: streak mechanics are "psychologically effective at producing consistency without producing proficiency," with documented streak anxiety and compulsive checking. For a personal-use tool it actively corrupts the signal | Neutral, honest progress: vocabulary retained, JLPT band trend, conversation minutes, error-rate-over-time. Recognition without coercion |
|
| 99 |
+
| **Free-form LLM grammar explanations** | Cheapest way to answer "why is it が not は?" | Documented failure mode: LLMs hallucinate Japanese grammar, assert ungrammatical particle combinations as natural, and are inconsistent when challenged. Also: LLMs over-correct learner text, rewriting correct-but-simple sentences into fluent ones the learner didn't intend. A tutor that teaches wrong Japanese is worse than no tutor | Curated grammar-point database (Bunpro-style, JLPT-tagged) as the source of truth; the LLM *selects and delivers* an entry in friendly wording but never authors the rule. Cite the grammar point ID in the UI |
|
| 100 |
+
| **True full-duplex barge-in (interrupt mid-sentence)** | Feels like "real conversation"; voice-agent industry treats it as standard | Requires VAD tuning, echo cancellation, prosody-aware end-of-turn models, and a co-located low-latency stack — unachievable on HF free CPU. Worse pedagogically: learners at N5-N4 *need* the full utterance, and half-heard sentences produce bad ASR and bad corrections | Half-duplex push-to-talk (or tap-to-talk) plus an explicit stop button. Frame it as a feature ("listen to the whole line, then reply") rather than a limitation. Revisit only if a realtime speech-to-speech backend is adopted |
|
| 101 |
+
| **Romaji-first interface** | Lowers the barrier for absolute beginners | Romaji is a documented crutch that delays kana acquisition and produces wrong pronunciation habits; serious Japanese apps treat it as beginner-only scaffolding | Kana/kanji + furigana by default; romaji available only as an on-demand per-word reveal, auto-disabled once the learner clears the kana module |
|
| 102 |
+
| **Photorealistic talking-head video generation** | "More impressive" avatar | Already Out of Scope. GPU-heavy, seconds of latency per utterance, and fails the free-tier constraint. Also uncanny-valley risk that stylized VRM avoids entirely | Stylized VRM at 60fps in-browser. Perceived responsiveness beats perceived realism for a tutor |
|
| 103 |
+
| **AI companion / romance / parasocial character drift** | The anime-avatar apps in this space (AITOMO, Novatar AI, Character.AI clones) monetize exactly this, and users will steer toward it | Destroys pedagogical credibility, wrecks the portfolio framing for a Google application, and creates content-moderation obligations on a public HF Space | Warm, encouraging *teacher* persona with explicit boundaries in the system prompt and a refusal path that redirects to a lesson. Personality yes; companion no |
|
| 104 |
+
| **Community/social features (friends, leaderboards, shared decks)** | Standard growth lever | Already Out of Scope. Needs moderation, spam handling, and a userbase this will not have | Optional single-player sharing: export a session summary or a progress card image |
|
| 105 |
+
| **Custom SRS algorithm** | "Our scheduler is tuned for Japanese" | Nobody will believe it and it will be worse. FSRS is open-source, trained on ~700M reviews, and cuts required reviews 20-30% vs SM-2 at 90% retention; it is Anki's default for new profiles from 23.12 | Adopt an FSRS implementation, expose the target-retention knob, and spend the saved effort on content quality |
|
| 106 |
+
| **Always-on microphone / auto-listening** | Most natural interaction | Privacy perception on a public HF Space is severe; also burns compute on ambient noise and produces phantom turns | Explicit push-to-talk with a visible recording indicator and a clear "audio is processed, not stored" statement |
|
| 107 |
+
| **JLPT N1 support** | Completeness | N1 needs 10,000+ words and 1,136 kanji, and its content (editorials, abstract argumentation) is exactly where a small open model produces subtly wrong Japanese. Scope explosion for a user who is nowhere near it | Ship N5→N2 well. N1 is a post-validation decision |
|
| 108 |
+
| **Kanji handwriting recognition** | Frequently requested by Japanese learners | High implementation cost (stroke-order model + canvas UX), orthogonal to a *speaking* product, and not on the JLPT (which has no writing section) | Recognition-only kanji drills (reading/meaning); defer production writing entirely |
|
| 109 |
+
| **Multi-language expansion / other languages** | Obvious market logic | Already Out of Scope, and correctly: furigana, pitch accent, politeness registers, morphological tokenization, and JLPT gating are all Japanese-specific. Generalizing would delete every differentiator | Single-language depth as the explicit positioning |
|
| 110 |
+
| **Auto-playing translation under every line** | "Helpful" | Kills comprehensible-input value; learners read English and stop processing Japanese | Translation always one click away, never shown unrequested |
|
| 111 |
+
|
| 112 |
+
---
|
| 113 |
+
|
| 114 |
+
## Feature Dependencies
|
| 115 |
+
|
| 116 |
+
```
|
| 117 |
+
[Japanese morphological tokenizer + dictionary] <- foundational, blocks a lot
|
| 118 |
+
├──enables──> [Furigana rendering]
|
| 119 |
+
├──enables──> [Click-any-word lookup]
|
| 120 |
+
├──enables──> [JLPT level gating of model output]
|
| 121 |
+
│ └──enables──> [Adaptive N5->N2 difficulty]
|
| 122 |
+
│ └──enables──> [Conversational level assessment]
|
| 123 |
+
└──enables──> [Vocabulary tracking / "words seen"]
|
| 124 |
+
└──enables──> [SRS review queue (FSRS)]
|
| 125 |
+
└──enables──> [Mistake-driven drill generation]
|
| 126 |
+
|
| 127 |
+
[Mic capture + Japanese ASR]
|
| 128 |
+
├──requires──> [Text input fallback] (degradation path, build together)
|
| 129 |
+
├──enables──> [Corrections of learner Japanese]
|
| 130 |
+
│ └──enables──> [End-of-session feedback report]
|
| 131 |
+
│ └──enables──> [Mistake-driven drill generation]
|
| 132 |
+
└──enables──> [Pronunciation scoring]
|
| 133 |
+
├──enables──> [Shadowing mode]
|
| 134 |
+
└──enables──> [Pitch-accent feedback] (also needs F0 extraction + accent dictionary)
|
| 135 |
+
|
| 136 |
+
[TTS with timing/viseme data]
|
| 137 |
+
└──requires──> [Lip-sync] HARD DEPENDENCY: no phoneme timings => no lip-sync => avatar reads as broken
|
| 138 |
+
└──enables──> [Speak-slower control]
|
| 139 |
+
└──enables──> [Post-session audio summary]
|
| 140 |
+
|
| 141 |
+
[VRM avatar rendering in-browser]
|
| 142 |
+
├──requires──> [Lip-sync] to be credible
|
| 143 |
+
└──enables──> [Pedagogically-driven expression system]
|
| 144 |
+
└──enhances──> [In-character immersion mode]
|
| 145 |
+
└──enhances──> [Latency perception] (thinking animation masks response time)
|
| 146 |
+
|
| 147 |
+
[HF OAuth accounts + server DB]
|
| 148 |
+
└──enables──> [Persistent progress]
|
| 149 |
+
├──enables──> [SRS scheduling across sessions]
|
| 150 |
+
├──enables──> [Avatar memory of the learner]
|
| 151 |
+
└──enables──> [Progress trend / JLPT band over time]
|
| 152 |
+
|
| 153 |
+
[Roleplay scenario schema (goal, register, level, exit condition)]
|
| 154 |
+
├──requires──> [JLPT level gating] (scenario must be renderable at the learner's level)
|
| 155 |
+
├──requires──> [Politeness register control]
|
| 156 |
+
└──enables──> [Scenario library]
|
| 157 |
+
└──enables──> [In-character immersion mode]
|
| 158 |
+
|
| 159 |
+
[Streaks/XP] ──conflicts──> [Honest progress metrics] (deliberately excluded)
|
| 160 |
+
[Full-duplex barge-in] ──conflicts──> [HF free-CPU latency budget]
|
| 161 |
+
[Romaji-first UI] ──conflicts──> [Kana acquisition module]
|
| 162 |
+
[LLM-authored grammar rules] ──conflicts──> [Correction trustworthiness]
|
| 163 |
+
```
|
| 164 |
+
|
| 165 |
+
### Dependency Notes
|
| 166 |
+
|
| 167 |
+
- **Tokenizer is the true foundation.** Furigana, word lookup, level gating, and vocabulary tracking all collapse without it. It should land in an early phase even though it is invisible on screen. It is also pure CPU — no model cost — so it is free capability on the HF free tier.
|
| 168 |
+
- **TTS choice constrains lip-sync, not the other way around.** If the selected TTS cannot emit phoneme/mora timings (or force-alignable audio), lip-sync degrades to amplitude-driven jaw flapping, which for Japanese's 5-vowel viseme set is a visible downgrade. Treat "does it give timings" as a hard TTS selection criterion, not a nice-to-have.
|
| 169 |
+
- **Level gating gates everything adaptive.** "Adaptive N5→N2" is not a model-prompting feature; it is a content-classification feature. Assessment, scenario selection, drill generation, and correction severity all read from the same level model.
|
| 170 |
+
- **Persistence gates the Core Value.** SRS, memory, and progress trend are all meaningless without server-side state — but the *anonymous* path must still work, so state has to be designed as "session state that can be adopted by an account on login," not "state requires an account."
|
| 171 |
+
- **Pronunciation scoring is one pipeline serving three features.** Build the F0/alignment pipeline once; shadowing, pitch-accent feedback, and per-mora pronunciation scores are three presentations of it. This changes the cost calculus — individually each looks expensive, together they amortize.
|
| 172 |
+
- **Expression system depends on tutor state being modeled explicitly.** If correction/confidence/level-fit are only implicit in prompt text, there is nothing for the avatar to react to. The conversation engine must emit structured state (`asr_confidence`, `error_severity`, `level_fit`, `turn_role`) as a first-class output.
|
| 173 |
+
|
| 174 |
+
---
|
| 175 |
+
|
| 176 |
+
## MVP Definition
|
| 177 |
+
|
| 178 |
+
### Launch With (v1)
|
| 179 |
+
|
| 180 |
+
The v1 test: *a stranger opens the Space, presses talk, says something in Japanese, and an anime character answers them at their level with a moving mouth, in under two seconds.* Everything that doesn't serve that sentence is v1.x.
|
| 181 |
+
|
| 182 |
+
- [ ] **VRM avatar rendering with idle/blink/gaze** — the product's face; a static model reads as a screenshot
|
| 183 |
+
- [ ] **TTS with lip-sync** — non-negotiable pairing; either both or neither
|
| 184 |
+
- [ ] **Push-to-talk mic → Japanese ASR → visible transcript** — the input half of the loop
|
| 185 |
+
- [ ] **Free conversation with level-appropriate replies (self-declared N5-N2)** — the core loop
|
| 186 |
+
- [ ] **Japanese tokenizer + furigana + click-to-lookup** — makes the transcript readable for the actual target user
|
| 187 |
+
- [ ] **JLPT level gating of output vocabulary/kanji** — what makes the small free model viable and the level claim honest
|
| 188 |
+
- [ ] **Corrections with grammar-DB-backed explanations** — the tutoring, not just the chatting
|
| 189 |
+
- [ ] **3-5 roleplay scenarios with goal + register + exit condition** — the immersion demo, and the most screenshot-able feature
|
| 190 |
+
- [ ] **HF OAuth account + server-persisted vocabulary/level/mistakes** — Core Value depends on memory
|
| 191 |
+
- [ ] **Basic expression mapping (thinking / listening / pleased / puzzled)** — cheap, and it is the whole reason for an avatar
|
| 192 |
+
- [ ] **Text-input fallback + mobile-responsive layout** — so the Space works for everyone who clicks the link
|
| 193 |
+
- [ ] **Anonymous zero-setup path** — hiring managers and casual visitors will not sign in
|
| 194 |
+
|
| 195 |
+
### Add After Validation (v1.x)
|
| 196 |
+
|
| 197 |
+
- [ ] **FSRS SRS review queue over saved vocabulary** — trigger: learner has accumulated 50+ words across sessions and asks "where do I review these?"
|
| 198 |
+
- [ ] **Mistake-driven drill generation** — trigger: error logs contain enough repeat patterns to be worth targeting (needs v1 persistence data first)
|
| 199 |
+
- [ ] **Structured kana/vocab/grammar drill modules** — trigger: conversation-only proves too unstructured for daily use
|
| 200 |
+
- [ ] **Conversational level assessment** — trigger: self-declared level proves inaccurate in practice (it will)
|
| 201 |
+
- [ ] **End-of-session feedback report + audio summary** — trigger: sessions get long enough that in-line corrections are forgotten
|
| 202 |
+
- [ ] **BYOK frontier model toggle** — trigger: free-model correction quality is visibly the bottleneck
|
| 203 |
+
- [ ] **Politeness register dial** — trigger: roleplay scenarios expose register as the recurring correction theme
|
| 204 |
+
- [ ] **Scenario library expansion (15-25 scenarios) with progressive difficulty** — trigger: the initial 3-5 get exhausted
|
| 205 |
+
- [ ] **Avatar memory of learner (interests, weak points)** — trigger: returning users expect continuity
|
| 206 |
+
- [ ] **Engineering transparency panel** — trigger: portfolio submission moment
|
| 207 |
+
- [ ] **Anki/CSV export** — trigger: first "can I get these into my deck?" request (or before showing it to any serious learner)
|
| 208 |
+
|
| 209 |
+
### Future Consideration (v2+)
|
| 210 |
+
|
| 211 |
+
- [ ] **Pitch-accent visualization and scoring** — defer: highest-effort differentiator; needs F0 pipeline + accent dictionary. Ship visualization-only first if it can be split
|
| 212 |
+
- [ ] **Shadowing mode** — defer: depends on the pronunciation pipeline landing
|
| 213 |
+
- [ ] **Per-mora pronunciation scoring** — defer: forced alignment on free CPU is a real research task, and imperfect scoring is worse than none (learners lose trust fast)
|
| 214 |
+
- [ ] **Full-body avatar animation / gestures / bowing** — defer: expression + lip-sync captures most of the value; gestures are polish
|
| 215 |
+
- [ ] **N1 content** — defer: needs both content scale and model quality this stack won't have
|
| 216 |
+
- [ ] **Listening-comprehension mode (avatar narrates, learner answers)** — defer: valuable but a separate loop
|
| 217 |
+
- [ ] **Multiple avatar characters / voices** — defer: pure polish until one character is excellent
|
| 218 |
+
|
| 219 |
+
---
|
| 220 |
+
|
| 221 |
+
## Feature Prioritization Matrix
|
| 222 |
+
|
| 223 |
+
| Feature | User Value | Implementation Cost | Priority |
|
| 224 |
+
|---------|------------|---------------------|----------|
|
| 225 |
+
| VRM avatar + lip-sync + TTS | HIGH | MEDIUM | P1 |
|
| 226 |
+
| Mic → ASR → transcript | HIGH | MEDIUM | P1 |
|
| 227 |
+
| Level-appropriate conversation replies | HIGH | MEDIUM | P1 |
|
| 228 |
+
| Tokenizer + furigana + word lookup | HIGH | MEDIUM | P1 |
|
| 229 |
+
| JLPT level gating of output | HIGH | HIGH | P1 |
|
| 230 |
+
| Corrections w/ grammar-DB backing | HIGH | MEDIUM | P1 |
|
| 231 |
+
| Latency under ~1.5s perceived | HIGH | HIGH | P1 |
|
| 232 |
+
| Roleplay scenarios (3-5, with goals) | HIGH | MEDIUM | P1 |
|
| 233 |
+
| Accounts + persistent progress | HIGH | MEDIUM | P1 |
|
| 234 |
+
| Basic expression mapping | MEDIUM | LOW | P1 |
|
| 235 |
+
| Anonymous zero-setup path | HIGH | MEDIUM | P1 |
|
| 236 |
+
| Text fallback + mobile responsive | MEDIUM | LOW | P1 |
|
| 237 |
+
| FSRS SRS queue | HIGH | MEDIUM | P2 |
|
| 238 |
+
| Mistake-driven drill generation | HIGH | MEDIUM | P2 |
|
| 239 |
+
| Structured kana/vocab/grammar drills | HIGH | MEDIUM | P2 |
|
| 240 |
+
| End-of-session feedback report | MEDIUM | MEDIUM | P2 |
|
| 241 |
+
| Conversational level assessment | MEDIUM | MEDIUM | P2 |
|
| 242 |
+
| Politeness register dial | MEDIUM | MEDIUM | P2 |
|
| 243 |
+
| BYOK frontier toggle | MEDIUM | MEDIUM | P2 |
|
| 244 |
+
| Avatar memory of learner | MEDIUM | MEDIUM | P2 |
|
| 245 |
+
| In-character immersion mode | MEDIUM | LOW | P2 |
|
| 246 |
+
| Engineering transparency panel | MEDIUM (portfolio: HIGH) | LOW | P2 |
|
| 247 |
+
| Anki/CSV export | LOW (trust: MEDIUM) | LOW | P2 |
|
| 248 |
+
| Post-session audio summary | LOW | LOW | P3 |
|
| 249 |
+
| Pitch-accent visualization | MEDIUM (portfolio: HIGH) | HIGH | P3 |
|
| 250 |
+
| Pitch-accent / mora-level scoring | MEDIUM | HIGH | P3 |
|
| 251 |
+
| Shadowing mode | MEDIUM | MEDIUM | P3 |
|
| 252 |
+
| Full-body gestures / bowing | LOW | MEDIUM | P3 |
|
| 253 |
+
| Multiple characters/voices | LOW | MEDIUM | P3 |
|
| 254 |
+
|
| 255 |
+
**Priority key:** P1 = must have for launch · P2 = should have, add when possible · P3 = nice to have, future
|
| 256 |
+
|
| 257 |
+
---
|
| 258 |
+
|
| 259 |
+
## Deep Dives on the Requested Dimensions
|
| 260 |
+
|
| 261 |
+
### Conversation practice UX (corrections, translations, hints)
|
| 262 |
+
|
| 263 |
+
The category has converged on a **dual-timing correction model**: non-intrusive signals during the conversation plus a comprehensive report after. Langua explicitly offers both, and lets the user choose written vs. verbal and subtle vs. explicit; Busuu corrects at conversation end; TalkPal explains the underlying rule in plain language when you err.
|
| 264 |
+
|
| 265 |
+
Recommended defaults for this project:
|
| 266 |
+
- **During conversation:** a subtle inline marker (underline on the learner's transcript line + a small count badge). No audio interruption, no modal. The avatar may *react* to an error with expression rather than words — a distinctive use of the avatar.
|
| 267 |
+
- **On demand:** tapping the marker expands correction → rule name → corrected form.
|
| 268 |
+
- **After session:** grouped report (particle errors, verb conjugation, register mismatches), each linked to a grammar-DB entry and each spawning an SRS card.
|
| 269 |
+
- **Hints:** a single "stuck" affordance (Langua uses a lightbulb) offering 2-3 level-appropriate suggested responses. Critically, also accept **English input mid-conversation** — Langua supports switching to your native language when stuck, and it prevents the conversation dying.
|
| 270 |
+
- **Translations:** always one click, never automatic. Word-level from dictionary lookup; sentence-level from the model.
|
| 271 |
+
|
| 272 |
+
### Drill / SRS systems
|
| 273 |
+
|
| 274 |
+
The Japanese-learning market has already settled this: WaniKani for kanji (radicals + mnemonics + SRS), Bunpro for JLPT-ordered grammar SRS with example sentences, Renshuu as the free all-rounder, Anki as the substrate. Learners currently run a **three-app stack** (grammar course + SRS + speaking tool). The strategic opportunity is not to out-SRS WaniKani — it is to be the one place where *the speaking tool feeds the SRS*.
|
| 275 |
+
|
| 276 |
+
Implementation stance: adopt **FSRS** (stability/difficulty/retrievability model, mean-reverting difficulty which eliminates SM-2's "ease hell", ~20-30% fewer reviews at 90% retention). Expose target retention; do not expose interval math. Card types worth having: recognition (kanji→reading/meaning), production (meaning→Japanese), audio (hear→meaning), and — uniquely enabled by this product — **spoken production** (prompt → say it aloud → ASR-checked).
|
| 277 |
+
|
| 278 |
+
### Role-play scenario design
|
| 279 |
+
|
| 280 |
+
Scenario-design research is consistent: define a **specific measurable goal**, set **explicit exit conditions** (resolution reached / N exchanges / outcome achieved), scale difficulty **progressively** (start cooperative, get harder as the learner handles early turns), and give feedback that explains *why* something worked rather than right/wrong. A scenario library beats a blank prompt box.
|
| 281 |
+
|
| 282 |
+
Concrete schema this implies:
|
| 283 |
+
|
| 284 |
+
```
|
| 285 |
+
scenario: { id, title_ja, title_en, jlpt_level_range,
|
| 286 |
+
setting, avatar_role, learner_role,
|
| 287 |
+
learner_goal (measurable), success_criteria,
|
| 288 |
+
required_register (casual/teineigo/keigo),
|
| 289 |
+
target_grammar_points[], target_vocab[],
|
| 290 |
+
exit_conditions[], difficulty_escalation_rules }
|
| 291 |
+
```
|
| 292 |
+
|
| 293 |
+
Starter set (N5→N2 spread): convenience store checkout, ordering at a restaurant, asking directions, train ticket purchase, clinic reception, apartment viewing, self-introduction at a new workplace, phone call to reschedule, complaining about a wrong order, casual chat with a friend about weekend plans. Note how register varies across that list — that is deliberate and is the politeness-dial hook.
|
| 294 |
+
|
| 295 |
+
### Level assessment and adaptation (JLPT-aligned)
|
| 296 |
+
|
| 297 |
+
Official JLPT competence definitions are the calibration anchor: N5 = hiragana/katakana/basic kanji, everyday conversation "spoken slowly"; N4 = familiar daily topics, slower speech; N3 = everyday materials, "near-natural speed"; N2 = newspaper/magazine articles, "nearly natural speed" across varied settings. Approximate content scale: N5 ~800 words/80 kanji → N4 ~1500/170 → N3 ~3750/370 → N2 ~6000/380.
|
| 298 |
+
|
| 299 |
+
Two adaptation axes, both needed:
|
| 300 |
+
1. **Content level** — vocabulary/kanji/grammar gating (deterministic, from lists).
|
| 301 |
+
2. **Delivery level** — TTS speed, sentence length, use of English scaffolding, correction verbosity. The JLPT definitions make speech rate an *explicit* level dimension, which most apps ignore.
|
| 302 |
+
|
| 303 |
+
Assessment stance: report an honest band, not a hard label (good JLPT estimators say "around N3, possibly N2"). Re-estimate continuously from conversational evidence — comprehension failures, error rates, response latency — rather than only at a placement gate. Note the honest limitation: a conversation-based estimate measures listening/speaking, while the JLPT measures reading/listening only. Say so in the UI rather than overclaiming.
|
| 304 |
+
|
| 305 |
+
### Pronunciation feedback
|
| 306 |
+
|
| 307 |
+
The state of the art (ELSA-class) is: ASR decomposes audio into phonemes, **forced alignment** against the expected native phoneme sequence, per-phoneme scoring on formant frequencies / voice onset time / spectral characteristics, plus prosodic scoring (intonation, stress, fluency) — surfaced as per-sound highlighting rather than one opaque score.
|
| 308 |
+
|
| 309 |
+
Japanese-specific adaptation: the segmental inventory is small and forgiving, so the discriminating signals are **mora timing**, **long/short vowel and geminate distinctions** (おばさん/おばあさん, きて/きって), and **pitch accent**. Pitch is the highest-value target and the least-served: OJAD's Suzuki-kun (University of Tokyo) is the established prior art — it converts arbitrary Japanese text to kana with a visualized F0 contour and expected accent-nucleus positions, and reads it aloud with those prosodic features. Mirroring that visualization for the avatar's line is achievable; scoring the learner's contour against it is the harder, later step.
|
| 310 |
+
|
| 311 |
+
Trust warning: an inaccurate pronunciation score is worse than no score. Ship a coarse, defensible signal (mora-timing + accent-pattern match) before attempting fine-grained per-phoneme grades.
|
| 312 |
+
|
| 313 |
+
### Progress tracking (without gamification)
|
| 314 |
+
|
| 315 |
+
Given streaks are excluded, progress must still be legible. Metrics that are honest and motivating: words in SRS by maturity, kanji known by JLPT level, estimated JLPT band over time, minutes of speaking, error-rate-by-category trend, scenarios completed at each level, and "grammar points mastered." Duolingo Lily's most defensible mechanic — *she calls you occasionally* — is worth noting as an opt-in reminder rather than a streak; it creates return pressure without loss-aversion mechanics.
|
| 316 |
+
|
| 317 |
+
### Avatar expressiveness
|
| 318 |
+
|
| 319 |
+
Duolingo's Video Call explicitly lists **expressive animations**, transcripts for post-call review, and memory of past calls as its headline enhancements — a useful confirmation of which avatar features actually matter. Research on LLM-driven pedagogical agents finds avatar guides significantly increase engagement versus chatbots or static UI.
|
| 320 |
+
|
| 321 |
+
Concrete expressiveness feature set, cheapest-first:
|
| 322 |
+
1. Idle life: blink, micro-saccades, breathing/sway. Without these the model looks dead in <3 seconds. Near-zero cost, highest perceived-quality return.
|
| 323 |
+
2. Viseme lip-sync from TTS timings (VRM presets `aa/ih/ou/ee/oh` map directly onto Japanese's five vowels — a structural advantage over English).
|
| 324 |
+
3. Tutor-state expressions: listening (attentive lean), thinking (during latency), pleased (correct answer), puzzled (low ASR confidence), gentle correction (slight head tilt).
|
| 325 |
+
4. Gaze targeting: look at the camera when speaking, away when thinking.
|
| 326 |
+
5. Register-linked posture: formal upright vs. relaxed — ties the politeness dial to the visual channel.
|
| 327 |
+
6. (Later) head nods as backchannel while the learner speaks — あいづち is culturally meaningful in Japanese and would be a genuinely distinctive touch.
|
| 328 |
+
|
| 329 |
+
---
|
| 330 |
+
|
| 331 |
+
## Competitor Feature Analysis
|
| 332 |
+
|
| 333 |
+
| Feature | Duolingo (Video Call/Lily) | Langua | Speak / TalkPal | Bunpro / WaniKani / Renshuu | Our Approach |
|
| 334 |
+
|---------|---------------------------|--------|-----------------|------------------------------|--------------|
|
| 335 |
+
| Avatar | 2D animated character (Rive), expressive, Max-tier only | None | None | None | 3D VRM, expression driven by tutor state, free |
|
| 336 |
+
| Voice conversation | Real-time, adapts to level, remembers past calls | Call mode, cloned native voices, slow-speed option | Real-time; Speak strong on structured drills | None | Push-to-talk half-duplex, streaming TTS |
|
| 337 |
+
| Corrections | In-call + transcript review | Written/verbal × subtle/explicit, + post-chat report | TalkPal explains the rule; Speak targets your error patterns | Bunpro: SRS-based grammar correctness | Inline marker + grammar-DB-backed report; avatar reacts expressively |
|
| 338 |
+
| Translation/reading aid | Transcript | Click-to-translate any word, context-aware saves | Varies | Furigana + readings native | Furigana (3-mode toggle) + click lookup + on-demand sentence translation |
|
| 339 |
+
| SRS | Weak/none for conversation vocab | Saves conversation words to SRS flashcards | Speak regenerates patterns from your errors | Core competency, JLPT-ordered | FSRS over conversation-sourced vocab + mistakes |
|
| 340 |
+
| Roleplay | Conversational, scenario-ish | Hundreds of roleplays/debates/chats | Roleplay + Free Talk + Tutor Q&A | None | Small curated set with goals, register, exit conditions, escalation |
|
| 341 |
+
| Level system | CEFR-ish internal | A1/A2 guided course | Level+topic lessons | JLPT-native | JLPT-native N5-N2 with deterministic gating |
|
| 342 |
+
| Pronunciation | Light | Light | Speak/ELSA-class scoring is the selling point | None | Mora-timing + pitch accent (Japanese-specific, later phase) |
|
| 343 |
+
| Gamification | Heavy (streaks, XP, leagues) | Minimal | Light | SRS-driven, mild | None by design; honest metrics only |
|
| 344 |
+
| Cost to start | Max subscription | Paid | Paid | Free tier (Renshuu/Bunpro) | Free, no signup, no key |
|
| 345 |
+
|
| 346 |
+
---
|
| 347 |
+
|
| 348 |
+
## Open Questions for Later Phases
|
| 349 |
+
|
| 350 |
+
- Does the selected Japanese TTS emit phoneme/mora timings? If not, lip-sync fidelity and the whole pitch-accent visualization path change materially. **Blocks:** lip-sync design, pitch features. (For STACK research.)
|
| 351 |
+
- Can Gradio host a custom three.js/VRM canvas cleanly, or does the avatar need a custom component / static HTML mount? This affects whether avatar↔conversation state sync is cheap or a build-out. (For ARCHITECTURE research.)
|
| 352 |
+
- What is the realistic end-to-end latency of ASR + small LLM + TTS on HF free CPU? If it exceeds ~3s, the "thinking animation" mitigation is insufficient and the interaction model may need to change (e.g., turn-based lesson framing rather than conversation framing).
|
| 353 |
+
- Licensing on JLPT-tagged wordlists, grammar-point databases, and pitch-accent dictionaries — several widely-used sources have restrictive or ambiguous terms. Needs resolution before the content spine is built.
|
| 354 |
+
- Whether an anonymous session's progress can be adopted into an account on later HF OAuth login, or whether progress is simply lost for anonymous users (acceptable, but must be stated in the UI).
|
| 355 |
+
|
| 356 |
+
---
|
| 357 |
+
|
| 358 |
+
## Sources
|
| 359 |
+
|
| 360 |
+
**Competitor feature documentation (HIGH confidence — vendor's own feature pages):**
|
| 361 |
+
- Langua AI tutor feature page — https://languatalk.com/ai-language-tutor
|
| 362 |
+
- Duolingo — AI behind Video Call — https://blog.duolingo.com/ai-and-video-call/
|
| 363 |
+
- Duolingo Video Call launch/expansion — https://investors.duolingo.com/news-releases/news-release-details/duolingo-launches-ai-powered-video-call-android
|
| 364 |
+
- Rive on Duolingo's Lily animation — https://rive.app/blog/duolingo-s-ai-powered-video-call-brings-lily-to-life
|
| 365 |
+
- Busuu Japanese placement test — https://www.busuu.com/en/japanese/placement-test
|
| 366 |
+
|
| 367 |
+
**Official standards (HIGH confidence):**
|
| 368 |
+
- JLPT official level summary N1-N5 — https://www.jlpt.jp/e/about/levelsummary.html
|
| 369 |
+
- OJAD / Suzuki-kun prosody tutor (University of Tokyo) — https://www.gavo.t.u-tokyo.ac.jp/ojad/eng/phrasing/index
|
| 370 |
+
- Minematsu et al., "Prosodic Reading Tutor of Japanese, Suzuki-kun" (ISCA) — https://www.isca-archive.org/ssw_2016/minematsu16_ssw.html
|
| 371 |
+
|
| 372 |
+
**Comparative reviews and market analysis (MEDIUM confidence — independent but commercially motivated):**
|
| 373 |
+
- Talkpal vs Langua comparison 2026 — https://www.borderset.com/blogs/posts/talkpal-vs-langua-best-ai-language-learning-app-2026
|
| 374 |
+
- Best AI speaking apps tested 2026 — https://lingtuitive.com/blog/best-ai-speaking-apps
|
| 375 |
+
- Speak app review 2026 — https://languatalk.com/blog/speak-app-review/
|
| 376 |
+
- Best Japanese learning apps 2026 (WaniKani/Bunpro/Renshuu) — https://www.shinobi-japanese.com/blog/best-apps-to-learn-japanese
|
| 377 |
+
- Bunpro review 2026 — https://languavibe.com/bunpro-review/
|
| 378 |
+
- Japanese app comparison — https://immit.co/blog/best-app-for-learning-japanese-in-2026-an-honest-comparison
|
| 379 |
+
|
| 380 |
+
**Technical/pedagogical evidence (MEDIUM-HIGH confidence):**
|
| 381 |
+
- FSRS in Anki — how it works and setup — https://medankigen.com/blog/fsrs-anki
|
| 382 |
+
- FSRS vs SM-2 — https://flica.app/article/fsrs-vs-sm2
|
| 383 |
+
- Voice AI barge-in and turn-taking implementation guide 2026 — https://futureagi.com/blog/voice-ai-barge-in-turn-taking-2026/
|
| 384 |
+
- Latency budgets for real-time voice — https://thepromptbench.com/voice-and-realtime/latency-budgets-for-realtime-voice/
|
| 385 |
+
- ELSA phoneme-level scoring methodology — https://www.buildfastwithai.com/ai-tools/elsa-speak
|
| 386 |
+
- ELSA API / scoring dimensions — https://elsaspeak.com/en/elsa-api/
|
| 387 |
+
- ELEVATE: Human-Centered GenAI Virtual Tutors (arXiv) — https://arxiv.org/pdf/2606.30662
|
| 388 |
+
- Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents (arXiv) — https://arxiv.org/pdf/2603.09536
|
| 389 |
+
- LLM-based intelligent tutoring in VR (MDPI Information) — https://doi.org/10.3390/info16070556
|
| 390 |
+
- TalkingHead (met4citizen) — real-time lip-sync reference implementation — https://github.com/met4citizen/talkinghead
|
| 391 |
+
|
| 392 |
+
**Anti-feature evidence (MEDIUM-HIGH confidence):**
|
| 393 |
+
- "The dangers of using AI to learn Japanese grammar: a case of hallucinating ChatGPT" — https://selftaughtjapanese.com/2025/07/10/the-dangers-of-using-ai-to-learn-japanese-grammar-a-case-of-hallucinating-chatgpt/
|
| 394 |
+
- "Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar" (arXiv) — https://arxiv.org/pdf/2505.19599
|
| 395 |
+
- "What LLMs Must Forget to Teach Effectively" — Japanese pedagogy (arXiv) — https://arxiv.org/pdf/2606.01410
|
| 396 |
+
- LLM grammatical error correction — over-correction findings — https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11524451/
|
| 397 |
+
- "Streak Creep: The perils of too much gamification" (The Decision Lab) — https://thedecisionlab.com/insights/consumer-insights/streak-creep-the-perils-of-too-much-gamification
|
| 398 |
+
- "The Duolingo Effect: Why $748M Revenue Doesn't Equal Best Pedagogy" — https://spellings.app/blog/duolingo-effect
|
| 399 |
+
|
| 400 |
+
**Scenario design (MEDIUM confidence):**
|
| 401 |
+
- AI roleplay practical playbook 2026 — https://www.outdoo.ai/blog/ai-roleplay-guide
|
| 402 |
+
- "How AI roleplay is redesigning language learning experiences" — https://medium.com/ai-product-design/how-ai-roleplay-is-redesigning-language-learning-experience-90522a8a68f1
|
| 403 |
+
|
| 404 |
+
**Japanese content scale (MEDIUM confidence — community consensus, JLPT does not publish official lists):**
|
| 405 |
+
- JLPT kanji counts by level (Kanshudo) — https://www.kanshudo.com/collections/jlpt_kanji
|
| 406 |
+
- JLPT study hours by level (Coto Academy) — https://cotoacademy.com/study-hours-needed-pass-jlpt-comparison-levels/
|
| 407 |
+
- JLPT study schedule guide (Migaku) — https://migaku.com/blog/japanese/jlpt-study-schedule-guide
|
| 408 |
+
|
| 409 |
+
**Adjacent-competitor scan (LOW confidence — thin, promotional sourcing):**
|
| 410 |
+
- Anime AI chat platforms roundup — https://findaichat.com/category/anime-ai-chat-bots
|
| 411 |
+
- Novatar AI (VRM import + language practice) — https://play.google.com/store/apps/details?id=com.storyverse.ai
|
| 412 |
+
|
| 413 |
+
---
|
| 414 |
+
*Feature research for: conversational AI Japanese tutor with 3D VRM avatar on Hugging Face Spaces*
|
| 415 |
+
*Researched: 2026-08-08*
|
.planning/research/PITFALLS.md
ADDED
|
@@ -0,0 +1,592 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Pitfalls Research
|
| 2 |
+
|
| 3 |
+
**Domain:** Avatar-based conversational language tutor (VRM + full-duplex Japanese voice + hybrid LLM + accounts) on Hugging Face Spaces
|
| 4 |
+
**Researched:** 2026-08-08
|
| 5 |
+
**Confidence:** MEDIUM-HIGH (platform limits and Gradio behavior verified against official docs / issue trackers; pedagogy and latency thresholds from research papers and industry benchmarks; a few Japanese-TTS quality numbers are vendor-reported and marked LOW)
|
| 6 |
+
|
| 7 |
+
> **Phase vocabulary used below.** The roadmap doesn't exist yet, so pitfalls are mapped to
|
| 8 |
+
> descriptive phase slots. Suggested ordering (derived from the dependency structure of these pitfalls):
|
| 9 |
+
> **P1 Voice+Avatar Loop Skeleton** → **P2 Japanese Language Core** → **P3 Tutoring Brain (LLM + guardrails)** →
|
| 10 |
+
> **P4 Accounts & Persistence** → **P5 Adaptive Level, Lessons & SRS** → **P6 Deploy / Portfolio Hardening**.
|
| 11 |
+
> The single most important structural recommendation in this document: **P1 must be a latency-and-integration
|
| 12 |
+
> spike that is allowed to fail, not a feature phase.**
|
| 13 |
+
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
## Critical Pitfalls
|
| 17 |
+
|
| 18 |
+
### Pitfall 1: Building the voice loop server-side on free CPU — latency makes it feel broken
|
| 19 |
+
|
| 20 |
+
**What goes wrong:**
|
| 21 |
+
A cascaded ASR → LLM → TTS pipeline runs on the Space's CPU. Each turn takes 5–15s. The avatar is beautiful and the conversation is unusable. This is the single most likely way this project dies: everything "works," and nobody wants to talk to it.
|
| 22 |
+
|
| 23 |
+
Industry latency thresholds for conversational voice: 200–500ms feels natural, 500–800ms is noticeable but acceptable, >800ms is clearly delayed, and **>1.5s users describe the conversation as "broken"** ([Prodinit](https://prodinit.com/blog/production-voice-ai-agents-latency-architecture), [Hamming AI](https://hamming.ai/resources/voice-ai-latency-whats-fast-whats-slow-how-to-fix-it)). Human turn-taking gaps are ~200ms ([AssemblyAI's "300ms rule"](https://www.assemblyai.com/blog/low-latency-voice-ai)). A well-funded cloud pipeline already spends 100–300ms ASR + 350–1000ms LLM + 90–200ms TTS + 50–200ms network. HF Spaces free tier is **2 vCPU / 16GB, no GPU** — every one of those numbers multiplies.
|
| 24 |
+
|
| 25 |
+
**Why it happens:**
|
| 26 |
+
The Gradio mental model ("write a Python function, wire it to a component") makes server-side inference the path of least resistance. Latency is treated as an optimization to do "later," but by then the architecture (Python-side ASR, Python-side TTS, blocking Gradio event handlers) is load-bearing.
|
| 27 |
+
|
| 28 |
+
**How to avoid:**
|
| 29 |
+
- Set an explicit **latency budget as a P1 acceptance gate**, not an aspiration. Suggested: p50 ≤ 2.0s and p95 ≤ 3.5s mic-release-to-first-audio for the free-tier default path. If a design can't hit it, the design is wrong.
|
| 30 |
+
- **Push inference to the browser wherever possible.** Client-side ASR (transformers.js Whisper/kotoba-whisper ONNX, or Web Speech API where available) and client-side or streaming TTS removes the server from the hot loop entirely and makes the free tier scale to N visitors for free. This also side-steps Pitfall 2 and Pitfall 11.
|
| 31 |
+
- **Overlap, don't sequence.** Start ASR on partial audio; stream LLM tokens; start TTS on the first sentence boundary, not the full response. Time-to-*first-audio* is the metric users feel, not total time.
|
| 32 |
+
- **Frame the interaction to hide latency.** Push-to-talk with an explicit "thinking" avatar animation buys ~1s of perceived tolerance that free-running full-duplex does not. Shipping push-to-talk in v1 and full-duplex in v2 is a legitimate, defensible plan.
|
| 33 |
+
- Build a **latency harness in P1** that logs per-stage timings (VAD end → ASR done → LLM first token → TTS first chunk → audio play) and prints them in a debug panel. You cannot optimize what you never measured.
|
| 34 |
+
|
| 35 |
+
**Warning signs:**
|
| 36 |
+
- No measured p50/p95 number exists by the end of the first phase.
|
| 37 |
+
- Timings are only ever measured on the dev laptop, never on the deployed Space.
|
| 38 |
+
- The word "optimize later" appears in a plan for anything in the audio path.
|
| 39 |
+
- Demo videos are edited to cut the pauses.
|
| 40 |
+
|
| 41 |
+
**Phase to address:** P1 (Voice+Avatar Loop Skeleton) — gate the phase on measured latency, on the deployed Space, not locally.
|
| 42 |
+
|
| 43 |
+
---
|
| 44 |
+
|
| 45 |
+
### Pitfall 2: Architecting the core loop on GPU (ZeroGPU) that visitors don't actually have
|
| 46 |
+
|
| 47 |
+
**What goes wrong:**
|
| 48 |
+
The obvious fix for Pitfall 1 is "use ZeroGPU." Then you discover ZeroGPU quota is **per-visitor and tiny**: roughly 2 minutes/day for anonymous visitors, ~3.5–5 minutes/day for free logged-in accounts, ~25 minutes/day for PRO ([HF forums](https://discuss.huggingface.co/t/free-account-zerogpu-quota-issue/175180), [HF ZeroGPU docs](https://huggingface.co/docs/hub/spaces-zerogpu)). A 10-minute Japanese conversation lesson burns a free visitor's entire daily quota in the first two exchanges and then hard-fails mid-lesson — the worst possible failure point.
|
| 49 |
+
|
| 50 |
+
**Why it happens:**
|
| 51 |
+
ZeroGPU is marketed as "free GPU," and quota is measured in GPU-seconds consumed by *the visitor's* account, which is easy to miss until a real user hits it. Local testing with the owner's PRO-ish quota never reproduces it.
|
| 52 |
+
|
| 53 |
+
**How to avoid:**
|
| 54 |
+
- Treat ZeroGPU as a **quality upgrade path, never the default path**. The free default must be CPU-or-browser only and must work with zero GPU seconds.
|
| 55 |
+
- If ZeroGPU is used at all, use it for **bounded, infrequent, non-conversational** work (e.g. a one-time placement-test scoring pass, or generating a lesson's audio assets), never for per-turn inference.
|
| 56 |
+
- Catch and degrade gracefully on quota exhaustion: the avatar should fall back to the CPU/browser path with a visible notice, not throw.
|
| 57 |
+
- The `.planning` decision "hybrid AI backend (small open model + BYOK)" should be explicit that **BYOK is the quality tier and the small model is the always-works tier** — and neither assumes GPU.
|
| 58 |
+
|
| 59 |
+
**Warning signs:**
|
| 60 |
+
- A requirements doc that says "ZeroGPU" without a quota arithmetic calculation next to it.
|
| 61 |
+
- No test where the GPU path is forcibly disabled.
|
| 62 |
+
|
| 63 |
+
**Phase to address:** P1 (architecture decision), verified again in P6 (deploy) with a quota-exhausted simulation.
|
| 64 |
+
|
| 65 |
+
---
|
| 66 |
+
|
| 67 |
+
### Pitfall 3: "Pronunciation scoring" built on Whisper — the model silently corrects the learner's mistakes
|
| 68 |
+
|
| 69 |
+
**What goes wrong:**
|
| 70 |
+
The requirement says "speech recognition transcribes **and scores**." A team wires up Whisper, gets clean Japanese transcripts, and ships a "pronunciation score." It is fiction. Whisper is a *language-model-conditioned* transcriber: it maps sloppy learner audio onto the most probable well-formed Japanese sentence. A learner who says こんにちわ with a flat, wrong pitch and a hard English `r` gets a perfect transcript and a 100% score.
|
| 71 |
+
|
| 72 |
+
This is documented, not speculative: research on Japanese speaking assessment found generic multilingual Whisper "not suitable for phonemic transcription tasks in Japanese speech assessment" and that it **"tended to ignore speaker errors commonly found in natural Japanese speech"** ([arXiv 2509.20655](https://arxiv.org/pdf/2509.20655)). Whisper accuracy also varies substantially by accent/L1 ([JASA Express Letters](https://pubs.aip.org/asa/jel/article/4/2/025206/3267247/Evaluating-OpenAI-s-Whisper-ASR-Performance)).
|
| 73 |
+
|
| 74 |
+
**Why it happens:**
|
| 75 |
+
ASR output *looks* like ground truth. The failure is invisible — the system never errors, it just flatters the learner. And a tutor that always says "great job!" is the definition of a demo that isn't a product, which is fatal given the stated dual purpose (the author actually wants to learn Japanese with it).
|
| 76 |
+
|
| 77 |
+
**How to avoid:**
|
| 78 |
+
- **Do not claim a pronunciation score you cannot defend.** Split the requirement into two: (a) *transcription* for conversation flow — Whisper-class is fine and excellent; (b) *pronunciation assessment* — a separate, explicitly-scoped mechanism.
|
| 79 |
+
- For (b), viable v1 approaches, in increasing order of effort: **constrained/forced decoding** against the expected target utterance (score = alignment cost against the expected mora sequence, so deviations can't be "corrected away"); **phoneme-level acoustic models** or forced alignment (e.g. a CTC phoneme model) rather than a seq2seq LM-conditioned model; goodness-of-pronunciation from frame posteriors. Whisper token log-probs are a weak, noisy proxy — usable only for a coarse "confident / unclear" signal, and should be labeled as such in the UI.
|
| 80 |
+
- **Constrain the task**: pronunciation scoring on a *known target sentence* (drill mode) is tractable; open-ended conversational scoring is a research problem. Ship the former.
|
| 81 |
+
- **Do not attempt pitch-accent scoring in v1** unless it's the explicit centerpiece — it is a specialized subfield with its own quality-assessment literature ([PASQA, arXiv 2606.20137](https://arxiv.org/pdf/2606.20137)).
|
| 82 |
+
- Sanity test: **deliberately mispronounce** 10 utterances (wrong mora, wrong pitch, English phonology) and confirm the score drops. If it doesn't, the feature is broken regardless of what the transcript says.
|
| 83 |
+
|
| 84 |
+
**Warning signs:**
|
| 85 |
+
- Scores cluster at the top for everyone.
|
| 86 |
+
- Nobody on the project can explain, in one sentence, *what signal* the score is computed from.
|
| 87 |
+
- The scoring code path reads `transcript == expected` string equality.
|
| 88 |
+
|
| 89 |
+
**Phase to address:** P2 (Japanese Language Core) — decide the assessment mechanism there; P5 if it feeds adaptive level. Flag P2 as needing **its own deeper research spike**.
|
| 90 |
+
|
| 91 |
+
---
|
| 92 |
+
|
| 93 |
+
### Pitfall 4: Whisper hallucination on silence/noise — the avatar answers things nobody said
|
| 94 |
+
|
| 95 |
+
**What goes wrong:**
|
| 96 |
+
Whisper generates fluent text from non-speech: room tone, breath, a cough, a 2-second silence become 「ご視聴ありがとうございました」 or similar training-data artifacts (this specific Japanese subtitle-boilerplate hallucination is notorious). The avatar then earnestly responds to a sentence the learner never uttered. In a *language tutor* this is worse than in a transcription tool, because a beginner cannot tell whether the weird exchange is their fault. Hallucination causes are documented in model design, training data, and input ambiguity ([Calm-Whisper, arXiv 2505.12969](https://arxiv.org/pdf/2505.12969)).
|
| 97 |
+
|
| 98 |
+
**Why it happens:**
|
| 99 |
+
No VAD gate. The pipeline sends every recorded chunk to ASR, including empty ones. Local testing happens in a quiet room with a good mic; real users are in cafés with laptop mics.
|
| 100 |
+
|
| 101 |
+
**How to avoid:**
|
| 102 |
+
- **VAD gate before ASR** (Silero VAD or equivalent, client-side). No speech energy → no ASR call. This also saves CPU (helps Pitfall 1).
|
| 103 |
+
- **Minimum duration + energy threshold**, and drop transcripts below a confidence/avg-logprob floor.
|
| 104 |
+
- **Blocklist the known hallucination phrases** for Japanese (subtitle/credits boilerplate) — cheap, high value.
|
| 105 |
+
- Prefer a Japanese-specialized distilled model (kotoba-whisper family, ~6.3× faster than large-v3 at comparable Japanese WER — [kotoba-tech/kotoba-whisper-v2.0](https://huggingface.co/kotoba-tech/kotoba-whisper-v2.0)) over generic multilingual Whisper: better Japanese accuracy *and* better latency.
|
| 106 |
+
- Test with 30s of pure silence and 30s of café noise. Expected result: zero avatar turns.
|
| 107 |
+
|
| 108 |
+
**Warning signs:** avatar occasionally replies to nothing; transcripts contain polite boilerplate the user would never say.
|
| 109 |
+
|
| 110 |
+
**Phase to address:** P1 (VAD is part of the audio loop skeleton, not a later polish item).
|
| 111 |
+
|
| 112 |
+
---
|
| 113 |
+
|
| 114 |
+
### Pitfall 5: Acoustic feedback loop — the avatar hears itself and interrupts itself
|
| 115 |
+
|
| 116 |
+
**What goes wrong:**
|
| 117 |
+
Full-duplex means the mic is open while the avatar speaks. Without acoustic echo cancellation, the avatar's own TTS bleeds into the mic, VAD fires, ASR transcribes the avatar's own Japanese, and the system treats it as a learner turn — **phantom interruptions on every single utterance** ([Deepgram](https://developers.deepgram.com/docs/voice-agent-echo-cancellation), [runedge](https://www.runedge.ai/blog/barge-in-interruption-handling-on-device-voice)). On laptop speakers this produces an infinite self-conversation loop. It is invisible during development because the developer wears headphones.
|
| 118 |
+
|
| 119 |
+
**Why it happens:**
|
| 120 |
+
"Full-duplex voice in v1" is listed as a key decision, but full-duplex is an *audio engineering* problem, not an LLM problem, and the audio engineering is usually discovered late.
|
| 121 |
+
|
| 122 |
+
**How to avoid:**
|
| 123 |
+
- Use the browser's built-in AEC: `getUserMedia({ audio: { echoCancellation: true, noiseSuppression: true, autoGainControl: true } })`. Free, and handles the common case — but **only if the TTS audio plays through the same device the browser knows about**; audio played via WebAudio into a separate context can defeat it.
|
| 124 |
+
- **Half-duplex gating as the safety net**: mute/ignore mic input while TTS is playing, with a short (~150–250ms) tail. Combine with explicit barge-in: allow interruption only when post-AEC residual energy clearly exceeds the echo profile.
|
| 125 |
+
- Ship **push-to-talk as the default** and full-duplex as an opt-in toggle ("use headphones for best results"). This converts a class of unfixable field bugs into a user setting.
|
| 126 |
+
- **Test on laptop speakers at high volume, no headphones.** This must be an explicit test case, written down, because every developer will default to headphones.
|
| 127 |
+
|
| 128 |
+
**Warning signs:** the avatar interrupts itself; conversation loops without user input; bugs that "only happen for other people."
|
| 129 |
+
|
| 130 |
+
**Phase to address:** P1. Do not defer — it changes the state machine design.
|
| 131 |
+
|
| 132 |
+
---
|
| 133 |
+
|
| 134 |
+
### Pitfall 6: Lip-sync driven by audio amplitude instead of Japanese phonemes
|
| 135 |
+
|
| 136 |
+
**What goes wrong:**
|
| 137 |
+
The quickest lip-sync is "map RMS volume to the `aa` blendshape." The mouth flaps. It reads as a cheap 2010-era chatbot, and for a *Japanese* tutor it actively destroys value: the learner should be able to watch the mouth to learn mora articulation, and a flapping jaw teaches nothing. Given this is the flagship portfolio piece, the avatar looking cheap is a direct hit to the stated goal.
|
| 138 |
+
|
| 139 |
+
The irony: **Japanese is the easiest major language to lip-sync well.** VRM defines exactly five vowel visemes — `aa / ih / ou / ee / oh` — which map 1:1 onto あ/い/う/え/お, and every Japanese mora reduces to one of them (plus ん and っ). A kana-driven viseme track is nearly free once the reading is known ([VRM viseme model overview](https://primeta.ai/blog/from-phonemes-to-mouth-shapes-animating-vrm-models)).
|
| 140 |
+
|
| 141 |
+
**Why it happens:**
|
| 142 |
+
Amplitude-driven lip-sync is one function; phoneme timing requires the TTS to emit (or the pipeline to derive) a timed phoneme/mora sequence, which most cloud TTS APIs don't return by default.
|
| 143 |
+
|
| 144 |
+
**How to avoid:**
|
| 145 |
+
- **Choose a TTS that can emit timing** (or that you run locally so you can extract it). OpenJTalk-based Japanese engines produce a labeled mora sequence; that timeline *is* your viseme track. Make "returns phoneme/mora timings" a hard selection criterion for TTS in P1 — retrofitting it later means changing TTS engines.
|
| 146 |
+
- Derive the viseme track from **kana**, not from romaji or from the raw kanji text — you need the reading anyway for furigana (Pitfall 8), so this is shared infrastructure.
|
| 147 |
+
- **Drive the viseme clock from `audio.currentTime`**, never from an independent `setTimeout`/frame counter, or you get progressive drift over a long utterance — drift is the classic failure and it grows with sentence length.
|
| 148 |
+
- Handle the streaming case: if audio streams chunk-by-chunk, the viseme timeline must be chunk-relative and re-anchored per chunk.
|
| 149 |
+
- Clamp small blendshape values (<0.001 → 0) and interpolate between visemes rather than snapping; keep an idle/breathing animation always running rather than switching between "idle" and "talking" state machines.
|
| 150 |
+
|
| 151 |
+
**Warning signs:** mouth movement identical for 「あ」 and 「お」; sync visibly degrades over the last third of long sentences; lip-sync works for pre-recorded audio but not streamed.
|
| 152 |
+
|
| 153 |
+
**Phase to address:** P1 for the timing architecture (`audio.currentTime`-driven, timing-capable TTS); P2 for kana-accurate viseme mapping.
|
| 154 |
+
|
| 155 |
+
---
|
| 156 |
+
|
| 157 |
+
### Pitfall 7: Gradio + three.js integration fights you the whole way (and Gradio 6 moves under you)
|
| 158 |
+
|
| 159 |
+
**What goes wrong:**
|
| 160 |
+
A real-time WebGL avatar is a stateful, long-lived, imperative canvas. Gradio is a declarative Python-driven UI that re-renders components. Predictable failures:
|
| 161 |
+
- WebGL context created inside a `gr.HTML` block gets destroyed/re-created when Gradio re-renders that component, losing the loaded VRM (multi-MB reload) and all animation state.
|
| 162 |
+
- Custom JS reaching into Gradio's DOM breaks on upgrade — the docs state plainly that **"the use of query selectors in custom JS and CSS is *not* guaranteed to work across Gradio versions"** ([Custom CSS and JS](https://gradio.app/guides/custom-CSS-and-JS)).
|
| 163 |
+
- Serving the `.vrm` file: local files must be explicitly allowed and referenced via the `/gradio_api/file=` mechanism, otherwise 403s that look like CORS errors.
|
| 164 |
+
- Writing a proper Gradio **custom component** means a Svelte/TypeScript build toolchain, `gradio cc build/publish`, and a frontend build in the Space — a large, ongoing tax; the component ecosystem itself is [under active reconsideration by the Gradio team](https://github.com/gradio-app/gradio/issues/12074).
|
| 165 |
+
- **Gradio 6 has many breaking changes** (Chatbot tuple format removed, `show_api` removed, `api_name=False` gone, event/API naming changes — [migration guide](https://gradio.app/main/guides/gradio-6-migration-guide)), with community friction ([#12921](https://github.com/gradio-app/gradio/issues/12921), [#12443](https://github.com/gradio-app/gradio/issues/12443)). Latest is 6.22.0 as of 2026-07-31.
|
| 166 |
+
|
| 167 |
+
**Plus the known-local gotcha:** this profile's existing Spaces are pinned in lockstep at `gradio==5.9.1 / huggingface_hub==0.26.5 / torch==2.5.1 / transformers==4.46.3` **specifically because newer versions broke on HF Spaces' Python 3.13 default (gradio/hub compat, `audioop` removal)** — documented at `hugginface_profile/docs/ai-model-updates-2026-07-02.md`, where 21 pins were audited and *refuted* for bumping without deploy testing. Note the recorded escape hatch: `dino-space` pins Python 3.11 in its README and successfully runs gradio 6.19.0.
|
| 168 |
+
|
| 169 |
+
**Why it happens:**
|
| 170 |
+
Gradio is chosen for the portfolio requirement ("the Gradio/Spaces demo slot"), then the avatar requirement demands something Gradio isn't designed for, and the seam is discovered mid-build.
|
| 171 |
+
|
| 172 |
+
**How to avoid:**
|
| 173 |
+
- **Decide the integration strategy in P1, explicitly, from three options**, and write down the choice:
|
| 174 |
+
1. `gr.HTML` + `head=` script injection with the avatar living in a **single, never-re-rendered** container, communicating via `window` events / a custom element — cheapest, most fragile to Gradio DOM changes.
|
| 175 |
+
2. A **proper Gradio custom component** — most robust and the most impressive on a portfolio, but a real frontend build pipeline.
|
| 176 |
+
3. **Avatar in an `<iframe>` hosting a self-contained static three.js app**, with `postMessage` as the sole Python↔avatar contract — strongest isolation from Gradio re-renders and version churn; the seam is a documented message protocol instead of DOM selectors. *(Recommended default unless the custom component itself is part of the portfolio story. Note the iframe needs `allow="microphone"` if mic capture lives inside it — see Pitfall 9.)*
|
| 177 |
+
- **Never rely on Gradio's DOM structure.** All JS hooks go through your own `elem_id`s and a narrow message API.
|
| 178 |
+
- **Pin `sdk_version` explicitly in README and pin Python to 3.11** in the Space config. Do not inherit the Spaces Python default. Match the `dino-space` precedent.
|
| 179 |
+
- Build a **thin smoke-test Space in P1** that loads the VRM, plays one canned TTS clip with lip-sync, and does nothing else — deployed, on the real Space, before any pedagogy work.
|
| 180 |
+
- Serve the `.vrm` via `allowed_paths` / static-path config, or better, from a Hub repo/CDN URL so the Space image stays small (Pitfall 12).
|
| 181 |
+
|
| 182 |
+
**Warning signs:** avatar disappears after any Gradio interaction; `document.querySelector` in your JS; the Space works locally and shows a blank canvas when deployed; a dependency bump breaks the build.
|
| 183 |
+
|
| 184 |
+
**Phase to address:** P1, as an explicit spike with a written decision record. Flag P1 as **needing deeper research**.
|
| 185 |
+
|
| 186 |
+
---
|
| 187 |
+
|
| 188 |
+
### Pitfall 8: LLM confidently hallucinating Japanese grammar explanations
|
| 189 |
+
|
| 190 |
+
**What goes wrong:**
|
| 191 |
+
The learner asks "why は here and not が?" and the model produces a fluent, authoritative, **wrong** explanation. This is documented specifically for Japanese ([Self Taught Japanese case study](https://selftaughtjapanese.com/2025/07/10/the-dangers-of-using-ai-to-learn-japanese-grammar-a-case-of-hallucinating-chatgpt/), where ChatGPT first marked a pattern incorrect and then implied it was natural). Japanese contains grammatical structures that are rare in multilingual training corpora, and SOTA models "frequently fail to respect nuanced aspects of Japanese grammar" ([arXiv 2505.19599](https://arxiv.org/pdf/2505.19599)). The **small open model** in the hybrid backend will be markedly worse than the frontier BYOK model here — meaning the free default tier, the one most visitors and the author will use, is the one that lies.
|
| 192 |
+
|
| 193 |
+
A beginner cannot detect the error. They memorize it. This is the single worst pedagogical failure mode, and it is silent.
|
| 194 |
+
|
| 195 |
+
**Why it happens:**
|
| 196 |
+
The LLM is treated as the knowledge source. It isn't; it's a fluent renderer.
|
| 197 |
+
|
| 198 |
+
**How to avoid:**
|
| 199 |
+
- **Separate knowledge from phrasing.** Build a curated grammar-point database (JLPT N5→N2 grammar points with meaning, formation, example sentences, common learner errors — this content is well-enumerated and finite, a few hundred points). The LLM's job is to *select and phrase* a retrieved entry, never to author grammar facts. This is a retrieval-grounded design and it is also a much better portfolio story than "I prompted a model."
|
| 200 |
+
- **Constrain the model's freedom:** structured output referencing a grammar-point ID; if no grammar point matches, the avatar says "I'm not sure — let's check that one" rather than inventing.
|
| 201 |
+
- **Tier-aware honesty:** if the small model is meaningfully worse at explanations, route *explanations* to retrieval+template and reserve free-form generation for *conversation*, where fluency matters and precision matters less.
|
| 202 |
+
- **Build a regression set**: ~50 known-tricky Japanese grammar questions (は/が, ~ている aspect, transitivity pairs, 敬語 levels, conditionals なら/たら/ば/と) with expert-correct answers. Run it against every model/prompt change. This is cheap and it's the only way you'll ever know quality moved.
|
| 203 |
+
- Add a visible, low-friction **"this looks wrong" report** control — the author is the primary user and will spot errors.
|
| 204 |
+
|
| 205 |
+
**Warning signs:** explanations vary between runs for the same question; the model hedges then contradicts itself; no test set exists; nobody has checked the small model's Japanese output against a textbook.
|
| 206 |
+
|
| 207 |
+
**Phase to address:** P3 (Tutoring Brain). Flag P3 as **needing deeper research**. The grammar-point database is P2 content work and is a *prerequisite* for P3.
|
| 208 |
+
|
| 209 |
+
---
|
| 210 |
+
|
| 211 |
+
### Pitfall 9: JLPT level drift — "adaptive N5→N2" that is actually one prompt string
|
| 212 |
+
|
| 213 |
+
**What goes wrong:**
|
| 214 |
+
The system prompt says "respond at JLPT N5 level." The first reply is N5. Four turns later the model is using N2 vocabulary and て-form chains because the conversation topic pulled it there. The learner is lost, and the "adaptive difficulty" requirement is unmet while appearing implemented. Level/difficulty **alignment drift in level-prompted LLMs is a documented, general problem** in tutoring research ([arXiv 2506.04072, controlling difficulty in conversation](https://arxiv.org/pdf/2506.04072)), and LLMs encode framework levels (CEFR/JLPT) only approximately.
|
| 215 |
+
|
| 216 |
+
**Why it happens:**
|
| 217 |
+
Level is expressed as a soft instruction in natural language, and soft instructions decay over long contexts. Nobody measures it, because measuring it requires level-tagged word lists that feel like busywork.
|
| 218 |
+
|
| 219 |
+
**How to avoid:**
|
| 220 |
+
- **Mechanically validate every generated Japanese utterance before it reaches the learner.** Tokenize the reply (Pitfall 10 infra), look each lemma up in JLPT-tagged vocabulary/kanji lists, and compute an out-of-level rate. If it exceeds a threshold, regenerate with the offending items named ("avoid: 経済, 〜に違いない"). This is a deterministic guardrail, not a hope.
|
| 221 |
+
- **Kanji gating is separate from vocabulary gating** — an N5 learner may know the *word* but not the *kanji*. Level state needs at least two axes (vocab level, kanji level), and furigana policy is a third.
|
| 222 |
+
- Re-inject the level constraint every turn and keep it late in the context, not just in a system prompt set once.
|
| 223 |
+
- **Instrument it:** log out-of-level token rate per turn. A dashboard number that is supposed to be near zero is worth more than any amount of prompt engineering.
|
| 224 |
+
|
| 225 |
+
**Warning signs:** no JLPT word list is in the repo; the level control is a single f-string; nobody can answer "what percentage of last session's tokens were above the learner's level?"
|
| 226 |
+
|
| 227 |
+
**Phase to address:** P3 (generation guardrail) built on P2 (tokenizer + JLPT lists); consumed by P5 (adaptive).
|
| 228 |
+
|
| 229 |
+
---
|
| 230 |
+
|
| 231 |
+
### Pitfall 10: Japanese text infrastructure treated as an afterthought (tokenization, readings, furigana)
|
| 232 |
+
|
| 233 |
+
**What goes wrong:**
|
| 234 |
+
Japanese has no spaces. Almost every downstream feature — furigana, level checking, vocabulary tracking, SRS card creation, viseme generation, "tap a word for its meaning" — requires **morphological analysis plus readings**. Teams bolt this on late and then discover:
|
| 235 |
+
- **Dictionary weight**: full UniDic is **770MB on disk** and needs a separate download step; `unidic-lite` is small but is a modified 2.1.2 ([fugashi docs](https://github.com/polm/fugashi)). Downloading UniDic at Space startup risks the **30-minute startup timeout** (Pitfall 12) and bloats the image.
|
| 236 |
+
- **Reproducibility**: "many different dictionaries for MeCab can give completely different results" — token boundaries and lemma IDs change with dictionary version, which silently invalidates stored vocabulary progress if the dictionary is ever swapped.
|
| 237 |
+
- **Reading ambiguity**: 「今日」→ きょう or こんにち; 「行った」→ いった or おこなった. Morphological analyzers resolve most from context, but not all.
|
| 238 |
+
- **Furigana rendering**: HTML `<ruby>/<rt>/<rp>` is required, and Gradio's Markdown rendering will sanitize/strip unexpected HTML depending on component and version — furigana silently vanishes or shows as literal text. Markdown-level furigana syntaxes exist (`[漢字]{かんじ}` style, [furigana-markdown-it](https://github.com/iltrof/furigana-markdown-it)), and a fallback with 【】 parentheses matters where ruby is unavailable.
|
| 239 |
+
|
| 240 |
+
**Why it happens:**
|
| 241 |
+
It reads like plumbing, not a feature, so it's deferred — but every feature depends on it.
|
| 242 |
+
|
| 243 |
+
**How to avoid:**
|
| 244 |
+
- **Make the Japanese text pipeline a first-class P2 deliverable** with one canonical function: `text → [{surface, reading, lemma, pos, jlpt_level, is_kanji}]`. Everything (furigana, level gate, SRS cards, visemes) consumes that one structure.
|
| 245 |
+
- **Pin the dictionary version and record it in the DB alongside stored vocabulary rows.** A schema field like `analyzer_version` costs nothing now and prevents a silent data corruption later.
|
| 246 |
+
- Prefer `unidic-lite` (or a pre-baked dictionary committed as an LFS asset / mounted repo volume) over downloading full UniDic at startup. Verify image size and startup time on the deployed Space.
|
| 247 |
+
- **Verify ruby rendering on the deployed Space early** with a one-line test string. Do not assume `gr.Markdown` passes `<ruby>` through; if it doesn't, use `gr.HTML` or render furigana in the avatar iframe.
|
| 248 |
+
- Provide a **manual reading-override table** for curated lesson content — a small YAML of `surface → reading` exceptions is far cheaper than fighting the analyzer.
|
| 249 |
+
|
| 250 |
+
**Warning signs:** furigana works locally, not deployed; vocabulary counts change after a dependency bump; the same word is tracked as two different entries.
|
| 251 |
+
|
| 252 |
+
**Phase to address:** P2 (Japanese Language Core), before P3 and P5 depend on it.
|
| 253 |
+
|
| 254 |
+
---
|
| 255 |
+
|
| 256 |
+
### Pitfall 11: Japanese TTS with wrong pitch accent and wrong kanji readings — the tutor teaches errors
|
| 257 |
+
|
| 258 |
+
**What goes wrong:**
|
| 259 |
+
A generic multilingual neural TTS pronounces 橋 (はし, L-H, bridge) and 箸 (はし, H-L, chopsticks) identically, reads 「今日」 wrongly, or produces flat, accentless Japanese. For a translation demo that's cosmetic; **for a tutor it is teaching wrong pronunciation to a learner who cannot detect it** — the audio equivalent of Pitfall 8. Even specialized commercial engines report ~91% pitch-accent accuracy on ambiguous pairs (vendor-reported, [AnySpeech](https://anyspeech.io/japanese-text-to-speech) — **LOW confidence**), i.e. ~1 in 11 ambiguous words wrong.
|
| 260 |
+
|
| 261 |
+
**Why it happens:**
|
| 262 |
+
TTS is selected on voice *pleasantness* and latency, not on linguistic correctness, and the developer (a learner) can't hear accent errors yet — that's the whole reason they're using the app.
|
| 263 |
+
|
| 264 |
+
**How to avoid:**
|
| 265 |
+
- **Prefer Japanese-specialized, pitch-accent-aware engines.** Rule-based OpenJTalk handles Japanese pitch-accent and intonation well, which is why the Style-Bert-VITS2 / AivisSpeech family remains actively used for Japanese local TTS ([lilting.ch 2026 OSS TTS survey](https://lilting.ch/en/articles/voxcpm2-tokenizer-free-local-tts) — MEDIUM confidence). Bonus: OpenJTalk-derived pipelines give you the **mora timing you need for lip-sync** (Pitfall 6). Selecting on both criteria at once collapses two problems into one decision.
|
| 266 |
+
- **Pre-generate and human-verify audio for curated content** (lesson vocabulary, drill sentences, model utterances). Free-form conversational TTS can tolerate an occasional accent error; a vocabulary drill card that mispronounces its own headword cannot.
|
| 267 |
+
- Maintain a **reading/accent override list** for the curated vocabulary set, keyed to the same lemma IDs as Pitfall 10.
|
| 268 |
+
- Add a native-speaker (or dictionary) spot-check to the definition of done for lesson content — NHK accent dictionary conventions are the reference.
|
| 269 |
+
|
| 270 |
+
**Warning signs:** homograph pairs sound identical; all sentences have the same flat contour; no one has ever listened to the drill audio with a reference.
|
| 271 |
+
|
| 272 |
+
**Phase to address:** P1 for engine selection (jointly with viseme timing), P2/P5 for curated-content audio verification.
|
| 273 |
+
|
| 274 |
+
---
|
| 275 |
+
|
| 276 |
+
### Pitfall 12: HF Spaces operational realities — ephemeral disk, cold starts, startup timeouts, and progress that vanishes
|
| 277 |
+
|
| 278 |
+
**What goes wrong:**
|
| 279 |
+
Multiple platform behaviors combine into "the portfolio piece is broken when the hiring manager opens it":
|
| 280 |
+
- **Disk is ephemeral.** HF's own docs: *"Every Space comes with a small amount of disk storage. This disk space is ephemeral, meaning its content will be lost if your Space restarts or is stopped."* Persistence now goes through **attached Storage Buckets** (mounted volumes) or an external DB ([spaces-storage docs](https://huggingface.co/docs/hub/en/spaces-storage), fetched 2026-08-08). Note this supersedes the older "$5/mo persistent storage add-on" framing still repeated in blog posts — **verify the current mechanism at build time**. A SQLite file on the Space disk means every `git push` silently deletes all user progress.
|
| 281 |
+
- **Free CPU Spaces sleep after 48h of inactivity** and `cpu-basic` **cannot configure a custom sleep time**. A portfolio Space gets sporadic traffic, so it is *usually asleep*, and the reviewer's first impression is a cold boot.
|
| 282 |
+
- **Startup timeout defaults to 30 minutes** (`startup_duration_timeout`), and *"model files in /tmp are temporary and redownload each session."* Downloading a Whisper model + a TTS model + UniDic at boot is minutes of cold start every wake ([HF forums](https://discuss.huggingface.co/t/launch-timed-out-space-was-not-healthy-after-30-min/63685), [Spaces config reference](https://huggingface.co/docs/hub/spaces-config-reference)).
|
| 283 |
+
- **Repo/file limits**: pre-receive hook rejects non-LFS files >10MiB; single LFS files capped at ~50GB. A `.vrm` (typically 10–60MB) and any bundled model weights need LFS or must be fetched from a Hub repo.
|
| 284 |
+
|
| 285 |
+
**Why it happens:**
|
| 286 |
+
Everything works during an active dev session, when the Space is warm and the disk hasn't been wiped yet. The failure mode only appears after a rebuild or after 48h.
|
| 287 |
+
|
| 288 |
+
**How to avoid:**
|
| 289 |
+
- **All user data goes to an external Postgres (or a mounted Storage Bucket) from day one.** Never SQLite-on-Space-disk, not even "temporarily" — see the tech-debt table.
|
| 290 |
+
- **Render the shell instantly.** The avatar + UI should paint and be interactive before any model is ready, with a visible "warming up" state. A cold-start reviewer should see a beautiful idle avatar in <3s, not a Gradio loading spinner.
|
| 291 |
+
- **Keep boot-time downloads near zero**: mount model repos as read-only volumes, or bake small models into the image, or (best) run ASR/TTS in the browser so the server has nothing to load.
|
| 292 |
+
- Explicitly set `sdk_version`, `python_version: "3.11"`, and `startup_duration_timeout` in the Space README.
|
| 293 |
+
- Consider a **keep-warm ping** for the portfolio window (cheap cron hitting the Space), and know that free `cpu-basic` will still sleep on its own schedule.
|
| 294 |
+
- **Test the actual cold path**: factory-rebuild the Space and time first-paint and first-conversation-turn as a P6 gate.
|
| 295 |
+
|
| 296 |
+
**Warning signs:** any `sqlite3.connect("./data.db")`; progress resets after a deploy; first visit takes >10s; Space "stuck on Starting."
|
| 297 |
+
|
| 298 |
+
**Phase to address:** P4 (persistence architecture), P6 (cold-start and rebuild verification).
|
| 299 |
+
|
| 300 |
+
---
|
| 301 |
+
|
| 302 |
+
### Pitfall 13: Gradio global state leaking across users — and BYOK API keys leaking with it
|
| 303 |
+
|
| 304 |
+
**What goes wrong:**
|
| 305 |
+
In Gradio, **any variable defined outside a function is shared across every user session**. Two learners hit the Space simultaneously and get each other's conversation history, level, or — catastrophically — each other's BYOK API keys. Worse, **the initial value of `gr.State` is itself shared across users and is not reset on refresh** ([#4558](https://github.com/gradio-app/gradio/issues/4558), [#2132](https://github.com/gradio-app/gradio/issues/2132), [#9983](https://github.com/gradio-app/gradio/issues/9983)) — the classic bug is assigning a UUID as a `gr.State` default and discovering every user shares one ID.
|
| 306 |
+
|
| 307 |
+
**Why it happens:**
|
| 308 |
+
Single-user local testing never reveals it. `global model_state` is the natural thing to write in a Python script.
|
| 309 |
+
|
| 310 |
+
**How to avoid:**
|
| 311 |
+
- **Zero mutable module-level globals** except read-only loaded models. Enforce it by review; it's a one-line grep.
|
| 312 |
+
- Per-user data lives in `gr.State` initialized via a **callable factory** (so each session constructs its own), keyed to the authenticated user, and reconciled with the DB.
|
| 313 |
+
- **BYOK keys**: hold in session state only; never log, never echo into the chat transcript, never persist to the DB, never include in error messages or analytics. Scrub them from exception traces. Use `type="password"` inputs. State clearly in the UI where the key goes and that it isn't stored.
|
| 314 |
+
- **Two-browser test as an explicit acceptance criterion**: open the deployed Space in two browsers, two accounts, and confirm zero cross-talk. Do this in P4 and again in P6.
|
| 315 |
+
|
| 316 |
+
**Warning signs:** someone else's Japanese sentence appears in your chat; a shared "current level" that changes when you didn't do anything; keys appearing in Space logs.
|
| 317 |
+
|
| 318 |
+
**Phase to address:** P4 (Accounts & Persistence) — but the discipline must start in P1.
|
| 319 |
+
|
| 320 |
+
---
|
| 321 |
+
|
| 322 |
+
### Pitfall 14: HF OAuth assumptions — works only in Spaces, and the local "mock" leaks the owner's token
|
| 323 |
+
|
| 324 |
+
**What goes wrong:**
|
| 325 |
+
Two distinct traps:
|
| 326 |
+
1. **OAuth features only work when the app runs in a Space**, and require `hf_oauth: true` in the Space README with `https://{SPACE_HOST}/login/callback` as redirect URI ([HF OAuth docs](https://huggingface.co/docs/hub/oauth)). Local dev of the logged-in experience is a different code path, so auth-dependent features get built untested against the real flow. There are also documented redirect-loop failures in Spaces ([gradio #8533](https://github.com/gradio-app/gradio/issues/8533)).
|
| 327 |
+
2. **Security advisory**: Gradio apps running *outside* Spaces auto-enable **mocked** OAuth routes when `gr.LoginButton` is used — visiting `/login/huggingface` makes the server fetch **its own** HF token via `huggingface_hub.get_token()` and place it in the visitor's session. If that dev server is network-reachable, any visitor steals the owner's HF token ([GHSA-h3h8-3v2v-rg7m](https://github.com/gradio-app/gradio/security/advisories/GHSA-h3h8-3v2v-rg7m)).
|
| 328 |
+
|
| 329 |
+
Also: HF OAuth tokens expire (`hf_oauth_expiration_minutes`); a tutor session that outlives the token must handle re-auth mid-lesson without losing the conversation.
|
| 330 |
+
|
| 331 |
+
**How to avoid:**
|
| 332 |
+
- Never run the app with `share=True` or bound to `0.0.0.0` during local OAuth development; keep it on `127.0.0.1`. Upgrade Gradio to a version with the advisory fix and pin it.
|
| 333 |
+
- Design the app so **anonymous use works fully** (guest mode with browser-local progress), and login is an upgrade that migrates local progress to the server. This decouples the demo from the auth path, protects the "reviewer opens it and it just works" case, and means an OAuth outage degrades rather than blocks.
|
| 334 |
+
- Handle token expiry: persist progress on every meaningful event, not at session end.
|
| 335 |
+
- Test the *real* OAuth flow on the deployed Space early — it cannot be validated locally.
|
| 336 |
+
|
| 337 |
+
**Warning signs:** login tested only via a mock; "works locally" claims about auth; progress lost when a session runs long.
|
| 338 |
+
|
| 339 |
+
**Phase to address:** P4, with the guest-mode decision made in P1 so it isn't a retrofit.
|
| 340 |
+
|
| 341 |
+
---
|
| 342 |
+
|
| 343 |
+
### Pitfall 15: The adaptive system that never adapts (and the progress schema that can't evolve)
|
| 344 |
+
|
| 345 |
+
**What goes wrong:**
|
| 346 |
+
Two linked failures that both look done:
|
| 347 |
+
- **Adaptation theater**: level is set by a one-time placement quiz and never moves again, or "adapts" via a rule nobody can articulate. The requirement "assesses and adjusts" is checked off; the learner is on N5 forever.
|
| 348 |
+
- **Derived-state-only schema**: the DB stores `user.level = "N4"` and `word.status = "learned"` — current state, no history. Later you want FSRS scheduling, retention analytics, or "words I keep failing," and the data to compute them was never recorded. FSRS-family algorithms are built on a **`ReviewLog` of individual review events** ([ts-fsrs data models](https://deepwiki.com/open-spaced-repetition/ts-fsrs/2.2-data-models)), and optimization needs on the order of **1,000+ reviews** before personalization beats defaults ([Deckbase FSRS guide](https://www.deckbase.co/resources/fsrs-guide)). If you didn't log from day one, you start that clock over.
|
| 349 |
+
|
| 350 |
+
**Why it happens:**
|
| 351 |
+
Adaptation is stated as a goal, not a specification. Schemas are designed around the current screen instead of around the events.
|
| 352 |
+
|
| 353 |
+
**How to avoid:**
|
| 354 |
+
- **Write the adaptation rule as an explicit spec before coding**: what signals (ASR-confirmed comprehension, drill accuracy, hesitation/response latency, requests for English, SRS retention), what thresholds, what promotion/demotion, what hysteresis (so it doesn't oscillate). If it can't be written in a paragraph, it can't be verified.
|
| 355 |
+
- **Append-only event log as the source of truth**: `(user_id, timestamp, event_type, item_id, payload_json)`. Derived state (current level, card scheduling) is a materialized projection that can be recomputed. This makes swapping SM-2 → FSRS a rewrite of a pure function instead of a data migration, and gives you the review history FSRS needs.
|
| 356 |
+
- **Don't optimize the scheduler early.** Ship stable defaults (FSRS default weights or SM-2), collect reviews, and only personalize after there's real data. Chasing high target retention (e.g. 97%) roughly doubles daily workload for marginal recall gain.
|
| 357 |
+
- **Test adaptation with a synthetic learner**: script a fake user who answers everything correctly for 50 turns and assert the level rises; one who fails everything and assert it falls. Without this the feature is unverifiable.
|
| 358 |
+
- Make it **legible to the learner** — "you moved to N4 because you got 85% on the last 40 N5 items." Legibility forces the rule to actually exist.
|
| 359 |
+
|
| 360 |
+
**Warning signs:** no table with a timestamp column; level only changes when the user manually changes it; nobody can state the promotion criterion.
|
| 361 |
+
|
| 362 |
+
**Phase to address:** P4 (schema: event log) and P5 (adaptation rule + synthetic-learner tests).
|
| 363 |
+
|
| 364 |
+
---
|
| 365 |
+
|
| 366 |
+
### Pitfall 16: Avatar polish eats the pedagogy budget (and the novelty wears off anyway)
|
| 367 |
+
|
| 368 |
+
**What goes wrong:**
|
| 369 |
+
The avatar is the fun part and it is infinitely polishable — hair physics, blinking, idle sway, camera work, outfits. Meanwhile the thing that determines whether the author actually learns Japanese (correct grammar explanations, a real drill loop, spaced repetition, level-appropriate content) stays a stub. Six months later there's a gorgeous character that teaches nothing.
|
| 370 |
+
|
| 371 |
+
This is compounded by the **novelty effect**: initial engagement from a novel interface reliably decays as it becomes routine unless the task itself becomes intrinsically meaningful ([novelty effect](https://en.wikipedia.org/wiki/Novelty_effect); language-app engagement decline is well documented). The avatar buys attention for week one. Pedagogy is what keeps week five.
|
| 372 |
+
|
| 373 |
+
There's a related, sharper research finding worth designing against: AI-supported students solved **48% more practice problems correctly but scored 17% lower on unassisted tests** — assistance that removes retrieval effort undermines learning ([pedagogy-driven ITS evaluation, arXiv 2510.22581](https://arxiv.org/html/2510.22581v1); AI tutors also show weak sensitivity to learner errors and over-direct feedback). A tutor that instantly supplies the word you're groping for feels amazing and teaches less than one that makes you struggle for three seconds.
|
| 374 |
+
|
| 375 |
+
**Why it happens:**
|
| 376 |
+
The portfolio framing ("would a Google hiring manager be impressed?") biases toward visual spectacle, and visual work has immediate, satisfying feedback. But the PROJECT.md explicitly says pedagogical quality matters equally.
|
| 377 |
+
|
| 378 |
+
**How to avoid:**
|
| 379 |
+
- **Time-box avatar polish per phase** and make each phase's exit criterion a *learning* outcome, not a visual one.
|
| 380 |
+
- **Define the "does it teach?" test early**: after N sessions, can the author produce X sentences / recall Y words unaided? Even a crude weekly self-quiz keeps the project honest.
|
| 381 |
+
- **Design for productive struggle**: delay hints, prefer prompting ("do you remember the て-form?") over supplying, require the learner to produce before showing the answer. Explicitly resist "helpful."
|
| 382 |
+
- Note the portfolio counter-argument: a hiring manager is *more* impressed by "I built a retrieval-grounded, level-validated tutoring pipeline with a measured latency budget and an event-sourced SRS" than by nicer hair physics. Depth of engineering beats surface polish for the stated audience — the avatar just has to be *good enough to be striking*, and it already is at ~80%.
|
| 383 |
+
- Ship the **retention loop** (review queue, streak of *learning*, "here's what you struggled with") in the same milestone as conversation — not as v2.
|
| 384 |
+
|
| 385 |
+
**Warning signs:** consecutive phases whose deliverables are all visual; conversation quality unchanged for weeks; the author stops using it after week two.
|
| 386 |
+
|
| 387 |
+
**Phase to address:** P5/P6, but enforced as a **roadmap-level constraint on every phase's exit criteria**.
|
| 388 |
+
|
| 389 |
+
---
|
| 390 |
+
|
| 391 |
+
## Technical Debt Patterns
|
| 392 |
+
|
| 393 |
+
| Shortcut | Immediate Benefit | Long-term Cost | When Acceptable |
|
| 394 |
+
|----------|-------------------|----------------|-----------------|
|
| 395 |
+
| SQLite file on the Space's ephemeral disk | Zero setup, no external service | **All user progress deleted on every rebuild/restart** — silent, total data loss | **Never.** External Postgres or a mounted Storage Bucket from the first line of persistence code |
|
| 396 |
+
| Amplitude-driven lip-sync ("good enough for now") | Ships in an hour | Rework touches the TTS choice, audio transport, and animation clock simultaneously | Only as a throwaway P1 spike, deleted before merge |
|
| 397 |
+
| Level control as a system-prompt string | Feature appears done immediately | Drift is invisible; "adaptive N5→N2" is unverifiable and unfixable without the tokenizer + JLPT lists that were skipped | Acceptable only in the P1 skeleton, with the validator scheduled in P3 |
|
| 398 |
+
| Storing only derived progress state (no event log) | Simpler schema, fewer rows | Cannot adopt FSRS, cannot compute retention, cannot explain adaptation; requires a migration + restarting data collection from zero | Never — the event table costs ~20 lines |
|
| 399 |
+
| Deriving furigana with a regex / romaji instead of morphological analysis | Skips the 770MB dictionary question | Wrong readings taught to learners; no lemma IDs to key SRS/vocab tracking to | Never for learner-facing text; fine for internal debug output |
|
| 400 |
+
| DOM `querySelector` into Gradio internals to mount the canvas | Works today, no build tooling | Breaks on any Gradio bump; explicitly unsupported | Acceptable behind an isolation layer (iframe/custom element) with one narrow selector, not scattered |
|
| 401 |
+
| Free-form LLM grammar explanations without retrieval grounding | Instantly impressive breadth | Teaches wrong Japanese, undetectably, to a beginner | Never for explanations; fine for conversational replies |
|
| 402 |
+
| Committing the `.vrm` and model weights into the Space repo | One deploy artifact | 10MiB push rejections, slow builds, long cold starts | Small assets only; use LFS/Hub repos/mounted volumes for anything large |
|
| 403 |
+
| Hard-pinning `gradio==5.9.1` lockstep to match the other Spaces | Known-good, matches the audited set | Locks out Gradio 6 features and the current custom-component tooling for a *new* project with no legacy | Not here — this is greenfield; pin Python 3.11 and use current Gradio (6.x), per the `dino-space` precedent |
|
| 404 |
+
|
| 405 |
+
---
|
| 406 |
+
|
| 407 |
+
## Integration Gotchas
|
| 408 |
+
|
| 409 |
+
| Integration | Common Mistake | Correct Approach |
|
| 410 |
+
|-------------|----------------|------------------|
|
| 411 |
+
| **HF Spaces (free CPU)** | Assuming the disk persists and the Space stays warm | Ephemeral disk + 48h sleep on `cpu-basic` (no configurable sleep time); external DB/bucket; instant-paint UI; verify by factory-rebuilding |
|
| 412 |
+
| **HF Spaces config** | Inheriting the platform's default Python and Gradio | Explicitly set `sdk_version`, `python_version: "3.11"`, `startup_duration_timeout` in README — the prior Python-3.13 `audioop`/hub breakage is documented in `docs/ai-model-updates-2026-07-02.md` |
|
| 413 |
+
| **HF OAuth** | Testing login locally and assuming it transfers | OAuth only works in Spaces (`hf_oauth: true`, `/login/callback`); local mock routes can leak the owner's HF token (GHSA-h3h8-3v2v-rg7m); test on the deployed Space |
|
| 414 |
+
| **Gradio custom JS/three.js** | Query-selecting Gradio's DOM; canvas re-created on re-render | Isolate the avatar (iframe + `postMessage`, or a real custom component); own `elem_id`s only; never assume DOM stability across versions |
|
| 415 |
+
| **Gradio static files** | Referencing a local `.vrm` by path | Serve via allowed paths / `/gradio_api/file=`, or from a Hub repo/CDN URL |
|
| 416 |
+
| **Browser microphone** | Testing on the direct `*.hf.space` URL only | Cross-origin iframes need `allow="microphone"` on the embed; the huggingface.co Space page and any personal-site embed are iframes. Classic symptom: mic works on the direct URL, `NotAllowedError` on the profile page. Also requires HTTPS + a user gesture |
|
| 417 |
+
| **FastRTC / WebRTC on Spaces** | Assuming peer connections just work behind the Space firewall | Cloud firewalls block direct connections — a **TURN server is required**; HF+Cloudflare gives 10GB/month free via HF token, with short-lived (~10min TTL) credentials ([fastrtc.org/deployment](https://fastrtc.org/deployment/)) |
|
| 418 |
+
| **Supabase / hosted Postgres** | Opening a connection per Gradio callback | The Space is a long-lived process — use one small pooled client (`(cores*2)+2` sizing, so ~5); use the pooler URL; know that transaction-mode pooling breaks prepared statements and session state |
|
| 419 |
+
| **Web Speech API (browser TTS/ASR)** | Assuming Japanese voices exist everywhere | Voice lists are OS/browser-specific, initially **empty** until the `voiceschanged` event fires on Chromium, and synthesis is throttled when the tab loses focus. Always feature-detect and have a server/neural fallback |
|
| 420 |
+
| **BYOK (Claude/OpenAI keys)** | Storing keys server-side or logging them | Session-scoped only, `type="password"`, never persisted/logged/echoed, scrubbed from tracebacks; explicit UI statement that keys aren't stored |
|
| 421 |
+
| **MeCab/fugashi dictionaries** | Downloading UniDic (770MB) at container start | Use `unidic-lite` or a pre-baked/mounted dictionary; pin and record the dictionary version alongside stored vocabulary rows |
|
| 422 |
+
|
| 423 |
+
---
|
| 424 |
+
|
| 425 |
+
## Performance Traps
|
| 426 |
+
|
| 427 |
+
| Trap | Symptoms | Prevention | When It Breaks |
|
| 428 |
+
|------|----------|------------|----------------|
|
| 429 |
+
| Server-side ASR+LLM+TTS on 2 vCPU | 5–15s turns; "thinking" spinner dominates the UX | Browser-side inference; streaming; sentence-level TTS chunking | Immediately — at **one** user |
|
| 430 |
+
| Gradio queue default `concurrency_limit=1` on a blocking inference fn | Second visitor waits for the first visitor's entire turn | Non-blocking handlers, browser-side inference, explicit `concurrency_limit`; note limits interact with `default_concurrency_limit` ([#7543](https://github.com/gradio-app/gradio/issues/7543)) | 2 concurrent visitors — i.e. the first time it's shared |
|
| 431 |
+
| Loading ASR/TTS/LLM models into RAM per session | OOM kill; Space restarts (wiping the ephemeral disk) | Load once at module scope, read-only; never per-request | 2–3 concurrent users on 16GB |
|
| 432 |
+
| Multi-MB `.vrm` re-downloaded on every Gradio re-render | Avatar flickers/reloads; slow interactions | Isolate the canvas so it never unmounts; cache-control headers; CDN/Hub-hosted asset | Every interaction |
|
| 433 |
+
| Cold start downloading models to ephemeral `/tmp` | Multi-minute first load after each sleep; occasional 30-min startup timeout | Bake/mount models; raise `startup_duration_timeout`; keep-warm ping | After any 48h idle period — i.e. most portfolio visits |
|
| 434 |
+
| Per-callback DB connections | `too many connections`; latency spikes | One pooled client for the process; pooler URL | Dozens of concurrent turns |
|
| 435 |
+
| Un-paginated progress/SRS queries (`SELECT *` over all review events) | Dashboard slows over months of use | Index `(user_id, timestamp)`; query projections, not raw events | ~10k review events — one dedicated user, one year |
|
| 436 |
+
| WebGL avatar + heavy JS on the main thread | Lip-sync stutter, audio glitching, mobile overheating | Web Workers for inference; keep render loop lean; cap DPR; test on mid-range mobile | Mid-range phones, immediately |
|
| 437 |
+
|
| 438 |
+
---
|
| 439 |
+
|
| 440 |
+
## Security Mistakes
|
| 441 |
+
|
| 442 |
+
| Mistake | Risk | Prevention |
|
| 443 |
+
|---------|------|------------|
|
| 444 |
+
| Module-level globals holding per-user data in Gradio | Cross-user leakage of conversations, progress, **and BYOK API keys** | No mutable globals; `gr.State` via factory; two-browser concurrent test as an acceptance gate |
|
| 445 |
+
| Relying on Gradio's `gr.State` default value for identity | The initial value is shared across users ([#4558](https://github.com/gradio-app/gradio/issues/4558)) — every user gets the same UUID | Generate identity inside the session-scoped callable; key everything to the authenticated HF user ID |
|
| 446 |
+
| Running a `gr.LoginButton` app outside Spaces on a reachable interface | Mocked OAuth hands the **server owner's HF token** to any visitor (GHSA-h3h8-3v2v-rg7m) | Bind to `127.0.0.1` locally, never `share=True` during auth work; patched Gradio version |
|
| 447 |
+
| Persisting or logging BYOK keys | Key theft; the visitor pays for someone else's usage; reputational damage on a public portfolio Space | Session-only, password input, scrubbed from logs/tracebacks/analytics, explicit UI promise |
|
| 448 |
+
| DB credentials in the repo instead of Space secrets | Public repo → public database | HF Space secrets/variables; never in `app.py`, never in the README |
|
| 449 |
+
| No row-level scoping on progress queries | User A reads/writes user B's progress | Every query filtered by authenticated user ID server-side (not client-supplied); RLS if using Supabase |
|
| 450 |
+
| Trusting client-supplied level/progress values | Trivial progress forgery (low stakes, but it corrupts the adaptation signal) | Server-side validation of state transitions |
|
| 451 |
+
| Rendering LLM-generated HTML (furigana ruby) unsanitized | XSS via prompt injection in a component that must allow `<ruby>` | Whitelist-sanitize to `<ruby>/<rt>/<rp>` only; better, render from the structured token list, never from model-authored HTML |
|
| 452 |
+
| Shipping a VRM whose license forbids this use | Takedown / licensing dispute on the flagship public portfolio piece | VRM files carry embedded license metadata; VRoid Hub models set per-model flags for personal vs **corporate commercial use, redistribution, and credit requirements** ([VRoid FAQ](https://vroid.pixiv.help/hc/en-us/articles/360016417013-About-VRoid-Hub-s-conditions-of-use-and-VRM-license)) — note CC0 cannot be set on VRoid Hub. Read the embedded license, keep a `LICENSES.md` with model + author + terms, display required credit, and prefer a self-made VRoid Studio character to remove the question entirely |
|
| 453 |
+
| Storing raw learner audio without disclosure | Privacy exposure; voice is biometric-adjacent | Process in memory, don't persist audio by default; if stored for scoring history, disclose and allow deletion |
|
| 454 |
+
|
| 455 |
+
---
|
| 456 |
+
|
| 457 |
+
## UX Pitfalls
|
| 458 |
+
|
| 459 |
+
| Pitfall | User Impact | Better Approach |
|
| 460 |
+
|---------|-------------|-----------------|
|
| 461 |
+
| Romaji shown as the default reading aid | Learners who lean on romaji for months find kana genuinely hard to switch to; romanization reliance "actively impedes reading fluency" ([research summary](https://www.researchgate.net/publication/237461322_CALL_Vocabulary_Learning_in_Japanese_Does_Romaji_Help_Beginners_Learn_More_Words)) | Kana-first with furigana; romaji available only as an explicit, temporary, decaying toggle (auto-off after the kana module) |
|
| 462 |
+
| Uniform furigana on everything, forever | Learner never has to read kanji; kanji progress stalls | Furigana gated by the learner's *kanji* level (separate axis from vocab level); tap-to-reveal rather than always-on |
|
| 463 |
+
| Silent dead air while the avatar "thinks" | 2–3s of nothing reads as a crash; user talks over it | Immediate acknowledgment: listening indicator, thinking pose, a short filler 「えーと…」 — perceived latency, not just real latency |
|
| 464 |
+
| Correcting every single learner error | Demoralizing; blocks fluency practice; not how tutors behave | Mode-dependent: conversation mode corrects sparingly and recasts naturally; drill mode corrects precisely. Make the mode visible |
|
| 465 |
+
| Instant help on any hesitation | Feels great, measurably reduces learning (48% better assisted, 17% worse unassisted) | Deliberate delay before hints; prompt for retrieval before supplying the answer |
|
| 466 |
+
| Corrections given in English by default at higher levels | Breaks immersion; caps ceiling | Explanation language follows level: English at N5, graded Japanese by N3+ |
|
| 467 |
+
| "Great job!" for everything | Feedback becomes noise; hides real problems (esp. with Pitfall 3's fake scores) | Specific, evidence-linked feedback: "your 「つ」 came out as 「ちゅ」" |
|
| 468 |
+
| Progress shown as a streak only | Streak anxiety; correlates with usage, not learning | Show *learning* evidence: words retained, level movement with the reason, sentences produced unaided |
|
| 469 |
+
| Avatar mouth/expression not matching the emotional content of speech | Uncanny; undermines the whole premise | Drive expressions from response metadata (a `tone` field the LLM emits) as well as visemes |
|
| 470 |
+
| No graceful degradation when mic is denied / voice unsupported | Blank screen on iOS or when permission is denied | Full text-chat fallback with the avatar still speaking; feature-detect and announce |
|
| 471 |
+
| Requiring login before anything works | Portfolio reviewers bounce; the hiring-manager path is broken | Guest mode fully functional; login upgrades and migrates local progress |
|
| 472 |
+
|
| 473 |
+
---
|
| 474 |
+
|
| 475 |
+
## "Looks Done But Isn't" Checklist
|
| 476 |
+
|
| 477 |
+
- [ ] **Voice conversation:** works in a quiet room with headphones — verify on **laptop speakers at volume, in a noisy room, on a phone**; confirm no self-triggering and no hallucinated turns from silence
|
| 478 |
+
- [ ] **Latency:** measured locally — verify **p50/p95 on the deployed Space, cold and warm**, with per-stage breakdown, on the free-tier default path
|
| 479 |
+
- [ ] **Lip-sync:** looks right on one short canned clip — verify on a **20-second streamed utterance** (drift), and that あ/い/う/え/お are visibly distinct
|
| 480 |
+
- [ ] **Pronunciation scoring:** returns a number — verify the number **drops for deliberately wrong pronunciation** (wrong mora, English phonology, wrong pitch)
|
| 481 |
+
- [ ] **Level adaptation:** placement quiz exists — verify the level **moves in both directions** under a scripted synthetic learner, and that out-of-level token rate is logged
|
| 482 |
+
- [ ] **Grammar explanations:** fluent and confident — verify against a **50-question expert-answered regression set**, run separately for the small model and the BYOK model
|
| 483 |
+
- [ ] **Furigana:** renders in dev — verify **on the deployed Space** (ruby sanitization), and that readings are correct for known-ambiguous words (今日/行った/人気)
|
| 484 |
+
- [ ] **Japanese TTS:** sounds natural — verify **homograph pairs** (橋/箸, 雨/飴) and drill-vocabulary accent against a reference
|
| 485 |
+
- [ ] **Accounts:** login works — verify **on the deployed Space** (not the local mock), with token expiry mid-session, and progress surviving a **factory rebuild**
|
| 486 |
+
- [ ] **Multi-user:** works for you — verify **two browsers / two accounts simultaneously**: no shared history, no shared level, no shared BYOK key
|
| 487 |
+
- [ ] **Cold start:** fast when warm — verify **first paint and first conversational turn after the Space has slept**
|
| 488 |
+
- [ ] **BYOK:** key field exists — verify the key **never appears** in Space logs, DB rows, error messages, or the chat transcript
|
| 489 |
+
- [ ] **Mic permissions:** works on `*.hf.space` — verify on the **huggingface.co Space page (iframe)** and any embedded context
|
| 490 |
+
- [ ] **Avatar asset:** loads — verify the **embedded VRM license** permits this use and required credit is displayed
|
| 491 |
+
- [ ] **Mobile:** responsive layout — verify **WebGL performance, mic capture, and audio playback** on a mid-range phone (iOS Safari specifically)
|
| 492 |
+
- [ ] **SRS:** cards appear — verify the **review event log** is being written and that a scheduler swap wouldn't require a data migration
|
| 493 |
+
|
| 494 |
+
---
|
| 495 |
+
|
| 496 |
+
## Recovery Strategies
|
| 497 |
+
|
| 498 |
+
| Pitfall | Recovery Cost | Recovery Steps |
|
| 499 |
+
|---------|---------------|----------------|
|
| 500 |
+
| Latency architecture wrong (server-side loop) | **HIGH** | Move ASR/TTS to the browser; restructure handlers to stream; may require re-choosing TTS (timing data) and re-doing the lip-sync clock. Effectively re-does P1 — which is why P1 must be a spike |
|
| 501 |
+
| Progress lost to ephemeral disk | **HIGH** (unrecoverable data) | Data is gone. Migrate to external DB immediately, add an export, and apologize in the changelog. Prevention is the only real strategy |
|
| 502 |
+
| Cross-user state leak discovered post-launch | MEDIUM | Audit for module-level mutables; convert to session factories; **rotate/invalidate any BYOK keys that may have leaked** and disclose |
|
| 503 |
+
| Pronunciation score is meaningless | MEDIUM | Relabel honestly ("transcription confidence") while a real assessment path is built; do not silently keep shipping a fake score |
|
| 504 |
+
| Level drift found late | MEDIUM | Add the post-generation validator as a middleware layer — this is retrofittable *if* the tokenizer and JLPT lists exist; if they don't, add P2 work first |
|
| 505 |
+
| Grammar hallucinations found in the wild | MEDIUM | Add the grammar-point DB + retrieval grounding; run the regression set; add user-reporting. Retrofittable because it wraps generation |
|
| 506 |
+
| Schema has no event log | MEDIUM-HIGH | Add the event table now and start collecting; historical data is unrecoverable and personalization restarts its ~1,000-review clock |
|
| 507 |
+
| Gradio upgrade breaks the avatar | LOW-MEDIUM | Pin `sdk_version`; if the seam was an iframe/`postMessage` contract, this is a non-event — the isolation choice *is* the recovery strategy |
|
| 508 |
+
| VRM license conflict | LOW-MEDIUM | Swap the model (the interface is a file path) and replace with a self-authored VRoid Studio character |
|
| 509 |
+
| Echo/self-interruption in the field | LOW | Ship the half-duplex gate + push-to-talk toggle; it's a state-machine change, not an architecture change — *if* the state machine was designed with a gate point |
|
| 510 |
+
|
| 511 |
+
---
|
| 512 |
+
|
| 513 |
+
## Pitfall-to-Phase Mapping
|
| 514 |
+
|
| 515 |
+
| Pitfall | Prevention Phase | Verification |
|
| 516 |
+
|---------|------------------|--------------|
|
| 517 |
+
| Latency makes voice feel broken (1) | **P1** | Measured p50/p95 mic-release→first-audio on the deployed free-tier Space, cold and warm |
|
| 518 |
+
| ZeroGPU/GPU dependency (2) | **P1** | Full conversation completes with GPU path forcibly disabled |
|
| 519 |
+
| Fake pronunciation scoring (3) | **P2** (research spike) | Deliberately-wrong utterances produce lower scores |
|
| 520 |
+
| ASR hallucination on silence (4) | **P1** | 30s silence + 30s café noise → zero avatar turns |
|
| 521 |
+
| Acoustic feedback / self-interruption (5) | **P1** | Laptop-speakers-at-volume test, no headphones, no self-triggered turns |
|
| 522 |
+
| Amplitude-only lip-sync (6) | **P1** arch, **P2** accuracy | Visemes differ per vowel; no drift over a 20s streamed utterance |
|
| 523 |
+
| Gradio↔three.js integration fragility (7) | **P1** (spike + written decision) | Avatar survives 20 Gradio interactions without reload; smoke Space deployed |
|
| 524 |
+
| Hallucinated grammar explanations (8) | **P3** (research spike) | 50-question regression set passes for **both** model tiers |
|
| 525 |
+
| JLPT level drift (9) | **P3** on **P2** infra | Out-of-level token rate logged and under threshold over a 30-turn session |
|
| 526 |
+
| Japanese text infra / furigana (10) | **P2** | One canonical tokenizer output feeds furigana + level gate + SRS; ruby verified on the deployed Space |
|
| 527 |
+
| TTS pitch accent / readings (11) | **P1** engine choice, **P2/P5** content | Homograph pairs audibly distinct; drill audio spot-checked |
|
| 528 |
+
| Spaces ephemeral disk / cold start (12) | **P4** persistence, **P6** verification | Factory rebuild → progress intact; timed cold-start first-paint |
|
| 529 |
+
| Gradio global state / BYOK leakage (13) | **P4** (discipline from P1) | Two-browser concurrent test; log grep for key material |
|
| 530 |
+
| HF OAuth quirks & mock-token leak (14) | **P4** | Real OAuth flow on the deployed Space; guest mode fully functional; expiry handled |
|
| 531 |
+
| Adaptation theater / rigid schema (15) | **P4** schema, **P5** rule | Synthetic learner moves the level both ways; event log populated |
|
| 532 |
+
| Avatar polish vs pedagogy (16) | **All phases** (exit-criteria constraint) | Every phase has at least one learning-outcome exit criterion, not only visual ones |
|
| 533 |
+
|
| 534 |
+
**Phases flagged as needing deeper, phase-specific research:** **P1** (voice-loop architecture + Gradio/three.js seam + TTS-with-timings selection — this is where the project is won or lost), **P2** (pronunciation-assessment mechanism; Japanese NLP stack), **P3** (grammar grounding + level-validation design).
|
| 535 |
+
|
| 536 |
+
---
|
| 537 |
+
|
| 538 |
+
## Sources
|
| 539 |
+
|
| 540 |
+
**Platform (HIGH confidence — official docs, fetched 2026-08-08)**
|
| 541 |
+
- [Disk usage on Spaces](https://huggingface.co/docs/hub/en/spaces-storage) — ephemeral disk, Storage Buckets as the persistence mechanism
|
| 542 |
+
- [Spaces Configuration Reference](https://huggingface.co/docs/hub/spaces-config-reference) — `startup_duration_timeout` (30 min default), port 7860, `python_version`, `sdk_version`
|
| 543 |
+
- [Spaces Overview](https://huggingface.co/docs/hub/en/spaces-overview) / [Manage your Space](https://huggingface.co/docs/huggingface_hub/v0.19.3/guides/manage-spaces) — 48h sleep, `cpu-basic` cannot set sleep time, 2 vCPU/16GB
|
| 544 |
+
- [Sign in with Hugging Face](https://huggingface.co/docs/hub/oauth) + [spaces-oauth.md](https://github.com/huggingface/hub-docs/blob/main/docs/hub/spaces-oauth.md) — Spaces-only OAuth, `hf_oauth: true`, callback URI, token expiry
|
| 545 |
+
- [Spaces ZeroGPU](https://huggingface.co/docs/hub/spaces-zerogpu) + [quota reference](https://github.com/huggingface/skills/blob/main/skills/huggingface-zerogpu/references/how-quota-works.md) — per-visitor daily quota
|
| 546 |
+
- [Storage limits](https://huggingface.co/docs/hub/main/en/storage-limits) / [repo recommendations](https://huggingface.co/docs/hub/repositories-recommendations) — 10MiB non-LFS push limit, ~50GB LFS file cap
|
| 547 |
+
- [FastRTC Deployment](https://fastrtc.org/deployment/) + [HF×Cloudflare TURN](https://huggingface.co/blog/fastrtc-cloudflare) — TURN required behind Spaces firewall, 10GB/mo free
|
| 548 |
+
|
| 549 |
+
**Gradio (HIGH — official docs + issue tracker/advisory)**
|
| 550 |
+
- [Custom CSS and JS](https://gradio.app/guides/custom-CSS-and-JS) — `head=`/`js=`, DOM selectors explicitly unsupported across versions
|
| 551 |
+
- [State in Blocks](https://gradio.app/guides/state-in-blocks) / [Queuing](https://gradio.app/guides/queuing) — global vs session state, `concurrency_limit`
|
| 552 |
+
- [Gradio 6 Migration Guide](https://gradio.app/main/guides/gradio-6-migration-guide); breaking-change friction [#12921](https://github.com/gradio-app/gradio/issues/12921), [#12443](https://github.com/gradio-app/gradio/issues/12443); custom components reconsidered [#12074](https://github.com/gradio-app/gradio/issues/12074)
|
| 553 |
+
- State leakage: [#4558](https://github.com/gradio-app/gradio/issues/4558), [#2132](https://github.com/gradio-app/gradio/issues/2132), [#9983](https://github.com/gradio-app/gradio/issues/9983); concurrency [#7543](https://github.com/gradio-app/gradio/issues/7543)
|
| 554 |
+
- Security advisory [GHSA-h3h8-3v2v-rg7m](https://github.com/gradio-app/gradio/security/advisories/GHSA-h3h8-3v2v-rg7m) — mocked OAuth leaks owner token
|
| 555 |
+
- OAuth redirect loops in Spaces [#8533](https://github.com/gradio-app/gradio/issues/8533)
|
| 556 |
+
|
| 557 |
+
**Speech / voice (MEDIUM-HIGH — peer-reviewed + industry benchmarks)**
|
| 558 |
+
- [Building Tailored Speech Recognizers for Japanese Speaking Assessment (arXiv 2509.20655)](https://arxiv.org/pdf/2509.20655) — Whisper unsuitable for Japanese phonemic assessment; ignores speaker errors
|
| 559 |
+
- [Calm-Whisper (arXiv 2505.12969)](https://arxiv.org/pdf/2505.12969) — hallucination on non-speech
|
| 560 |
+
- [Evaluating Whisper across accents, JASA Express Letters](https://pubs.aip.org/asa/jel/article/4/2/025206/3267247/Evaluating-OpenAI-s-Whisper-ASR-Performance)
|
| 561 |
+
- [Whisper for L2 speech scoring (Springer / preprint)](https://taylorarnold.org/pdf/2024-whisperl2.pdf)
|
| 562 |
+
- [kotoba-whisper-v2.0](https://huggingface.co/kotoba-tech/kotoba-whisper-v2.0) — 6.3× faster than large-v3 for Japanese
|
| 563 |
+
- Latency thresholds: [Prodinit](https://prodinit.com/blog/production-voice-ai-agents-latency-architecture), [Hamming AI](https://hamming.ai/resources/voice-ai-latency-whats-fast-whats-slow-how-to-fix-it), [AssemblyAI 300ms rule](https://www.assemblyai.com/blog/low-latency-voice-ai)
|
| 564 |
+
- Echo/barge-in: [Deepgram AEC](https://developers.deepgram.com/docs/voice-agent-echo-cancellation), [Deepgram preprocessing](https://developers.deepgram.com/guides/deep-dives/audio-preprocessing-barge-in), [runedge on-device barge-in](https://www.runedge.ai/blog/barge-in-interruption-handling-on-device-voice), [FireRedChat full-duplex (arXiv 2509.06502)](https://arxiv.org/pdf/2509.06502)
|
| 565 |
+
|
| 566 |
+
**Japanese NLP / TTS (MEDIUM; vendor accuracy figures LOW)**
|
| 567 |
+
- [fugashi](https://github.com/polm/fugashi) + [How to Tokenize Japanese in Python](https://www.dampfkraft.com/nlp/how-to-tokenize-japanese.html) — UniDic 770MB, dictionary-version sensitivity
|
| 568 |
+
- [PASQA: pitch-accent-focused speech quality assessment (arXiv 2606.20137)](https://arxiv.org/pdf/2606.20137)
|
| 569 |
+
- [VoxCPM2 and OSS TTS in 2026 — Japanese fine-tune notes](https://lilting.ch/en/articles/voxcpm2-tokenizer-free-local-tts) — OpenJTalk pitch-accent strength; Style-Bert-VITS2/AivisSpeech
|
| 570 |
+
- [AnySpeech Japanese TTS](https://anyspeech.io/japanese-text-to-speech) — 91% pitch-accent accuracy claim on NHK-dictionary ambiguous pairs; **vendor-reported, LOW confidence**
|
| 571 |
+
- [furigana-markdown-it](https://github.com/iltrof/furigana-markdown-it), [showdown-kanji](https://github.com/joeellis/showdown-kanji) — ruby rendering approaches
|
| 572 |
+
|
| 573 |
+
**Avatar / VRM (MEDIUM)**
|
| 574 |
+
- [From Phonemes to Mouth Shapes: Animating VRM Models](https://primeta.ai/blog/from-phonemes-to-mouth-shapes-animating-vrm-models) — five VRM vowel visemes, clamping, always-on idle animation
|
| 575 |
+
- [VRChat Visemes wiki](https://wiki.vrchat.com/wiki/Visemes)
|
| 576 |
+
- [VRoid Hub conditions of use and VRM license](https://vroid.pixiv.help/hc/en-us/articles/360016417013-About-VRoid-Hub-s-conditions-of-use-and-VRM-license), [commercial use](https://vroid.pixiv.help/hc/en-us/articles/360014192953-About-commercial-use-on-VRoid-Hub), [license data in VRM files](https://vroid.pixiv.help/hc/en-us/articles/360014193033-About-license-data-for-VRM-files)
|
| 577 |
+
|
| 578 |
+
**Pedagogy (MEDIUM-HIGH — peer-reviewed)**
|
| 579 |
+
- [Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation (arXiv 2506.04072)](https://arxiv.org/pdf/2506.04072) — level/difficulty drift
|
| 580 |
+
- [Inconsistent Tokenizations Cause LMs to be Perplexed by Japanese Grammar (arXiv 2505.19599)](https://arxiv.org/pdf/2505.19599)
|
| 581 |
+
- [The dangers of using AI to learn Japanese grammar (case study)](https://selftaughtjapanese.com/2025/07/10/the-dangers-of-using-ai-to-learn-japanese-grammar-a-case-of-hallucinating-chatgpt/)
|
| 582 |
+
- [Pedagogy-driven Evaluation of GenAI Intelligent Tutoring Systems (arXiv 2510.22581)](https://arxiv.org/html/2510.22581v1) — weak error sensitivity; 48%/-17% assisted-vs-unassisted finding
|
| 583 |
+
- [Novelty effect](https://en.wikipedia.org/wiki/Novelty_effect); language-app engagement decline
|
| 584 |
+
- Romaji: [CALL Vocabulary Learning in Japanese: Does Romaji Help?](https://www.researchgate.net/publication/237461322_CALL_Vocabulary_Learning_in_Japanese_Does_Romaji_Help_Beginners_Learn_More_Words), [Why You Should Never Learn Japanese with Romaji](https://unseen-japan.com/japanese-learning-no-romaji/)
|
| 585 |
+
- SRS: [ts-fsrs data models](https://deepwiki.com/open-spaced-repetition/ts-fsrs/2.2-data-models), [FSRS fundamentals](https://github.com/open-spaced-repetition/fsrs4anki/wiki/The-fundamental-of-FSRS/8c793cefb3ec361cd6fa6ab8f750e31c3da57e8e), [FSRS guide / common mistakes](https://www.deckbase.co/resources/fsrs-guide)
|
| 586 |
+
|
| 587 |
+
**Prior local experience (HIGH — first-party, this profile)**
|
| 588 |
+
- `hugginface_profile/docs/ai-model-updates-2026-07-02.md` — the `gradio==5.9.1 / huggingface_hub==0.26.5 / torch==2.5.1 / transformers==4.46.3` lockstep pins exist because newer versions broke on HF Spaces' **Python 3.13 default** (gradio/hub compat, `audioop` removal); 21 pins audited and refuted for blind bumping; `dino-space` demonstrates the escape hatch (pin Python 3.11, run Gradio 6.x); `mechspec-qwen` 5.9.1→6.19.0 verified only by building a py3.11 venv and launching
|
| 589 |
+
|
| 590 |
+
---
|
| 591 |
+
*Pitfalls research for: avatar-based conversational Japanese tutor on HF Spaces*
|
| 592 |
+
*Researched: 2026-08-08*
|
.planning/research/STACK.md
ADDED
|
@@ -0,0 +1,369 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Stack Research
|
| 2 |
+
|
| 3 |
+
**Domain:** Avatar-based conversational language tutor on Hugging Face Spaces (Gradio + in-browser VRM + Japanese full-duplex voice + hybrid LLM)
|
| 4 |
+
**Researched:** 2026-08-08
|
| 5 |
+
**Confidence:** HIGH on platform/hosting and avatar rendering (official docs verified), MEDIUM on TTS/ASR/LLM model selection (benchmarks are third-party), MEDIUM on licensing nuances (VOICEVOX character terms need a human read).
|
| 6 |
+
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
## ⚠️ Read This First: Two Findings That Reshape the Project
|
| 10 |
+
|
| 11 |
+
### 1. "Free CPU tier" is no longer the free path for a Gradio Space
|
| 12 |
+
|
| 13 |
+
`PROJECT.md` assumes "free CPU tier for baseline experience." **That assumption is now wrong.** Per the official Spaces docs:
|
| 14 |
+
|
| 15 |
+
> "Static Spaces are free for everyone. Gradio and Docker Spaces run on compute and **require a paid plan to create: PRO for personal accounts**… Free personal accounts in good standing can still host **up to 2 Gradio Spaces running on ZeroGPU**."
|
| 16 |
+
|
| 17 |
+
So the free hosting path for this project is **a Gradio Space on ZeroGPU**, not CPU Basic. Requirements to qualify: verified email, account older than 30 days, max 2 such Spaces.
|
| 18 |
+
|
| 19 |
+
**This is good news.** ZeroGPU is backed by **NVIDIA RTX Pro 6000 Blackwell** (48 GB VRAM at the default `large` size). The free path just got a datacenter GPU instead of 2 vCPUs. Model choices below assume this.
|
| 20 |
+
|
| 21 |
+
Confidence: **HIGH** — [Spaces Overview](https://huggingface.co/docs/hub/spaces-overview), [Spaces ZeroGPU](https://huggingface.co/docs/hub/spaces-zerogpu).
|
| 22 |
+
|
| 23 |
+
### 2. ZeroGPU quota is consumed by the *visitor*, not the owner
|
| 24 |
+
|
| 25 |
+
| Visitor account type | Included daily GPU quota | Queue priority |
|
| 26 |
+
|---|---|---|
|
| 27 |
+
| Unauthenticated | **2 minutes/day** | Low |
|
| 28 |
+
| Free HF account | **5 minutes/day** | Medium |
|
| 29 |
+
| PRO account | 40 minutes/day (extensible) | Highest |
|
| 30 |
+
|
| 31 |
+
Quota is only burned *inside* `@spaces.GPU`-decorated functions. This is the single hardest architectural constraint in the project: **an anonymous visitor gets ~120 seconds of GPU per day.** A naive design that runs ASR + LLM + TTS on GPU for every conversational turn would give a recruiter about four turns before hitting a quota wall.
|
| 32 |
+
|
| 33 |
+
**The stack below is designed around minimizing GPU-seconds per turn:** TTS runs on CPU (zero quota), ASR can run in the visitor's browser (zero quota), and only the LLM reliably touches `@spaces.GPU`. It also creates a natural, honest upsell for the BYOK path (bring your own key → zero GPU quota consumed).
|
| 34 |
+
|
| 35 |
+
Confidence: **HIGH** — official ZeroGPU docs usage-tier table.
|
| 36 |
+
|
| 37 |
+
---
|
| 38 |
+
|
| 39 |
+
## Recommended Stack
|
| 40 |
+
|
| 41 |
+
### Core Technologies
|
| 42 |
+
|
| 43 |
+
| Technology | Version | Purpose | Why Recommended |
|
| 44 |
+
|---|---|---|---|
|
| 45 |
+
| **Gradio** | `6.22.0` (pin exactly) | App framework + Space SDK | Gradio 6 is the only maintained major line; the team has stated only v6 gets future releases. Crucially, **ZeroGPU is exclusively compatible with the Gradio SDK** — no Docker, no Static. Gradio 6 also ships the `gr.HTML` templating system (below) that makes the three.js avatar viable without a build step. |
|
| 46 |
+
| **Python** | `3.12.12` (declare in README) | Runtime | ZeroGPU provides **only** Python `3.10.13` and `3.12.12`. Pick 3.12.12. **Do not use 3.13** — this is the recurring "Python 3.13 / Gradio pin" gotcha from prior Spaces work, and here it's a hard platform limit, not a soft one (`spaces` requires `>=3.10,<3.15`, but ZeroGPU only builds those two images). |
|
| 47 |
+
| **`spaces`** | `0.51.1` | ZeroGPU allocation | Provides `@spaces.GPU(duration=...)`. Effect-free off-ZeroGPU, so local dev is unaffected. Use **dynamic durations** (`duration=callable`) — shorter declared durations improve queue priority for your visitors. |
|
| 48 |
+
| **three.js** | `0.185.1` (MIT) | 3D renderer | The only realistic in-browser 3D engine with first-class VRM support. Pin the exact version — three.js makes breaking renderer changes at a ~6-week cadence and `three-vrm` tracks it loosely. |
|
| 49 |
+
| **@pixiv/three-vrm** | `3.5.5` (MIT) | VRM avatar loading, humanoid bones, expressions, spring-bone physics | The reference implementation from the company that runs VRoid Hub. Supports both VRM 0.0 and VRM 1.0, so community VRoid models load without conversion. Peer dep is `three >= 0.137`, so 0.185.1 is safe. |
|
| 50 |
+
| **VOICEVOX CORE** | `0.16.4` (`+cpu` abi3 wheel) | Japanese TTS **with per-mora phoneme timing** | See "The lip-sync decision" below. This is the highest-leverage choice in the whole stack. |
|
| 51 |
+
| **whisper-large-v3-turbo** | via `transformers` 5.x on ZeroGPU | Japanese ASR | Best measured accuracy/speed balance for Japanese conversational audio in the Feb-2026 benchmark: CER 0.184, **RTF 0.013**. MIT-licensed. |
|
| 52 |
+
| **Qwen3.5-4B** | Apache-2.0, released 2026-03-02 | Free-default tutor LLM | 4B dense, **201 languages incl. Japanese**, 256K context, hybrid thinking (disable thinking for latency). Apache-2.0 means no license story to explain to a hiring manager. Fits easily in 48 GB with room for KV cache. |
|
| 53 |
+
| **`transformers`** | `5.14.1` | Model loading | v5 is current. Note: v5 was a breaking release — do not copy v4-era snippets from blogs. |
|
| 54 |
+
| **`huggingface_hub`** | `1.27.0` | Hub + Inference Providers client | v1.x. Provides `InferenceClient` for the BYOK/routed paths. |
|
| 55 |
+
| **Neon Serverless Postgres** | Free plan | User accounts + progress persistence | See "The database decision" below. |
|
| 56 |
+
|
| 57 |
+
---
|
| 58 |
+
|
| 59 |
+
### The lip-sync decision (do not skip this)
|
| 60 |
+
|
| 61 |
+
VRM 1.0 defines exactly five lip-sync expression presets: **`aa`, `ih`, `ou`, `ee`, `oh`** — which are precisely the five Japanese vowels あいうえお. Japanese is mora-timed, and every mora is `(consonant +) vowel`. This means **Japanese is the single easiest language on earth to lip-sync on a VRM avatar**, provided you can get per-mora timings.
|
| 62 |
+
|
| 63 |
+
**VOICEVOX gives you exactly that, for free, on CPU.** Its `AudioQuery` structure returns `accent_phrases[] → moras[]`, where every mora carries `text`, `consonant`, `consonant_length`, `vowel`, `vowel_length`, and `pitch`. You get a deterministic phoneme timeline *before* synthesis, which you ship to the browser alongside the WAV.
|
| 64 |
+
|
| 65 |
+
Consequences:
|
| 66 |
+
- **No audio-amplitude analysis.** ChatVRM and most AITuber kits drive the `aa` blendshape from a Web Audio `AnalyserNode` RMS — the mouth just flaps open and shut. Mora-timed visemes are visibly better and are the thing that makes the demo look expensive.
|
| 67 |
+
- **No forced alignment, no viseme ML model, no GPU.**
|
| 68 |
+
- **Zero ZeroGPU quota consumed for speech output**, which is what makes the free tier survivable.
|
| 69 |
+
- The same mora timeline doubles as **pedagogical data**: mora-by-mora pitch accent is already in the payload, so "here is where your pitch accent was wrong" is nearly free.
|
| 70 |
+
|
| 71 |
+
Recommended implementation: build a `visemes: [{t_start, t_end, shape}]` array server-side from the `AudioQuery`, send it with the base64 audio to the `gr.HTML` avatar component, and drive `vrm.expressionManager.setValue(shape, weight)` off `audioElement.currentTime` in the render loop.
|
| 72 |
+
|
| 73 |
+
Confidence: **HIGH** on the VRM preset names ([VRM 1.0 expressions spec](https://github.com/vrm-c/vrm-specification/blob/master/specification/VRMC_vrm-1.0/expressions.md)) and on the AudioQuery structure ([VOICEVOX engine API reference](https://deepwiki.com/VOICEVOX/voicevox_engine/4.1-tts-pipeline-api)). The "AudioQuery enables audio-analysis-free, lightweight and accurate LipSync" technique is documented practice in the Japanese dev community.
|
| 74 |
+
|
| 75 |
+
**⚠️ Licensing action item (MEDIUM confidence — requires a human decision):**
|
| 76 |
+
- `voicevox_core` is dual-licensed **LGPL v3** + a commercial no-source-disclosure license. Dynamic linking via the Python wheel keeps you clear of copyleft on your own code.
|
| 77 |
+
- **Each character voice has its own terms.** The general rule is *free for commercial and non-commercial use provided you display credit* in the form `VOICEVOX:キャラクター名`. Without credit, per-character paid licensing is reportedly ¥400,000/character.
|
| 78 |
+
- **Read the specific character's terms before shipping**, and render the credit string persistently in the UI (a footer line next to the avatar). Prefer a character with permissive, well-documented terms (ずんだもん / VOICEVOX Nemo voices are the usual safe picks).
|
| 79 |
+
- Also verify the terms permit *offering synthesis to third parties via a web app*, not merely using generated audio — this is the one clause a portfolio Space most plausibly trips.
|
| 80 |
+
|
| 81 |
+
---
|
| 82 |
+
|
| 83 |
+
### The avatar-embedding decision
|
| 84 |
+
|
| 85 |
+
Gradio 6's `gr.HTML` gained `html_template`, `css_template`, `head`, `js_on_load`, `server_functions`, `watch()`, `trigger()` and `upload()`. This turns `gr.HTML` into a full custom-component system **in a single Python file, with no npm build step**. HF's own launch blog for this feature showcases "a full Three.js viewport inside a Gradio app."
|
| 86 |
+
|
| 87 |
+
```python
|
| 88 |
+
avatar = gr.HTML(
|
| 89 |
+
value={"audio": None, "visemes": [], "emotion": "neutral"},
|
| 90 |
+
html_template="<div id='vrm-stage'></div>",
|
| 91 |
+
css_template="#vrm-stage { width: 100%; aspect-ratio: 3/4; }",
|
| 92 |
+
head="""
|
| 93 |
+
<script type="importmap">
|
| 94 |
+
{"imports": {
|
| 95 |
+
"three": "https://cdn.jsdelivr.net/npm/three@0.185.1/build/three.module.js",
|
| 96 |
+
"three/addons/": "https://cdn.jsdelivr.net/npm/three@0.185.1/examples/jsm/",
|
| 97 |
+
"@pixiv/three-vrm": "https://cdn.jsdelivr.net/npm/@pixiv/three-vrm@3.5.5/lib/three-vrm.module.js"
|
| 98 |
+
}}
|
| 99 |
+
</script>
|
| 100 |
+
""",
|
| 101 |
+
js_on_load="""
|
| 102 |
+
const { VRMLoaderPlugin } = await import('@pixiv/three-vrm');
|
| 103 |
+
// ... build scene, load VRM, start render loop
|
| 104 |
+
watch('value', () => playUtterance(props.value)); // audio + viseme timeline
|
| 105 |
+
""",
|
| 106 |
+
server_functions=[synthesize], # JS can call Python directly
|
| 107 |
+
)
|
| 108 |
+
```
|
| 109 |
+
|
| 110 |
+
**Why this over the alternatives:**
|
| 111 |
+
|
| 112 |
+
| Approach | Verdict |
|
| 113 |
+
|---|---|
|
| 114 |
+
| `gr.HTML` with `html_template`/`js_on_load` | ✅ **Recommended.** No build step, no npm publish, hot-reloads with the app, `server_functions` gives direct JS→Python calls, `watch('value')` gives Python→JS push. Ships as one `app.py`. |
|
| 115 |
+
| Full Gradio custom component (`gradio cc`) | ❌ Requires Svelte, Node 18+, npm 9+, a build/publish cycle, and a separate package on PyPI. Justified only if you intend to publish the VRM component for others — a nice-to-have, not a v1 need. |
|
| 116 |
+
| Raw `<iframe>` + `postMessage` | ❌ You hand-roll the entire bridge, fight Space iframe sandboxing and cookie issues, and lose Gradio's event system. |
|
| 117 |
+
| Server-rendered talking-head video | ❌ Already correctly in `PROJECT.md` Out of Scope. Would also incinerate ZeroGPU quota. |
|
| 118 |
+
|
| 119 |
+
⚠️ `js_on_load` event listeners attach **once** on first render — use event delegation for dynamically created elements. And never interpolate untrusted user text into `html_template`/`js_on_load` (XSS); a Japanese-learning app will absolutely have user-supplied text flowing near the avatar (subtitles, corrections).
|
| 120 |
+
|
| 121 |
+
Confidence: **HIGH** — [Gradio custom HTML components guide](https://gradio.app/main/guides/custom-HTML-components), [HF blog: One-Shot Any Web App with gr.HTML](https://huggingface.co/blog/gradio-html-one-shot-apps).
|
| 122 |
+
|
| 123 |
+
---
|
| 124 |
+
|
| 125 |
+
### The database decision
|
| 126 |
+
|
| 127 |
+
| Option | Verdict |
|
| 128 |
+
|---|---|
|
| 129 |
+
| **Neon serverless Postgres (free plan)** | ✅ **Recommended.** 0.5 GB storage + 100 compute-hours/project/month, up to 100 projects, no credit card. Compute **scales to zero after 5 min idle and auto-resumes in a few hundred milliseconds on the next query** — no manual intervention, ever. |
|
| 130 |
+
| Supabase (free plan) | ⚠️ Only if you need its Auth/Storage too. **Free projects are auto-paused after 7 days of inactivity and require a manual dashboard unpause.** For a portfolio piece whose entire value is "a hiring manager clicks the link and it works," a silent 7-day dead-man's-switch on your user database is a project-killing failure mode. |
|
| 131 |
+
| SQLite on Space disk | ❌ Space disk is **ephemeral** — wiped on every restart/rebuild, and Spaces rebuild on every git push. You would lose user progress every deploy. |
|
| 132 |
+
| HF Storage Buckets (S3-like, Xet-backed) | ⚠️ Great for **assets** (VRM files, cached TTS audio, lesson media) — mountable as a Space volume, mutable, free allowance. **Not** a transactional database; SQLite over a network-mounted bucket invites corruption under concurrency. |
|
| 133 |
+
| HF Dataset repo + `CommitScheduler` | ⚠️ The classic HF persistence pattern, fine for **append-only analytics/telemetry**. Wrong shape for per-user read-modify-write of progress records (git commits, no transactions, no concurrent-write story). |
|
| 134 |
+
|
| 135 |
+
Recommended split: **Neon Postgres for user/progress state; HF Storage Bucket for binary assets; HF Dataset repo (private) for anonymized interaction logs** — the last one is also the honest, portfolio-legible answer to "where would your fine-tuning data come from?"
|
| 136 |
+
|
| 137 |
+
Confidence: **HIGH** on Neon/Supabase mechanics (vendor docs + multiple 2026 sources agree), **HIGH** on Space disk ephemerality and Storage Buckets ([Spaces disk usage](https://huggingface.co/docs/hub/spaces-storage), [Storage Buckets](https://huggingface.co/docs/hub/storage-buckets)).
|
| 138 |
+
|
| 139 |
+
---
|
| 140 |
+
|
| 141 |
+
### Supporting Libraries
|
| 142 |
+
|
| 143 |
+
| Library | Version | Purpose | When to Use |
|
| 144 |
+
|---|---|---|---|
|
| 145 |
+
| **SudachiPy** | `0.6.11` (Apache-2.0) | Japanese tokenization + `reading_form()` for furigana + lemmas | Primary Japanese NLP engine. Chosen over MeCab/fugashi for the clean Apache-2.0 story and pure-wheel install (no MeCab system dependency to fight in a Space build). Three splitting modes (A/B/C) map neatly onto "vocab item" vs "phrase" granularity for a tutor. |
|
| 146 |
+
| **SudachiDict-core** | `20260723` (Apache-2.0) | Dictionary for SudachiPy | Required. `-core` is the right size/coverage tradeoff; `-full` bloats the image. |
|
| 147 |
+
| **jaconv** | `0.5.0` (MIT) | hiragana ⇄ katakana ⇄ half-width conversion | Normalizing learner input, generating kana drills, feeding the viseme mapper. Tiny, zero-dependency, MIT. |
|
| 148 |
+
| **cutlet** | `0.5.2` | Romaji generation | N5 learners need romaji scaffolding. Built on fugashi; use only if you also pull fugashi, otherwise derive romaji from SudachiPy readings + `jaconv`. |
|
| 149 |
+
| **`@huggingface/transformers`** (transformers.js) | `4.2.0` (Apache-2.0) | **Browser-side** Whisper ASR fallback | v4 (Feb 2026) shipped a C++-rewritten WebGPU runtime (~4× faster, 53% smaller bundles). Running ASR here costs **zero ZeroGPU quota** — this is how anonymous visitors get more than four conversational turns. |
|
| 150 |
+
| **`anthropic`** | `0.121.0` | BYOK frontier path (Claude) | Visitor pastes their key; nothing is stored server-side. |
|
| 151 |
+
| **`openai`** | `2.53.0` | BYOK frontier path (OpenAI) — and reused as the client for HF's OpenAI-compatible router (`https://router.huggingface.co/v1`) | One client library covers two providers. |
|
| 152 |
+
| **`psycopg[binary]`** | 3.x | Postgres driver | Neon over TLS. Use a pooled connection string; a Space restart should not leak connections. |
|
| 153 |
+
| **SQLAlchemy** | 2.x | ORM / schema | Optional but recommended once progress modelling gets past ~4 tables (level, vocab-seen SRS state, mistakes, sessions). |
|
| 154 |
+
| **`onnxruntime`** | pulled by voicevox_core downloader | VOICEVOX inference backend | CPU build. Managed by the VOICEVOX downloader, do not hand-install. |
|
| 155 |
+
| **`fastrtc`** | `0.0.34` | *Deferred.* True barge-in / VAD / turn-taking over WebRTC | HF's real-time library, Gradio-team maintained, with built-in VAD and turn detection + Cloudflare TURN. **Still `0.0.x`** — pre-1.0 API churn is not what you want under a flagship portfolio piece in v1. Revisit for a "real-time conversation" milestone once turn-based voice is solid. |
|
| 156 |
+
|
| 157 |
+
---
|
| 158 |
+
|
| 159 |
+
### Development Tools
|
| 160 |
+
|
| 161 |
+
| Tool | Purpose | Notes |
|
| 162 |
+
|---|---|---|
|
| 163 |
+
| **uv** | Dependency resolution + venv | Fast, and its lockfile makes the "works locally, breaks on Space rebuild" class of bug reproducible. Export to `requirements.txt` for the Space (Spaces do not read `uv.lock`). |
|
| 164 |
+
| **ruff** | Lint + format | Pre-commit hook. Prior GSD work in this account has been blocked by accumulated ruff debt — enforce from commit #1 here. |
|
| 165 |
+
| **pytest** | Unit tests | Focus on the viseme-timeline builder and the JLPT-level content selector — both are pure functions with high bug density. |
|
| 166 |
+
| **Playwright** | E2E | The avatar is a canvas; assert on `trigger()` events and audio playback state, not pixels. Screenshot-diff the VRM render at your own risk. |
|
| 167 |
+
| **`hf` CLI** | Space + bucket management | `hf buckets sync` for asset deploys. |
|
| 168 |
+
| **Space README front-matter** | Config | Must pin `sdk: gradio`, `sdk_version: 6.22.0`, `python_version: 3.12.12`, `hf_oauth: true`. |
|
| 169 |
+
|
| 170 |
+
---
|
| 171 |
+
|
| 172 |
+
## Authentication
|
| 173 |
+
|
| 174 |
+
Gradio's built-in HF OAuth is the right call and is nearly free to implement:
|
| 175 |
+
|
| 176 |
+
```yaml
|
| 177 |
+
# README.md front-matter
|
| 178 |
+
sdk: gradio
|
| 179 |
+
sdk_version: 6.22.0
|
| 180 |
+
python_version: 3.12.12
|
| 181 |
+
hf_oauth: true
|
| 182 |
+
hf_oauth_expiration_minutes: 43200 # 30 days = max; a daily-practice app wants the max
|
| 183 |
+
# no extra scopes needed for progress tracking — "openid profile" is always included
|
| 184 |
+
```
|
| 185 |
+
|
| 186 |
+
```python
|
| 187 |
+
def load_progress(profile: gr.OAuthProfile | None):
|
| 188 |
+
if profile is None:
|
| 189 |
+
return anonymous_session()
|
| 190 |
+
return db.get_or_create(hf_user=profile.username) # stable subject id
|
| 191 |
+
```
|
| 192 |
+
|
| 193 |
+
Notes:
|
| 194 |
+
- Request **only** `openid profile` (always included) since progress lives in your own DB. Every extra scope is a consent-screen deterrent.
|
| 195 |
+
- Add the `inference-api` scope **only** if you implement the optional "spend my own HF credits" path (lets you call Inference Providers on the visitor's behalf — free users get ~$0.10/month of credits, PRO $2.00, so treat it as a nicety, not a tier).
|
| 196 |
+
- `hf_oauth_expiration_minutes` maxes at 43200 (30 days). Use it — re-authing daily kills a daily-habit app.
|
| 197 |
+
- ⚠️ **`gr.LogoutButton` was removed in Gradio 6.** Build your own sign-out. (`gr.LoginButton` / `gr.OAuthProfile` / `gr.OAuthToken` are believed intact — **verify on first spike**, MEDIUM confidence.)
|
| 198 |
+
- ⚠️ Use `target="_blank"` on the sign-in button, per HF docs, or third-party-cookie policies break the flow inside the Space iframe.
|
| 199 |
+
|
| 200 |
+
Confidence: **HIGH** on the OAuth mechanism ([Spaces OAuth docs](https://huggingface.co/docs/hub/spaces-oauth)); **MEDIUM** on which Gradio 6 OAuth components survived the v6 API sweep.
|
| 201 |
+
|
| 202 |
+
---
|
| 203 |
+
|
| 204 |
+
## Installation
|
| 205 |
+
|
| 206 |
+
```txt
|
| 207 |
+
# requirements.txt (Hugging Face Space)
|
| 208 |
+
gradio==6.22.0
|
| 209 |
+
spaces==0.51.1
|
| 210 |
+
|
| 211 |
+
# --- inference ---
|
| 212 |
+
torch==2.8.0 # ZeroGPU supports 2.8.0 → 2.11.0 only
|
| 213 |
+
transformers==5.14.1
|
| 214 |
+
huggingface_hub==1.27.0
|
| 215 |
+
accelerate
|
| 216 |
+
|
| 217 |
+
# --- Japanese TTS (CPU, zero GPU quota) ---
|
| 218 |
+
https://github.com/VOICEVOX/voicevox_core/releases/download/0.16.4/voicevox_core-0.16.4+cpu-cp310-abi3-manylinux_2_34_x86_64.whl
|
| 219 |
+
|
| 220 |
+
# --- Japanese NLP ---
|
| 221 |
+
SudachiPy==0.6.11
|
| 222 |
+
SudachiDict-core==20260723
|
| 223 |
+
jaconv==0.5.0
|
| 224 |
+
|
| 225 |
+
# --- persistence ---
|
| 226 |
+
psycopg[binary]==3.*
|
| 227 |
+
SQLAlchemy==2.*
|
| 228 |
+
|
| 229 |
+
# --- BYOK ---
|
| 230 |
+
anthropic==0.121.0
|
| 231 |
+
openai==2.53.0
|
| 232 |
+
```
|
| 233 |
+
|
| 234 |
+
```bash
|
| 235 |
+
# One-time, in the Space build (voicevox_core needs runtime + dict + voice models)
|
| 236 |
+
# The official downloader fetches onnxruntime, the Open JTalk dictionary, and .vvm voice models.
|
| 237 |
+
curl -sSfL https://github.com/VOICEVOX/voicevox_core/releases/latest/download/download-linux-x64 \
|
| 238 |
+
-o download && chmod +x download && ./download
|
| 239 |
+
```
|
| 240 |
+
|
| 241 |
+
> ⚠️ Fetching the VOICEVOX runtime at build time makes your Space build depend on GitHub availability. **Better:** vendor `onnxruntime` + Open JTalk dict + the chosen `.vvm` into an **HF Storage Bucket** mounted read-only at a fixed path. Deterministic, fast, and offline-safe.
|
| 242 |
+
|
| 243 |
+
Browser-side dependencies are loaded from CDN via the `head` importmap shown earlier — **there is no `package.json` in this project**, which is exactly the point of the `gr.HTML` approach.
|
| 244 |
+
|
| 245 |
+
---
|
| 246 |
+
|
| 247 |
+
## Alternatives Considered
|
| 248 |
+
|
| 249 |
+
| Recommended | Alternative | When to Use Alternative |
|
| 250 |
+
|---|---|---|
|
| 251 |
+
| **VOICEVOX** (TTS) | **Qwen3-TTS-12Hz-0.6B / 1.7B** (Apache-2.0, Jan 2026, JA among 10 langs, voice cloning from 3s) | If VOICEVOX character terms turn out to prohibit third-party web synthesis, or you want a custom/cloned tutor voice. Cost: runs on GPU (burns visitor quota), and **gives you no phoneme timings** — you fall back to amplitude-driven lip-sync or add a forced aligner. |
|
| 252 |
+
| **VOICEVOX** (TTS) | **Style-Bert-VITS2** (`2.5.0`) | If you want maximum expressive/emotional Japanese prosody. Cost: **AGPL-3.0**, which would infect your whole app unless isolated behind a separate network service. For a public portfolio repo, that is a conversation you do not want to have in an interview. |
|
| 253 |
+
| **VOICEVOX** (TTS) | **Kokoro-82M** (Apache-2.0, `kokoro-js` 1.2.1 runs in-browser) | English-first fallback, or if you need TTS with literally zero server cost. **Rejected for Japanese**: its own VOICES.md grades the five JA voices **C- to C+** on <10 hours of training data and warns "support for non-English languages may be absent or thin due to weak G2P." A pronunciation-teaching app cannot ship C-grade Japanese. |
|
| 254 |
+
| **whisper-large-v3-turbo** (ASR) | **Qwen3-ASR-1.7B** | Best measured Japanese accuracy (CER 0.140 vs 0.184) — use if transcription errors are visibly hurting the correction/scoring feature. Cost: RTF 0.036 vs 0.013 (~3× slower), which matters when GPU-seconds are the visitor's scarcest resource. |
|
| 255 |
+
| **whisper-large-v3-turbo** (ASR) | **transformers.js Whisper in-browser (WebGPU)** | **Recommend shipping both.** Browser-side = zero quota, works while the Space is queued, and is a genuinely impressive "runs on your GPU" demo. Server-side = the accuracy fallback for weak devices / no-WebGPU browsers. |
|
| 256 |
+
| **whisper-large-v3-turbo** (ASR) | **nvidia/parakeet-tdt-0.6b-v3** | Only if latency is the binding constraint (RTF 0.003) and you can tolerate CER 0.321 — too error-prone to grade a learner's pronunciation against. |
|
| 257 |
+
| **Qwen3.5-4B** (LLM) | **Nemotron Nano 9B JP** (NVIDIA) | Ranked **#1 in the sub-10B category on Nejumi Leaderboard 4** (the Japanese LLM leaderboard). Swap in if Qwen3.5-4B's Japanese pedagogy/naturalness disappoints in evaluation. Cost: 2× the size (more GPU-seconds per turn) and the NVIDIA Open Model License rather than Apache-2.0. |
|
| 258 |
+
| **Qwen3.5-4B** (LLM) | **Qwen3.5-2B / 0.8B** | If per-turn GPU-seconds prove to be the binding UX constraint. Cheapest lever for stretching an anonymous visitor's 2-minute daily quota. |
|
| 259 |
+
| **Qwen3.5-4B** (LLM) | **LLM-jp-4** (32B MoE, 3.8B active, Apache-2.0, NII) | The Japanese-sovereign, research-credible choice — reportedly MT-Bench JA 7.82 vs GPT-4o's 7.29. MoE with 3.8B active params is surprisingly viable on a 48 GB Blackwell. Strong "I evaluated Japanese-specific models" story for a Google interview. Cost: 32B of weights to load = slow cold starts. |
|
| 260 |
+
| **Neon** (DB) | **Supabase** | If you later want its Auth, Realtime, or Storage as a bundle and will pay for Pro (which removes the 7-day pause). |
|
| 261 |
+
| **Gradio `gr.HTML`** | **Gradio custom component (`gradio cc`)** | If you decide to publish `gradio-vrm-avatar` as a reusable package — genuinely good portfolio surface area, but a Phase-N+1 concern. |
|
| 262 |
+
|
| 263 |
+
---
|
| 264 |
+
|
| 265 |
+
## What NOT to Use
|
| 266 |
+
|
| 267 |
+
| Avoid | Why | Use Instead |
|
| 268 |
+
|---|---|---|
|
| 269 |
+
| **CPU Basic as the free hosting plan** | Gradio Spaces on compute now require PRO for personal accounts. CPU Basic is not the free path, and its 2 vCPU could not run a 4B LLM anyway. | ZeroGPU Gradio Space (free for up to 2 Spaces on an account in good standing). |
|
| 270 |
+
| **Python 3.13** | ZeroGPU images provide **only** 3.10.13 and 3.12.12. This is the known "Python 3.13 / gradio pin" gotcha, and on ZeroGPU it is a hard wall, not a warning. | `python_version: 3.12.12` in README front-matter. |
|
| 271 |
+
| **Unpinned `sdk_version`** | Spaces silently rebuild on push. An unpinned Gradio floats you into a breaking release while you sleep — and Gradio 6 removed/renamed a lot (`show_api`→`api_visibility`, tuple chatbot messages, `gr.LogoutButton`, `mirror_webcam`, `cache_examples="lazy"`). | Pin `sdk_version: 6.22.0` and bump deliberately. |
|
| 272 |
+
| **Gradio 5-era tutorials and code** | Gradio 6 moved `theme`/`css`/`js`/`head` from the `Blocks()` constructor to `Blocks.launch()`, consolidated all `show_*_button` params into `buttons`, and **removed tuple-format chatbot messages entirely**. Copy-pasted v5 snippets will fail in non-obvious ways. | The [Gradio 6 migration guide](https://gradio.app/main/guides/gradio-6-migration-guide) as the primary reference. |
|
| 273 |
+
| **`torch.compile` on ZeroGPU** | Explicitly unsupported. | PyTorch **ahead-of-time** compilation (`torch >= 2.8`), documented in HF's zerogpu-aoti guide. Meaningful latency win = meaningful quota win. |
|
| 274 |
+
| **Lazy `.to('cuda')` inside `@spaces.GPU`** | Docs are explicit that CUDA transfers are optimized for startup placement; lazy loading is "significantly less efficient." Every wasted second is billed to your visitor's daily quota. | Load models to `cuda` at **module level** (a CUDA emulation layer makes this work outside GPU context). |
|
| 275 |
+
| **`pykakasi`** | **GPL-3.0-or-later.** For kana/romaji conversion — a trivially replaceable function — it would impose GPL on a flagship public portfolio repo. | `SudachiPy` (Apache-2.0) readings + `jaconv` (MIT) + `cutlet` for romaji. |
|
| 276 |
+
| **Style-Bert-VITS2 embedded in-process** | **AGPL-3.0.** §13 means network use triggers source disclosure for the embedding server. | VOICEVOX (LGPL + credit) or Qwen3-TTS (Apache-2.0). |
|
| 277 |
+
| **`edge-tts` / unofficial Microsoft endpoints** | Excellent Japanese voices, but it's an undocumented consumer endpoint used against its ToS, and it breaks without notice. A hiring manager finding a ToS-violating dependency in a flagship repo is a pure downside. | VOICEVOX. |
|
| 278 |
+
| **Web Speech API (`SpeechRecognition`) as primary ASR** | Chromium-only in practice, silently ships the learner's audio to Google, and returns no usable confidence/timing data — which is exactly the data pronunciation scoring needs. | transformers.js Whisper (browser) + whisper-large-v3-turbo (ZeroGPU). Web Speech is acceptable as a last-resort tertiary fallback only. |
|
| 279 |
+
| **kotoba-whisper (v1/v2)** | Fast (6.3× large-v3) and great in-domain, but the Feb-2026 benchmark flags it as struggling with **unscripted/natural conversation** — which is 100% of this app's input. | whisper-large-v3-turbo. |
|
| 280 |
+
| **SQLite on Space disk** | Space disk is ephemeral; every `git push` rebuilds and wipes it. | Neon Postgres. |
|
| 281 |
+
| **Supabase free tier for the user DB** | Auto-pauses after 7 days idle, requires manual dashboard unpause. Guaranteed to be paused exactly when someone finally clicks your portfolio link. | Neon (scale-to-zero, auto-resume in ~hundreds of ms). |
|
| 282 |
+
| **`fastrtc` in v1** | `0.0.34` — pre-1.0, active API churn, and it wraps the Gradio version you must pin. | Turn-based `gr.Audio(sources=["microphone"])`; adopt fastrtc in a later real-time milestone. |
|
| 283 |
+
| **AI talking-head video generation** | Already Out of Scope in PROJECT.md — and it would consume a visitor's entire daily GPU quota in a single utterance. | VRM + mora-timed visemes. |
|
| 284 |
+
|
| 285 |
+
---
|
| 286 |
+
|
| 287 |
+
## Stack Patterns by Variant
|
| 288 |
+
|
| 289 |
+
**If the account is a free personal account (assume this for v1):**
|
| 290 |
+
- Hardware: **ZeroGPU**, max 2 such Spaces on the account — budget that slot deliberately against the other 4 HF-profile projects.
|
| 291 |
+
- Design every turn to fit inside ~3–6 GPU-seconds so an anonymous visitor gets 20+ turns from 2 minutes.
|
| 292 |
+
- Push ASR to the browser by default; use ZeroGPU ASR only as fallback.
|
| 293 |
+
- Free Spaces sleep after inactivity → **expect a cold start on the recruiter's first click.** Make the loading state part of the show: render the VRM avatar immediately (it's client-side and needs no backend) with a "waking up" idle animation while the Python backend boots.
|
| 294 |
+
|
| 295 |
+
**If the account upgrades to PRO ($9/mo):**
|
| 296 |
+
- Up to 10 ZeroGPU Spaces; the owner's own testing gets 40 min/day at highest queue priority.
|
| 297 |
+
- $2.00/month Inference Provider credits make an **Inference Providers–routed default LLM** viable, removing cold-start and quota concerns entirely for light traffic.
|
| 298 |
+
- Visitor quotas are unchanged — PRO does not fix the anonymous visitor's 2 minutes. Design for free-tier visitors regardless.
|
| 299 |
+
|
| 300 |
+
**If VOICEVOX character licensing blocks shipping:**
|
| 301 |
+
- Swap to **Qwen3-TTS-12Hz-0.6B** (Apache-2.0) on ZeroGPU.
|
| 302 |
+
- You lose free phoneme timings → recover lip-sync by deriving the mora sequence from **SudachiPy readings → katakana → vowel sequence** and distributing it proportionally across the audio duration, refined by a Web Audio RMS envelope. Noticeably worse than VOICEVOX timings, still far better than pure amplitude flapping.
|
| 303 |
+
|
| 304 |
+
**If ZeroGPU queueing makes conversation feel slow:**
|
| 305 |
+
- Move the default LLM to **HF Inference Providers** routed with the *visitor's* OAuth token (`inference-api` scope) — their credits, their latency, no queue.
|
| 306 |
+
- Keep BYOK (Anthropic/OpenAI) as the premium path. This is the honest three-tier story: free/queued → your-HF-credits → your-frontier-key.
|
| 307 |
+
|
| 308 |
+
---
|
| 309 |
+
|
| 310 |
+
## Version Compatibility
|
| 311 |
+
|
| 312 |
+
| Package A | Compatible With | Notes |
|
| 313 |
+
|---|---|---|
|
| 314 |
+
| `gradio==6.22.0` | Python `>=3.10` (3.10–3.13 classifiers) | But **ZeroGPU narrows this to 3.10.13 / 3.12.12**. The platform is stricter than the package. |
|
| 315 |
+
| `spaces==0.51.1` | Python `>=3.10,<3.15` | Consistent with the above. |
|
| 316 |
+
| ZeroGPU | `torch` **2.8.0 → 2.11.0**, Gradio **4+** | Pin torch explicitly; a transitive bump outside this range breaks GPU allocation. |
|
| 317 |
+
| ZeroGPU | **Gradio SDK only** | Not Docker, not Static. This makes Gradio non-negotiable, which conveniently matches the portfolio requirement. |
|
| 318 |
+
| `@pixiv/three-vrm@3.5.5` | `three >= 0.137` (peer) | Verified against `three@0.185.1`. Load both from the *same* CDN origin via importmap — mismatched three.js instances cause the classic "multiple instances of three.js imported" breakage. |
|
| 319 |
+
| `@pixiv/three-vrm-animation@3.5.5` | same peer range | Only needed if you use `.vrma` animation clips for idle/gesture motion. Recommended for idle breathing/blinking. |
|
| 320 |
+
| `transformers==5.14.1` | `huggingface_hub>=1.x` | Both are v-major releases from the last cycle; **v4-era `transformers` snippets will not run**. |
|
| 321 |
+
| `voicevox_core` 0.16.4 `cp310-abi3` wheel | Python `>=3.10` incl. 3.12 | `abi3` wheels are forward-compatible across minor versions — the `cp310` tag is not a 3.10-only restriction. Linux wheel: `manylinux_2_34_x86_64`. |
|
| 322 |
+
| `SudachiPy 0.6.11` | `SudachiDict-core 20260723` | Dict packages are date-versioned and must match the SudachiPy major line; pin both together. |
|
| 323 |
+
| Neon Postgres | `psycopg[binary]` 3.x | Requires TLS + SNI. Use the pooled endpoint; Spaces restart often and unpooled connections leak. |
|
| 324 |
+
|
| 325 |
+
---
|
| 326 |
+
|
| 327 |
+
## Open Questions for the Roadmap
|
| 328 |
+
|
| 329 |
+
1. **Does the WolfDavid account qualify for free ZeroGPU hosting** (verified email, >30 days old, <2 existing ZeroGPU Spaces)? Blocking, and cheap to check. If not, this is a $9/mo PRO decision to make before Phase 1.
|
| 330 |
+
2. **VOICEVOX character terms** — needs a human to read the specific character's 利用規約 and confirm third-party web synthesis is permitted with credit. Gate this before building the TTS phase; the fallback (Qwen3-TTS) changes the lip-sync design.
|
| 331 |
+
3. **Which VRM model?** Licensing on VRoid Hub models varies per author (some forbid commercial/redistribution). Needs the same treatment as the voice.
|
| 332 |
+
4. **Do `gr.LoginButton` / `gr.OAuthProfile` survive Gradio 6's API sweep unchanged?** MEDIUM confidence — verify in a 20-minute spike, since auth sits under the whole progress-tracking requirement.
|
| 333 |
+
5. **Real per-turn GPU-second cost** of Qwen3.5-4B on RTX Pro 6000 Blackwell is unmeasured. The entire free-tier UX budget depends on it. Measure early; it may force 2B, AoT compilation, or shorter generation caps.
|
| 334 |
+
6. **Pronunciation scoring method** is unresolved at the stack level — Whisper transcripts give you *what* was said, not *how well*. Options (forced alignment + GOP scoring, or CER against the target text) deserve their own research pass at that phase.
|
| 335 |
+
|
| 336 |
+
---
|
| 337 |
+
|
| 338 |
+
## Sources
|
| 339 |
+
|
| 340 |
+
**HIGH confidence (official documentation, fetched 2026-08-08):**
|
| 341 |
+
- [HF Spaces Overview](https://huggingface.co/docs/hub/spaces-overview) — hardware tiers, PRO requirement for Gradio/Docker Spaces, ephemeral disk, sleep behavior, built-in env vars
|
| 342 |
+
- [HF Spaces ZeroGPU](https://huggingface.co/docs/hub/spaces-zerogpu) — RTX Pro 6000 Blackwell, Gradio-only compatibility, torch 2.8–2.11, Python 3.10.13/3.12.12, per-visitor daily quotas, `@spaces.GPU`, no `torch.compile`
|
| 343 |
+
- [HF Spaces OAuth](https://huggingface.co/docs/hub/spaces-oauth) — `hf_oauth` metadata, scopes, expiration limits, `target=_blank` caveat
|
| 344 |
+
- [HF Spaces disk usage](https://huggingface.co/docs/hub/spaces-storage) + [Storage Buckets](https://huggingface.co/docs/hub/storage-buckets) — ephemerality, bucket volumes
|
| 345 |
+
- [HF Inference Providers pricing](https://huggingface.co/docs/inference-providers/pricing) — $0.10 free / $2.00 PRO monthly credits, OpenAI-compatible router
|
| 346 |
+
- [HF pricing](https://huggingface.co/pricing) — PRO $9/mo
|
| 347 |
+
- [Gradio 6 migration guide](https://gradio.app/main/guides/gradio-6-migration-guide) — breaking changes, removals
|
| 348 |
+
- [Gradio custom HTML components](https://gradio.app/main/guides/custom-HTML-components) + [HF blog: gr.HTML one-shot apps](https://huggingface.co/blog/gradio-html-one-shot-apps) — `html_template`, `head`, `js_on_load`, `server_functions`, three.js precedent
|
| 349 |
+
- [VRM 1.0 expressions spec](https://github.com/vrm-c/vrm-specification/blob/master/specification/VRMC_vrm-1.0/expressions.md) — `aa`/`ih`/`ou`/`ee`/`oh` presets
|
| 350 |
+
- [VOICEVOX core releases](https://github.com/VOICEVOX/voicevox_core/releases) — 0.16.4, manylinux abi3 CPU wheel
|
| 351 |
+
- [Kokoro-82M VOICES.md](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md) — Japanese voice grades C- to C+
|
| 352 |
+
- [Qwen3.5-4B model card](https://huggingface.co/Qwen/Qwen3.5-4B) — Apache-2.0, 201 languages, 256K context, thinking mode
|
| 353 |
+
- PyPI/npm registry APIs queried directly for every pinned version in this document
|
| 354 |
+
|
| 355 |
+
**MEDIUM confidence (third-party, single-source or benchmark-dependent):**
|
| 356 |
+
- [Japanese ASR benchmark, Feb 2026 (Neosophie)](https://neosophie.com/en/blog/20260226-japanese-asr-benchmark) — CER/WER/RTF table; single benchmark, not independently replicated
|
| 357 |
+
- [Japanese LLMs compared, Apr 2026 (lilting.ch)](https://lilting.ch/en/articles/japanese-llm-options-compared) — LLM-jp-4, Nemotron Nano 9B JP, PLaMo, Swallow, Namazu; Nejumi 4 rankings quoted secondhand
|
| 358 |
+
- [VOICEVOX engine API reference (DeepWiki)](https://deepwiki.com/VOICEVOX/voicevox_engine/4.1-tts-pipeline-api) — AudioQuery/mora structure
|
| 359 |
+
- VOICEVOX commercial-use and credit requirements — Japanese-language secondary sources; **must be confirmed against the official 利用規約 per character**
|
| 360 |
+
- [Transformers.js v4 release](https://huggingface.co/blog/transformersjs-v4) — WebGPU runtime rewrite
|
| 361 |
+
- Neon vs Supabase free-tier behavior — multiple 2026 sources agree, vendor docs corroborate scale-to-zero
|
| 362 |
+
|
| 363 |
+
**Notable correction made during research:** a widely-cited 2026 pricing blog stated PRO gets "25 minutes of daily H200 ZeroGPU." Official docs say **40 minutes** on **RTX Pro 6000 Blackwell**. Third-party pricing content on HF is stale by roughly one hardware generation — trust `huggingface.co/docs` only.
|
| 364 |
+
|
| 365 |
+
**Tooling note:** Context7, Exa, Firecrawl and Brave Search were all disabled in `.planning/config.json` for this run, so verification relied on WebFetch against official docs plus direct PyPI/npm registry queries. Every version number in this document was read from a package registry or official doc page, not from training data.
|
| 366 |
+
|
| 367 |
+
---
|
| 368 |
+
*Stack research for: avatar-based Japanese language tutor on Hugging Face Spaces*
|
| 369 |
+
*Researched: 2026-08-08*
|
.planning/research/SUMMARY.md
ADDED
|
@@ -0,0 +1,186 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Project Research Summary
|
| 2 |
+
|
| 3 |
+
**Project:** Japanese Learning Avatar
|
| 4 |
+
**Domain:** Avatar-based conversational language tutor (VRM 3D avatar + Japanese full-duplex-feel voice + hybrid LLM + persistent accounts) delivered as a public Hugging Face Space
|
| 5 |
+
**Researched:** 2026-08-08
|
| 6 |
+
**Confidence:** MEDIUM-HIGH — platform mechanics and avatar rendering are HIGH (official docs, verified 2026-08-08); TTS/ASR/LLM model choice, pedagogical patterns, and latency numbers are MEDIUM; a handful of licensing and benchmark claims are explicitly flagged LOW below.
|
| 7 |
+
|
| 8 |
+
## Executive Summary
|
| 9 |
+
|
| 10 |
+
This is a real-time, avatar-fronted conversational tutoring product wrapped in a portfolio-piece Gradio Space, and the four research passes converge tightly on how experts actually build this shape of system: the browser owns rendering and the audio clock, the server owns cognition, and a single narrow message contract (an AvatarDirective carrying audio + a precomputed viseme timeline + expression + subtitle) crosses the boundary. Gradio 6's gr.HTML custom-component API (head/html_template/js_on_load/server_functions) makes this possible with zero npm build step, which is the single most load-bearing platform finding — it is what keeps the whole project inside sdk: gradio instead of forcing a Docker SDK and a frontend build pipeline. The other structurally decisive finding is Japanese-specific and unusually favorable: Japanese is mora-timed with exactly five vowels, and VRM 1.0 defines exactly five matching viseme presets (aa/ih/ou/ee/oh). A TTS engine that emits per-mora timings (VOICEVOX, or any OpenJTalk-family engine) turns lip-sync into a nearly-free, deterministic, GPU-free feature instead of an ML problem — this is the highest-leverage decision in the entire stack and every research file independently arrived at it.
|
| 11 |
+
|
| 12 |
+
The recommended approach is turn-based voice first, hybrid free/BYOK LLM, deterministic guardrails everywhere an LLM would otherwise be trusted as ground truth. Concretely: VOICEVOX (CPU, mora timings, LGPL+credit) for TTS; whisper-large-v3-turbo or a Japanese-specialized ASR model for transcription only (never as a pronunciation scorer — Whisper actively "corrects" learner mispronunciation into fluent text, which silently teaches nothing); a small open LLM (Qwen3.5-4B, Apache-2.0) as the always-available floor with BYOK Anthropic/OpenAI as the quality upsell; a curated, JLPT-tagged grammar database as the source of truth for grammar explanations with the LLM only selecting/phrasing entries (LLMs hallucinate Japanese grammar convincingly and the small free-tier model will be the worst offender); and a deterministic post-generation OutputGuard that tokenizes every tutor reply and rejects/regenerates anything above the learner's JLPT level, because prompt-only level control provably drifts within a few turns — worse on small open models, which is exactly the default tier. SudachiPy tokenization (Apache-2.0) is the shared foundation under furigana, word lookup, level gating, and vocabulary/SRS tracking, and should land early even though it's invisible on screen.
|
| 13 |
+
|
| 14 |
+
The single biggest platform correction this research surfaced — and the main risk to manage — is hosting. PROJECT.md and the PITFALLS pass both assumed a free CPU tier (2 vCPU / 16 GB) with the expected 5-15s cascade latency; STACK research (fetched the same day, 2026-08-08) found that free Gradio/Docker Spaces on personal accounts now require a paid PRO plan to create at all — the actual free path is a Gradio Space on ZeroGPU (RTX Pro 6000 Blackwell, 48 GB VRAM), which is real datacenter GPU, not CPU. This is good news for LLM inference but introduces a harder constraint than "slow CPU": ZeroGPU quota is burned per visitor, not per owner — an anonymous visitor gets roughly 2 minutes of GPU-seconds per day, a logged-in free visitor ~5 minutes. All four research files converge on the same mitigation despite disagreeing on the underlying platform: push everything that can run without GPU off the GPU path. TTS (VOICEVOX) runs on CPU for free; ASR can run in-browser via transformers.js/WebGPU for zero quota; only the LLM turn reliably needs @spaces.GPU, and even that should be designed to fit in single-digit GPU-seconds per turn. The practical effect is that the "free CPU tier" framing in PROJECT.md should be read as "free CPU-or-browser path, with ZeroGPU as a scarce, quota-metered upgrade for the LLM" — not as 2 vCPUs running everything. Either way, the architectural discipline is identical: minimize server-side, per-turn, blocking inference, and design the interaction (push-to-talk, visible "thinking" avatar state, streaming-first-sentence playback) so that whichever hosting reality applies, turns feel responsive rather than broken.
|
| 15 |
+
|
| 16 |
+
## Key Findings
|
| 17 |
+
|
| 18 |
+
### Recommended Stack
|
| 19 |
+
|
| 20 |
+
Full detail: .planning/research/STACK.md. The stack is chosen almost entirely to minimize both GPU-seconds and licensing risk on a public portfolio repo, with Gradio 6 + gr.HTML as the mechanism that lets a full three.js/VRM avatar live inside a single-file Gradio app.
|
| 21 |
+
|
| 22 |
+
Core technologies:
|
| 23 |
+
- Gradio 6.22.0 (pinned exactly): app framework + Space SDK — the only maintained major line, and ZeroGPU is exclusively Gradio-SDK-compatible.
|
| 24 |
+
- Python 3.12.12 (pinned in README): ZeroGPU images provide only 3.10.13/3.12.12; Python 3.13 is a hard platform wall here, echoing this account's known Gradio/Python-3.13 pin gotcha.
|
| 25 |
+
- three.js 0.185.1 + @pixiv/three-vrm 3.5.5 (MIT): the only realistic in-browser VRM renderer, loaded via CDN importmap inside gr.HTML's head — no npm build.
|
| 26 |
+
- VOICEVOX CORE 0.16.4+cpu: Japanese TTS that emits per-mora phoneme timings via AudioQuery — the lip-sync decision. Runs on CPU (zero GPU quota). LGPL+per-character-credit licensing needs a human read before shipping (action item).
|
| 27 |
+
- whisper-large-v3-turbo (server, ZeroGPU) + transformers.js Whisper (browser, WebGPU, zero quota): dual-path ASR — browser-first for quota economy, server as accuracy fallback.
|
| 28 |
+
- Qwen3.5-4B (Apache-2.0): free-default tutor LLM, 201 languages incl. Japanese, fits comfortably in 48 GB with GBNF/schema-constrained decoding for structured output.
|
| 29 |
+
- SudachiPy 0.6.11 + SudachiDict-core (Apache-2.0): Japanese tokenization, readings, furigana, lemma/JLPT tagging — the shared foundation under nearly every Japanese-specific feature.
|
| 30 |
+
- Neon Serverless Postgres (free plan): the only viable persistence choice — Space disk is ephemeral, and Supabase's free-tier 7-day auto-pause is a project-killing failure mode for a portfolio link opened sporadically.
|
| 31 |
+
- spaces 0.51.1: @spaces.GPU(duration=...) allocation, dynamic durations for queue priority.
|
| 32 |
+
|
| 33 |
+
Explicitly avoided: CPU Basic as "the free tier" (no longer free for Gradio/Docker Spaces on personal accounts), pykakasi (GPL-3.0), Style-Bert-VITS2 embedded in-process (AGPL-3.0), edge-tts (ToS-violating undocumented endpoint), SQLite on Space disk, Supabase for the user DB, fastrtc in v1 (pre-1.0, Gradio-6 compatibility unverified).
|
| 34 |
+
|
| 35 |
+
### Expected Features
|
| 36 |
+
|
| 37 |
+
Full detail: .planning/research/FEATURES.md. The competitive landscape shows a genuine gap: AI speaking tutors (Langua, TalkPal) have deep pedagogy but no avatar; avatar/companion apps (Duolingo's Lily, Novatar AI) have expressive faces but shallow or absent JLPT grounding. This project's differentiation is a serious JLPT-grounded pedagogy behind an expressive VRM face — nobody currently owns that intersection.
|
| 38 |
+
|
| 39 |
+
Must have (table stakes) for v1:
|
| 40 |
+
- Mic capture → visible ASR transcript, TTS with mora-timed lip-sync, perceived turn latency under ~1.5s, replay/speak-slower/text-input fallback
|
| 41 |
+
- Furigana (toggleable) + click-any-word lookup + on-demand sentence translation, JLPT-level-appropriate output (deterministic gating, not prompt-only)
|
| 42 |
+
- Corrections backed by a grammar database (not free-form LLM authorship), level selection at onboarding, 3-5 roleplay scenarios with goal/register/exit condition
|
| 43 |
+
- HF OAuth + server-persisted progress, anonymous zero-setup path (a hiring manager must reach a talking avatar with no signup), mobile-responsive layout
|
| 44 |
+
|
| 45 |
+
Should have (differentiators): pedagogically-driven avatar expression (state-of-tutor, not sentiment-of-text — the actual novel idea and cheap to build), hard JLPT-level gating that makes a small open model usable, pitch-accent visualization (target contour first, scoring later), mistake-driven drill generation connecting conversation to SRS, conversational (no-quiz) level assessment, politeness-register dial, engineering transparency panel (latency/model breakdown — directly targets the hiring-manager audience).
|
| 46 |
+
|
| 47 |
+
Defer (v2+): FSRS SRS review queue, structured drill modules, end-of-session feedback reports, BYOK toggle, pitch-accent scoring (visualization ships first), shadowing mode, avatar memory, N1 content, gestures/multiple characters.
|
| 48 |
+
|
| 49 |
+
Explicitly anti-features: streaks/XP/gamification, free-form LLM grammar explanations, true full-duplex barge-in, romaji-first UI, photorealistic video avatars, companion/romance framing, custom SRS algorithms (use FSRS), always-on mic.
|
| 50 |
+
|
| 51 |
+
### Architecture Approach
|
| 52 |
+
|
| 53 |
+
Full detail: .planning/research/ARCHITECTURE.md. Five layers with one hard boundary: browser (rendering, animation clock, mic/playback) ↔ Gradio app shell ↔ voice pipeline + tutor core ↔ shared services (LLM gateway, Japanese analyzer, assessment engine) ↔ persistence (Neon Postgres + repo-based content). The closest verified prior art, HF-staff's victor/gemma-avatar Space, validates this layer split while being deliberately diverged from (it uses Docker SDK + cloud LLM + ARKit visemes; this project stays on Gradio SDK + VRM's native 5-vowel visemes).
|
| 54 |
+
|
| 55 |
+
Major components:
|
| 56 |
+
1. VRM Stage + Lip-Sync Driver + Audio Queue (browser, avatar/avatar.js as the sole facade) — owns the render loop and plays a precomputed VisemeTimeline off AudioContext.currentTime; never receives per-frame instructions from Python.
|
| 57 |
+
2. SessionOrchestrator + Modes (Free/Drill/RolePlay behind one ABC) — turn lifecycle; all three modes share LevelPolicy, OutputGuard, LLMGateway, AssessmentEngine by construction rather than reimplementing difficulty three times.
|
| 58 |
+
3. LLMGateway — one chat(messages, schema, stream) interface resolving BYOK key → visitor's HF OAuth credits → local/ZeroGPU model, so modes never know which provider answered.
|
| 59 |
+
4. JapaneseAnalyzer (SudachiPy) + OutputGuard — deterministic, LLM-free tokenization/level-checking that both grades learner input and validates tutor output before TTS.
|
| 60 |
+
5. Persistence Repository (Neon Postgres, append-only turns/event log) — the only module that speaks SQL; schema deliberately event-sourced so FSRS/level-recompute never requires a migration.
|
| 61 |
+
|
| 62 |
+
Suggested build order (risk-retirement, not feature-glamour): Space skeleton/deploy smoke test → avatar stage with a canned WAV (zero AI, highest novelty x risk) → voice out (TTS+viseme) → voice in (ASR echo bot, measure real latency) → LLM gateway + structured TurnResult → Japanese analyzer + LevelPolicy + OutputGuard → persistence + FSRS → content/scenarios → hybrid routing/BYOK/polish → (optional) streaming transport.
|
| 63 |
+
|
| 64 |
+
### Critical Pitfalls
|
| 65 |
+
|
| 66 |
+
Full detail: .planning/research/PITFALLS.md (16 pitfalls, mapped to a 6-phase vocabulary: P1 Voice+Avatar Skeleton → P2 Japanese Language Core → P3 Tutoring Brain → P4 Accounts & Persistence → P5 Adaptive Level/Lessons/SRS → P6 Deploy Hardening). Top five, reconciled against the STACK hosting correction:
|
| 67 |
+
|
| 68 |
+
1. Voice-loop latency makes the product feel broken — whether the bottleneck is 2-vCPU CPU or per-visitor ZeroGPU queueing, the fix is the same: push ASR/TTS off the hot server path (browser ASR, CPU TTS), stream/overlap stages, and gate P1 on a measured p50/p95 latency number on the deployed Space, not a local dev-machine guess.
|
| 69 |
+
2. Architecting the core loop on GPU quota visitors don't actually have — STACK confirms ZeroGPU is now the free hosting mechanism itself, which sharpens rather than removes this pitfall: an anonymous visitor's ~2 minutes/day must not be the only path to a working conversation. Treat GPU as a quality upgrade for the LLM call only, never the default for TTS/ASR, and test the full loop with the GPU path forcibly disabled.
|
| 70 |
+
3. "Pronunciation scoring" built on Whisper is fiction — Whisper is LM-conditioned and silently "corrects" mispronunciation into fluent, confident transcripts. Split transcription (fine) from pronunciation assessment (a separate, constrained-decoding or forced-alignment mechanism, ship later and scoped to known-target-sentence drills only).
|
| 71 |
+
4. LLM hallucinates confident, wrong Japanese grammar — worst on exactly the small free-tier model most visitors will use. Ground explanations in a curated, JLPT-tagged grammar database; the LLM selects and phrases, never authors rules.
|
| 72 |
+
5. JLPT level drift — a system-prompt-only level instruction decays within a handful of turns, markedly worse on small open models. Requires the deterministic OutputGuard (tokenize every tutor reply, reject/regenerate anything above the learner's level) as a P3 architectural component, not a prompt-engineering fix.
|
| 73 |
+
|
| 74 |
+
Two additional pitfalls carry outsized weight for this specific stack: Gradio↔three.js integration fragility (query-selecting Gradio's DOM is explicitly unsupported across versions; isolate the avatar behind its own elem_ids/facade or an iframe+postMessage contract) and Gradio global-state/BYOK-key leakage (module-level variables and gr.State default values are shared across all users in Gradio — a documented, easy-to-hit bug class that would leak one visitor's BYOK API key to another).
|
| 75 |
+
|
| 76 |
+
## Implications for Roadmap
|
| 77 |
+
|
| 78 |
+
Based on combined research, suggested phase structure (using PITFALLS' phase vocabulary, which STACK, FEATURES, and ARCHITECTURE all independently support):
|
| 79 |
+
|
| 80 |
+
### Phase 1: Voice + Avatar Loop Skeleton (spike, allowed to fail)
|
| 81 |
+
Rationale: Highest novelty × highest architectural risk, and it needs zero AI — if VRM-in-Gradio via gr.HTML doesn't work cleanly, everything downstream is void. This is also where the hosting-model decision (ZeroGPU free-tier confirmation, quota budget per turn) must be made concrete before any pedagogy work starts.
|
| 82 |
+
Delivers: A deployed Space with a VRM avatar (idle/blink/breathe), TTS+mora-timed lip-sync on a canned utterance, ASR wired as an echo bot, and a measured p50/p95 latency number on the real deployed Space (cold and warm).
|
| 83 |
+
Addresses: Table-stakes voice-conversation-loop features (mic capture, TTS, lip-sync, replay, speak-slower, text fallback) from FEATURES.md.
|
| 84 |
+
Avoids: Pitfalls 1 (latency), 2 (GPU-quota dependency), 4 (ASR hallucination on silence — VAD gate), 5 (acoustic feedback loop — AEC/half-duplex), 6 (amplitude-only lip-sync), 7 (Gradio↔three.js fragility — write the isolation-strategy decision down here).
|
| 85 |
+
|
| 86 |
+
### Phase 2: Japanese Language Core
|
| 87 |
+
Rationale: SudachiPy tokenization is the true foundation — furigana, word lookup, level gating, and vocabulary tracking all collapse without it, and it is pure CPU (free capability regardless of hosting tier). Must land before any level-gating or grammar-grounding work in Phase 3.
|
| 88 |
+
Delivers: Canonical text → [{surface, reading, lemma, pos, jlpt_level}] pipeline; furigana rendering verified on the deployed Space (ruby sanitization is a documented trap); a curated JLPT-tagged grammar-point database; the pronunciation-assessment mechanism decision (research spike — Whisper-as-scorer is out).
|
| 89 |
+
Uses: SudachiPy/SudachiDict-core, jaconv (STACK.md); JapaneseAnalyzer component (ARCHITECTURE.md).
|
| 90 |
+
Implements: Deterministic, LLM-free Japanese analysis layer shared by input grading and output guarding.
|
| 91 |
+
|
| 92 |
+
### Phase 3: Tutoring Brain (LLM + guardrails)
|
| 93 |
+
Rationale: Only viable once Phase 2's tokenizer/JLPT lists exist — the OutputGuard and grammar-grounding both consume that infrastructure directly.
|
| 94 |
+
Delivers: LLMGateway with provider resolution chain (BYOK → visitor OAuth credits → local/ZeroGPU floor), structured TurnResult schema (never free prose), deterministic level guard rejecting/regenerating out-of-level replies, grammar-DB-grounded corrections, a 50-question expert-answered regression set run against both the free and BYOK model tiers.
|
| 95 |
+
Delivers (features): Free-form conversation with level-appropriate replies, corrections with grammar-DB explanations (FEATURES.md P1 items).
|
| 96 |
+
Avoids: Pitfalls 8 (grammar hallucination), 9 (level drift), 3 revisited (constrained scoring, not open scoring).
|
| 97 |
+
|
| 98 |
+
### Phase 4: Accounts & Persistence
|
| 99 |
+
Rationale: Deliberately after the tutor loop — the schema depends on what the assessment engine actually emits (event-sourced, not derived-state-only, per Pitfall 15). Designing the DB before the tutor core exists guarantees a migration.
|
| 100 |
+
Delivers: HF OAuth (with guest-mode fully functional as the default, login as an upgrade/migration path — not a gate), Neon Postgres with an append-only turns/event-log schema, per-user session isolation discipline (zero mutable module-level globals; two-browser concurrent test as an acceptance gate), BYOK key handling (session-only, never persisted/logged).
|
| 101 |
+
Addresses: Persistent progress across sessions/devices, zero-setup anonymous path (FEATURES.md table stakes).
|
| 102 |
+
Avoids: Pitfalls 12 (ephemeral disk), 13 (Gradio state leakage / BYOK leakage), 14 (OAuth-only-works-in-Spaces, mock-token security advisory).
|
| 103 |
+
|
| 104 |
+
### Phase 5: Adaptive Level, Lessons & SRS
|
| 105 |
+
Rationale: Needs both the mode ABC (Phase 3) and item-level scheduling infrastructure (Phase 4) to exist first.
|
| 106 |
+
Delivers: Roleplay scenario library (3-5 scenarios with goal/register/exit condition), FSRS-based SRS over conversation-sourced vocabulary, mistake-driven drill generation, a written adaptation-rule spec verified against a synthetic scripted learner (level must move in both directions), politeness-register dial.
|
| 107 |
+
Addresses: Structured drills, roleplay scenarios, SRS, mistake-driven drills, conversational level assessment (FEATURES.md differentiators/P2).
|
| 108 |
+
Avoids: Pitfall 15 (adaptation theater / schema that can't evolve — this is why Phase 4's event log is a prerequisite, not this phase's own work).
|
| 109 |
+
|
| 110 |
+
### Phase 6: Deploy / Portfolio Hardening
|
| 111 |
+
Rationale: Final gate before treating the Space as launch-ready; several pitfalls are only detectable on the deployed Space, not locally.
|
| 112 |
+
Delivers: Cold-start verification (factory rebuild timing, first-paint under load), quota-exhausted-path simulation, multi-user isolation re-verification, VRM/voice-license documentation (LICENSES.md), engineering transparency panel, mobile responsiveness pass.
|
| 113 |
+
Addresses: Engineering transparency panel, Anki/CSV export (FEATURES.md P2 polish).
|
| 114 |
+
Avoids: Pitfall 16 (avatar polish eating the pedagogy budget) as a cross-phase constraint enforced here as a final audit.
|
| 115 |
+
|
| 116 |
+
### Phase Ordering Rationale
|
| 117 |
+
|
| 118 |
+
- Risk retirement over feature glamour: the avatar/voice transport (Phase 1) has the least prior art in a Gradio context and cannot be worked around if it fails, so it is proven first with zero AI dependency, exactly as ARCHITECTURE.md's build order and PITFALLS' "P1 must be a spike, not a feature phase" both independently conclude.
|
| 119 |
+
- Tokenizer-first dependency chain: FEATURES.md's dependency graph and ARCHITECTURE.md's component list agree that SudachiPy/JapaneseAnalyzer blocks nearly everything Japanese-specific (furigana, level gating, vocabulary tracking) — it lands in Phase 2 even though it's invisible on screen.
|
| 120 |
+
- Guardrails before adaptivity: the deterministic OutputGuard (Phase 3) must exist before "adaptive N5→N2" (Phase 5) can be honestly claimed, because prompt-only level control is documented to drift — this ordering directly avoids Pitfall 9.
|
| 121 |
+
- Schema before scheduling: persistence (Phase 4) is deliberately placed after the tutor core so the event-sourced schema reflects what AssessmentEngine actually produces, avoiding a Phase-5 migration (Pitfall 15).
|
| 122 |
+
- Hosting-model discipline threads through every phase: regardless of whether the deployed reality is CPU-constrained or ZeroGPU-quota-constrained, every phase's exit criteria should include "does this still work with GPU forcibly disabled / quota exhausted" — the two research files that disagreed on the platform still agreed on this design response.
|
| 123 |
+
|
| 124 |
+
### Research Flags
|
| 125 |
+
|
| 126 |
+
Phases likely needing deeper research during planning:
|
| 127 |
+
- Phase 1: Gradio 6 gr.HTML custom-component specifics under real deployment (vs. local dev), VRM/three-vrm CDN-import behavior, TTS-engine-with-timings selection and its licensing footprint, and confirming actual ZeroGPU eligibility/quota arithmetic for the WolfDavid account before committing the phase's design.
|
| 128 |
+
- Phase 2: Pronunciation-assessment mechanism (Whisper is explicitly disqualified as a scorer; forced-alignment/constrained-decoding approaches need their own research pass) and confirming furigana/ruby rendering survives Gradio's Markdown sanitization on the deployed Space.
|
| 129 |
+
- Phase 3: Grammar-grounding design (retrieval-vs-generation split) and level-validation guardrail implementation — flagged independently by both ARCHITECTURE.md and PITFALLS.md as needing a dedicated spike.
|
| 130 |
+
- Phase 9-equivalent (streaming, if pursued post-v1): FastRTC × Gradio 6 compatibility is explicitly unverified (last FastRTC release predates Gradio 6); do not plan a phase around it without an empirical spike first.
|
| 131 |
+
|
| 132 |
+
Phases with standard, well-documented patterns (skip dedicated research-phase):
|
| 133 |
+
- Phase 4: HF OAuth + Neon Postgres integration is officially documented and has clear vendor guidance; the main work is discipline (session isolation, event-log schema) rather than novel research.
|
| 134 |
+
- Phase 5: FSRS is a mature, well-documented library (py-fsrs) with an established data model; roleplay scenario schema design has consistent, converging guidance across sources.
|
| 135 |
+
- Phase 6: Deploy hardening follows a documented checklist (cold-start, LFS limits, secrets) rather than open research questions.
|
| 136 |
+
|
| 137 |
+
## Confidence Assessment
|
| 138 |
+
|
| 139 |
+
| Area | Confidence | Notes |
|
| 140 |
+
|------|------------|-------|
|
| 141 |
+
| Stack | HIGH on platform/hosting mechanics and avatar rendering (official HF/Gradio docs fetched 2026-08-08); MEDIUM on TTS/ASR/LLM model selection (third-party benchmarks); MEDIUM on VOICEVOX character licensing nuances (needs a human read of specific terms) |
|
| 142 |
+
| Features | MEDIUM-HIGH — vendor feature pages and the official JLPT spec are HIGH; competitive-landscape/"gap in the market" framing and engagement research are MEDIUM; adjacent-competitor scan (anime companion apps) is LOW-MEDIUM (thin, promotional sourcing) |
|
| 143 |
+
| Architecture | MEDIUM-HIGH — platform APIs verified against official Gradio/HF docs; pedagogical-layer patterns (alignment drift, structured turn results) verified against 2025-2026 papers; the latency budget table is explicitly LOW confidence (extrapolated, not measured on this stack) |
|
| 144 |
+
| Pitfalls | MEDIUM-HIGH — platform limits and Gradio behavior verified against official docs and the GitHub issue tracker/security advisory; pedagogy and latency thresholds from research papers and industry benchmarks; a few Japanese-TTS pitch-accent accuracy numbers are vendor-reported and marked LOW |
|
| 145 |
+
|
| 146 |
+
**Overall confidence:** MEDIUM-HIGH. The platform mechanics (Gradio, HF Spaces, ZeroGPU, OAuth, VRM/three.js integration) rest on official documentation fetched the same day as this research and are trustworthy. The pedagogical and model-selection layers rest on a mix of peer-reviewed findings (alignment drift, Whisper-as-scorer failure, LLM Japanese-grammar hallucination) and third-party benchmarks that should be treated as directional, not final — several open questions below are flagged for empirical validation early in the build rather than trusted as settled.
|
| 147 |
+
|
| 148 |
+
### Gaps to Address
|
| 149 |
+
|
| 150 |
+
- The hosting-model conflict itself (STACK vs. PITFALLS) needs one resolving decision, not two parallel plans. Confirm during Phase 1 planning: does the WolfDavid HF account currently qualify for free ZeroGPU (verified email, >30 days old, <2 existing ZeroGPU Spaces)? If not, PRO ($9/mo) becomes a pre-Phase-1 decision. Either way, design every turn to survive a forcibly-disabled GPU path.
|
| 151 |
+
- VOICEVOX character licensing — needs a human to read the specific character's terms and confirm third-party web synthesis (not just generated-audio use) is permitted with credit, before the TTS phase is built around it. The fallback (Qwen3-TTS) changes the lip-sync design (loses free phoneme timings), so this gates a design branch, not just a legal checkbox.
|
| 152 |
+
- VRM model licensing — VRoid Hub models carry per-author license flags (commercial/redistribution/credit vary); resolve before committing to a specific avatar asset, or default to a self-authored VRoid Studio character to remove the question.
|
| 153 |
+
- Real per-turn GPU-second cost of the LLM on RTX Pro 6000 Blackwell is unmeasured — the entire free-tier UX budget depends on this number; measure in Phase 1/3 rather than assuming.
|
| 154 |
+
- Pronunciation scoring mechanism is unresolved at the stack level (Whisper transcripts show what was said, not how well); this is explicitly flagged as needing its own research pass at Phase 2, not a default implementation.
|
| 155 |
+
- Whether gr.LoginButton/gr.OAuthProfile/gr.OAuthToken survived Gradio 6's API sweep unchanged — MEDIUM confidence per STACK.md; verify in a short spike since auth underlies the whole progress-tracking requirement.
|
| 156 |
+
- FastRTC × Gradio 6 compatibility — unverified; only relevant if/when a post-v1 streaming-transport phase is considered.
|
| 157 |
+
|
| 158 |
+
## Sources
|
| 159 |
+
|
| 160 |
+
### Primary (HIGH confidence)
|
| 161 |
+
- HF Spaces Overview / ZeroGPU / OAuth / Disk usage / Storage Buckets / Spaces Config Reference docs (huggingface.co/docs/hub) — hosting tiers, quota mechanics, ephemeral disk, OAuth mechanism, startup limits (all fetched 2026-08-08)
|
| 162 |
+
- Gradio 6 Migration Guide / Custom HTML Components / Custom CSS and JS / State in Blocks (gradio.app) — breaking changes, gr.HTML custom-component API, DOM-selector unsupported guarantee, global-vs-session state
|
| 163 |
+
- VRM 1.0 expressions spec (github.com/vrm-c/vrm-specification) / @pixiv/three-vrm — five viseme presets, expressionManager API
|
| 164 |
+
- VOICEVOX core releases / AudioQuery-Mora model (deepwiki.com/VOICEVOX) — mora-timing structure
|
| 165 |
+
- JLPT official level summary (jlpt.jp) — N1-N5 competence definitions, calibration anchor
|
| 166 |
+
- Gradio security advisory GHSA-h3h8-3v2v-rg7m and state-leakage issues #4558, #2132, #9983
|
| 167 |
+
- victor/gemma-avatar Space (huggingface.co) — verified prior-art reference implementation
|
| 168 |
+
- hugginface_profile/docs/ai-model-updates-2026-07-02.md — first-party prior local experience on Gradio/Python version pinning gotchas on this account
|
| 169 |
+
|
| 170 |
+
### Secondary (MEDIUM confidence)
|
| 171 |
+
- Japanese ASR benchmark, Feb 2026 (Neosophie) — CER/WER/RTF comparisons, single benchmark
|
| 172 |
+
- Japanese LLMs compared, Apr 2026 (lilting.ch) — Nejumi leaderboard rankings quoted secondhand
|
| 173 |
+
- Langua/Duolingo/Busuu vendor feature pages — competitive feature landscape
|
| 174 |
+
- Alignment Drift in CEFR-prompted LLMs, BEA 2025 (ACL Anthology) — level-control decay finding
|
| 175 |
+
- Building Tailored Speech Recognizers for Japanese Speaking Assessment, arXiv 2509.20655 — Whisper unsuitable for phonemic/pronunciation assessment
|
| 176 |
+
- Neon vs Supabase free-tier behavior comparisons (multiple 2026 sources agree, vendor docs corroborate)
|
| 177 |
+
- FSRS vs SM-2 comparisons — ts-fsrs data models (deepwiki.com)
|
| 178 |
+
|
| 179 |
+
### Tertiary (LOW confidence)
|
| 180 |
+
- AnySpeech Japanese TTS pitch-accent accuracy claim (91%) — vendor-reported, single source
|
| 181 |
+
- Adjacent anime-companion-app competitor scan (Novatar AI, AITOMO) — thin, promotional sourcing
|
| 182 |
+
- Free-Space sleep interval (~48h) — widely reported but not re-verified against current docs at research time
|
| 183 |
+
|
| 184 |
+
---
|
| 185 |
+
*Research completed: 2026-08-08*
|
| 186 |
+
*Ready for roadmap: yes*
|