Spaces:
Running on Zero
Running on Zero
docs(02-05): complete canonical analyzer and fixed sample sentence set plan
Browse files- 02-05-SUMMARY.md: warm-up per component, 0.14-0.19 ms warm analysis,
the four accepted edges (#14, #15, #17, #27), the three list-level
corrections (日本語 / ない / こんにちは) pinned by JMdict id, quick loop
293 passed / 30.7 s, commit SHAs, self-check
- STATE.md: plan 6 of 11, progress 16/22, metric, four decisions, session
- ROADMAP.md: Phase 02 plan progress 5/11
.planning/ROADMAP.md
CHANGED
|
@@ -58,7 +58,7 @@ Decimal phases appear between their surrounding integers in numeric order.
|
|
| 58 |
- [x] 02-02-PLAN.md — Tokenizer singleton, whole-word units (rules A0–A7), okurigana ruby alignment, reading-override table (TDD) [wave 2]
|
| 59 |
- [x] 02-03-PLAN.md — Ranked JMdict lookup + JLPT word/kanji level derivation (N1+, name) (TDD) [wave 2]
|
| 60 |
- [x] 02-04-PLAN.md — OPUS-MT ja-en → CTranslate2 int8 under LFS, conversion script, translate() runtime (mt marker) [wave 2]
|
| 61 |
-
- [
|
| 62 |
- [ ] 02-06-PLAN.md — tokens on the directive; analyze / translate / language_info server functions; warm-up on load; turn-loop + facade surface; parity [wave 4]
|
| 63 |
- [ ] 02-07-PLAN.md — Furigana: transcript.js renderer, three-mode control + level picker, persistence, three-layer numbers (JPN-02) [wave 5]
|
| 64 |
- [ ] 02-08-PLAN.md — Word-lookup popover with published numbers, desktop + Pixel 7 rows (JPN-03) [wave 6]
|
|
|
|
| 58 |
- [x] 02-02-PLAN.md — Tokenizer singleton, whole-word units (rules A0–A7), okurigana ruby alignment, reading-override table (TDD) [wave 2]
|
| 59 |
- [x] 02-03-PLAN.md — Ranked JMdict lookup + JLPT word/kanji level derivation (N1+, name) (TDD) [wave 2]
|
| 60 |
- [x] 02-04-PLAN.md — OPUS-MT ja-en → CTranslate2 int8 under LFS, conversion script, translate() runtime (mt marker) [wave 2]
|
| 61 |
+
- [x] 02-05-PLAN.md — The canonical analyze(text) + the 28-sentence golden fixture set (SC-1) [wave 3]
|
| 62 |
- [ ] 02-06-PLAN.md — tokens on the directive; analyze / translate / language_info server functions; warm-up on load; turn-loop + facade surface; parity [wave 4]
|
| 63 |
- [ ] 02-07-PLAN.md — Furigana: transcript.js renderer, three-mode control + level picker, persistence, three-layer numbers (JPN-02) [wave 5]
|
| 64 |
- [ ] 02-08-PLAN.md — Word-lookup popover with published numbers, desktop + Pixel 7 rows (JPN-03) [wave 6]
|
.planning/STATE.md
CHANGED
|
@@ -3,13 +3,13 @@ gsd_state_version: 1.0
|
|
| 3 |
milestone: v1.0
|
| 4 |
milestone_name: milestone
|
| 5 |
status: unknown
|
| 6 |
-
stopped_at: "Completed 02-
|
| 7 |
-
last_updated: "2026-09-
|
| 8 |
progress:
|
| 9 |
total_phases: 6
|
| 10 |
completed_phases: 1
|
| 11 |
total_plans: 22
|
| 12 |
-
completed_plans:
|
| 13 |
---
|
| 14 |
|
| 15 |
# Project State
|
|
@@ -24,7 +24,7 @@ See: .planning/PROJECT.md (updated 2026-08-08)
|
|
| 24 |
## Current Position
|
| 25 |
|
| 26 |
Phase: 02 (japanese-language-core) — EXECUTING
|
| 27 |
-
Plan:
|
| 28 |
|
| 29 |
## Performance Metrics
|
| 30 |
|
|
@@ -65,6 +65,7 @@ Plan: 5 of 11
|
|
| 65 |
| Phase 02 P02 | 12 | 3 tasks | 8 files |
|
| 66 |
| Phase 02 P03 | 11 | 2 tasks | 5 files |
|
| 67 |
| Phase 02 P04 | 17 | 2 tasks | 13 files |
|
|
|
|
| 68 |
|
| 69 |
## Accumulated Context
|
| 70 |
|
|
@@ -140,6 +141,10 @@ Recent decisions affecting current work:
|
|
| 140 |
- [Phase 02]: The CTranslate2 converter writes shared_vocabulary.json/config.json with the platform newline; convert_mt.py normalises them to LF because the repo has core.autocrlf=input and hashing a CRLF working copy would fail test_mt_files_match_readme on every fresh clone (the Space). Research's 233 B / 1,269,578 B were Windows-CRLF sizes; committed bytes are 223 B / 1,208,862 B
|
| 141 |
- [Phase 02]: data/mt/README.md carries run measurements (read-back ms, conversion s) and stays idempotent because the script leaves it untouched when every recorded hash matches; the conversion proves sentencepiece-only pieces == MarianTokenizer on 7 sentences before recording, so the Space imports neither transformers nor torch
|
| 142 |
- [Phase 02]: JPN-04 left Pending after 02-04: 02-VALIDATION.md binds it to 02-10's deployed test_translation_reveal + docs/LATENCY.md § Translation row; translate() local baseline 0.35 s warm-up, 22-286 ms per contract sentence at intra_threads=2, MT_INTRA_THREADS is the contention lever
|
|
|
|
|
|
|
|
|
|
|
|
|
| 143 |
|
| 144 |
### Pending Todos
|
| 145 |
|
|
@@ -161,8 +166,8 @@ None yet.
|
|
| 161 |
|
| 162 |
## Session Continuity
|
| 163 |
|
| 164 |
-
Last session: 2026-09-
|
| 165 |
-
Stopped at: Completed 02-
|
| 166 |
Resume file: None
|
| 167 |
|
| 168 |
**Roadmap-level open decision:** the Phase 1 latency trigger (p50 7.1 s, ~90 % server synthesis) fired; recorded in `docs/LATENCY.md` and in 02-CONTEXT.md § Deferred. Phase 2 adds a CPU translation model to the same container — measure on the Space. Consider `/gsd:insert-phase` for streaming/synthesis relocation before Phase 3 lengthens turns.
|
|
|
|
| 3 |
milestone: v1.0
|
| 4 |
milestone_name: milestone
|
| 5 |
status: unknown
|
| 6 |
+
stopped_at: "Completed 02-05-PLAN.md (wave 3): analyze(text) canonical token record (14 keys, 0.14-0.19 ms warm), levels/ruby is_kanji reconciled, 28-sentence golden fixture + refusal-guard generator, 14 analyzer tests; quick loop 293 passed / 30.7 s. Next 02-06."
|
| 7 |
+
last_updated: "2026-09-07T00:08:47.920Z"
|
| 8 |
progress:
|
| 9 |
total_phases: 6
|
| 10 |
completed_phases: 1
|
| 11 |
total_plans: 22
|
| 12 |
+
completed_plans: 16
|
| 13 |
---
|
| 14 |
|
| 15 |
# Project State
|
|
|
|
| 24 |
## Current Position
|
| 25 |
|
| 26 |
Phase: 02 (japanese-language-core) — EXECUTING
|
| 27 |
+
Plan: 6 of 11
|
| 28 |
|
| 29 |
## Performance Metrics
|
| 30 |
|
|
|
|
| 65 |
| Phase 02 P02 | 12 | 3 tasks | 8 files |
|
| 66 |
| Phase 02 P03 | 11 | 2 tasks | 5 files |
|
| 67 |
| Phase 02 P04 | 17 | 2 tasks | 13 files |
|
| 68 |
+
| Phase 02 P05 | 12 | 2 tasks | 7 files |
|
| 69 |
|
| 70 |
## Accumulated Context
|
| 71 |
|
|
|
|
| 141 |
- [Phase 02]: The CTranslate2 converter writes shared_vocabulary.json/config.json with the platform newline; convert_mt.py normalises them to LF because the repo has core.autocrlf=input and hashing a CRLF working copy would fail test_mt_files_match_readme on every fresh clone (the Space). Research's 233 B / 1,269,578 B were Windows-CRLF sizes; committed bytes are 223 B / 1,208,862 B
|
| 142 |
- [Phase 02]: data/mt/README.md carries run measurements (read-back ms, conversion s) and stays idempotent because the script leaves it untouched when every recorded hash matches; the conversion proves sentencepiece-only pieces == MarianTokenizer on 7 sentences before recording, so the Space imports neither transformers nor torch
|
| 143 |
- [Phase 02]: JPN-04 left Pending after 02-04: 02-VALIDATION.md binds it to 02-10's deployed test_translation_reveal + docs/LATENCY.md § Translation row; translate() local baseline 0.35 s warm-up, 22-286 ms per contract sentence at intra_threads=2, MT_INTRA_THREADS is the contention lever
|
| 144 |
+
- [Phase 02]: analyze(text) is the ONE canonical text -> tokens function: records built from TOKEN_KEYS (14 keys, fixed order), level from the level_key entry and gloss from the lemma entry, non-tappable units carry jlpt None / gloss [] / ruby [[surface, None]]; warm analysis 0.14-0.19 ms per 36-mora sentence
|
| 145 |
+
- [Phase 02]: Three plan-assumed word levels corrected to what the pinned lists say and pinned by JMdict id, not patched: 日本語 1464530 is N1+ (on no list; kanji all N5 - the two-axis example), ない/なかった resolve to 無い 1529520 which no list carries (N1+), こんにちは 1289400 is N3; 人気 read ひとけ is the separate entry 1367020 (N1+). A level-alias table is the honest fix if the product wants otherwise (02-10 review), never a weaker lookup
|
| 146 |
+
- [Phase 02]: levels.is_kanji IS ruby.is_kanji (02-03's parallel-wave copy removed): one predicate decides both ruby and the kanji-axis gate, so supplementary-plane kanji (𠮷) reach the axis as unlisted instead of being skipped
|
| 147 |
+
- [Phase 02]: The fixed sample sentence set (SC-1) follows the golden_timeline pattern: 28 sentences / 115 records generated by the real pipeline, 18 EXPECTED facts checked before writing (exit 1, nothing written otherwise), frozen LF byte-for-byte with data pins read from the environment; JPN-01 still bound to 02-10's rows
|
| 148 |
|
| 149 |
### Pending Todos
|
| 150 |
|
|
|
|
| 166 |
|
| 167 |
## Session Continuity
|
| 168 |
|
| 169 |
+
Last session: 2026-09-07T00:08:47.915Z
|
| 170 |
+
Stopped at: Completed 02-05-PLAN.md (wave 3): analyze(text) canonical token record (14 keys, 0.14-0.19 ms warm), levels/ruby is_kanji reconciled, 28-sentence golden fixture + refusal-guard generator, 14 analyzer tests; quick loop 293 passed / 30.7 s. Next 02-06.
|
| 171 |
Resume file: None
|
| 172 |
|
| 173 |
**Roadmap-level open decision:** the Phase 1 latency trigger (p50 7.1 s, ~90 % server synthesis) fired; recorded in `docs/LATENCY.md` and in 02-CONTEXT.md § Deferred. Phase 2 adds a CPU translation model to the same container — measure on the Space. Consider `/gsd:insert-phase` for streaming/synthesis relocation before Phase 3 lengthens turns.
|
.planning/phases/02-japanese-language-core/02-05-SUMMARY.md
ADDED
|
@@ -0,0 +1,187 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
phase: 02-japanese-language-core
|
| 3 |
+
plan: 05
|
| 4 |
+
subsystem: nlp
|
| 5 |
+
tags: [analyzer, token-record, golden-fixture, sudachipy, jmdict, jlpt, furigana, tdd, pytest]
|
| 6 |
+
|
| 7 |
+
# Dependency graph
|
| 8 |
+
requires:
|
| 9 |
+
- phase: 02-japanese-language-core
|
| 10 |
+
plan: 02
|
| 11 |
+
provides: "tokenizer.morphemes() plain dicts; units.build_units() (A0-A7); ruby.align_ruby() / is_kanji(); overrides.apply_overrides() + data/nlp/reading_overrides.json"
|
| 12 |
+
- phase: 02-japanese-language-core
|
| 13 |
+
plan: 03
|
| 14 |
+
provides: "jmdict.load()/lookup()/glosses_for() ranked headword lookup; levels.vocab_levels() (7,748 ids) / kanji_levels() (2,211) / derive_level()"
|
| 15 |
+
- phase: 01-voice-avatar-loop-skeleton
|
| 16 |
+
provides: "tests/fixtures/make_golden_timeline.py refuse-to-write pattern; voice/tts.py warmup() shape"
|
| 17 |
+
provides:
|
| 18 |
+
- "src/japanese_avatar/nlp/analyzer.py: analyze(text) -> list[dict] with exactly TOKEN_KEYS (14 keys, fixed order), warmup() -> {tokenizer_s, jmdict_s, jmdict_entries, vocab_ids, kanji}; the ONE canonical text -> tokens function, no LLM anywhere on the path"
|
| 19 |
+
- "tests/fixtures/sentences.json: the fixed sample sentence set of success criterion 1 - 28 sentences, 115 golden token records, with the data pins that produced them"
|
| 20 |
+
- "tests/fixtures/make_sentences.py: regenerator that checks 18 EXPECTED facts before writing and exits 1 with nothing written otherwise; byte-idempotent"
|
| 21 |
+
- "tests/test_analyzer.py (14 tests): D-10/D-11/D-12/D-13 as named tests, the override, the ambiguous readings, exact golden equality, contract coverage, pin check, the honest list gaps"
|
| 22 |
+
- "tests/conftest.py: session-scoped `analyzer` fixture - one dictionary open + one compact-JMdict load per session"
|
| 23 |
+
- "levels.py now uses ruby.is_kanji (02-03's parallel-wave copy removed)"
|
| 24 |
+
affects: [02-06, 02-07, 02-08, 02-10, 02-11, phase-3-level-guard, phase-5-vocab-tracking]
|
| 25 |
+
|
| 26 |
+
# Tech tracking
|
| 27 |
+
tech-stack:
|
| 28 |
+
added: []
|
| 29 |
+
patterns:
|
| 30 |
+
- "One canonical analyze(text): every consumer reads the same 14-key plain-dict record; engine objects never leave tokenizer.py; the dict is built from TOKEN_KEYS so key order is a contract, not an accident"
|
| 31 |
+
- "Level from the level_key entry, gloss from the lemma entry (research § Q2 step 5) - one ranked lookup when the two keys coincide, which is the common case"
|
| 32 |
+
- "Golden language fixture = golden_timeline pattern: generated once by the real pipeline, guarded by explicit EXPECTED facts the generator refuses to violate, frozen byte-for-byte (LF, indent=1, ensure_ascii=False) with the data pins that produced it"
|
| 33 |
+
- "Upstream data quirks are pinned BY JMDICT ID in a named test (test_known_list_gaps_are_honest_n1_plus) so a future level alias is a conscious edit to the test, never a weaker lookup"
|
| 34 |
+
|
| 35 |
+
key-files:
|
| 36 |
+
created:
|
| 37 |
+
- src/japanese_avatar/nlp/analyzer.py
|
| 38 |
+
- tests/fixtures/make_sentences.py
|
| 39 |
+
- tests/fixtures/sentences.json
|
| 40 |
+
- tests/test_analyzer.py
|
| 41 |
+
modified:
|
| 42 |
+
- src/japanese_avatar/nlp/levels.py
|
| 43 |
+
- tests/conftest.py
|
| 44 |
+
- tests/test_levels.py
|
| 45 |
+
|
| 46 |
+
key-decisions:
|
| 47 |
+
- "The plan's three assumed word levels were corrected to what the pinned lists say and pinned honestly under D-10 rather than patched: 日本語 (1464530) is N1+ - it is on none of the five yomitan-jlpt-vocab CSVs (日本 itself is N3 there); ない / なかった resolve correctly to 無い 1529520, which no list carries, so they badge N1+ (the plan's #5 'ない tappable N5' and research's 'badge shows N5 (無い)' were guesses); こんにちは 1289400 is on N3 and N1, so it badges N3. The lookup was not weakened for any of them"
|
| 48 |
+
- "人気 read ひとけ is the separate JMdict entry 1367020 ('sign of life'), unlisted -> N1+, while にんき is 1367010 N3: the override changes the reading and the reading-first rank then picks the right WORD, which is exactly what a learner tapping 人気のない通り should see"
|
| 49 |
+
- "Non-tappable units carry ruby [[surface, None]] rather than align_ruby's output: a punctuation 'reading' is the mark itself and would otherwise become an rt over 。"
|
| 50 |
+
- "tests/test_levels.py (not in the plan's file list) had to change: its test_is_kanji_ranges_local_copy asserted on levels._is_kanji, which the plan requires removing; it is now test_is_kanji_is_rubys and pins identity with ruby.is_kanji plus the supplementary-plane consequence (𠮷野家 -> {𠮷: None, 野: N4, 家: N4})"
|
| 51 |
+
- "JPN-01 NOT marked complete: 02-VALIDATION.md binds it to 02-10's rows (test_fixture_sentences here + the deployed test_language_assets_loaded); this plan delivers the quick-loop half"
|
| 52 |
+
|
| 53 |
+
patterns-established:
|
| 54 |
+
- "Session-scoped warm fixture that prints its per-component seconds, so the quick loop reports the language core's real cost every run (0.03 s dictionary, ~2.0-2.4 s compact JMdict)"
|
| 55 |
+
- "A regenerator records the versions it ran under by READING them (sudachipy.__version__, importlib.metadata, the compact file's meta.version, the build script's pinned tag/commit), and a test compares them with the installed environment so a stale golden file and a stale venv are told apart"
|
| 56 |
+
|
| 57 |
+
requirements-completed: [] # JPN-01 is marked only in 02-10 on its bound rows (02-VALIDATION.md)
|
| 58 |
+
|
| 59 |
+
# Metrics
|
| 60 |
+
duration: 12min
|
| 61 |
+
completed: 2026-09-06
|
| 62 |
+
---
|
| 63 |
+
|
| 64 |
+
# Phase 02 Plan 05: The Canonical Analyzer and the Fixed Sample Sentence Set Summary
|
| 65 |
+
|
| 66 |
+
**One function, `analyze(text)`, turns any Japanese string into 14-key token records - whole-word surfaces, hiragana readings after the override table, dictionary forms, word level by JMdict-id join, kanji level per character, okurigana-aligned ruby spans, up to 3x3 English glosses and offsets - with no LLM on the path, in 0.14-0.19 ms per 36-mora sentence once warm; and success criterion 1 is now a committed file: 28 sentences / 115 golden records that the analyzer reproduces byte-for-byte and that the generator refuses to overwrite with wrong 今日 / 行った / 人気 readings.**
|
| 67 |
+
|
| 68 |
+
## Performance
|
| 69 |
+
|
| 70 |
+
- **Duration:** 12 min (started 2026-09-06T23:53:18Z, last task commit 2026-09-07T00:05:16Z)
|
| 71 |
+
- **Tasks:** 2 (Task 1 TDD: RED + GREEN; Task 2 feature)
|
| 72 |
+
- **Files:** 7 (4 created, 3 modified)
|
| 73 |
+
- **Warm-up, per component** (three fresh processes): tokenizer (Sudachi core) **0.030 / 0.032 / 0.033 s**; compact JMdict **1.97 / 2.32 / 2.40 s**; `vocab_levels()` + `kanji_levels()` are inside the noise (< 0.05 s). `warmup()` returns `{'tokenizer_s', 'jmdict_s', 'jmdict_entries': 218672, 'vocab_ids': 7748, 'kanji': 2211}`.
|
| 74 |
+
- **Warm analysis of the 36-mora sentence** (今日はいい天気ですから、公園を散歩してから、買い物に行きました。, 16 units): **mean 0.142 ms / 0.194 ms over 100 runs** in two test sessions (plan expectation < 5 ms; assertion budget 50 ms).
|
| 75 |
+
- **Quick loop** (`pytest tests --ignore=tests/e2e -m "not mt"`): after Task 1 **289 passed, 14 deselected, 28.8 s pytest / 30.7 s wall**; after Task 2 **293 passed, 14 deselected, 28.8 s / 30.7 s wall** (budget 40 s; was 283 before this plan). The session fixture makes the JMdict load a one-time ~2.3 s.
|
| 76 |
+
- `tests/test_analyzer.py` alone: 14 passed in 2.5-2.6 s (dominated by the one JMdict load).
|
| 77 |
+
|
| 78 |
+
## Accomplishments
|
| 79 |
+
|
| 80 |
+
- **`analyzer.py`** composes the pure functions of 02-02 and 02-03 in the plan's order: `morphemes -> build_units -> apply_overrides -> (per tappable unit) align_ruby, lookup x1-2, derive_level, kanji_levels_for, glosses_for`. Each record is built from `TOKEN_KEYS`, so `tuple(u.keys()) == TOKEN_KEYS` for every unit of every sentence (asserted in three places). Non-tappable units (D-11) carry `jlpt None`, `jmdict_id None`, `gloss []`, `kanji_levels {}`, `ruby [[surface, None]]`. `grep -c sudachipy analyzer.py` is 0.
|
| 81 |
+
- **Level and gloss from different entries when they must be:** 話せる -> gloss from its own entry 1562360 ("to be able to speak"), level N5 from level_key 話す; 勉強しています -> lemma 勉強する (no such headword), level and gloss from level_key 勉強 (1512670, N5); 三人 -> level_key `"3"` matches no headword, lemma 三 gives 1579350 N5.
|
| 82 |
+
- **The reconciliation the wave-2 plans left:** `levels.py` imports `is_kanji` from `ruby.py`; `def _is_kanji` is gone (grep 0, `import is_kanji` 1). Consequence made visible: a supplementary-plane kanji (𠮷) now reaches the kanji axis as unlisted (`None`) instead of being skipped. `test_kanji_axis` passes unchanged.
|
| 83 |
+
- **The fixed sample sentence set** (`tests/fixtures/sentences.json`, 59,029 bytes, LF only, `filter: unspecified` - text, not LFS): 28 sentences in the plan's order, 115 units, `analyzer` block `{sudachipy 0.6.11, sudachidict_core 20260723, jmdict 3.6.2+20260831182826, jlpt_vocab 2025.08.01.0, kanji_data 00fd7079c3890f430759536f91aa5e854ec0ca4f}` - every pin read from the environment, none typed. The generator ran three times to the identical SHA-256 `8d97fd7683e4de6fae430ca1523a2cdc31a3188511ae33b0774a69275adfe89a`; after the commit a further run left `git status --porcelain tests/fixtures/sentences.json` empty.
|
| 84 |
+
- **Spot-checked by eye** (the plan's list #2, #3, #4, #5, #6, #7, #11, #12, #13, #15, #16, #17, #20, #27, #28 plus #1, #9, #14, #21, #22, #23, #24): 今日 きょう N5; 行った いった/行く N5 vs おこなった/行う N4; 人気 ひとけ (1367020) vs にんき N3 (1367010); 食べました one unit `[[食,た],[べました,None]]` N5 1358280 "to eat"; 美味しかった|です with `[[美味,おい],[しかった,None]]`; きれい|じゃ|なかった tappable T/F/T; 田中さん name (no entry, gloss []) and 東京 name (entry, gloss present); 話せる N5 1562360; 学生 N5 + です; 待ち合わせ `[[待,ま],[ち,None],[合,あ],[わせ,None]]` N1; 買い物 `[[買,か],[い,None],[物,もの]]` N5; 取り扱い `[[取,と],[り,None],[扱,あつか],[い,None]]` N1; 大阪駅 `[[大阪駅,おおさかえき]]` name, kanji {大 N5, 阪 None, 駅 N4}; コーヒー no rt, reading こーひー, N5; 美味しくない ONE unit, lemma 美味しい, N5. All eight okurigana spans equal research § Q1's table.
|
| 85 |
+
- **Tests:** 14 `def test_` in `tests/test_analyzer.py` (plan minimum 10 then 13): `test_record_keys_exact`, `test_reading_disambiguation`, `test_hitoke_override`, `test_tappability` (D-11), `test_level_derivation` (D-12, D-10 via 瀟洒 - a JMdict entry on no list), `test_kanji_axis_separate_from_word_axis` (D-13), `test_gloss_shipped_with_token` (D-07), `test_offsets_cover_text`, `test_analyze_is_fast`, `test_empty_text`, `test_known_list_gaps_are_honest_n1_plus`, `test_fixture_sentences` (exact equality; failure text names the regeneration command and says REVIEW the diff), `test_fixture_set_covers_the_contract` (also pins the five ambiguous readings as readings), `test_fixture_pins_match_installed`.
|
| 86 |
+
|
| 87 |
+
## The four accepted edges (observed units, now frozen in the fixture)
|
| 88 |
+
|
| 89 |
+
| # | Sentence | Units (surface [reading | jlpt]) | The edge |
|
| 90 |
+
|---|---|---|---|
|
| 91 |
+
| 14 | 三人で行きました。 | 三人 [さんにん \| **N5**] · で · 行きました [いきました \| N5] · 。 | 三人 is one unit (A6); `level_key` is Sudachi's `"3"`, which matches no JMdict headword, so both level and gloss come from lemma 三 (1579350, "three"); `kanji_levels {三: N5, 人: N5}`; ruby `[[三人, さんにん]]`. で is 助動詞 (lemma だ), non-tappable |
|
| 92 |
+
| 15 | お願いします | お願い [おねがい \| **N3**] · します [します \| N5] | Both tappable. お願い's head is 願い (動詞,非自立可能), so `lemma` is 願う -> 1217950 "to desire; to wish", N3; ruby `[[お,None],[願,ねが],[い,None]]`; します -> lemma する, level_key 為る, 1157170 N5. The card for お願い therefore names 願う, not お願い (1001720) - a consequence of A7 + the head rule, recorded for 02-08/02-10's review |
|
| 93 |
+
| 17 | 行かなければならない | 行かなけれ [いかなけれ \| N5] · ば · ならない [ならない \| **N3**] | 02-02's observed split. 行かなけれ -> 行く 1578850 N5, ruby `[[行,い],[かなけれ,None]]`; ば breaks the chain (non-tappable); ならない -> lemma なる, level_key 成る -> 1375610, **N3** by the lists (the なる quirk 02-03 recorded: n5.csv's なる row is the archaic copula 2138260) |
|
| 94 |
+
| 27 | 鬱陶しい天気だ。 | 鬱陶しい [うっとうしい \| **N1**] · 天気 [てんき \| N5] · だ · 。 | 鬱陶しい IS on the N1 list (1568430), so its word level is N1 not N1+; `kanji_levels {鬱: None, 陶: N1}` - the unlisted kanji is on the kanji axis; ruby `[[鬱陶,うっとう],[しい,None]]` |
|
| 95 |
+
|
| 96 |
+
Other observed levels worth knowing (all in the golden file): こんにちは **N3** (1289400 is on N3 and N1), 人気/ひとけ **N1+** (1367020 unlisted), ない / なかった **N1+** (無い 1529520 unlisted), 日本語 **N1+** (1464530 unlisted) with kanji all N5, なりたい **N3** (なる), はじめまして N4, よろしく N3, 待ち合わせ / 取り扱い N1, 注意してください N4 (lemma 注意する), 遅れた N4.
|
| 97 |
+
|
| 98 |
+
## Task Commits
|
| 99 |
+
|
| 100 |
+
| Task | Commit | Type |
|
| 101 |
+
|---|---|---|
|
| 102 |
+
| 1 RED - failing analyzer behaviour tests + session fixture | `a09270c` | test |
|
| 103 |
+
| 1 GREEN - analyze(text), warmup(), levels.py reconciliation, test_levels follow-up | `46459e0` | feat |
|
| 104 |
+
| 2 - generator with refusal guards, golden file, golden tests | `b0413c1` | feat |
|
| 105 |
+
|
| 106 |
+
All commits `--no-verify`, explicit paths only; nothing pushed; no `git add -A`.
|
| 107 |
+
|
| 108 |
+
## Files Created/Modified
|
| 109 |
+
|
| 110 |
+
- `src/japanese_avatar/nlp/analyzer.py` (new, ~150 lines) - `TOKEN_KEYS`, `_lookup_pair(unit, level_of)`, `_token(unit, level_of)`, `analyze(text)`, `warmup()`, a dev-only `_self_time()`; docstring carries the record's field-by-field meaning and the pipeline order
|
| 111 |
+
- `src/japanese_avatar/nlp/levels.py` - `from japanese_avatar.nlp.ruby import is_kanji`; the local `_KANJI_RANGES` / `_is_kanji` block removed; `kanji_levels_for` unchanged otherwise
|
| 112 |
+
- `tests/conftest.py` - session-scoped `analyzer` fixture beside `synth_meta` (`--space-url` untouched); `scope="session"` count is now 7
|
| 113 |
+
- `tests/test_analyzer.py` (new) - 14 tests as listed; `_golden()` cached loader; `REGENERATE` message constant
|
| 114 |
+
- `tests/test_levels.py` - `test_is_kanji_ranges_local_copy` -> `test_is_kanji_is_rubys` (identity with `ruby.is_kanji`, `not hasattr(levels, "_is_kanji")`, supplementary plane, 𠮷野家 levels)
|
| 115 |
+
- `tests/fixtures/make_sentences.py` (new) - `SENTENCES` (28 x (text, what it pins)), 18 `EXPECTED` check functions (the plan's 15 facts + tiling + tappable<->jlpt + exact keys), `data_pins()`, LF writer, per-sentence review print
|
| 116 |
+
- `tests/fixtures/sentences.json` (new) - the golden file
|
| 117 |
+
|
| 118 |
+
## Decisions Made
|
| 119 |
+
|
| 120 |
+
See `key-decisions` in the frontmatter. The load-bearing one: three word levels the plan assumed (日本語 N5, ない N5, こんにちは implied N5) are not what the pinned lists say; each was verified against the CSVs and the lookup (the lookup picks the right entry in every case), then pinned honestly as N1+ / N3 with the JMdict id in a named test. No lookup or data was changed to make a number come out "nicer".
|
| 121 |
+
|
| 122 |
+
## Deviations from Plan
|
| 123 |
+
|
| 124 |
+
### Auto-fixed Issues
|
| 125 |
+
|
| 126 |
+
**1. [Rule 1 - Bug in the spec] 日本語 is `N1+` by the pinned lists, not the plan's `N5`**
|
| 127 |
+
- **Found during:** Task 1 GREEN, `test_kanji_axis_separate_from_word_axis` failed on `nihongo["jlpt"] == "N5"` (`'N1+'`).
|
| 128 |
+
- **Issue:** `analyze` resolves 日本語 to JMdict 1464530 (kanji 日本語, kana にほんご / にっぽんご, common) - the right entry - and 1464530 appears in none of `data/jlpt/n1..n5.csv` (grep: no `日本語` and no `1464530` row; 日本 1582710 is on N3). D-10 says a word on no list is "N1+", never a guess.
|
| 129 |
+
- **Fix:** The test pins `jmdict_id == 1464530` and `jlpt == "N1+"` with the reason, and `kanji_levels == {日: N5, 本: N5, 語: N5}` - which makes it the clearest possible two-axis example. The `SENTENCES` pin text for #9 says the same.
|
| 130 |
+
- **Files modified:** `tests/test_analyzer.py`, `tests/fixtures/make_sentences.py`
|
| 131 |
+
- **Committed in:** `46459e0`, `b0413c1`
|
| 132 |
+
|
| 133 |
+
**2. [Rule 1 - Bug in the spec] ない / なかった are `N1+`, こんにちは is `N3`, 人気(ひとけ) is `N1+`**
|
| 134 |
+
- **Found during:** Task 2, reading the generator's review print (the plan's `EXPECTED` list does not cover these levels, so the guard did not fire; the by-eye review the plan mandates did).
|
| 135 |
+
- **Issue:** Plan table #5 says "ない tappable N5" and research § Q1 says "badge shows N5 (無い)" for なかった. Measured: `lookup("ない", "無い", "ない")` -> 無い 1529520 (the only entry with kanji 無い; reading matches), and **no entry read ない is on any list**. Plan #1 says only "kana headword 1289400"; 1289400 is on the N3 and N1 lists -> N3. 人気 after the override reads ひとけ, so the reading-first rank selects 1367020 ("sign of life", unlisted) -> N1+, the correct word.
|
| 136 |
+
- **Fix:** None to the code - the lookup is right and the lists are what they are. Pinned by id in `test_known_list_gaps_are_honest_n1_plus` and in the golden file; pin text of #5 corrected. Flagged below for 02-10's override review: if the product wants ない / 日本語 / なる / こんにちは at learner-intuitive levels, that is a **level-alias table** (a new, human-judged data file like `reading_overrides.json`), not a change to the join.
|
| 137 |
+
- **Files modified:** `tests/test_analyzer.py`, `tests/fixtures/make_sentences.py`
|
| 138 |
+
- **Committed in:** `b0413c1`
|
| 139 |
+
|
| 140 |
+
**3. [Rule 3 - Blocking] `tests/test_levels.py` had to change with `levels.py`**
|
| 141 |
+
- **Found during:** Task 1 GREEN design - the plan lists `levels.py` but not `test_levels.py`, yet 02-03's `test_is_kanji_ranges_local_copy` asserts on `levels._is_kanji`, which the plan's acceptance requires to be gone (`grep -c "def _is_kanji" == 0`).
|
| 142 |
+
- **Fix:** Rewritten as `test_is_kanji_is_rubys`: `levels.is_kanji is ruby.is_kanji`, `not hasattr(levels, "_is_kanji")`, the same range checks, the supplementary plane (𠮷), and the measured `kanji_levels_for("𠮷野家") == {𠮷: None, 野: N4, 家: N4}` (家 is N4 on the pinned kanji list - I would have guessed N5, which is why it is asserted from a measurement). `test_kanji_axis` unchanged and green.
|
| 143 |
+
- **Files modified:** `tests/test_levels.py`
|
| 144 |
+
- **Committed in:** `46459e0`
|
| 145 |
+
|
| 146 |
+
### Scope notes (not deviations)
|
| 147 |
+
|
| 148 |
+
- A 14th test, `test_known_list_gaps_are_honest_n1_plus`, was added beyond the plan's 13 so the honest-N1+ facts are pinned by JMdict id and not only by the opaque golden comparison. Cost: ~3 analyses in an already-warm session.
|
| 149 |
+
- `test_offsets_cover_text` iterates the golden file's texts (the plan's "every fixture sentence") rather than a duplicate inline list.
|
| 150 |
+
- The `REGENERATE` message constant was re-wrapped so the literal `REVIEW the diff` sits in one string (the plan's `grep -c` is a literal contract; the runtime message was already correct).
|
| 151 |
+
|
| 152 |
+
---
|
| 153 |
+
|
| 154 |
+
**Total deviations:** 3 auto-fixed (2 Rule 1 spec corrections proven on the pinned data, 1 Rule 3 file the plan omitted). No architectural change; no scope creep; no lookup or data weakened; nothing pushed.
|
| 155 |
+
|
| 156 |
+
## Issues Encountered
|
| 157 |
+
|
| 158 |
+
- The Edit tool could not match a `test_levels.py` block containing U+9FFF / U+FAFF glyphs even though `Read` showed the text verbatim; the block was replaced by line range with a short Python script (UTF-8 in, LF out). No other file needed this.
|
| 159 |
+
- `ruff check --fix && ruff format` as an `&&` chain never reaches `format` while E501s remain; `format` had to run first (it wraps the long Japanese-dense tuples itself), then `check` was clean.
|
| 160 |
+
- `levels.py` in the working copy was CRLF (02-03's `Path.write_text`); Git normalised the blob to LF on commit as before (warning only, no content change).
|
| 161 |
+
- `pytest -q` over `addopts="-q"` hides the pass count; counts here were taken with `-o addopts=""` / `-o addopts="--strict-markers"`.
|
| 162 |
+
- `ruff` printed its usual `.ruff_cache` "Access is denied" cache-write warning once; the checks passed.
|
| 163 |
+
|
| 164 |
+
## Known Stubs
|
| 165 |
+
|
| 166 |
+
None. Every field of every record is produced from the committed data (SudachiDict-core, the compact JMdict, the JLPT CSVs, the kanji map, the override table); `[]` / `None` / `{}` values are the D-11 contract for non-tappable units and the honest "no entry" for names without a JMdict entry (田中), never placeholders.
|
| 167 |
+
|
| 168 |
+
## User Setup Required
|
| 169 |
+
|
| 170 |
+
None. Nothing pushed to the `space` remote; no external service touched.
|
| 171 |
+
|
| 172 |
+
## Next Phase Readiness
|
| 173 |
+
|
| 174 |
+
- **02-06 (server functions / directive):** `from japanese_avatar.nlp.analyzer import TOKEN_KEYS, analyze, warmup`. `warmup()` returns the dict above - log `jmdict_s` next to `warm_synthesizer` on `Blocks.load`; budget +~310 MB (02-03's measurement) for the JMdict. Records are JSON-plain (str / bool / int / None / list / dict); the 36-mora line (16 units) serialises to the size recorded in the self-check below with `ensure_ascii=False`. `analyze("")` is `[]`; whitespace yields non-tappable 空白 units - the `{error}` guard for empty text belongs in the server function, not here.
|
| 175 |
+
- **02-07 (furigana):** per unit use `ruby` (already per kanji run, `rt` None over kana) and `kanji_levels` for the "above my level" gate (`None` = unlisted = above every level); non-tappable units are `[[surface, None]]` and need no gate. `tappable` decides `.tok`.
|
| 176 |
+
- **02-08 (popover):** card = `reading`, `gloss` (<= 3 x 3), `jlpt`, `lemma`; `jlpt` is `"name"` / `"N1+"` / `"N5".."N1"` for every tappable unit (asserted: tappable <-> jlpt not None). Note the お願い card names 願う (edge #15).
|
| 177 |
+
- **02-10 (override review row + requirements):** `test_fixture_sentences` is JPN-01's quick-loop row and is green; mark JPN-01 only with the deployed `test_language_assets_loaded`. **Product question for the review:** ない / なかった (N1+), 日本語 (N1+), なる forms (N3) and こんにちは (N3) are correct by the pinned lists and surprising to a learner; a level-alias override table (id -> level, with `why`) is the honest fix if wanted - the reading-override table is the pattern.
|
| 178 |
+
- **02-11 (docs/LANGUAGE.md):** the regeneration command and the REVIEW-the-diff rule are in `make_sentences.py`'s docstring; the record's field meanings are in `analyzer.py`'s docstring; the known list gaps are in `test_known_list_gaps_are_honest_n1_plus`'s docstring.
|
| 179 |
+
- **Quick-loop budget:** 30.7 s of 40 s; this plan added ~2.5 s (one JMdict load + 14 tests). Any new module that loads JMdict should use the `analyzer` fixture or `jmdict.load()` (cached), never a second process.
|
| 180 |
+
|
| 181 |
+
---
|
| 182 |
+
*Phase: 02-japanese-language-core*
|
| 183 |
+
*Completed: 2026-09-06*
|
| 184 |
+
|
| 185 |
+
## Self-Check: PASSED
|
| 186 |
+
|
| 187 |
+
All 4 created files (`analyzer.py`, `make_sentences.py`, `sentences.json`, `test_analyzer.py`) and 3 modified files (`levels.py`, `conftest.py`, `test_levels.py`) present on disk; commits `a09270c`, `46459e0`, `b0413c1` present in `git log`; `analyzer.py` is 160 lines (plan minimum 60); `sentences.json` contains ひとけ (2), にんき (2), おこなっ (1), `"name"` (3) and `git status --porcelain` on it is empty after a post-commit regeneration; `tests/test_analyzer.py` 14 passed; quick loop 293 passed / 14 deselected / 30.7 s wall; whole-repo `ruff check . && ruff format --check .` clean (53 files). Measured for 02-06: `analyze()` of the 36-mora line = 16 units, **4,843 bytes** as JSON with `ensure_ascii=False` (5,521 escaped). Nothing pushed.
|