WolfDavid commited on
Commit
16f28bd
·
1 Parent(s): b0413c1

docs(02-05): complete canonical analyzer and fixed sample sentence set plan

Browse files

- 02-05-SUMMARY.md: warm-up per component, 0.14-0.19 ms warm analysis,
the four accepted edges (#14, #15, #17, #27), the three list-level
corrections (日本語 / ない / こんにちは) pinned by JMdict id, quick loop
293 passed / 30.7 s, commit SHAs, self-check
- STATE.md: plan 6 of 11, progress 16/22, metric, four decisions, session
- ROADMAP.md: Phase 02 plan progress 5/11

.planning/ROADMAP.md CHANGED
@@ -58,7 +58,7 @@ Decimal phases appear between their surrounding integers in numeric order.
58
  - [x] 02-02-PLAN.md — Tokenizer singleton, whole-word units (rules A0–A7), okurigana ruby alignment, reading-override table (TDD) [wave 2]
59
  - [x] 02-03-PLAN.md — Ranked JMdict lookup + JLPT word/kanji level derivation (N1+, name) (TDD) [wave 2]
60
  - [x] 02-04-PLAN.md — OPUS-MT ja-en → CTranslate2 int8 under LFS, conversion script, translate() runtime (mt marker) [wave 2]
61
- - [ ] 02-05-PLAN.md — The canonical analyze(text) + the 28-sentence golden fixture set (SC-1) [wave 3]
62
  - [ ] 02-06-PLAN.md — tokens on the directive; analyze / translate / language_info server functions; warm-up on load; turn-loop + facade surface; parity [wave 4]
63
  - [ ] 02-07-PLAN.md — Furigana: transcript.js renderer, three-mode control + level picker, persistence, three-layer numbers (JPN-02) [wave 5]
64
  - [ ] 02-08-PLAN.md — Word-lookup popover with published numbers, desktop + Pixel 7 rows (JPN-03) [wave 6]
 
58
  - [x] 02-02-PLAN.md — Tokenizer singleton, whole-word units (rules A0–A7), okurigana ruby alignment, reading-override table (TDD) [wave 2]
59
  - [x] 02-03-PLAN.md — Ranked JMdict lookup + JLPT word/kanji level derivation (N1+, name) (TDD) [wave 2]
60
  - [x] 02-04-PLAN.md — OPUS-MT ja-en → CTranslate2 int8 under LFS, conversion script, translate() runtime (mt marker) [wave 2]
61
+ - [x] 02-05-PLAN.md — The canonical analyze(text) + the 28-sentence golden fixture set (SC-1) [wave 3]
62
  - [ ] 02-06-PLAN.md — tokens on the directive; analyze / translate / language_info server functions; warm-up on load; turn-loop + facade surface; parity [wave 4]
63
  - [ ] 02-07-PLAN.md — Furigana: transcript.js renderer, three-mode control + level picker, persistence, three-layer numbers (JPN-02) [wave 5]
64
  - [ ] 02-08-PLAN.md — Word-lookup popover with published numbers, desktop + Pixel 7 rows (JPN-03) [wave 6]
.planning/STATE.md CHANGED
@@ -3,13 +3,13 @@ gsd_state_version: 1.0
3
  milestone: v1.0
4
  milestone_name: milestone
5
  status: unknown
6
- stopped_at: "Completed 02-04-PLAN.md (wave 2): OPUS-MT ja-en CT2 int8 under LFS (convert_mt.py, 5 hash rows, Apache-2.0 + NOTICE), translate.py runtime (14 mt tests, 0.35 s warm-up, 22-286 ms/sentence, no transformers/torch). Wave 2 complete; next 02-05."
7
- last_updated: "2026-09-06T23:48:29.726Z"
8
  progress:
9
  total_phases: 6
10
  completed_phases: 1
11
  total_plans: 22
12
- completed_plans: 15
13
  ---
14
 
15
  # Project State
@@ -24,7 +24,7 @@ See: .planning/PROJECT.md (updated 2026-08-08)
24
  ## Current Position
25
 
26
  Phase: 02 (japanese-language-core) — EXECUTING
27
- Plan: 5 of 11
28
 
29
  ## Performance Metrics
30
 
@@ -65,6 +65,7 @@ Plan: 5 of 11
65
  | Phase 02 P02 | 12 | 3 tasks | 8 files |
66
  | Phase 02 P03 | 11 | 2 tasks | 5 files |
67
  | Phase 02 P04 | 17 | 2 tasks | 13 files |
 
68
 
69
  ## Accumulated Context
70
 
@@ -140,6 +141,10 @@ Recent decisions affecting current work:
140
  - [Phase 02]: The CTranslate2 converter writes shared_vocabulary.json/config.json with the platform newline; convert_mt.py normalises them to LF because the repo has core.autocrlf=input and hashing a CRLF working copy would fail test_mt_files_match_readme on every fresh clone (the Space). Research's 233 B / 1,269,578 B were Windows-CRLF sizes; committed bytes are 223 B / 1,208,862 B
141
  - [Phase 02]: data/mt/README.md carries run measurements (read-back ms, conversion s) and stays idempotent because the script leaves it untouched when every recorded hash matches; the conversion proves sentencepiece-only pieces == MarianTokenizer on 7 sentences before recording, so the Space imports neither transformers nor torch
142
  - [Phase 02]: JPN-04 left Pending after 02-04: 02-VALIDATION.md binds it to 02-10's deployed test_translation_reveal + docs/LATENCY.md § Translation row; translate() local baseline 0.35 s warm-up, 22-286 ms per contract sentence at intra_threads=2, MT_INTRA_THREADS is the contention lever
 
 
 
 
143
 
144
  ### Pending Todos
145
 
@@ -161,8 +166,8 @@ None yet.
161
 
162
  ## Session Continuity
163
 
164
- Last session: 2026-09-06T23:48:29.720Z
165
- Stopped at: Completed 02-04-PLAN.md (wave 2): OPUS-MT ja-en CT2 int8 under LFS (convert_mt.py, 5 hash rows, Apache-2.0 + NOTICE), translate.py runtime (14 mt tests, 0.35 s warm-up, 22-286 ms/sentence, no transformers/torch). Wave 2 complete; next 02-05.
166
  Resume file: None
167
 
168
  **Roadmap-level open decision:** the Phase 1 latency trigger (p50 7.1 s, ~90 % server synthesis) fired; recorded in `docs/LATENCY.md` and in 02-CONTEXT.md § Deferred. Phase 2 adds a CPU translation model to the same container — measure on the Space. Consider `/gsd:insert-phase` for streaming/synthesis relocation before Phase 3 lengthens turns.
 
3
  milestone: v1.0
4
  milestone_name: milestone
5
  status: unknown
6
+ stopped_at: "Completed 02-05-PLAN.md (wave 3): analyze(text) canonical token record (14 keys, 0.14-0.19 ms warm), levels/ruby is_kanji reconciled, 28-sentence golden fixture + refusal-guard generator, 14 analyzer tests; quick loop 293 passed / 30.7 s. Next 02-06."
7
+ last_updated: "2026-09-07T00:08:47.920Z"
8
  progress:
9
  total_phases: 6
10
  completed_phases: 1
11
  total_plans: 22
12
+ completed_plans: 16
13
  ---
14
 
15
  # Project State
 
24
  ## Current Position
25
 
26
  Phase: 02 (japanese-language-core) — EXECUTING
27
+ Plan: 6 of 11
28
 
29
  ## Performance Metrics
30
 
 
65
  | Phase 02 P02 | 12 | 3 tasks | 8 files |
66
  | Phase 02 P03 | 11 | 2 tasks | 5 files |
67
  | Phase 02 P04 | 17 | 2 tasks | 13 files |
68
+ | Phase 02 P05 | 12 | 2 tasks | 7 files |
69
 
70
  ## Accumulated Context
71
 
 
141
  - [Phase 02]: The CTranslate2 converter writes shared_vocabulary.json/config.json with the platform newline; convert_mt.py normalises them to LF because the repo has core.autocrlf=input and hashing a CRLF working copy would fail test_mt_files_match_readme on every fresh clone (the Space). Research's 233 B / 1,269,578 B were Windows-CRLF sizes; committed bytes are 223 B / 1,208,862 B
142
  - [Phase 02]: data/mt/README.md carries run measurements (read-back ms, conversion s) and stays idempotent because the script leaves it untouched when every recorded hash matches; the conversion proves sentencepiece-only pieces == MarianTokenizer on 7 sentences before recording, so the Space imports neither transformers nor torch
143
  - [Phase 02]: JPN-04 left Pending after 02-04: 02-VALIDATION.md binds it to 02-10's deployed test_translation_reveal + docs/LATENCY.md § Translation row; translate() local baseline 0.35 s warm-up, 22-286 ms per contract sentence at intra_threads=2, MT_INTRA_THREADS is the contention lever
144
+ - [Phase 02]: analyze(text) is the ONE canonical text -> tokens function: records built from TOKEN_KEYS (14 keys, fixed order), level from the level_key entry and gloss from the lemma entry, non-tappable units carry jlpt None / gloss [] / ruby [[surface, None]]; warm analysis 0.14-0.19 ms per 36-mora sentence
145
+ - [Phase 02]: Three plan-assumed word levels corrected to what the pinned lists say and pinned by JMdict id, not patched: 日本語 1464530 is N1+ (on no list; kanji all N5 - the two-axis example), ない/なかった resolve to 無い 1529520 which no list carries (N1+), こんにちは 1289400 is N3; 人気 read ひとけ is the separate entry 1367020 (N1+). A level-alias table is the honest fix if the product wants otherwise (02-10 review), never a weaker lookup
146
+ - [Phase 02]: levels.is_kanji IS ruby.is_kanji (02-03's parallel-wave copy removed): one predicate decides both ruby and the kanji-axis gate, so supplementary-plane kanji (𠮷) reach the axis as unlisted instead of being skipped
147
+ - [Phase 02]: The fixed sample sentence set (SC-1) follows the golden_timeline pattern: 28 sentences / 115 records generated by the real pipeline, 18 EXPECTED facts checked before writing (exit 1, nothing written otherwise), frozen LF byte-for-byte with data pins read from the environment; JPN-01 still bound to 02-10's rows
148
 
149
  ### Pending Todos
150
 
 
166
 
167
  ## Session Continuity
168
 
169
+ Last session: 2026-09-07T00:08:47.915Z
170
+ Stopped at: Completed 02-05-PLAN.md (wave 3): analyze(text) canonical token record (14 keys, 0.14-0.19 ms warm), levels/ruby is_kanji reconciled, 28-sentence golden fixture + refusal-guard generator, 14 analyzer tests; quick loop 293 passed / 30.7 s. Next 02-06.
171
  Resume file: None
172
 
173
  **Roadmap-level open decision:** the Phase 1 latency trigger (p50 7.1 s, ~90 % server synthesis) fired; recorded in `docs/LATENCY.md` and in 02-CONTEXT.md § Deferred. Phase 2 adds a CPU translation model to the same container — measure on the Space. Consider `/gsd:insert-phase` for streaming/synthesis relocation before Phase 3 lengthens turns.
.planning/phases/02-japanese-language-core/02-05-SUMMARY.md ADDED
@@ -0,0 +1,187 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ phase: 02-japanese-language-core
3
+ plan: 05
4
+ subsystem: nlp
5
+ tags: [analyzer, token-record, golden-fixture, sudachipy, jmdict, jlpt, furigana, tdd, pytest]
6
+
7
+ # Dependency graph
8
+ requires:
9
+ - phase: 02-japanese-language-core
10
+ plan: 02
11
+ provides: "tokenizer.morphemes() plain dicts; units.build_units() (A0-A7); ruby.align_ruby() / is_kanji(); overrides.apply_overrides() + data/nlp/reading_overrides.json"
12
+ - phase: 02-japanese-language-core
13
+ plan: 03
14
+ provides: "jmdict.load()/lookup()/glosses_for() ranked headword lookup; levels.vocab_levels() (7,748 ids) / kanji_levels() (2,211) / derive_level()"
15
+ - phase: 01-voice-avatar-loop-skeleton
16
+ provides: "tests/fixtures/make_golden_timeline.py refuse-to-write pattern; voice/tts.py warmup() shape"
17
+ provides:
18
+ - "src/japanese_avatar/nlp/analyzer.py: analyze(text) -> list[dict] with exactly TOKEN_KEYS (14 keys, fixed order), warmup() -> {tokenizer_s, jmdict_s, jmdict_entries, vocab_ids, kanji}; the ONE canonical text -> tokens function, no LLM anywhere on the path"
19
+ - "tests/fixtures/sentences.json: the fixed sample sentence set of success criterion 1 - 28 sentences, 115 golden token records, with the data pins that produced them"
20
+ - "tests/fixtures/make_sentences.py: regenerator that checks 18 EXPECTED facts before writing and exits 1 with nothing written otherwise; byte-idempotent"
21
+ - "tests/test_analyzer.py (14 tests): D-10/D-11/D-12/D-13 as named tests, the override, the ambiguous readings, exact golden equality, contract coverage, pin check, the honest list gaps"
22
+ - "tests/conftest.py: session-scoped `analyzer` fixture - one dictionary open + one compact-JMdict load per session"
23
+ - "levels.py now uses ruby.is_kanji (02-03's parallel-wave copy removed)"
24
+ affects: [02-06, 02-07, 02-08, 02-10, 02-11, phase-3-level-guard, phase-5-vocab-tracking]
25
+
26
+ # Tech tracking
27
+ tech-stack:
28
+ added: []
29
+ patterns:
30
+ - "One canonical analyze(text): every consumer reads the same 14-key plain-dict record; engine objects never leave tokenizer.py; the dict is built from TOKEN_KEYS so key order is a contract, not an accident"
31
+ - "Level from the level_key entry, gloss from the lemma entry (research § Q2 step 5) - one ranked lookup when the two keys coincide, which is the common case"
32
+ - "Golden language fixture = golden_timeline pattern: generated once by the real pipeline, guarded by explicit EXPECTED facts the generator refuses to violate, frozen byte-for-byte (LF, indent=1, ensure_ascii=False) with the data pins that produced it"
33
+ - "Upstream data quirks are pinned BY JMDICT ID in a named test (test_known_list_gaps_are_honest_n1_plus) so a future level alias is a conscious edit to the test, never a weaker lookup"
34
+
35
+ key-files:
36
+ created:
37
+ - src/japanese_avatar/nlp/analyzer.py
38
+ - tests/fixtures/make_sentences.py
39
+ - tests/fixtures/sentences.json
40
+ - tests/test_analyzer.py
41
+ modified:
42
+ - src/japanese_avatar/nlp/levels.py
43
+ - tests/conftest.py
44
+ - tests/test_levels.py
45
+
46
+ key-decisions:
47
+ - "The plan's three assumed word levels were corrected to what the pinned lists say and pinned honestly under D-10 rather than patched: 日本語 (1464530) is N1+ - it is on none of the five yomitan-jlpt-vocab CSVs (日本 itself is N3 there); ない / なかった resolve correctly to 無い 1529520, which no list carries, so they badge N1+ (the plan's #5 'ない tappable N5' and research's 'badge shows N5 (無い)' were guesses); こんにちは 1289400 is on N3 and N1, so it badges N3. The lookup was not weakened for any of them"
48
+ - "人気 read ひとけ is the separate JMdict entry 1367020 ('sign of life'), unlisted -> N1+, while にんき is 1367010 N3: the override changes the reading and the reading-first rank then picks the right WORD, which is exactly what a learner tapping 人気のない通り should see"
49
+ - "Non-tappable units carry ruby [[surface, None]] rather than align_ruby's output: a punctuation 'reading' is the mark itself and would otherwise become an rt over 。"
50
+ - "tests/test_levels.py (not in the plan's file list) had to change: its test_is_kanji_ranges_local_copy asserted on levels._is_kanji, which the plan requires removing; it is now test_is_kanji_is_rubys and pins identity with ruby.is_kanji plus the supplementary-plane consequence (𠮷野家 -> {𠮷: None, 野: N4, 家: N4})"
51
+ - "JPN-01 NOT marked complete: 02-VALIDATION.md binds it to 02-10's rows (test_fixture_sentences here + the deployed test_language_assets_loaded); this plan delivers the quick-loop half"
52
+
53
+ patterns-established:
54
+ - "Session-scoped warm fixture that prints its per-component seconds, so the quick loop reports the language core's real cost every run (0.03 s dictionary, ~2.0-2.4 s compact JMdict)"
55
+ - "A regenerator records the versions it ran under by READING them (sudachipy.__version__, importlib.metadata, the compact file's meta.version, the build script's pinned tag/commit), and a test compares them with the installed environment so a stale golden file and a stale venv are told apart"
56
+
57
+ requirements-completed: [] # JPN-01 is marked only in 02-10 on its bound rows (02-VALIDATION.md)
58
+
59
+ # Metrics
60
+ duration: 12min
61
+ completed: 2026-09-06
62
+ ---
63
+
64
+ # Phase 02 Plan 05: The Canonical Analyzer and the Fixed Sample Sentence Set Summary
65
+
66
+ **One function, `analyze(text)`, turns any Japanese string into 14-key token records - whole-word surfaces, hiragana readings after the override table, dictionary forms, word level by JMdict-id join, kanji level per character, okurigana-aligned ruby spans, up to 3x3 English glosses and offsets - with no LLM on the path, in 0.14-0.19 ms per 36-mora sentence once warm; and success criterion 1 is now a committed file: 28 sentences / 115 golden records that the analyzer reproduces byte-for-byte and that the generator refuses to overwrite with wrong 今日 / 行った / 人気 readings.**
67
+
68
+ ## Performance
69
+
70
+ - **Duration:** 12 min (started 2026-09-06T23:53:18Z, last task commit 2026-09-07T00:05:16Z)
71
+ - **Tasks:** 2 (Task 1 TDD: RED + GREEN; Task 2 feature)
72
+ - **Files:** 7 (4 created, 3 modified)
73
+ - **Warm-up, per component** (three fresh processes): tokenizer (Sudachi core) **0.030 / 0.032 / 0.033 s**; compact JMdict **1.97 / 2.32 / 2.40 s**; `vocab_levels()` + `kanji_levels()` are inside the noise (< 0.05 s). `warmup()` returns `{'tokenizer_s', 'jmdict_s', 'jmdict_entries': 218672, 'vocab_ids': 7748, 'kanji': 2211}`.
74
+ - **Warm analysis of the 36-mora sentence** (今日はいい天気ですから、公園を散歩してから、買い物に行きました。, 16 units): **mean 0.142 ms / 0.194 ms over 100 runs** in two test sessions (plan expectation < 5 ms; assertion budget 50 ms).
75
+ - **Quick loop** (`pytest tests --ignore=tests/e2e -m "not mt"`): after Task 1 **289 passed, 14 deselected, 28.8 s pytest / 30.7 s wall**; after Task 2 **293 passed, 14 deselected, 28.8 s / 30.7 s wall** (budget 40 s; was 283 before this plan). The session fixture makes the JMdict load a one-time ~2.3 s.
76
+ - `tests/test_analyzer.py` alone: 14 passed in 2.5-2.6 s (dominated by the one JMdict load).
77
+
78
+ ## Accomplishments
79
+
80
+ - **`analyzer.py`** composes the pure functions of 02-02 and 02-03 in the plan's order: `morphemes -> build_units -> apply_overrides -> (per tappable unit) align_ruby, lookup x1-2, derive_level, kanji_levels_for, glosses_for`. Each record is built from `TOKEN_KEYS`, so `tuple(u.keys()) == TOKEN_KEYS` for every unit of every sentence (asserted in three places). Non-tappable units (D-11) carry `jlpt None`, `jmdict_id None`, `gloss []`, `kanji_levels {}`, `ruby [[surface, None]]`. `grep -c sudachipy analyzer.py` is 0.
81
+ - **Level and gloss from different entries when they must be:** 話せる -> gloss from its own entry 1562360 ("to be able to speak"), level N5 from level_key 話す; 勉強しています -> lemma 勉強する (no such headword), level and gloss from level_key 勉強 (1512670, N5); 三人 -> level_key `"3"` matches no headword, lemma 三 gives 1579350 N5.
82
+ - **The reconciliation the wave-2 plans left:** `levels.py` imports `is_kanji` from `ruby.py`; `def _is_kanji` is gone (grep 0, `import is_kanji` 1). Consequence made visible: a supplementary-plane kanji (𠮷) now reaches the kanji axis as unlisted (`None`) instead of being skipped. `test_kanji_axis` passes unchanged.
83
+ - **The fixed sample sentence set** (`tests/fixtures/sentences.json`, 59,029 bytes, LF only, `filter: unspecified` - text, not LFS): 28 sentences in the plan's order, 115 units, `analyzer` block `{sudachipy 0.6.11, sudachidict_core 20260723, jmdict 3.6.2+20260831182826, jlpt_vocab 2025.08.01.0, kanji_data 00fd7079c3890f430759536f91aa5e854ec0ca4f}` - every pin read from the environment, none typed. The generator ran three times to the identical SHA-256 `8d97fd7683e4de6fae430ca1523a2cdc31a3188511ae33b0774a69275adfe89a`; after the commit a further run left `git status --porcelain tests/fixtures/sentences.json` empty.
84
+ - **Spot-checked by eye** (the plan's list #2, #3, #4, #5, #6, #7, #11, #12, #13, #15, #16, #17, #20, #27, #28 plus #1, #9, #14, #21, #22, #23, #24): 今日 きょう N5; 行った いった/行く N5 vs おこなった/行う N4; 人気 ひとけ (1367020) vs にんき N3 (1367010); 食べました one unit `[[食,た],[べました,None]]` N5 1358280 "to eat"; 美味しかった|です with `[[美味,おい],[しかった,None]]`; きれい|じゃ|なかった tappable T/F/T; 田中さん name (no entry, gloss []) and 東京 name (entry, gloss present); 話せる N5 1562360; 学生 N5 + です; 待ち合わせ `[[待,ま],[ち,None],[合,あ],[わせ,None]]` N1; 買い物 `[[買,か],[い,None],[物,もの]]` N5; 取り扱い `[[取,と],[り,None],[扱,あつか],[い,None]]` N1; 大阪駅 `[[大阪駅,おおさかえき]]` name, kanji {大 N5, 阪 None, 駅 N4}; コーヒー no rt, reading こーひー, N5; 美味しくない ONE unit, lemma 美味しい, N5. All eight okurigana spans equal research § Q1's table.
85
+ - **Tests:** 14 `def test_` in `tests/test_analyzer.py` (plan minimum 10 then 13): `test_record_keys_exact`, `test_reading_disambiguation`, `test_hitoke_override`, `test_tappability` (D-11), `test_level_derivation` (D-12, D-10 via 瀟洒 - a JMdict entry on no list), `test_kanji_axis_separate_from_word_axis` (D-13), `test_gloss_shipped_with_token` (D-07), `test_offsets_cover_text`, `test_analyze_is_fast`, `test_empty_text`, `test_known_list_gaps_are_honest_n1_plus`, `test_fixture_sentences` (exact equality; failure text names the regeneration command and says REVIEW the diff), `test_fixture_set_covers_the_contract` (also pins the five ambiguous readings as readings), `test_fixture_pins_match_installed`.
86
+
87
+ ## The four accepted edges (observed units, now frozen in the fixture)
88
+
89
+ | # | Sentence | Units (surface [reading | jlpt]) | The edge |
90
+ |---|---|---|---|
91
+ | 14 | 三人で行きました。 | 三人 [さんにん \| **N5**] · で · 行きました [いきました \| N5] · 。 | 三人 is one unit (A6); `level_key` is Sudachi's `"3"`, which matches no JMdict headword, so both level and gloss come from lemma 三 (1579350, "three"); `kanji_levels {三: N5, 人: N5}`; ruby `[[三人, さんにん]]`. で is 助動詞 (lemma だ), non-tappable |
92
+ | 15 | お願いします | お願い [おねがい \| **N3**] · します [します \| N5] | Both tappable. お願い's head is 願い (動詞,非自立可能), so `lemma` is 願う -> 1217950 "to desire; to wish", N3; ruby `[[お,None],[願,ねが],[い,None]]`; します -> lemma する, level_key 為る, 1157170 N5. The card for お願い therefore names 願う, not お願い (1001720) - a consequence of A7 + the head rule, recorded for 02-08/02-10's review |
93
+ | 17 | 行かなければならない | 行かなけれ [いかなけれ \| N5] · ば · ならない [ならない \| **N3**] | 02-02's observed split. 行かなけれ -> 行く 1578850 N5, ruby `[[行,い],[かなけれ,None]]`; ば breaks the chain (non-tappable); ならない -> lemma なる, level_key 成る -> 1375610, **N3** by the lists (the なる quirk 02-03 recorded: n5.csv's なる row is the archaic copula 2138260) |
94
+ | 27 | 鬱陶しい天気だ。 | 鬱陶しい [うっとうしい \| **N1**] · 天気 [てんき \| N5] · だ · 。 | 鬱陶しい IS on the N1 list (1568430), so its word level is N1 not N1+; `kanji_levels {鬱: None, 陶: N1}` - the unlisted kanji is on the kanji axis; ruby `[[鬱陶,うっとう],[しい,None]]` |
95
+
96
+ Other observed levels worth knowing (all in the golden file): こんにちは **N3** (1289400 is on N3 and N1), 人気/ひとけ **N1+** (1367020 unlisted), ない / なかった **N1+** (無い 1529520 unlisted), 日本語 **N1+** (1464530 unlisted) with kanji all N5, なりたい **N3** (なる), はじめまして N4, よろしく N3, 待ち合わせ / 取り扱い N1, 注意してください N4 (lemma 注意する), 遅れた N4.
97
+
98
+ ## Task Commits
99
+
100
+ | Task | Commit | Type |
101
+ |---|---|---|
102
+ | 1 RED - failing analyzer behaviour tests + session fixture | `a09270c` | test |
103
+ | 1 GREEN - analyze(text), warmup(), levels.py reconciliation, test_levels follow-up | `46459e0` | feat |
104
+ | 2 - generator with refusal guards, golden file, golden tests | `b0413c1` | feat |
105
+
106
+ All commits `--no-verify`, explicit paths only; nothing pushed; no `git add -A`.
107
+
108
+ ## Files Created/Modified
109
+
110
+ - `src/japanese_avatar/nlp/analyzer.py` (new, ~150 lines) - `TOKEN_KEYS`, `_lookup_pair(unit, level_of)`, `_token(unit, level_of)`, `analyze(text)`, `warmup()`, a dev-only `_self_time()`; docstring carries the record's field-by-field meaning and the pipeline order
111
+ - `src/japanese_avatar/nlp/levels.py` - `from japanese_avatar.nlp.ruby import is_kanji`; the local `_KANJI_RANGES` / `_is_kanji` block removed; `kanji_levels_for` unchanged otherwise
112
+ - `tests/conftest.py` - session-scoped `analyzer` fixture beside `synth_meta` (`--space-url` untouched); `scope="session"` count is now 7
113
+ - `tests/test_analyzer.py` (new) - 14 tests as listed; `_golden()` cached loader; `REGENERATE` message constant
114
+ - `tests/test_levels.py` - `test_is_kanji_ranges_local_copy` -> `test_is_kanji_is_rubys` (identity with `ruby.is_kanji`, `not hasattr(levels, "_is_kanji")`, supplementary plane, 𠮷野家 levels)
115
+ - `tests/fixtures/make_sentences.py` (new) - `SENTENCES` (28 x (text, what it pins)), 18 `EXPECTED` check functions (the plan's 15 facts + tiling + tappable<->jlpt + exact keys), `data_pins()`, LF writer, per-sentence review print
116
+ - `tests/fixtures/sentences.json` (new) - the golden file
117
+
118
+ ## Decisions Made
119
+
120
+ See `key-decisions` in the frontmatter. The load-bearing one: three word levels the plan assumed (日本語 N5, ない N5, こんにちは implied N5) are not what the pinned lists say; each was verified against the CSVs and the lookup (the lookup picks the right entry in every case), then pinned honestly as N1+ / N3 with the JMdict id in a named test. No lookup or data was changed to make a number come out "nicer".
121
+
122
+ ## Deviations from Plan
123
+
124
+ ### Auto-fixed Issues
125
+
126
+ **1. [Rule 1 - Bug in the spec] 日本語 is `N1+` by the pinned lists, not the plan's `N5`**
127
+ - **Found during:** Task 1 GREEN, `test_kanji_axis_separate_from_word_axis` failed on `nihongo["jlpt"] == "N5"` (`'N1+'`).
128
+ - **Issue:** `analyze` resolves 日本語 to JMdict 1464530 (kanji 日本語, kana にほんご / にっぽんご, common) - the right entry - and 1464530 appears in none of `data/jlpt/n1..n5.csv` (grep: no `日本語` and no `1464530` row; 日本 1582710 is on N3). D-10 says a word on no list is "N1+", never a guess.
129
+ - **Fix:** The test pins `jmdict_id == 1464530` and `jlpt == "N1+"` with the reason, and `kanji_levels == {日: N5, 本: N5, 語: N5}` - which makes it the clearest possible two-axis example. The `SENTENCES` pin text for #9 says the same.
130
+ - **Files modified:** `tests/test_analyzer.py`, `tests/fixtures/make_sentences.py`
131
+ - **Committed in:** `46459e0`, `b0413c1`
132
+
133
+ **2. [Rule 1 - Bug in the spec] ない / なかった are `N1+`, こんにちは is `N3`, 人気(ひとけ) is `N1+`**
134
+ - **Found during:** Task 2, reading the generator's review print (the plan's `EXPECTED` list does not cover these levels, so the guard did not fire; the by-eye review the plan mandates did).
135
+ - **Issue:** Plan table #5 says "ない tappable N5" and research § Q1 says "badge shows N5 (無い)" for なかった. Measured: `lookup("ない", "無い", "ない")` -> 無い 1529520 (the only entry with kanji 無い; reading matches), and **no entry read ない is on any list**. Plan #1 says only "kana headword 1289400"; 1289400 is on the N3 and N1 lists -> N3. 人気 after the override reads ひとけ, so the reading-first rank selects 1367020 ("sign of life", unlisted) -> N1+, the correct word.
136
+ - **Fix:** None to the code - the lookup is right and the lists are what they are. Pinned by id in `test_known_list_gaps_are_honest_n1_plus` and in the golden file; pin text of #5 corrected. Flagged below for 02-10's override review: if the product wants ない / 日本語 / なる / こんにちは at learner-intuitive levels, that is a **level-alias table** (a new, human-judged data file like `reading_overrides.json`), not a change to the join.
137
+ - **Files modified:** `tests/test_analyzer.py`, `tests/fixtures/make_sentences.py`
138
+ - **Committed in:** `b0413c1`
139
+
140
+ **3. [Rule 3 - Blocking] `tests/test_levels.py` had to change with `levels.py`**
141
+ - **Found during:** Task 1 GREEN design - the plan lists `levels.py` but not `test_levels.py`, yet 02-03's `test_is_kanji_ranges_local_copy` asserts on `levels._is_kanji`, which the plan's acceptance requires to be gone (`grep -c "def _is_kanji" == 0`).
142
+ - **Fix:** Rewritten as `test_is_kanji_is_rubys`: `levels.is_kanji is ruby.is_kanji`, `not hasattr(levels, "_is_kanji")`, the same range checks, the supplementary plane (𠮷), and the measured `kanji_levels_for("𠮷野家") == {𠮷: None, 野: N4, 家: N4}` (家 is N4 on the pinned kanji list - I would have guessed N5, which is why it is asserted from a measurement). `test_kanji_axis` unchanged and green.
143
+ - **Files modified:** `tests/test_levels.py`
144
+ - **Committed in:** `46459e0`
145
+
146
+ ### Scope notes (not deviations)
147
+
148
+ - A 14th test, `test_known_list_gaps_are_honest_n1_plus`, was added beyond the plan's 13 so the honest-N1+ facts are pinned by JMdict id and not only by the opaque golden comparison. Cost: ~3 analyses in an already-warm session.
149
+ - `test_offsets_cover_text` iterates the golden file's texts (the plan's "every fixture sentence") rather than a duplicate inline list.
150
+ - The `REGENERATE` message constant was re-wrapped so the literal `REVIEW the diff` sits in one string (the plan's `grep -c` is a literal contract; the runtime message was already correct).
151
+
152
+ ---
153
+
154
+ **Total deviations:** 3 auto-fixed (2 Rule 1 spec corrections proven on the pinned data, 1 Rule 3 file the plan omitted). No architectural change; no scope creep; no lookup or data weakened; nothing pushed.
155
+
156
+ ## Issues Encountered
157
+
158
+ - The Edit tool could not match a `test_levels.py` block containing U+9FFF / U+FAFF glyphs even though `Read` showed the text verbatim; the block was replaced by line range with a short Python script (UTF-8 in, LF out). No other file needed this.
159
+ - `ruff check --fix && ruff format` as an `&&` chain never reaches `format` while E501s remain; `format` had to run first (it wraps the long Japanese-dense tuples itself), then `check` was clean.
160
+ - `levels.py` in the working copy was CRLF (02-03's `Path.write_text`); Git normalised the blob to LF on commit as before (warning only, no content change).
161
+ - `pytest -q` over `addopts="-q"` hides the pass count; counts here were taken with `-o addopts=""` / `-o addopts="--strict-markers"`.
162
+ - `ruff` printed its usual `.ruff_cache` "Access is denied" cache-write warning once; the checks passed.
163
+
164
+ ## Known Stubs
165
+
166
+ None. Every field of every record is produced from the committed data (SudachiDict-core, the compact JMdict, the JLPT CSVs, the kanji map, the override table); `[]` / `None` / `{}` values are the D-11 contract for non-tappable units and the honest "no entry" for names without a JMdict entry (田中), never placeholders.
167
+
168
+ ## User Setup Required
169
+
170
+ None. Nothing pushed to the `space` remote; no external service touched.
171
+
172
+ ## Next Phase Readiness
173
+
174
+ - **02-06 (server functions / directive):** `from japanese_avatar.nlp.analyzer import TOKEN_KEYS, analyze, warmup`. `warmup()` returns the dict above - log `jmdict_s` next to `warm_synthesizer` on `Blocks.load`; budget +~310 MB (02-03's measurement) for the JMdict. Records are JSON-plain (str / bool / int / None / list / dict); the 36-mora line (16 units) serialises to the size recorded in the self-check below with `ensure_ascii=False`. `analyze("")` is `[]`; whitespace yields non-tappable 空白 units - the `{error}` guard for empty text belongs in the server function, not here.
175
+ - **02-07 (furigana):** per unit use `ruby` (already per kanji run, `rt` None over kana) and `kanji_levels` for the "above my level" gate (`None` = unlisted = above every level); non-tappable units are `[[surface, None]]` and need no gate. `tappable` decides `.tok`.
176
+ - **02-08 (popover):** card = `reading`, `gloss` (<= 3 x 3), `jlpt`, `lemma`; `jlpt` is `"name"` / `"N1+"` / `"N5".."N1"` for every tappable unit (asserted: tappable <-> jlpt not None). Note the お願い card names 願う (edge #15).
177
+ - **02-10 (override review row + requirements):** `test_fixture_sentences` is JPN-01's quick-loop row and is green; mark JPN-01 only with the deployed `test_language_assets_loaded`. **Product question for the review:** ない / なかった (N1+), 日本語 (N1+), なる forms (N3) and こんにちは (N3) are correct by the pinned lists and surprising to a learner; a level-alias override table (id -> level, with `why`) is the honest fix if wanted - the reading-override table is the pattern.
178
+ - **02-11 (docs/LANGUAGE.md):** the regeneration command and the REVIEW-the-diff rule are in `make_sentences.py`'s docstring; the record's field meanings are in `analyzer.py`'s docstring; the known list gaps are in `test_known_list_gaps_are_honest_n1_plus`'s docstring.
179
+ - **Quick-loop budget:** 30.7 s of 40 s; this plan added ~2.5 s (one JMdict load + 14 tests). Any new module that loads JMdict should use the `analyzer` fixture or `jmdict.load()` (cached), never a second process.
180
+
181
+ ---
182
+ *Phase: 02-japanese-language-core*
183
+ *Completed: 2026-09-06*
184
+
185
+ ## Self-Check: PASSED
186
+
187
+ All 4 created files (`analyzer.py`, `make_sentences.py`, `sentences.json`, `test_analyzer.py`) and 3 modified files (`levels.py`, `conftest.py`, `test_levels.py`) present on disk; commits `a09270c`, `46459e0`, `b0413c1` present in `git log`; `analyzer.py` is 160 lines (plan minimum 60); `sentences.json` contains ひとけ (2), にんき (2), おこなっ (1), `"name"` (3) and `git status --porcelain` on it is empty after a post-commit regeneration; `tests/test_analyzer.py` 14 passed; quick loop 293 passed / 14 deselected / 30.7 s wall; whole-repo `ruff check . && ruff format --check .` clean (53 files). Measured for 02-06: `analyze()` of the 36-mora line = 16 units, **4,843 bytes** as JSON with `ensure_ascii=False` (5,521 escaped). Nothing pushed.