WolfDavid commited on
Commit
f7e73d2
·
1 Parent(s): adfe806

feat(02-02): okurigana-aligned ruby spans

Browse files

- ruby.py: KANA class, is_kanji (incl. 々 and the supplementary planes), kanji_runs,
align_ruby with the whole-unit fallback; pure over strings, no Sudachi

Files changed (1) hide show
  1. src/japanese_avatar/nlp/ruby.py +88 -0
src/japanese_avatar/nlp/ruby.py ADDED
@@ -0,0 +1,88 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Okurigana alignment: a unit's surface + its hiragana reading -> per-kanji-run ruby spans.
2
+
3
+ The reading comes from Sudachi (``units.build_units`` already converted it to hiragana with
4
+ ``jaconv.kata2hira``); this module only ALIGNS it - it never guesses a reading (research
5
+ § Anti-Patterns: no furigana by regex or by romaji). The standard trick: split the surface into
6
+ kana runs and non-kana runs, force every kana run to match the reading verbatim, and let each
7
+ non-kana run absorb whatever lies between with a non-greedy group. ``<rt>`` then sits over 食,
8
+ not over 食べ. Verified on the 23 pairs in ``tests/test_ruby.py``.
9
+
10
+ Ruby is per kanji run: 待ち合わせ -> [待,ま] [ち,-] [合,あ] [わせ,-]. A run's reading cannot be
11
+ split further (大阪駅 -> one span おおさかえき), which is also why the kanji-axis gate (02-05)
12
+ is evaluated per run.
13
+
14
+ Fallback is honest: when the kana runs cannot be matched (a reading that does not fit the
15
+ surface) the whole unit gets one ruby span carrying the whole reading. Never raises.
16
+
17
+ ``jaconv.normalize`` must NOT be used on readings - it folds ー and small kana. Pure function
18
+ over strings; no Sudachi import.
19
+ """
20
+
21
+ from __future__ import annotations
22
+
23
+ import re
24
+
25
+ import jaconv
26
+
27
+ #: Character class body for kana: hiragana ぁ-ゖ, the hiragana iteration marks ゝゞ, katakana
28
+ #: ァ-ヶ, the prolonged-sound mark ー and the katakana iteration marks ヽヾ. Everything else -
29
+ #: kanji, 々, digits, Latin, punctuation - is a "non-kana run" and receives ruby.
30
+ KANA = "ぁ-ゖゝゞァ-ヶーヽヾ"
31
+
32
+ _RUNS = re.compile(f"[{KANA}]+|[^{KANA}]+")
33
+ _KANA_START = re.compile(f"[{KANA}]")
34
+
35
+ #: Kanji code point ranges: CJK Unified Ideographs Extension A, CJK Unified Ideographs, CJK
36
+ #: Compatibility Ideographs, the supplementary planes (Extensions B-F, e.g. 𠮷), and U+3005 々,
37
+ #: the iteration mark, which repeats the kanji before it and therefore belongs to its run.
38
+ _KANJI_RANGES = (
39
+ (0x3400, 0x4DBF),
40
+ (0x4E00, 0x9FFF),
41
+ (0xF900, 0xFAFF),
42
+ (0x20000, 0x2FA1F),
43
+ )
44
+ _ITERATION_MARK = 0x3005
45
+
46
+
47
+ def is_kanji(ch: str) -> bool:
48
+ """True for a single kanji character (including 々). False for kana, Latin, digits, marks."""
49
+ cp = ord(ch)
50
+ return cp == _ITERATION_MARK or any(lo <= cp <= hi for lo, hi in _KANJI_RANGES)
51
+
52
+
53
+ def _runs(surface: str) -> list[tuple[str, bool]]:
54
+ """``[(text, is_kana), ...]`` - the surface split into maximal kana / non-kana runs."""
55
+ return [(m.group(0), bool(_KANA_START.match(m.group(0)))) for m in _RUNS.finditer(surface)]
56
+
57
+
58
+ def kanji_runs(surface: str) -> list[str]:
59
+ """The non-kana runs of ``surface`` that contain at least one kanji, in order.
60
+
61
+ These are the spans that receive ruby and that the kanji-axis gate evaluates: if any kanji in
62
+ a run is above the learner's level, the whole run shows its reading.
63
+ """
64
+ return [text for text, kana in _runs(surface) if not kana and any(is_kanji(c) for c in text)]
65
+
66
+
67
+ def align_ruby(surface: str, reading_hira: str) -> list[list[str | None]]:
68
+ """Split ``surface`` into ``[text, rt]`` spans: ``rt`` is None over kana, hiragana otherwise.
69
+
70
+ ``reading_hira`` must already be hiragana (the caller converts Sudachi's katakana with
71
+ ``jaconv.kata2hira``). Returns one bare span for an all-kana surface, one whole-unit ruby
72
+ span when the reading cannot be aligned, and never raises.
73
+ """
74
+ runs = _runs(surface)
75
+ if all(kana for _, kana in runs):
76
+ return [[surface, None]]
77
+ if not reading_hira:
78
+ return [[surface, reading_hira]]
79
+ pattern = "".join(
80
+ "(" + re.escape(jaconv.kata2hira(text)) + ")" if kana else "(.+?)" for text, kana in runs
81
+ )
82
+ match = re.fullmatch(pattern, reading_hira)
83
+ if not match:
84
+ return [[surface, reading_hira]] # honest fallback: whole-unit ruby
85
+ return [
86
+ [text, None if kana else group]
87
+ for (text, kana), group in zip(runs, match.groups(), strict=True)
88
+ ]