Spaces:
Running on Zero
Running on Zero
Harden hybrid network inference recovery
Browse files- README.md +13 -7
- app.py +96 -3
- production.py +577 -38
- tests/test_production.py +519 -48
- tests/test_quality_runtime.py +2 -2
- tests/test_release_pins.py +81 -0
README.md
CHANGED
|
@@ -53,7 +53,7 @@ Barbet 另固定在 `6fcd7ce4aa37f2250a3242995bef0fbc3b026ba8`,
|
|
| 53 |
| Sparse completion-headroom policy | every fourth retry: 4.2 CJK / 3.6 ASCII-mixed units/sec + 1 latent step |
|
| 54 |
| Short-text guidance | minimum CFG 3.0 at no more than 6 speech units |
|
| 55 |
| Generation guidance | fixed mixed-CFG assignment within the bounded cascade; public NFE fixed at 10 |
|
| 56 |
-
| Email / URL frontend |
|
| 57 |
| Stop policy | 0.50 → 0.05 from 75% to 95% predicted progress, 1 hit |
|
| 58 |
| Endpoint cue | append terminal punctuation for model input when missing, except very short text |
|
| 59 |
| Hard stop | native-pace target steps, independent of playback pace |
|
|
@@ -70,7 +70,7 @@ Barbet 另固定在 `6fcd7ce4aa37f2250a3242995bef0fbc3b026ba8`,
|
|
| 70 |
| Final output gate | re-verify joined/faded/RMS-matched/speed-adjusted whole waveform; fail closed |
|
| 71 |
| Runtime budget | at most 20 generated TTS chunks and 800 generated speech units per request; NFE fixed at 10 |
|
| 72 |
| Maximum chunk | ordinary text 80 speech units; URL/email-bearing generation chunks 36 units |
|
| 73 |
-
| Minimum chunk | 12 speech units
|
| 74 |
| Crossfade / internal edge fade | semantic boundary 80 ms / 80 ms; proven URL/email internal boundary 5 ms / 5 ms with no inserted pause |
|
| 75 |
| Chunk RMS adjustment | at most 4 dB |
|
| 76 |
| Generated continuation context | 0 sec |
|
|
@@ -137,13 +137,19 @@ transcript。
|
|
| 137 |
日期、24 小時制時間、百分比、常見單位與大寫 acronym/model code 會先轉成保守的
|
| 138 |
zh-TW spoken form,例如 `2026/07/16`、`15:30`、`12.5%` 與 `10 km`。ASR 會先統一
|
| 139 |
繁簡字形再評分,避免把正確的台灣華語輸出誤判為內容錯誤。
|
| 140 |
-
Email 與 URL 會以可辨識
|
| 141 |
-
|
| 142 |
-
讀成「點、
|
| 143 |
完整 identifier 仍保留為 joined、full-large-v3 與 final ASR 的 exact protected target;只有模型
|
| 144 |
generation 會在由原始 ASCII grammar 證明的 scheme/domain/path/query/email component 邊界切段,
|
| 145 |
-
以 32 units 為目標、36 units 為硬上限
|
| 146 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 147 |
這能降低模型把不常見 TLD 自動補成 `.com` 的風險;一般英文句子不會套用這個規則。
|
| 148 |
URL 與後續英文 prose 應以空白或中文標點分隔;未分隔的 RFC path punctuation 會視為 URL
|
| 149 |
本身的一部分並納入 exact gate。Quoted email local-part 暫不支援,輸入時會直接 fail closed。
|
|
|
|
| 53 |
| Sparse completion-headroom policy | every fourth retry: 4.2 CJK / 3.6 ASCII-mixed units/sec + 1 latent step |
|
| 54 |
| Short-text guidance | minimum CFG 3.0 at no more than 6 speech units |
|
| 55 |
| Generation guidance | fixed mixed-CFG assignment within the bounded cascade; public NFE fixed at 10 |
|
| 56 |
+
| Email / URL frontend | hybrid lexical labels; scheme, all-uppercase atoms and short suffixes use explicit ASCII letter tokens; digits are read one by one and separators remain audible |
|
| 57 |
| Stop policy | 0.50 → 0.05 from 75% to 95% predicted progress, 1 hit |
|
| 58 |
| Endpoint cue | append terminal punctuation for model input when missing, except very short text |
|
| 59 |
| Hard stop | native-pace target steps, independent of playback pace |
|
|
|
|
| 70 |
| Final output gate | re-verify joined/faded/RMS-matched/speed-adjusted whole waveform; fail closed |
|
| 71 |
| Runtime budget | at most 20 generated TTS chunks and 800 generated speech units per request; NFE fixed at 10 |
|
| 72 |
| Maximum chunk | ordinary text 80 speech units; URL/email-bearing generation chunks 36 units |
|
| 73 |
+
| Minimum chunk | ordinary text 12 speech units; URL/email-bearing chunks 8 units; genuine short non-network requests/natural short sentence boundaries are preserved |
|
| 74 |
| Crossfade / internal edge fade | semantic boundary 80 ms / 80 ms; proven URL/email internal boundary 5 ms / 5 ms with no inserted pause |
|
| 75 |
| Chunk RMS adjustment | at most 4 dB |
|
| 76 |
| Generated continuation context | 0 sec |
|
|
|
|
| 137 |
日期、24 小時制時間、百分比、常見單位與大寫 acronym/model code 會先轉成保守的
|
| 138 |
zh-TW spoken form,例如 `2026/07/16`、`15:30`、`12.5%` 與 `10 km`。ASR 會先統一
|
| 139 |
繁簡字形再評分,避免把正確的台灣華語輸出誤判為內容錯誤。
|
| 140 |
+
Email 與 URL 會以混合式可辨識讀法展開:一般 local/domain/path label 保留 lexical ASCII
|
| 141 |
+
(例如 `tour`、`help`),scheme、全大寫 atom 與短 suffix 使用明確 ASCII letter tokens,數字逐位
|
| 142 |
+
朗讀,分隔符也明確朗讀(例如 `.tw` 讀成「點、T、W」)。
|
| 143 |
完整 identifier 仍保留為 joined、full-large-v3 與 final ASR 的 exact protected target;只有模型
|
| 144 |
generation 會在由原始 ASCII grammar 證明的 scheme/domain/path/query/email component 邊界切段,
|
| 145 |
+
以 32 units 為目標、36 units 為硬上限、8 units 為 network minimum。Email 第一輪必須在 `小老鼠`
|
| 146 |
+
之前切開;有 scheme 的 URL 第一輪必須在 `冒號 斜線 斜線` 之後切開,第一輪 DP 不得跨越這兩類
|
| 147 |
+
grammar boundary。只有像 `a@b.co`、`https://a.tw` 這類 identifier 的某一側本身不足 8 units 且
|
| 148 |
+
第一輪無解時,才放寬該短 identifier 的 preferred cut 並重跑相同 component-safe DP;其他強制切點
|
| 149 |
+
不變。Identifier 內部邊界不插入 pause,只使用 5 ms fade/crossfade;component proof 無法完整
|
| 150 |
+
重建、非 ASCII IRI 或單一不可拆 component 超限時會 fail closed。
|
| 151 |
+
Local fragment gate 只接受 range-bound exact proof,並檢查所有最佳 edit alignment;重複的普通文字
|
| 152 |
+
不能借用受保護片段的 proof。Joined、whole 與 final gate 不套用這項 local-only canonicalization。
|
| 153 |
這能降低模型把不常見 TLD 自動補成 `.com` 的風險;一般英文句子不會套用這個規則。
|
| 154 |
URL 與後續英文 prose 應以空白或中文標點分隔;未分隔的 RFC path punctuation 會視為 URL
|
| 155 |
本身的一部分並納入 exact gate。Quoted email local-part 暫不支援,輸入時會直接 fail closed。
|
app.py
CHANGED
|
@@ -19,8 +19,10 @@ from bluemagpie import BlueMagpieModel
|
|
| 19 |
from production import (
|
| 20 |
StopHysteresisController,
|
| 21 |
GenerationChunkSpec,
|
|
|
|
| 22 |
active_pace_correction_speed,
|
| 23 |
apply_loudness_floor,
|
|
|
|
| 24 |
coalesce_text_chunks,
|
| 25 |
contains_network_identifier,
|
| 26 |
count_speech_units,
|
|
@@ -115,6 +117,7 @@ MIN_CHUNK_CHARS = 12
|
|
| 115 |
CROSSFADE_MS = 80.0
|
| 116 |
CHUNK_EDGE_FADE_MS = 80.0
|
| 117 |
CHUNK_RMS_MATCH_DB = 4.0
|
|
|
|
| 118 |
NETWORK_GENERATION_TARGET_UNITS = 32
|
| 119 |
NETWORK_GENERATION_MAX_UNITS = 36
|
| 120 |
NETWORK_INTERNAL_FADE_MS = 5.0
|
|
@@ -442,13 +445,26 @@ def _verify_trajectory_audio(
|
|
| 442 |
release_speaker_gate: bool = False,
|
| 443 |
transcriber=transcribe_whisper,
|
| 444 |
semantic_only: bool = False,
|
|
|
|
|
|
|
|
|
|
| 445 |
):
|
| 446 |
if len(trajectory) != len(chunks):
|
| 447 |
return verify_trajectory(())
|
|
|
|
|
|
|
|
|
|
|
|
|
| 448 |
encoder = None
|
| 449 |
observations: list[CandidateObservation] = []
|
| 450 |
artifacts: list[ChunkCandidateArtifact] = []
|
| 451 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 452 |
prepared = prepare_candidate_audio(
|
| 453 |
audio,
|
| 454 |
SR,
|
|
@@ -475,6 +491,19 @@ def _verify_trajectory_audio(
|
|
| 475 |
except ValueError:
|
| 476 |
duration = 0.0
|
| 477 |
transcript = prepared.transcript_text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 478 |
speaker_similarity = None
|
| 479 |
begin_similarity = None
|
| 480 |
end_similarity = None
|
|
@@ -611,6 +640,37 @@ def _verify_independent_whole_audio(
|
|
| 611 |
)
|
| 612 |
|
| 613 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 614 |
def _assemble_trajectory_audio(
|
| 615 |
trajectory: tuple[np.ndarray, ...],
|
| 616 |
chunks: tuple[str, ...],
|
|
@@ -742,6 +802,7 @@ def _qualify_candidate_trajectory_audio(
|
|
| 742 |
chunks,
|
| 743 |
anchor,
|
| 744 |
playback_speed,
|
|
|
|
| 745 |
)
|
| 746 |
if not local_verification.passed:
|
| 747 |
return CandidateVerification(local_verification)
|
|
@@ -802,6 +863,7 @@ def _verify_refill_candidate_trajectory_audio(
|
|
| 802 |
chunks: tuple[str, ...],
|
| 803 |
anchor: np.ndarray,
|
| 804 |
playback_speed: float,
|
|
|
|
| 805 |
):
|
| 806 |
"""Apply the unchanged strict local gates to one safe-duration refill."""
|
| 807 |
|
|
@@ -812,6 +874,7 @@ def _verify_refill_candidate_trajectory_audio(
|
|
| 812 |
chunks,
|
| 813 |
anchor,
|
| 814 |
playback_speed,
|
|
|
|
| 815 |
)
|
| 816 |
|
| 817 |
|
|
@@ -922,6 +985,7 @@ def _synthesize(
|
|
| 922 |
raw_text,
|
| 923 |
text,
|
| 924 |
min_units=MIN_CHUNK_CHARS,
|
|
|
|
| 925 |
target_units=NETWORK_GENERATION_TARGET_UNITS,
|
| 926 |
network_max_units=NETWORK_GENERATION_MAX_UNITS,
|
| 927 |
ordinary_max_units=CHUNK_CHARS,
|
|
@@ -948,6 +1012,30 @@ def _synthesize(
|
|
| 948 |
request_seed = resolve_request_seed(request_seed, secrets.randbelow)
|
| 949 |
anchor = _speaker_anchor_array(centroid)
|
| 950 |
independent_cache = WholeWaveformVerificationCache()
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 951 |
|
| 952 |
def candidate_cfg(candidate_ordinal: int) -> float:
|
| 953 |
return generation_cfg_for_candidate_offset(
|
|
@@ -1032,6 +1120,10 @@ def _synthesize(
|
|
| 1032 |
) -> tuple[np.ndarray, ...]:
|
| 1033 |
if generation_context.seed != seed:
|
| 1034 |
raise ValueError("generation context seed does not match the request seed")
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1035 |
ordinals = generation_context.chunk_candidate_ordinals
|
| 1036 |
if not ordinals or len(set(ordinals)) != 1:
|
| 1037 |
raise ValueError("one generation call must use one candidate ordinal")
|
|
@@ -1113,7 +1205,7 @@ def _synthesize(
|
|
| 1113 |
speed,
|
| 1114 |
independent_cache,
|
| 1115 |
candidate_seed=seed,
|
| 1116 |
-
chunk_specs=
|
| 1117 |
),
|
| 1118 |
lambda trajectory, candidate_chunks, seed: (
|
| 1119 |
_qualify_candidate_trajectory_audio(
|
|
@@ -1124,7 +1216,7 @@ def _synthesize(
|
|
| 1124 |
speed,
|
| 1125 |
independent_cache,
|
| 1126 |
candidate_seed=seed,
|
| 1127 |
-
chunk_specs=
|
| 1128 |
)
|
| 1129 |
if len(chunks) == 1
|
| 1130 |
else _verify_refill_candidate_trajectory_audio(
|
|
@@ -1132,6 +1224,7 @@ def _synthesize(
|
|
| 1132 |
candidate_chunks,
|
| 1133 |
anchor,
|
| 1134 |
speed,
|
|
|
|
| 1135 |
)
|
| 1136 |
),
|
| 1137 |
sequence_final_verifier=lambda sequence_result, candidate_chunks: (
|
|
|
|
| 19 |
from production import (
|
| 20 |
StopHysteresisController,
|
| 21 |
GenerationChunkSpec,
|
| 22 |
+
NetworkFragmentProof,
|
| 23 |
active_pace_correction_speed,
|
| 24 |
apply_loudness_floor,
|
| 25 |
+
canonicalize_asr_network_fragments,
|
| 26 |
coalesce_text_chunks,
|
| 27 |
contains_network_identifier,
|
| 28 |
count_speech_units,
|
|
|
|
| 117 |
CROSSFADE_MS = 80.0
|
| 118 |
CHUNK_EDGE_FADE_MS = 80.0
|
| 119 |
CHUNK_RMS_MATCH_DB = 4.0
|
| 120 |
+
NETWORK_GENERATION_MIN_UNITS = 8
|
| 121 |
NETWORK_GENERATION_TARGET_UNITS = 32
|
| 122 |
NETWORK_GENERATION_MAX_UNITS = 36
|
| 123 |
NETWORK_INTERNAL_FADE_MS = 5.0
|
|
|
|
| 445 |
release_speaker_gate: bool = False,
|
| 446 |
transcriber=transcribe_whisper,
|
| 447 |
semantic_only: bool = False,
|
| 448 |
+
network_fragment_proofs: (
|
| 449 |
+
tuple[tuple[NetworkFragmentProof, ...], ...] | None
|
| 450 |
+
) = None,
|
| 451 |
):
|
| 452 |
if len(trajectory) != len(chunks):
|
| 453 |
return verify_trajectory(())
|
| 454 |
+
if network_fragment_proofs is not None and len(network_fragment_proofs) != len(
|
| 455 |
+
chunks
|
| 456 |
+
):
|
| 457 |
+
raise ValueError("network fragment proofs must align with local chunks")
|
| 458 |
encoder = None
|
| 459 |
observations: list[CandidateObservation] = []
|
| 460 |
artifacts: list[ChunkCandidateArtifact] = []
|
| 461 |
+
proof_rows = network_fragment_proofs or ((),) * len(chunks)
|
| 462 |
+
for chunk, audio, fragment_proofs in zip(
|
| 463 |
+
chunks,
|
| 464 |
+
trajectory,
|
| 465 |
+
proof_rows,
|
| 466 |
+
strict=True,
|
| 467 |
+
):
|
| 468 |
prepared = prepare_candidate_audio(
|
| 469 |
audio,
|
| 470 |
SR,
|
|
|
|
| 491 |
except ValueError:
|
| 492 |
duration = 0.0
|
| 493 |
transcript = prepared.transcript_text
|
| 494 |
+
if fragment_proofs:
|
| 495 |
+
fragment_evidence = canonicalize_asr_network_fragments(
|
| 496 |
+
transcript,
|
| 497 |
+
chunk,
|
| 498 |
+
fragment_proofs,
|
| 499 |
+
)
|
| 500 |
+
# A range-bound identifier mismatch is a hard local semantic
|
| 501 |
+
# failure even when its contribution to whole-chunk CER is small.
|
| 502 |
+
transcript = (
|
| 503 |
+
fragment_evidence.transcript_text
|
| 504 |
+
if fragment_evidence.passed
|
| 505 |
+
else ""
|
| 506 |
+
)
|
| 507 |
speaker_similarity = None
|
| 508 |
begin_similarity = None
|
| 509 |
end_similarity = None
|
|
|
|
| 640 |
)
|
| 641 |
|
| 642 |
|
| 643 |
+
def _network_fragment_proof_rows(
|
| 644 |
+
chunks: tuple[str, ...],
|
| 645 |
+
chunk_specs: tuple[GenerationChunkSpec, ...] | None,
|
| 646 |
+
) -> tuple[tuple[NetworkFragmentProof, ...], ...] | None:
|
| 647 |
+
"""Validate explicit planner provenance before any local ASR relaxation."""
|
| 648 |
+
|
| 649 |
+
if chunk_specs is None:
|
| 650 |
+
return None
|
| 651 |
+
if len(chunk_specs) != len(chunks):
|
| 652 |
+
raise ValueError("network fragment provenance must align with chunks")
|
| 653 |
+
rows: list[tuple[NetworkFragmentProof, ...]] = []
|
| 654 |
+
for chunk, spec in zip(chunks, chunk_specs, strict=True):
|
| 655 |
+
proofs = spec.network_fragment_proofs
|
| 656 |
+
if spec.text != chunk:
|
| 657 |
+
raise ValueError("network fragment provenance text does not match")
|
| 658 |
+
if spec.network_conditioned:
|
| 659 |
+
if (
|
| 660 |
+
not proofs
|
| 661 |
+
or len(proofs) != len(spec.network_span_indices)
|
| 662 |
+
or tuple(proof.span_index for proof in proofs)
|
| 663 |
+
!= spec.network_span_indices
|
| 664 |
+
or tuple(proof.full_spoken_proof for proof in proofs)
|
| 665 |
+
!= spec.network_full_spoken_proofs
|
| 666 |
+
):
|
| 667 |
+
raise ValueError("network-conditioned chunk lacks exact fragment proof")
|
| 668 |
+
elif proofs or spec.network_full_spoken_proofs:
|
| 669 |
+
raise ValueError("ordinary chunk carries network fragment proof")
|
| 670 |
+
rows.append(proofs)
|
| 671 |
+
return tuple(rows)
|
| 672 |
+
|
| 673 |
+
|
| 674 |
def _assemble_trajectory_audio(
|
| 675 |
trajectory: tuple[np.ndarray, ...],
|
| 676 |
chunks: tuple[str, ...],
|
|
|
|
| 802 |
chunks,
|
| 803 |
anchor,
|
| 804 |
playback_speed,
|
| 805 |
+
network_fragment_proofs=_network_fragment_proof_rows(chunks, chunk_specs),
|
| 806 |
)
|
| 807 |
if not local_verification.passed:
|
| 808 |
return CandidateVerification(local_verification)
|
|
|
|
| 863 |
chunks: tuple[str, ...],
|
| 864 |
anchor: np.ndarray,
|
| 865 |
playback_speed: float,
|
| 866 |
+
chunk_specs: tuple[GenerationChunkSpec, ...] | None = None,
|
| 867 |
):
|
| 868 |
"""Apply the unchanged strict local gates to one safe-duration refill."""
|
| 869 |
|
|
|
|
| 874 |
chunks,
|
| 875 |
anchor,
|
| 876 |
playback_speed,
|
| 877 |
+
network_fragment_proofs=_network_fragment_proof_rows(chunks, chunk_specs),
|
| 878 |
)
|
| 879 |
|
| 880 |
|
|
|
|
| 985 |
raw_text,
|
| 986 |
text,
|
| 987 |
min_units=MIN_CHUNK_CHARS,
|
| 988 |
+
network_min_units=NETWORK_GENERATION_MIN_UNITS,
|
| 989 |
target_units=NETWORK_GENERATION_TARGET_UNITS,
|
| 990 |
network_max_units=NETWORK_GENERATION_MAX_UNITS,
|
| 991 |
ordinary_max_units=CHUNK_CHARS,
|
|
|
|
| 1012 |
request_seed = resolve_request_seed(request_seed, secrets.randbelow)
|
| 1013 |
anchor = _speaker_anchor_array(centroid)
|
| 1014 |
independent_cache = WholeWaveformVerificationCache()
|
| 1015 |
+
generation_context_by_seed: dict[int, CandidateGenerationContext] = {}
|
| 1016 |
+
|
| 1017 |
+
def generation_chunk_specs(
|
| 1018 |
+
seed: int,
|
| 1019 |
+
candidate_chunks: tuple[str, ...],
|
| 1020 |
+
) -> tuple[GenerationChunkSpec, ...] | None:
|
| 1021 |
+
if chunk_specs is None:
|
| 1022 |
+
return None
|
| 1023 |
+
context = generation_context_by_seed.get(seed)
|
| 1024 |
+
if context is None or len(context.chunk_indices) != len(candidate_chunks):
|
| 1025 |
+
raise ValueError("candidate verification lacks generation provenance")
|
| 1026 |
+
selected: list[GenerationChunkSpec] = []
|
| 1027 |
+
for chunk_index, chunk in zip(
|
| 1028 |
+
context.chunk_indices,
|
| 1029 |
+
candidate_chunks,
|
| 1030 |
+
strict=True,
|
| 1031 |
+
):
|
| 1032 |
+
if not 0 <= chunk_index < len(chunk_specs):
|
| 1033 |
+
raise ValueError("candidate verification provenance is out of range")
|
| 1034 |
+
spec = chunk_specs[chunk_index]
|
| 1035 |
+
if spec.text != chunk:
|
| 1036 |
+
raise ValueError("candidate verification provenance text does not match")
|
| 1037 |
+
selected.append(spec)
|
| 1038 |
+
return tuple(selected)
|
| 1039 |
|
| 1040 |
def candidate_cfg(candidate_ordinal: int) -> float:
|
| 1041 |
return generation_cfg_for_candidate_offset(
|
|
|
|
| 1120 |
) -> tuple[np.ndarray, ...]:
|
| 1121 |
if generation_context.seed != seed:
|
| 1122 |
raise ValueError("generation context seed does not match the request seed")
|
| 1123 |
+
previous_context = generation_context_by_seed.get(seed)
|
| 1124 |
+
if previous_context is not None and previous_context != generation_context:
|
| 1125 |
+
raise ValueError("one candidate seed cannot carry two generation contexts")
|
| 1126 |
+
generation_context_by_seed[seed] = generation_context
|
| 1127 |
ordinals = generation_context.chunk_candidate_ordinals
|
| 1128 |
if not ordinals or len(set(ordinals)) != 1:
|
| 1129 |
raise ValueError("one generation call must use one candidate ordinal")
|
|
|
|
| 1205 |
speed,
|
| 1206 |
independent_cache,
|
| 1207 |
candidate_seed=seed,
|
| 1208 |
+
chunk_specs=generation_chunk_specs(seed, candidate_chunks),
|
| 1209 |
),
|
| 1210 |
lambda trajectory, candidate_chunks, seed: (
|
| 1211 |
_qualify_candidate_trajectory_audio(
|
|
|
|
| 1216 |
speed,
|
| 1217 |
independent_cache,
|
| 1218 |
candidate_seed=seed,
|
| 1219 |
+
chunk_specs=generation_chunk_specs(seed, candidate_chunks),
|
| 1220 |
)
|
| 1221 |
if len(chunks) == 1
|
| 1222 |
else _verify_refill_candidate_trajectory_audio(
|
|
|
|
| 1224 |
candidate_chunks,
|
| 1225 |
anchor,
|
| 1226 |
speed,
|
| 1227 |
+
generation_chunk_specs(seed, candidate_chunks),
|
| 1228 |
)
|
| 1229 |
),
|
| 1230 |
sequence_final_verifier=lambda sequence_result, candidate_chunks: (
|
production.py
CHANGED
|
@@ -1030,7 +1030,7 @@ def _canonicalize_target_proven_frequency_scales(
|
|
| 1030 |
|
| 1031 |
|
| 1032 |
def _zh_network_text(value: str) -> str:
|
| 1033 |
-
"""
|
| 1034 |
|
| 1035 |
output: list[str] = []
|
| 1036 |
buffer: list[str] = []
|
|
@@ -1042,7 +1042,10 @@ def _zh_network_text(value: str) -> str:
|
|
| 1042 |
if token.isdigit():
|
| 1043 |
output.append(_zh_digit_sequence(token))
|
| 1044 |
elif token.isascii() and token.isalpha():
|
| 1045 |
-
|
|
|
|
|
|
|
|
|
|
| 1046 |
else:
|
| 1047 |
output.append(token)
|
| 1048 |
buffer.clear()
|
|
@@ -1070,10 +1073,10 @@ def _zh_network_text(value: str) -> str:
|
|
| 1070 |
|
| 1071 |
|
| 1072 |
def _zh_network_letters(value: str) -> str:
|
| 1073 |
-
"""Return
|
| 1074 |
|
| 1075 |
return " ".join(
|
| 1076 |
-
|
| 1077 |
for character in value
|
| 1078 |
if not character.isspace()
|
| 1079 |
)
|
|
@@ -2660,6 +2663,25 @@ def split_text_for_tts(text: str, max_chars: int = 80, min_chunk_chars: int = 12
|
|
| 2660 |
return chunks or [text]
|
| 2661 |
|
| 2662 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2663 |
@dataclass(frozen=True)
|
| 2664 |
class GenerationChunkSpec:
|
| 2665 |
"""One generation-only slice with immutable full-identifier provenance."""
|
|
@@ -2670,6 +2692,7 @@ class GenerationChunkSpec:
|
|
| 2670 |
network_span_indices: tuple[int, ...] = ()
|
| 2671 |
network_component_indices: tuple[tuple[int, int], ...] = ()
|
| 2672 |
network_full_spoken_proofs: tuple[str, ...] = ()
|
|
|
|
| 2673 |
boundary_after: str = "none"
|
| 2674 |
|
| 2675 |
@property
|
|
@@ -2683,6 +2706,7 @@ class _NetworkGenerationSpan:
|
|
| 2683 |
end: int
|
| 2684 |
spoken_proof: str
|
| 2685 |
component_ranges: tuple[tuple[int, int], ...]
|
|
|
|
| 2686 |
|
| 2687 |
|
| 2688 |
def _raw_network_generation_matches(
|
|
@@ -2762,6 +2786,50 @@ def _network_component_ranges(
|
|
| 2762 |
return tuple(ranges)
|
| 2763 |
|
| 2764 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2765 |
def _network_generation_spans(
|
| 2766 |
raw_text: str,
|
| 2767 |
normalized_text: str,
|
|
@@ -2810,12 +2878,28 @@ def _network_generation_spans(
|
|
| 2810 |
absolute_start=start,
|
| 2811 |
hard_max_units=hard_max_units,
|
| 2812 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2813 |
spans.append(
|
| 2814 |
_NetworkGenerationSpan(
|
| 2815 |
start=start,
|
| 2816 |
end=end,
|
| 2817 |
spoken_proof=spoken,
|
| 2818 |
component_ranges=components,
|
|
|
|
| 2819 |
)
|
| 2820 |
)
|
| 2821 |
cursor = end
|
|
@@ -2832,6 +2916,7 @@ def _network_generation_group_plan(
|
|
| 2832 |
original_boundaries: Sequence[int],
|
| 2833 |
spans: Sequence[_NetworkGenerationSpan],
|
| 2834 |
minimum: int,
|
|
|
|
| 2835 |
target: int,
|
| 2836 |
network_maximum: int,
|
| 2837 |
ordinary_maximum: int,
|
|
@@ -2862,6 +2947,13 @@ def _network_generation_group_plan(
|
|
| 2862 |
for offset in component
|
| 2863 |
}
|
| 2864 |
preferred_cuts.update(component_cuts)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2865 |
network_interiors = {
|
| 2866 |
offset
|
| 2867 |
for span in group_spans
|
|
@@ -2881,39 +2973,68 @@ def _network_generation_group_plan(
|
|
| 2881 |
}
|
| 2882 |
|
| 2883 |
ordered = sorted(offset for offset in cuts if group_start <= offset <= group_end)
|
| 2884 |
-
|
| 2885 |
-
|
| 2886 |
-
|
| 2887 |
-
|
| 2888 |
-
|
| 2889 |
-
|
| 2890 |
-
|
| 2891 |
-
|
| 2892 |
-
|
| 2893 |
-
|
| 2894 |
-
|
| 2895 |
-
|
| 2896 |
-
|
| 2897 |
-
|
| 2898 |
-
|
| 2899 |
-
|
| 2900 |
-
|
| 2901 |
-
|
| 2902 |
-
|
| 2903 |
-
|
| 2904 |
-
|
| 2905 |
-
|
| 2906 |
-
|
| 2907 |
-
|
| 2908 |
-
|
| 2909 |
-
|
| 2910 |
-
|
| 2911 |
-
|
| 2912 |
-
|
| 2913 |
-
|
| 2914 |
-
|
| 2915 |
-
|
| 2916 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2917 |
if plan is None:
|
| 2918 |
raise ValueError("network-bearing sentence cannot satisfy generation limits")
|
| 2919 |
output: list[tuple[int, int]] = []
|
|
@@ -3019,6 +3140,7 @@ def plan_generation_chunks(
|
|
| 3019 |
normalized_text: str,
|
| 3020 |
*,
|
| 3021 |
min_units: int = 12,
|
|
|
|
| 3022 |
target_units: int = 32,
|
| 3023 |
network_max_units: int = 36,
|
| 3024 |
ordinary_max_units: int = 80,
|
|
@@ -3034,10 +3156,17 @@ def plan_generation_chunks(
|
|
| 3034 |
if network_identifier_has_ambiguous_iri(raw_text):
|
| 3035 |
raise ValueError("network generation does not support non-ASCII IRI")
|
| 3036 |
minimum = max(1, int(min_units))
|
|
|
|
| 3037 |
target = max(minimum, int(target_units))
|
| 3038 |
network_maximum = max(minimum, int(network_max_units))
|
| 3039 |
ordinary_maximum = max(minimum, int(ordinary_max_units))
|
| 3040 |
-
if not
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3041 |
raise ValueError("generation chunk limits are inconsistent")
|
| 3042 |
|
| 3043 |
normalized = normalize_tts_text(normalized_text)
|
|
@@ -3088,6 +3217,7 @@ def plan_generation_chunks(
|
|
| 3088 |
original_boundaries=(),
|
| 3089 |
spans=spans,
|
| 3090 |
minimum=minimum,
|
|
|
|
| 3091 |
target=target,
|
| 3092 |
network_maximum=network_maximum,
|
| 3093 |
ordinary_maximum=ordinary_maximum,
|
|
@@ -3141,6 +3271,19 @@ def plan_generation_chunks(
|
|
| 3141 |
)
|
| 3142 |
if left < end and right > start
|
| 3143 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3144 |
specs.append(
|
| 3145 |
GenerationChunkSpec(
|
| 3146 |
text=chunk,
|
|
@@ -3151,6 +3294,7 @@ def plan_generation_chunks(
|
|
| 3151 |
network_full_spoken_proofs=tuple(
|
| 3152 |
spans[span_index].spoken_proof for span_index in overlapping
|
| 3153 |
),
|
|
|
|
| 3154 |
boundary_after=(
|
| 3155 |
"network_internal"
|
| 3156 |
if end in internal_boundaries
|
|
@@ -3166,6 +3310,63 @@ def plan_generation_chunks(
|
|
| 3166 |
for spec in specs
|
| 3167 |
):
|
| 3168 |
raise ValueError("generation chunk exceeds its provenance-specific limit")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3169 |
return tuple(specs)
|
| 3170 |
|
| 3171 |
|
|
@@ -3710,6 +3911,344 @@ def _levenshtein_alignment(source: str, hypothesis: str) -> list[tuple[str, int,
|
|
| 3710 |
return operations
|
| 3711 |
|
| 3712 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3713 |
@dataclass(frozen=True)
|
| 3714 |
class AsrComparison:
|
| 3715 |
"""Orthographic whole-text and acoustic endpoint evidence for one candidate."""
|
|
|
|
| 1030 |
|
| 1031 |
|
| 1032 |
def _zh_network_text(value: str) -> str:
|
| 1033 |
+
"""Read lexical labels naturally while keeping typed atoms explicit."""
|
| 1034 |
|
| 1035 |
output: list[str] = []
|
| 1036 |
buffer: list[str] = []
|
|
|
|
| 1042 |
if token.isdigit():
|
| 1043 |
output.append(_zh_digit_sequence(token))
|
| 1044 |
elif token.isascii() and token.isalpha():
|
| 1045 |
+
if token.casefold() == "www" or (len(token) > 1 and token.isupper()):
|
| 1046 |
+
output.append(_zh_network_letters(token))
|
| 1047 |
+
else:
|
| 1048 |
+
output.append(token)
|
| 1049 |
else:
|
| 1050 |
output.append(token)
|
| 1051 |
buffer.clear()
|
|
|
|
| 1073 |
|
| 1074 |
|
| 1075 |
def _zh_network_letters(value: str) -> str:
|
| 1076 |
+
"""Return model-native spaced uppercase letters for typed ASCII atoms."""
|
| 1077 |
|
| 1078 |
return " ".join(
|
| 1079 |
+
character.upper()
|
| 1080 |
for character in value
|
| 1081 |
if not character.isspace()
|
| 1082 |
)
|
|
|
|
| 2663 |
return chunks or [text]
|
| 2664 |
|
| 2665 |
|
| 2666 |
+
@dataclass(frozen=True)
|
| 2667 |
+
class NetworkFragmentProof:
|
| 2668 |
+
"""Range-bound local slice of one planner-proven network identifier.
|
| 2669 |
+
|
| 2670 |
+
``chunk_start``/``chunk_end`` are offsets in the exact generation chunk.
|
| 2671 |
+
``parent_start``/``parent_end`` are offsets in ``full_spoken_proof``. The
|
| 2672 |
+
two slices must be byte-for-byte equal after NFC normalization. Keeping
|
| 2673 |
+
both coordinate systems prevents a repeated label elsewhere in the chunk
|
| 2674 |
+
from borrowing a parent URL/email proof.
|
| 2675 |
+
"""
|
| 2676 |
+
|
| 2677 |
+
span_index: int
|
| 2678 |
+
chunk_start: int
|
| 2679 |
+
chunk_end: int
|
| 2680 |
+
parent_start: int
|
| 2681 |
+
parent_end: int
|
| 2682 |
+
full_spoken_proof: str
|
| 2683 |
+
|
| 2684 |
+
|
| 2685 |
@dataclass(frozen=True)
|
| 2686 |
class GenerationChunkSpec:
|
| 2687 |
"""One generation-only slice with immutable full-identifier provenance."""
|
|
|
|
| 2692 |
network_span_indices: tuple[int, ...] = ()
|
| 2693 |
network_component_indices: tuple[tuple[int, int], ...] = ()
|
| 2694 |
network_full_spoken_proofs: tuple[str, ...] = ()
|
| 2695 |
+
network_fragment_proofs: tuple[NetworkFragmentProof, ...] = ()
|
| 2696 |
boundary_after: str = "none"
|
| 2697 |
|
| 2698 |
@property
|
|
|
|
| 2706 |
end: int
|
| 2707 |
spoken_proof: str
|
| 2708 |
component_ranges: tuple[tuple[int, int], ...]
|
| 2709 |
+
mandatory_cut_offsets: tuple[int, ...]
|
| 2710 |
|
| 2711 |
|
| 2712 |
def _raw_network_generation_matches(
|
|
|
|
| 2786 |
return tuple(ranges)
|
| 2787 |
|
| 2788 |
|
| 2789 |
+
def _network_mandatory_cut_offsets(
|
| 2790 |
+
spoken: str,
|
| 2791 |
+
*,
|
| 2792 |
+
kind: str,
|
| 2793 |
+
absolute_start: int,
|
| 2794 |
+
) -> tuple[int, ...]:
|
| 2795 |
+
"""Return grammar boundaries that no generation chunk may cross.
|
| 2796 |
+
|
| 2797 |
+
Hybrid lexical labels are substantially easier for the model than spelling
|
| 2798 |
+
every letter, but a whole email/URL still asks one trajectory to preserve a
|
| 2799 |
+
long opaque identifier. Split email local/domain parts and URL
|
| 2800 |
+
scheme/remainder parts independently so coverage selection can combine
|
| 2801 |
+
exact fragments without relaxing whole-output verification.
|
| 2802 |
+
"""
|
| 2803 |
+
|
| 2804 |
+
token_matches = tuple(re.finditer(r"\S+", spoken))
|
| 2805 |
+
tokens = tuple(
|
| 2806 |
+
_SIMPLIFIED_NETWORK_SYMBOL_READINGS.get(match.group(0), match.group(0))
|
| 2807 |
+
for match in token_matches
|
| 2808 |
+
)
|
| 2809 |
+
offsets: list[int] = []
|
| 2810 |
+
if kind == "email":
|
| 2811 |
+
separators = [
|
| 2812 |
+
match.start()
|
| 2813 |
+
for token, match in zip(tokens, token_matches, strict=True)
|
| 2814 |
+
if token == "小老鼠"
|
| 2815 |
+
]
|
| 2816 |
+
if len(separators) != 1:
|
| 2817 |
+
raise ValueError("email generation proof lacks one separator")
|
| 2818 |
+
offsets.append(separators[0])
|
| 2819 |
+
elif kind == "url":
|
| 2820 |
+
scheme_ends = [
|
| 2821 |
+
token_matches[index].end()
|
| 2822 |
+
for index in range(2, len(tokens))
|
| 2823 |
+
if tokens[index - 2 : index + 1] == ("冒號", "斜線", "斜線")
|
| 2824 |
+
]
|
| 2825 |
+
if len(scheme_ends) > 1:
|
| 2826 |
+
raise ValueError("URL generation proof has ambiguous scheme separators")
|
| 2827 |
+
offsets.extend(scheme_ends)
|
| 2828 |
+
else:
|
| 2829 |
+
raise ValueError("network generation proof has an invalid kind")
|
| 2830 |
+
return tuple(absolute_start + offset for offset in offsets)
|
| 2831 |
+
|
| 2832 |
+
|
| 2833 |
def _network_generation_spans(
|
| 2834 |
raw_text: str,
|
| 2835 |
normalized_text: str,
|
|
|
|
| 2878 |
absolute_start=start,
|
| 2879 |
hard_max_units=hard_max_units,
|
| 2880 |
)
|
| 2881 |
+
mandatory_cuts = _network_mandatory_cut_offsets(
|
| 2882 |
+
spoken,
|
| 2883 |
+
kind=kind,
|
| 2884 |
+
absolute_start=start,
|
| 2885 |
+
)
|
| 2886 |
+
component_boundaries = {
|
| 2887 |
+
boundary
|
| 2888 |
+
for component in components
|
| 2889 |
+
for boundary in component
|
| 2890 |
+
}
|
| 2891 |
+
if any(
|
| 2892 |
+
cut <= start or cut >= end or cut not in component_boundaries
|
| 2893 |
+
for cut in mandatory_cuts
|
| 2894 |
+
):
|
| 2895 |
+
raise ValueError("mandatory network cut lacks component provenance")
|
| 2896 |
spans.append(
|
| 2897 |
_NetworkGenerationSpan(
|
| 2898 |
start=start,
|
| 2899 |
end=end,
|
| 2900 |
spoken_proof=spoken,
|
| 2901 |
component_ranges=components,
|
| 2902 |
+
mandatory_cut_offsets=mandatory_cuts,
|
| 2903 |
)
|
| 2904 |
)
|
| 2905 |
cursor = end
|
|
|
|
| 2916 |
original_boundaries: Sequence[int],
|
| 2917 |
spans: Sequence[_NetworkGenerationSpan],
|
| 2918 |
minimum: int,
|
| 2919 |
+
network_minimum: int,
|
| 2920 |
target: int,
|
| 2921 |
network_maximum: int,
|
| 2922 |
ordinary_maximum: int,
|
|
|
|
| 2947 |
for offset in component
|
| 2948 |
}
|
| 2949 |
preferred_cuts.update(component_cuts)
|
| 2950 |
+
mandatory_cuts = {
|
| 2951 |
+
offset
|
| 2952 |
+
for span in group_spans
|
| 2953 |
+
for offset in span.mandatory_cut_offsets
|
| 2954 |
+
}
|
| 2955 |
+
if not mandatory_cuts.issubset(component_cuts):
|
| 2956 |
+
raise ValueError("mandatory network cuts lack component provenance")
|
| 2957 |
network_interiors = {
|
| 2958 |
offset
|
| 2959 |
for span in group_spans
|
|
|
|
| 2973 |
}
|
| 2974 |
|
| 2975 |
ordered = sorted(offset for offset in cuts if group_start <= offset <= group_end)
|
| 2976 |
+
def solve(
|
| 2977 |
+
active_mandatory_cuts: set[int],
|
| 2978 |
+
) -> tuple[int, int, int, int, tuple[int, ...]] | None:
|
| 2979 |
+
best: dict[int, tuple[int, int, int, int, tuple[int, ...]]] = {
|
| 2980 |
+
group_end: (0, 0, 0, 0, ())
|
| 2981 |
+
}
|
| 2982 |
+
for start in reversed(ordered[:-1]):
|
| 2983 |
+
selected: tuple[int, int, int, int, tuple[int, ...]] | None = None
|
| 2984 |
+
for end in ordered:
|
| 2985 |
+
if end <= start:
|
| 2986 |
+
continue
|
| 2987 |
+
if any(start < cut < end for cut in active_mandatory_cuts):
|
| 2988 |
+
continue
|
| 2989 |
+
chunk = text[start:end]
|
| 2990 |
+
units = count_speech_units(chunk)
|
| 2991 |
+
network_conditioned = any(
|
| 2992 |
+
span.start < end and span.end > start for span in group_spans
|
| 2993 |
+
)
|
| 2994 |
+
maximum = (
|
| 2995 |
+
network_maximum if network_conditioned else ordinary_maximum
|
| 2996 |
+
)
|
| 2997 |
+
effective_minimum = (
|
| 2998 |
+
network_minimum if network_conditioned else minimum
|
| 2999 |
+
)
|
| 3000 |
+
if units > maximum:
|
| 3001 |
+
continue
|
| 3002 |
+
if units < effective_minimum and not (
|
| 3003 |
+
start == group_start and end == group_end
|
| 3004 |
+
):
|
| 3005 |
+
continue
|
| 3006 |
+
remainder = best.get(end)
|
| 3007 |
+
if remainder is None:
|
| 3008 |
+
continue
|
| 3009 |
+
candidate = (
|
| 3010 |
+
1 + remainder[0],
|
| 3011 |
+
int(end != group_end and end not in preferred_cuts)
|
| 3012 |
+
+ remainder[1],
|
| 3013 |
+
max(units, remainder[2]),
|
| 3014 |
+
abs(target - units) + remainder[3],
|
| 3015 |
+
(end,) + remainder[4],
|
| 3016 |
+
)
|
| 3017 |
+
if selected is None or candidate < selected:
|
| 3018 |
+
selected = candidate
|
| 3019 |
+
if selected is not None:
|
| 3020 |
+
best[start] = selected
|
| 3021 |
+
return best.get(group_start)
|
| 3022 |
+
|
| 3023 |
+
plan = solve(mandatory_cuts)
|
| 3024 |
+
if plan is None:
|
| 3025 |
+
# A tiny common identifier such as ``https://a.tw`` cannot place both
|
| 3026 |
+
# sides of the preferred scheme cut above the network minimum. Retry
|
| 3027 |
+
# only after proving which identifier-local side is intrinsically short;
|
| 3028 |
+
# all other mandatory cuts and every grammar component boundary remain.
|
| 3029 |
+
short_identifier_cuts = {
|
| 3030 |
+
cut
|
| 3031 |
+
for span in group_spans
|
| 3032 |
+
for cut in span.mandatory_cut_offsets
|
| 3033 |
+
if count_speech_units(text[span.start:cut]) < network_minimum
|
| 3034 |
+
or count_speech_units(text[cut:span.end]) < network_minimum
|
| 3035 |
+
}
|
| 3036 |
+
if short_identifier_cuts:
|
| 3037 |
+
plan = solve(mandatory_cuts - short_identifier_cuts)
|
| 3038 |
if plan is None:
|
| 3039 |
raise ValueError("network-bearing sentence cannot satisfy generation limits")
|
| 3040 |
output: list[tuple[int, int]] = []
|
|
|
|
| 3140 |
normalized_text: str,
|
| 3141 |
*,
|
| 3142 |
min_units: int = 12,
|
| 3143 |
+
network_min_units: int = 8,
|
| 3144 |
target_units: int = 32,
|
| 3145 |
network_max_units: int = 36,
|
| 3146 |
ordinary_max_units: int = 80,
|
|
|
|
| 3156 |
if network_identifier_has_ambiguous_iri(raw_text):
|
| 3157 |
raise ValueError("network generation does not support non-ASCII IRI")
|
| 3158 |
minimum = max(1, int(min_units))
|
| 3159 |
+
network_minimum = max(1, int(network_min_units))
|
| 3160 |
target = max(minimum, int(target_units))
|
| 3161 |
network_maximum = max(minimum, int(network_max_units))
|
| 3162 |
ordinary_maximum = max(minimum, int(ordinary_max_units))
|
| 3163 |
+
if not (
|
| 3164 |
+
network_minimum
|
| 3165 |
+
<= minimum
|
| 3166 |
+
<= target
|
| 3167 |
+
<= network_maximum
|
| 3168 |
+
<= ordinary_maximum
|
| 3169 |
+
):
|
| 3170 |
raise ValueError("generation chunk limits are inconsistent")
|
| 3171 |
|
| 3172 |
normalized = normalize_tts_text(normalized_text)
|
|
|
|
| 3217 |
original_boundaries=(),
|
| 3218 |
spans=spans,
|
| 3219 |
minimum=minimum,
|
| 3220 |
+
network_minimum=network_minimum,
|
| 3221 |
target=target,
|
| 3222 |
network_maximum=network_maximum,
|
| 3223 |
ordinary_maximum=ordinary_maximum,
|
|
|
|
| 3271 |
)
|
| 3272 |
if left < end and right > start
|
| 3273 |
)
|
| 3274 |
+
fragment_proofs = tuple(
|
| 3275 |
+
NetworkFragmentProof(
|
| 3276 |
+
span_index=span_index,
|
| 3277 |
+
chunk_start=max(start, spans[span_index].start) - start,
|
| 3278 |
+
chunk_end=min(end, spans[span_index].end) - start,
|
| 3279 |
+
parent_start=max(start, spans[span_index].start)
|
| 3280 |
+
- spans[span_index].start,
|
| 3281 |
+
parent_end=min(end, spans[span_index].end)
|
| 3282 |
+
- spans[span_index].start,
|
| 3283 |
+
full_spoken_proof=spans[span_index].spoken_proof,
|
| 3284 |
+
)
|
| 3285 |
+
for span_index in overlapping
|
| 3286 |
+
)
|
| 3287 |
specs.append(
|
| 3288 |
GenerationChunkSpec(
|
| 3289 |
text=chunk,
|
|
|
|
| 3294 |
network_full_spoken_proofs=tuple(
|
| 3295 |
spans[span_index].spoken_proof for span_index in overlapping
|
| 3296 |
),
|
| 3297 |
+
network_fragment_proofs=fragment_proofs,
|
| 3298 |
boundary_after=(
|
| 3299 |
"network_internal"
|
| 3300 |
if end in internal_boundaries
|
|
|
|
| 3310 |
for spec in specs
|
| 3311 |
):
|
| 3312 |
raise ValueError("generation chunk exceeds its provenance-specific limit")
|
| 3313 |
+
for spec in specs:
|
| 3314 |
+
proofs = spec.network_fragment_proofs
|
| 3315 |
+
if spec.network_conditioned:
|
| 3316 |
+
if not (
|
| 3317 |
+
len(proofs)
|
| 3318 |
+
== len(spec.network_span_indices)
|
| 3319 |
+
== len(spec.network_full_spoken_proofs)
|
| 3320 |
+
):
|
| 3321 |
+
raise ValueError("network fragment provenance is incomplete")
|
| 3322 |
+
elif proofs or spec.network_full_spoken_proofs:
|
| 3323 |
+
raise ValueError("ordinary chunk carries network fragment provenance")
|
| 3324 |
+
previous_chunk_end = 0
|
| 3325 |
+
for span_index, full_proof, proof in zip(
|
| 3326 |
+
spec.network_span_indices,
|
| 3327 |
+
spec.network_full_spoken_proofs,
|
| 3328 |
+
proofs,
|
| 3329 |
+
strict=True,
|
| 3330 |
+
):
|
| 3331 |
+
if (
|
| 3332 |
+
proof.span_index != span_index
|
| 3333 |
+
or proof.full_spoken_proof != full_proof
|
| 3334 |
+
or not 0 <= proof.chunk_start < proof.chunk_end <= len(spec.text)
|
| 3335 |
+
or not 0 <= proof.parent_start < proof.parent_end <= len(full_proof)
|
| 3336 |
+
or proof.chunk_start < previous_chunk_end
|
| 3337 |
+
or spec.text[proof.chunk_start : proof.chunk_end]
|
| 3338 |
+
!= full_proof[proof.parent_start : proof.parent_end]
|
| 3339 |
+
):
|
| 3340 |
+
raise ValueError("network fragment provenance is inconsistent")
|
| 3341 |
+
previous_chunk_end = proof.chunk_end
|
| 3342 |
+
|
| 3343 |
+
# The local fragments are not merely hints: taken across the plan they must
|
| 3344 |
+
# form an exact, non-overlapping partition of every parent identifier.
|
| 3345 |
+
for span_index, span in enumerate(spans):
|
| 3346 |
+
partition = sorted(
|
| 3347 |
+
(
|
| 3348 |
+
proof.parent_start,
|
| 3349 |
+
proof.parent_end,
|
| 3350 |
+
proof.full_spoken_proof,
|
| 3351 |
+
)
|
| 3352 |
+
for spec in specs
|
| 3353 |
+
for proof in spec.network_fragment_proofs
|
| 3354 |
+
if proof.span_index == span_index
|
| 3355 |
+
)
|
| 3356 |
+
if (
|
| 3357 |
+
not partition
|
| 3358 |
+
or partition[0][0] != 0
|
| 3359 |
+
or partition[-1][1] != len(span.spoken_proof)
|
| 3360 |
+
or any(
|
| 3361 |
+
left_end != right_start
|
| 3362 |
+
for (_, left_end, _), (right_start, _, _) in zip(
|
| 3363 |
+
partition,
|
| 3364 |
+
partition[1:],
|
| 3365 |
+
)
|
| 3366 |
+
)
|
| 3367 |
+
or any(full_proof != span.spoken_proof for _, _, full_proof in partition)
|
| 3368 |
+
):
|
| 3369 |
+
raise ValueError("network fragment provenance does not partition its parent")
|
| 3370 |
return tuple(specs)
|
| 3371 |
|
| 3372 |
|
|
|
|
| 3911 |
return operations
|
| 3912 |
|
| 3913 |
|
| 3914 |
+
def _protected_ranges_are_exact_in_all_optimal_alignments(
|
| 3915 |
+
source: str,
|
| 3916 |
+
hypothesis: str,
|
| 3917 |
+
protected_ranges: Sequence[tuple[int, int]],
|
| 3918 |
+
) -> bool:
|
| 3919 |
+
"""Require protected source ranges to survive every optimal edit path.
|
| 3920 |
+
|
| 3921 |
+
A deterministic Levenshtein backtrace is insufficient when ordinary text
|
| 3922 |
+
duplicates a protected fragment: one optimal path can match the protected
|
| 3923 |
+
occurrence while another deletes it and matches the ordinary occurrence.
|
| 3924 |
+
Prefix/suffix distances let us inspect every edge that belongs to at least
|
| 3925 |
+
one globally optimal path. Protected characters may only traverse equal
|
| 3926 |
+
diagonal edges, and insertions may not be anchored inside the closed range
|
| 3927 |
+
(including immediately before its first or after its last character).
|
| 3928 |
+
"""
|
| 3929 |
+
|
| 3930 |
+
ranges = tuple(protected_ranges)
|
| 3931 |
+
if not ranges:
|
| 3932 |
+
return True
|
| 3933 |
+
if any(
|
| 3934 |
+
isinstance(start, bool)
|
| 3935 |
+
or isinstance(end, bool)
|
| 3936 |
+
or not isinstance(start, int)
|
| 3937 |
+
or not isinstance(end, int)
|
| 3938 |
+
or not 0 <= start < end <= len(source)
|
| 3939 |
+
for start, end in ranges
|
| 3940 |
+
):
|
| 3941 |
+
raise ValueError("protected alignment range is invalid")
|
| 3942 |
+
|
| 3943 |
+
source_protected = [False] * len(source)
|
| 3944 |
+
insertion_protected = [False] * (len(source) + 1)
|
| 3945 |
+
for start, end in ranges:
|
| 3946 |
+
for source_index in range(start, end):
|
| 3947 |
+
source_protected[source_index] = True
|
| 3948 |
+
for source_offset in range(start, end + 1):
|
| 3949 |
+
insertion_protected[source_offset] = True
|
| 3950 |
+
|
| 3951 |
+
source_length = len(source)
|
| 3952 |
+
hypothesis_length = len(hypothesis)
|
| 3953 |
+
forward = [
|
| 3954 |
+
[0] * (hypothesis_length + 1) for _ in range(source_length + 1)
|
| 3955 |
+
]
|
| 3956 |
+
for source_index in range(source_length + 1):
|
| 3957 |
+
forward[source_index][0] = source_index
|
| 3958 |
+
for hypothesis_index in range(hypothesis_length + 1):
|
| 3959 |
+
forward[0][hypothesis_index] = hypothesis_index
|
| 3960 |
+
for source_index in range(1, source_length + 1):
|
| 3961 |
+
for hypothesis_index in range(1, hypothesis_length + 1):
|
| 3962 |
+
forward[source_index][hypothesis_index] = min(
|
| 3963 |
+
forward[source_index - 1][hypothesis_index] + 1,
|
| 3964 |
+
forward[source_index][hypothesis_index - 1] + 1,
|
| 3965 |
+
forward[source_index - 1][hypothesis_index - 1]
|
| 3966 |
+
+ (
|
| 3967 |
+
source[source_index - 1]
|
| 3968 |
+
!= hypothesis[hypothesis_index - 1]
|
| 3969 |
+
),
|
| 3970 |
+
)
|
| 3971 |
+
|
| 3972 |
+
backward = [
|
| 3973 |
+
[0] * (hypothesis_length + 1) for _ in range(source_length + 1)
|
| 3974 |
+
]
|
| 3975 |
+
for source_index in range(source_length + 1):
|
| 3976 |
+
backward[source_index][hypothesis_length] = source_length - source_index
|
| 3977 |
+
for hypothesis_index in range(hypothesis_length + 1):
|
| 3978 |
+
backward[source_length][hypothesis_index] = (
|
| 3979 |
+
hypothesis_length - hypothesis_index
|
| 3980 |
+
)
|
| 3981 |
+
for source_index in range(source_length - 1, -1, -1):
|
| 3982 |
+
for hypothesis_index in range(hypothesis_length - 1, -1, -1):
|
| 3983 |
+
backward[source_index][hypothesis_index] = min(
|
| 3984 |
+
backward[source_index + 1][hypothesis_index] + 1,
|
| 3985 |
+
backward[source_index][hypothesis_index + 1] + 1,
|
| 3986 |
+
backward[source_index + 1][hypothesis_index + 1]
|
| 3987 |
+
+ (source[source_index] != hypothesis[hypothesis_index]),
|
| 3988 |
+
)
|
| 3989 |
+
|
| 3990 |
+
optimal_distance = forward[source_length][hypothesis_length]
|
| 3991 |
+
for source_index in range(source_length + 1):
|
| 3992 |
+
for hypothesis_index in range(hypothesis_length + 1):
|
| 3993 |
+
prefix_distance = forward[source_index][hypothesis_index]
|
| 3994 |
+
if (
|
| 3995 |
+
source_index < source_length
|
| 3996 |
+
and source_protected[source_index]
|
| 3997 |
+
and prefix_distance
|
| 3998 |
+
+ 1
|
| 3999 |
+
+ backward[source_index + 1][hypothesis_index]
|
| 4000 |
+
== optimal_distance
|
| 4001 |
+
):
|
| 4002 |
+
return False
|
| 4003 |
+
if (
|
| 4004 |
+
source_index < source_length
|
| 4005 |
+
and hypothesis_index < hypothesis_length
|
| 4006 |
+
and source_protected[source_index]
|
| 4007 |
+
and source[source_index] != hypothesis[hypothesis_index]
|
| 4008 |
+
and prefix_distance
|
| 4009 |
+
+ 1
|
| 4010 |
+
+ backward[source_index + 1][hypothesis_index + 1]
|
| 4011 |
+
== optimal_distance
|
| 4012 |
+
):
|
| 4013 |
+
return False
|
| 4014 |
+
if (
|
| 4015 |
+
hypothesis_index < hypothesis_length
|
| 4016 |
+
and insertion_protected[source_index]
|
| 4017 |
+
and prefix_distance
|
| 4018 |
+
+ 1
|
| 4019 |
+
+ backward[source_index][hypothesis_index + 1]
|
| 4020 |
+
== optimal_distance
|
| 4021 |
+
):
|
| 4022 |
+
return False
|
| 4023 |
+
return True
|
| 4024 |
+
|
| 4025 |
+
|
| 4026 |
+
@dataclass(frozen=True)
|
| 4027 |
+
class NetworkFragmentAsrEvidence:
|
| 4028 |
+
"""Result of exact, range-bound local network-fragment verification."""
|
| 4029 |
+
|
| 4030 |
+
transcript_text: str
|
| 4031 |
+
protected_range_count: int
|
| 4032 |
+
passed: bool
|
| 4033 |
+
|
| 4034 |
+
|
| 4035 |
+
_NETWORK_FRAGMENT_IDENTIFIER_CHARACTERS = (
|
| 4036 |
+
"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789"
|
| 4037 |
+
+ "".join(_ZH_NETWORK_SYMBOL_READINGS)
|
| 4038 |
+
)
|
| 4039 |
+
_NETWORK_FRAGMENT_READING_TO_ASCII = {
|
| 4040 |
+
reading: symbol for symbol, reading in _ZH_NETWORK_SYMBOL_READINGS.items()
|
| 4041 |
+
}
|
| 4042 |
+
|
| 4043 |
+
|
| 4044 |
+
def _network_fragment_token_ascii(token: str) -> str | None:
|
| 4045 |
+
"""Invert one typed frontend network atom without prose heuristics."""
|
| 4046 |
+
|
| 4047 |
+
canonical = _SIMPLIFIED_NETWORK_SYMBOL_READINGS.get(token, token)
|
| 4048 |
+
letter = _network_letter_token(canonical)
|
| 4049 |
+
if letter is not None:
|
| 4050 |
+
return letter
|
| 4051 |
+
if canonical.isascii() and canonical.isalnum():
|
| 4052 |
+
# A hybrid frontend may intentionally retain a planner-bound opaque
|
| 4053 |
+
# label (``tour``/``islandmuseum``) while spelling only its separators
|
| 4054 |
+
# and short suffix. It is still an exact typed atom, never prose: the
|
| 4055 |
+
# range object below binds its precise occurrence to the parent URL.
|
| 4056 |
+
return canonical.casefold()
|
| 4057 |
+
if canonical and all(character in _ZH_DIGITS for character in canonical):
|
| 4058 |
+
return "".join(str(_ZH_DIGITS.index(character)) for character in canonical)
|
| 4059 |
+
return _NETWORK_FRAGMENT_READING_TO_ASCII.get(canonical)
|
| 4060 |
+
|
| 4061 |
+
|
| 4062 |
+
def _network_fragment_ascii(spoken_fragment: str) -> str | None:
|
| 4063 |
+
"""Return the exact ASCII atom sequence for one planner-bound fragment."""
|
| 4064 |
+
|
| 4065 |
+
output: list[str] = []
|
| 4066 |
+
for token in normalize_tts_text(spoken_fragment).split():
|
| 4067 |
+
value = _network_fragment_token_ascii(token)
|
| 4068 |
+
if value is None:
|
| 4069 |
+
return None
|
| 4070 |
+
output.append(value)
|
| 4071 |
+
rendered = "".join(output)
|
| 4072 |
+
return rendered if rendered and any(character.isalnum() for character in rendered) else None
|
| 4073 |
+
|
| 4074 |
+
|
| 4075 |
+
def _network_fragment_alignment_text(value: str) -> str:
|
| 4076 |
+
"""Composable orthographic key used only by the exact fragment gate."""
|
| 4077 |
+
|
| 4078 |
+
normalized = _T2S_CONVERTER.convert(
|
| 4079 |
+
normalize_tts_eval_text(unicodedata.normalize("NFC", str(value or ""))).casefold()
|
| 4080 |
+
)
|
| 4081 |
+
return "".join(character for character in normalized if character.isalnum())
|
| 4082 |
+
|
| 4083 |
+
|
| 4084 |
+
def canonicalize_asr_network_fragments(
|
| 4085 |
+
text: str,
|
| 4086 |
+
target_text: str,
|
| 4087 |
+
fragment_proofs: Sequence[NetworkFragmentProof],
|
| 4088 |
+
) -> NetworkFragmentAsrEvidence:
|
| 4089 |
+
"""Canonicalize and exact-gate local ASCII URL/email fragments.
|
| 4090 |
+
|
| 4091 |
+
This helper is deliberately separate from :func:`compare_asr_text`. It is
|
| 4092 |
+
valid only for generation chunks carrying the range objects emitted by
|
| 4093 |
+
:func:`plan_generation_chunks`; whole/joined/final verification must not
|
| 4094 |
+
call it. Each literal replacement must equal the *entire* chunk-local
|
| 4095 |
+
parent slice. Substrings of a full URL cannot borrow its proof.
|
| 4096 |
+
"""
|
| 4097 |
+
|
| 4098 |
+
raw_text = unicodedata.normalize("NFC", str(text or ""))
|
| 4099 |
+
target = unicodedata.normalize("NFC", str(target_text or ""))
|
| 4100 |
+
if isinstance(fragment_proofs, (str, bytes)):
|
| 4101 |
+
raise ValueError("network fragment proofs must be range objects")
|
| 4102 |
+
try:
|
| 4103 |
+
proofs = tuple(fragment_proofs)
|
| 4104 |
+
except TypeError as error:
|
| 4105 |
+
raise ValueError("network fragment proofs must be iterable") from error
|
| 4106 |
+
if not proofs:
|
| 4107 |
+
return NetworkFragmentAsrEvidence(raw_text, 0, True)
|
| 4108 |
+
if not target:
|
| 4109 |
+
raise ValueError("network fragment target must be non-empty")
|
| 4110 |
+
|
| 4111 |
+
previous_chunk_end = 0
|
| 4112 |
+
typed_fragments: dict[str, str] = {}
|
| 4113 |
+
for proof in proofs:
|
| 4114 |
+
if not isinstance(proof, NetworkFragmentProof):
|
| 4115 |
+
raise ValueError("network fragment proof has an invalid type")
|
| 4116 |
+
integer_fields = (
|
| 4117 |
+
proof.span_index,
|
| 4118 |
+
proof.chunk_start,
|
| 4119 |
+
proof.chunk_end,
|
| 4120 |
+
proof.parent_start,
|
| 4121 |
+
proof.parent_end,
|
| 4122 |
+
)
|
| 4123 |
+
if any(isinstance(value, bool) or not isinstance(value, int) for value in integer_fields):
|
| 4124 |
+
raise ValueError("network fragment proof offsets must be integers")
|
| 4125 |
+
full_proof = unicodedata.normalize("NFC", str(proof.full_spoken_proof or ""))
|
| 4126 |
+
if (
|
| 4127 |
+
proof.span_index < 0
|
| 4128 |
+
or not 0 <= proof.chunk_start < proof.chunk_end <= len(target)
|
| 4129 |
+
or not 0 <= proof.parent_start < proof.parent_end <= len(full_proof)
|
| 4130 |
+
or proof.chunk_start < previous_chunk_end
|
| 4131 |
+
or target[proof.chunk_start : proof.chunk_end]
|
| 4132 |
+
!= full_proof[proof.parent_start : proof.parent_end]
|
| 4133 |
+
or full_proof not in network_protected_spoken_spans(full_proof)
|
| 4134 |
+
):
|
| 4135 |
+
raise ValueError("network fragment proof does not bind its target range")
|
| 4136 |
+
previous_chunk_end = proof.chunk_end
|
| 4137 |
+
fragment = target[proof.chunk_start : proof.chunk_end]
|
| 4138 |
+
ascii_fragment = _network_fragment_ascii(fragment)
|
| 4139 |
+
if ascii_fragment is None:
|
| 4140 |
+
# Unmapped atoms remain eligible only in their canonical spoken
|
| 4141 |
+
# form. A literal approximation is never inferred.
|
| 4142 |
+
continue
|
| 4143 |
+
old_fragment = typed_fragments.get(ascii_fragment.casefold())
|
| 4144 |
+
if old_fragment is not None and old_fragment != fragment:
|
| 4145 |
+
raise ValueError("network fragment ASCII proof is ambiguous")
|
| 4146 |
+
typed_fragments[ascii_fragment.casefold()] = fragment
|
| 4147 |
+
|
| 4148 |
+
# Protect longer exact atoms first, normalize surrounding transcript
|
| 4149 |
+
# punctuation, then restore their planner-proven spoken forms. A typed atom
|
| 4150 |
+
# is eligible only at a conservative prose boundary. In particular, an
|
| 4151 |
+
# unknown adjacent symbol (backslash, angle bracket, emoji, and so on) must
|
| 4152 |
+
# remain visible rather than letting a global replacement erase its
|
| 4153 |
+
# provenance before the exact range alignment.
|
| 4154 |
+
canonicalized = raw_text
|
| 4155 |
+
placeholder_prefix = "藍鵲片段保護佔位符"
|
| 4156 |
+
while placeholder_prefix in raw_text or placeholder_prefix in target:
|
| 4157 |
+
placeholder_prefix += "號"
|
| 4158 |
+
fragment_placeholders: list[tuple[str, str]] = []
|
| 4159 |
+
unsafe_typed_literal = False
|
| 4160 |
+
ascii_sentence_boundaries = frozenset(",.;:!?\"'()[]{}")
|
| 4161 |
+
|
| 4162 |
+
def literal_boundary_kind(value: str, offset: int) -> str:
|
| 4163 |
+
if offset < 0 or offset >= len(value):
|
| 4164 |
+
return "safe"
|
| 4165 |
+
character = value[offset]
|
| 4166 |
+
if character.isascii() and character.isalnum():
|
| 4167 |
+
return "identifier"
|
| 4168 |
+
if (
|
| 4169 |
+
character.isspace()
|
| 4170 |
+
or _is_cjk(character)
|
| 4171 |
+
or character in _NETWORK_NONASCII_BOUNDARY_PUNCTUATION
|
| 4172 |
+
or character in ascii_sentence_boundaries
|
| 4173 |
+
):
|
| 4174 |
+
return "safe"
|
| 4175 |
+
return "unsafe"
|
| 4176 |
+
|
| 4177 |
+
for ascii_fragment, spoken_fragment in sorted(
|
| 4178 |
+
typed_fragments.items(),
|
| 4179 |
+
key=lambda item: (-len(item[0]), item[0]),
|
| 4180 |
+
):
|
| 4181 |
+
character_pattern = r"\s*".join(
|
| 4182 |
+
re.escape(character) for character in ascii_fragment
|
| 4183 |
+
)
|
| 4184 |
+
pattern = re.compile(
|
| 4185 |
+
character_pattern,
|
| 4186 |
+
flags=re.IGNORECASE | re.ASCII,
|
| 4187 |
+
)
|
| 4188 |
+
|
| 4189 |
+
def protect_literal(match: re.Match[str]) -> str:
|
| 4190 |
+
nonlocal unsafe_typed_literal
|
| 4191 |
+
boundary_kinds = (
|
| 4192 |
+
literal_boundary_kind(match.string, match.start() - 1),
|
| 4193 |
+
literal_boundary_kind(match.string, match.end()),
|
| 4194 |
+
)
|
| 4195 |
+
if "identifier" in boundary_kinds:
|
| 4196 |
+
return match.group(0)
|
| 4197 |
+
if "unsafe" in boundary_kinds:
|
| 4198 |
+
unsafe_typed_literal = True
|
| 4199 |
+
return match.group(0)
|
| 4200 |
+
placeholder = (
|
| 4201 |
+
f"{placeholder_prefix}"
|
| 4202 |
+
f"{_zh_integer(len(fragment_placeholders) + 1)}結束"
|
| 4203 |
+
)
|
| 4204 |
+
fragment_placeholders.append((placeholder, spoken_fragment))
|
| 4205 |
+
return placeholder
|
| 4206 |
+
|
| 4207 |
+
canonicalized = pattern.sub(protect_literal, canonicalized)
|
| 4208 |
+
|
| 4209 |
+
canonicalized = normalize_tts_text(canonicalized)
|
| 4210 |
+
for placeholder, spoken_fragment in fragment_placeholders:
|
| 4211 |
+
canonicalized = canonicalized.replace(placeholder, spoken_fragment)
|
| 4212 |
+
|
| 4213 |
+
target_key = _network_fragment_alignment_text(target)
|
| 4214 |
+
transcript_key = _network_fragment_alignment_text(canonicalized)
|
| 4215 |
+
protected_ranges: list[tuple[int, int]] = []
|
| 4216 |
+
rebuilt_target_key: list[str] = []
|
| 4217 |
+
cursor = 0
|
| 4218 |
+
key_cursor = 0
|
| 4219 |
+
for proof in proofs:
|
| 4220 |
+
prefix_key = _network_fragment_alignment_text(
|
| 4221 |
+
target[cursor : proof.chunk_start]
|
| 4222 |
+
)
|
| 4223 |
+
fragment_key = _network_fragment_alignment_text(
|
| 4224 |
+
target[proof.chunk_start : proof.chunk_end]
|
| 4225 |
+
)
|
| 4226 |
+
rebuilt_target_key.extend((prefix_key, fragment_key))
|
| 4227 |
+
key_cursor += len(prefix_key)
|
| 4228 |
+
if not fragment_key:
|
| 4229 |
+
raise ValueError("network fragment proof has no comparable content")
|
| 4230 |
+
protected_ranges.append((key_cursor, key_cursor + len(fragment_key)))
|
| 4231 |
+
key_cursor += len(fragment_key)
|
| 4232 |
+
cursor = proof.chunk_end
|
| 4233 |
+
rebuilt_target_key.append(_network_fragment_alignment_text(target[cursor:]))
|
| 4234 |
+
if "".join(rebuilt_target_key) != target_key:
|
| 4235 |
+
raise ValueError("network fragment comparison ranges are not composable")
|
| 4236 |
+
|
| 4237 |
+
passed = (
|
| 4238 |
+
not unsafe_typed_literal
|
| 4239 |
+
and _protected_ranges_are_exact_in_all_optimal_alignments(
|
| 4240 |
+
target_key,
|
| 4241 |
+
transcript_key,
|
| 4242 |
+
protected_ranges,
|
| 4243 |
+
)
|
| 4244 |
+
)
|
| 4245 |
+
return NetworkFragmentAsrEvidence(
|
| 4246 |
+
transcript_text=canonicalized,
|
| 4247 |
+
protected_range_count=len(protected_ranges),
|
| 4248 |
+
passed=bool(passed),
|
| 4249 |
+
)
|
| 4250 |
+
|
| 4251 |
+
|
| 4252 |
@dataclass(frozen=True)
|
| 4253 |
class AsrComparison:
|
| 4254 |
"""Orthographic whole-text and acoustic endpoint evidence for one candidate."""
|
tests/test_production.py
CHANGED
|
@@ -6,10 +6,12 @@ import torch
|
|
| 6 |
from torch import nn
|
| 7 |
|
| 8 |
from production import (
|
|
|
|
| 9 |
StopHysteresisController,
|
| 10 |
active_pace_correction_speed,
|
| 11 |
candidate_local_score,
|
| 12 |
candidate_transition_score,
|
|
|
|
| 13 |
coalesce_text_chunks,
|
| 14 |
compare_asr_text,
|
| 15 |
contains_network_identifier,
|
|
@@ -40,8 +42,8 @@ from production import (
|
|
| 40 |
)
|
| 41 |
|
| 42 |
|
| 43 |
-
_SPOKEN_HTTPS = "
|
| 44 |
-
_SPOKEN_PATH = "
|
| 45 |
|
| 46 |
|
| 47 |
def test_eval_only_pronoun_homophones_do_not_change_model_input_text():
|
|
@@ -629,14 +631,14 @@ def test_non_ascii_iri_is_explicitly_unsupported_and_fails_closed(transcript):
|
|
| 629 |
(
|
| 630 |
(
|
| 631 |
"前方甲乙X,H T T P S 冒號 斜線 斜線 orange 點 "
|
| 632 |
-
"
|
| 633 |
"後方丙丁。",
|
| 634 |
1.0 / 6.0,
|
| 635 |
0.0,
|
| 636 |
),
|
| 637 |
(
|
| 638 |
"前方甲乙,H T T P S 冒號 斜線 斜線 orange 點 "
|
| 639 |
-
"
|
| 640 |
"X後方丙丁。",
|
| 641 |
0.0,
|
| 642 |
1.0 / 6.0,
|
|
@@ -875,20 +877,18 @@ def test_frontend_clause_split_matrix_for_structured_and_network_holdouts():
|
|
| 875 |
"H06": (
|
| 876 |
"若要更換導覽場次,請寄信到 tour.help@islandmuseum.tw。",
|
| 877 |
(
|
| 878 |
-
"若要更換導覽場次,請寄信到,"
|
| 879 |
-
"
|
| 880 |
-
"艾爾 欸 恩 迪 艾姆 優 艾斯 伊 優 艾姆 點 踢 達不溜。",
|
| 881 |
),
|
| 882 |
-
(
|
| 883 |
),
|
| 884 |
"H07": (
|
| 885 |
"潮汐預報可查詢 https://coastwatch.example.tw/tide。",
|
| 886 |
(
|
| 887 |
-
"潮汐預報可查詢,
|
| 888 |
-
"
|
| 889 |
-
"批 艾爾 伊 點 踢 達不溜 斜線 踢 愛 迪 伊。",
|
| 890 |
),
|
| 891 |
-
(
|
| 892 |
),
|
| 893 |
"H11": (
|
| 894 |
raw_h11,
|
|
@@ -897,18 +897,17 @@ def test_frontend_clause_split_matrix_for_structured_and_network_holdouts():
|
|
| 897 |
"七點十五分 集合,",
|
| 898 |
"先用定位器 A X 五二零 核對座標,再分組檢查木棧道、"
|
| 899 |
"里程牌與飲水站。",
|
| 900 |
-
"若氣象網站,
|
| 901 |
-
"
|
| 902 |
-
"欸 艾姆 批 艾爾 伊 點 踢 達不溜,",
|
| 903 |
"顯示降雨機率超過 百分之六十五,",
|
| 904 |
"領隊就取消高海拔路線,改走較短的林間環線。",
|
| 905 |
-
"途中若發現落石或樹枝阻斷通行,
|
| 906 |
-
"
|
| 907 |
-
"
|
| 908 |
"所有隊員回到登山口後,還要清點無線電與急救包,"
|
| 909 |
"確認沒有任何人落單,才結束當天的巡查。",
|
| 910 |
),
|
| 911 |
-
(30, 29,
|
| 912 |
),
|
| 913 |
}
|
| 914 |
|
|
@@ -944,16 +943,22 @@ def test_generation_only_network_component_planner_holdout_matrix():
|
|
| 944 |
matrix = {
|
| 945 |
"H06": (
|
| 946 |
"若要更換導覽場次,請寄信到 tour.help@islandmuseum.tw。",
|
| 947 |
-
(
|
|
|
|
| 948 |
),
|
| 949 |
"H07": (
|
| 950 |
"潮汐預報可查詢 https://coastwatch.example.tw/tide。",
|
| 951 |
-
(
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 952 |
),
|
| 953 |
-
"H11": (raw_h11, (30, 29, 18, 35, 33, 28, 34, 38)),
|
| 954 |
}
|
| 955 |
|
| 956 |
-
for raw, expected_units in matrix.values():
|
| 957 |
normalized = normalize_spoken_forms(raw)
|
| 958 |
specs = plan_generation_chunks(raw, normalized)
|
| 959 |
|
|
@@ -964,14 +969,384 @@ def test_generation_only_network_component_planner_holdout_matrix():
|
|
| 964 |
for spec, following in zip(specs, specs[1:])
|
| 965 |
)
|
| 966 |
assert all(
|
| 967 |
-
|
|
|
|
|
|
|
| 968 |
for spec in specs
|
| 969 |
)
|
| 970 |
-
assert
|
|
|
|
|
|
|
|
|
|
|
|
|
| 971 |
for spec in specs:
|
| 972 |
assert len(spec.network_span_indices) == len(
|
| 973 |
spec.network_full_spoken_proofs
|
| 974 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 975 |
|
| 976 |
|
| 977 |
def test_generation_network_planner_keeps_ordinary_semantic_chunks_unchanged():
|
|
@@ -1007,25 +1382,124 @@ def test_generation_network_planner_fails_closed_on_invalid_provenance():
|
|
| 1007 |
with pytest.raises(ValueError, match="non-ASCII IRI"):
|
| 1008 |
plan_generation_chunks(emoji_iri, normalize_spoken_forms(emoji_iri))
|
| 1009 |
|
| 1010 |
-
oversized = "請查詢 https://" + "a" *
|
| 1011 |
with pytest.raises(ValueError, match="indivisible network component"):
|
| 1012 |
plan_generation_chunks(oversized, normalize_spoken_forms(oversized))
|
| 1013 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1014 |
|
| 1015 |
-
|
| 1016 |
-
|
| 1017 |
-
|
| 1018 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1019 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1020 |
raw = f"先念 {spoken},再查 https://example.tw。"
|
| 1021 |
normalized = normalize_spoken_forms(raw)
|
| 1022 |
|
| 1023 |
specs = plan_generation_chunks(raw, normalized)
|
| 1024 |
|
| 1025 |
assert "".join(spec.text for spec in specs) == normalized
|
| 1026 |
-
assert [spec.network_conditioned for spec in specs] == [
|
| 1027 |
-
assert
|
| 1028 |
-
assert specs[
|
|
|
|
|
|
|
|
|
|
| 1029 |
assert all(
|
| 1030 |
normalized[spec.source_start : spec.source_end] == spec.text
|
| 1031 |
for spec in specs
|
|
@@ -1277,10 +1751,10 @@ def test_ascii_text_receives_a_wider_generation_window():
|
|
| 1277 |
assert select_generation_cps("AI TTS 測試", cjk_cps=5.2, ascii_cps=4.6) == 4.6
|
| 1278 |
|
| 1279 |
|
| 1280 |
-
def
|
| 1281 |
normalized = normalize_spoken_forms("https://example.tw/path")
|
| 1282 |
|
| 1283 |
-
assert
|
| 1284 |
assert network_protected_spoken_spans(normalized)
|
| 1285 |
assert select_generation_cps(normalized, cjk_cps=5.2, ascii_cps=4.6) == 4.6
|
| 1286 |
|
|
@@ -1426,9 +1900,8 @@ def test_spoken_form_normalizer_expands_clock_time_and_spells_urls():
|
|
| 1426 |
"無效 24點30分、15點60分。"
|
| 1427 |
)
|
| 1428 |
assert normalize_spoken_forms("IP 192.168.1.1,網址 https://example.test:30/path") == (
|
| 1429 |
-
"I P 192.168.1.1,網址,
|
| 1430 |
-
"
|
| 1431 |
-
"冒號 三零 斜線 批 欸 踢 艾取"
|
| 1432 |
)
|
| 1433 |
assert normalize_spoken_forms(
|
| 1434 |
"音量15點05分貝,區間15點30分鐘,比分15點05分。"
|
|
@@ -1864,14 +2337,12 @@ def test_spoken_form_normalizer_expands_network_identifiers_and_versions():
|
|
| 1864 |
assert normalize_spoken_forms(
|
| 1865 |
"網址 https://api.example.com/v1/items?q=RTX-5090&n=2。"
|
| 1866 |
) == (
|
| 1867 |
-
"網址,
|
| 1868 |
-
"
|
| 1869 |
-
"
|
| 1870 |
-
"橫線 五零九零 和 恩 等於 二。"
|
| 1871 |
)
|
| 1872 |
assert normalize_spoken_forms("信箱 USER.name+tts@example.com。") == (
|
| 1873 |
-
"信箱,
|
| 1874 |
-
"小老鼠 伊 艾克斯 欸 艾姆 批 艾爾 伊 點 西 歐 艾姆。"
|
| 1875 |
)
|
| 1876 |
assert normalize_spoken_forms(
|
| 1877 |
"版本 v1.2.3,候選 version 10.4.0-beta.1+build.5。"
|
|
@@ -1916,7 +2387,7 @@ def test_expanded_url_rejects_raw_symbol_inserted_inside_path(symbol):
|
|
| 1916 |
target = normalize_spoken_forms("https://example.tw/path")
|
| 1917 |
transcript = target.replace(
|
| 1918 |
_SPOKEN_PATH,
|
| 1919 |
-
f"
|
| 1920 |
)
|
| 1921 |
comparison = compare_asr_text(
|
| 1922 |
target,
|
|
@@ -1935,8 +2406,8 @@ def test_expanded_url_rejects_raw_symbol_inserted_inside_path(symbol):
|
|
| 1935 |
def test_expanded_email_rejects_raw_symbol_inserted_inside_local_part(symbol):
|
| 1936 |
target = normalize_spoken_forms("museum@example.tw")
|
| 1937 |
transcript = target.replace(
|
| 1938 |
-
"
|
| 1939 |
-
f"
|
| 1940 |
)
|
| 1941 |
comparison = compare_asr_text(
|
| 1942 |
target,
|
|
@@ -2093,7 +2564,7 @@ def test_expanded_www_url_keeps_an_exact_protected_span():
|
|
| 2093 |
)
|
| 2094 |
comparison = compare_asr_text(
|
| 2095 |
target,
|
| 2096 |
-
target.replace(_SPOKEN_PATH, "
|
| 2097 |
max_cer=0.20,
|
| 2098 |
max_prefix_cer=1.0,
|
| 2099 |
max_suffix_cer=1.0,
|
|
|
|
| 6 |
from torch import nn
|
| 7 |
|
| 8 |
from production import (
|
| 9 |
+
NetworkFragmentProof,
|
| 10 |
StopHysteresisController,
|
| 11 |
active_pace_correction_speed,
|
| 12 |
candidate_local_score,
|
| 13 |
candidate_transition_score,
|
| 14 |
+
canonicalize_asr_network_fragments,
|
| 15 |
coalesce_text_chunks,
|
| 16 |
compare_asr_text,
|
| 17 |
contains_network_identifier,
|
|
|
|
| 42 |
)
|
| 43 |
|
| 44 |
|
| 45 |
+
_SPOKEN_HTTPS = "H T T P S"
|
| 46 |
+
_SPOKEN_PATH = "path"
|
| 47 |
|
| 48 |
|
| 49 |
def test_eval_only_pronoun_homophones_do_not_change_model_input_text():
|
|
|
|
| 631 |
(
|
| 632 |
(
|
| 633 |
"前方甲乙X,H T T P S 冒號 斜線 斜線 orange 點 "
|
| 634 |
+
"example 點 T W 斜線 road,"
|
| 635 |
"後方丙丁。",
|
| 636 |
1.0 / 6.0,
|
| 637 |
0.0,
|
| 638 |
),
|
| 639 |
(
|
| 640 |
"前方甲乙,H T T P S 冒號 斜線 斜線 orange 點 "
|
| 641 |
+
"example 點 T W 斜線 road,"
|
| 642 |
"X後方丙丁。",
|
| 643 |
0.0,
|
| 644 |
1.0 / 6.0,
|
|
|
|
| 877 |
"H06": (
|
| 878 |
"若要更換導覽場次,請寄信到 tour.help@islandmuseum.tw。",
|
| 879 |
(
|
| 880 |
+
"若要更換導覽場次,請寄信到,tour 點 help 小老鼠 "
|
| 881 |
+
"islandmuseum 點 T W。",
|
|
|
|
| 882 |
),
|
| 883 |
+
(24,),
|
| 884 |
),
|
| 885 |
"H07": (
|
| 886 |
"潮汐預報可查詢 https://coastwatch.example.tw/tide。",
|
| 887 |
(
|
| 888 |
+
"潮汐預報可查詢,H T T P S 冒號 斜線 斜線 coastwatch 點 "
|
| 889 |
+
"example 點 T W 斜線 tide。",
|
|
|
|
| 890 |
),
|
| 891 |
+
(30,),
|
| 892 |
),
|
| 893 |
"H11": (
|
| 894 |
raw_h11,
|
|
|
|
| 897 |
"七點十五分 集合,",
|
| 898 |
"先用定位器 A X 五二零 核對座標,再分組檢查木棧道、"
|
| 899 |
"里程牌與飲水站。",
|
| 900 |
+
"若氣象網站,H T T P S 冒號 斜線 斜線 trailweather 點 "
|
| 901 |
+
"example 點 T W,",
|
|
|
|
| 902 |
"顯示降雨機率超過 百分之六十五,",
|
| 903 |
"領隊就取消高海拔路線,改走較短的林間環線。",
|
| 904 |
+
"途中若發現落石或樹枝阻斷通行,",
|
| 905 |
+
"請拍照並寄到,patrol 小老鼠 forestmail 點 T W,"
|
| 906 |
+
"不要自行搬動大型障礙物。",
|
| 907 |
"所有隊員回到登山口後,還要清點無線電與急救包,"
|
| 908 |
"確認沒有任何人落單,才結束當天的巡查。",
|
| 909 |
),
|
| 910 |
+
(30, 29, 25, 14, 19, 14, 28, 38),
|
| 911 |
),
|
| 912 |
}
|
| 913 |
|
|
|
|
| 943 |
matrix = {
|
| 944 |
"H06": (
|
| 945 |
"若要更換導覽場次,請寄信到 tour.help@islandmuseum.tw。",
|
| 946 |
+
(15, 9),
|
| 947 |
+
(0,),
|
| 948 |
),
|
| 949 |
"H07": (
|
| 950 |
"潮汐預報可查詢 https://coastwatch.example.tw/tide。",
|
| 951 |
+
(18, 12),
|
| 952 |
+
(0,),
|
| 953 |
+
),
|
| 954 |
+
"H11": (
|
| 955 |
+
raw_h11,
|
| 956 |
+
(30, 29, 16, 23, 19, 22, 20, 38),
|
| 957 |
+
(2, 5),
|
| 958 |
),
|
|
|
|
| 959 |
}
|
| 960 |
|
| 961 |
+
for raw, expected_units, mandatory_after in matrix.values():
|
| 962 |
normalized = normalize_spoken_forms(raw)
|
| 963 |
specs = plan_generation_chunks(raw, normalized)
|
| 964 |
|
|
|
|
| 969 |
for spec, following in zip(specs, specs[1:])
|
| 970 |
)
|
| 971 |
assert all(
|
| 972 |
+
(8 if spec.network_conditioned else 12)
|
| 973 |
+
<= count_speech_units(spec.text)
|
| 974 |
+
<= (36 if spec.network_conditioned else 80)
|
| 975 |
for spec in specs
|
| 976 |
)
|
| 977 |
+
assert tuple(
|
| 978 |
+
index
|
| 979 |
+
for index, spec in enumerate(specs)
|
| 980 |
+
if spec.boundary_after == "network_internal"
|
| 981 |
+
) == mandatory_after
|
| 982 |
for spec in specs:
|
| 983 |
assert len(spec.network_span_indices) == len(
|
| 984 |
spec.network_full_spoken_proofs
|
| 985 |
)
|
| 986 |
+
assert len(spec.network_span_indices) == len(
|
| 987 |
+
spec.network_fragment_proofs
|
| 988 |
+
)
|
| 989 |
+
for proof in spec.network_fragment_proofs:
|
| 990 |
+
assert (
|
| 991 |
+
spec.text[proof.chunk_start : proof.chunk_end]
|
| 992 |
+
== proof.full_spoken_proof[
|
| 993 |
+
proof.parent_start : proof.parent_end
|
| 994 |
+
]
|
| 995 |
+
)
|
| 996 |
+
|
| 997 |
+
for span_index in {
|
| 998 |
+
index for spec in specs for index in spec.network_span_indices
|
| 999 |
+
}:
|
| 1000 |
+
ranges = sorted(
|
| 1001 |
+
(proof.parent_start, proof.parent_end, proof.full_spoken_proof)
|
| 1002 |
+
for spec in specs
|
| 1003 |
+
for proof in spec.network_fragment_proofs
|
| 1004 |
+
if proof.span_index == span_index
|
| 1005 |
+
)
|
| 1006 |
+
assert ranges[0][0] == 0
|
| 1007 |
+
assert ranges[-1][1] == len(ranges[0][2])
|
| 1008 |
+
assert all(
|
| 1009 |
+
left_end == right_start
|
| 1010 |
+
for (_, left_end, _), (right_start, _, _) in zip(
|
| 1011 |
+
ranges,
|
| 1012 |
+
ranges[1:],
|
| 1013 |
+
)
|
| 1014 |
+
)
|
| 1015 |
+
|
| 1016 |
+
|
| 1017 |
+
def _tour_help_fragment_case(prefix: str = "請核對,", suffix: str = ",完成。"):
|
| 1018 |
+
fragment = "tour 點 help"
|
| 1019 |
+
full_proof = normalize_spoken_forms("tour.help@islandmuseum.tw")
|
| 1020 |
+
parent_start = full_proof.index(fragment)
|
| 1021 |
+
target = f"{prefix}{fragment}{suffix}"
|
| 1022 |
+
chunk_start = len(prefix)
|
| 1023 |
+
return target, NetworkFragmentProof(
|
| 1024 |
+
span_index=0,
|
| 1025 |
+
chunk_start=chunk_start,
|
| 1026 |
+
chunk_end=chunk_start + len(fragment),
|
| 1027 |
+
parent_start=parent_start,
|
| 1028 |
+
parent_end=parent_start + len(fragment),
|
| 1029 |
+
full_spoken_proof=full_proof,
|
| 1030 |
+
)
|
| 1031 |
+
|
| 1032 |
+
|
| 1033 |
+
@pytest.mark.parametrize("literal", ("tour.help", "Tour.Help", "t o u r . h e l p"))
|
| 1034 |
+
def test_local_network_fragment_canonicalizes_only_exact_typed_ascii(literal):
|
| 1035 |
+
target, proof = _tour_help_fragment_case()
|
| 1036 |
+
transcript = f"請核對,{literal},完成。"
|
| 1037 |
+
|
| 1038 |
+
evidence = canonicalize_asr_network_fragments(transcript, target, (proof,))
|
| 1039 |
+
|
| 1040 |
+
assert evidence.passed is True
|
| 1041 |
+
assert evidence.protected_range_count == 1
|
| 1042 |
+
assert "tour 點 help" in evidence.transcript_text
|
| 1043 |
+
assert compare_asr_text(target, evidence.transcript_text, max_cer=0.20).passed
|
| 1044 |
+
|
| 1045 |
+
|
| 1046 |
+
def test_local_network_fragment_normalizes_ascii_sentence_boundaries_first():
|
| 1047 |
+
target, proof = _tour_help_fragment_case()
|
| 1048 |
+
|
| 1049 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1050 |
+
"請核對,tour.help,完成.",
|
| 1051 |
+
target,
|
| 1052 |
+
(proof,),
|
| 1053 |
+
)
|
| 1054 |
+
|
| 1055 |
+
assert evidence.transcript_text == target
|
| 1056 |
+
assert evidence.passed is True
|
| 1057 |
+
|
| 1058 |
+
|
| 1059 |
+
@pytest.mark.parametrize(
|
| 1060 |
+
"literal",
|
| 1061 |
+
(
|
| 1062 |
+
"tour.herp", # wrong content
|
| 1063 |
+
"tourhelp", # missing separator
|
| 1064 |
+
"tour.hel", # missing suffix letter
|
| 1065 |
+
"help.tour", # reordered labels
|
| 1066 |
+
"xtour.help", # extra prefix content
|
| 1067 |
+
"tour.help/path", # extra suffix content
|
| 1068 |
+
"tour.help.help", # repeated content
|
| 1069 |
+
),
|
| 1070 |
+
)
|
| 1071 |
+
def test_local_network_fragment_rejects_wrong_missing_reordered_or_extra_content(
|
| 1072 |
+
literal,
|
| 1073 |
+
):
|
| 1074 |
+
target, proof = _tour_help_fragment_case()
|
| 1075 |
+
transcript = f"請核對,{literal},完成。"
|
| 1076 |
+
|
| 1077 |
+
evidence = canonicalize_asr_network_fragments(transcript, target, (proof,))
|
| 1078 |
+
|
| 1079 |
+
assert evidence.passed is False
|
| 1080 |
+
assert evidence.protected_range_count == 1
|
| 1081 |
+
|
| 1082 |
+
|
| 1083 |
+
@pytest.mark.parametrize("symbol", ("\\", "<", ">", "§", "🙂"))
|
| 1084 |
+
@pytest.mark.parametrize("side", ("prefix", "suffix"))
|
| 1085 |
+
def test_local_network_fragment_rejects_adjacent_unknown_symbols(symbol, side):
|
| 1086 |
+
target, proof = _tour_help_fragment_case()
|
| 1087 |
+
literal = f"{symbol}tour.help" if side == "prefix" else f"tour.help{symbol}"
|
| 1088 |
+
|
| 1089 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1090 |
+
f"請核對,{literal},完成。",
|
| 1091 |
+
target,
|
| 1092 |
+
(proof,),
|
| 1093 |
+
)
|
| 1094 |
+
|
| 1095 |
+
assert evidence.protected_range_count == 1
|
| 1096 |
+
assert evidence.passed is False
|
| 1097 |
+
|
| 1098 |
+
|
| 1099 |
+
def _pure_label_fragment_case(label: str):
|
| 1100 |
+
fragment = f"{label} "
|
| 1101 |
+
full_proof = normalize_spoken_forms(f"{label}@islandmuseum.tw")
|
| 1102 |
+
parent_start = full_proof.index(fragment)
|
| 1103 |
+
prefix = "請核對,"
|
| 1104 |
+
target = f"{prefix}{fragment},完成。"
|
| 1105 |
+
chunk_start = len(prefix)
|
| 1106 |
+
return target, NetworkFragmentProof(
|
| 1107 |
+
span_index=0,
|
| 1108 |
+
chunk_start=chunk_start,
|
| 1109 |
+
chunk_end=chunk_start + len(fragment),
|
| 1110 |
+
parent_start=parent_start,
|
| 1111 |
+
parent_end=parent_start + len(fragment),
|
| 1112 |
+
full_spoken_proof=full_proof,
|
| 1113 |
+
)
|
| 1114 |
+
|
| 1115 |
+
|
| 1116 |
+
@pytest.mark.parametrize("label", ("tour", "a"))
|
| 1117 |
+
@pytest.mark.parametrize("side", ("prefix", "suffix"))
|
| 1118 |
+
@pytest.mark.parametrize(
|
| 1119 |
+
"symbol",
|
| 1120 |
+
(
|
| 1121 |
+
"/",
|
| 1122 |
+
"\\",
|
| 1123 |
+
"@",
|
| 1124 |
+
"-",
|
| 1125 |
+
"_",
|
| 1126 |
+
"%",
|
| 1127 |
+
"$",
|
| 1128 |
+
"~",
|
| 1129 |
+
"|",
|
| 1130 |
+
"=",
|
| 1131 |
+
"&",
|
| 1132 |
+
"#",
|
| 1133 |
+
"+",
|
| 1134 |
+
"*",
|
| 1135 |
+
"<",
|
| 1136 |
+
">",
|
| 1137 |
+
"§",
|
| 1138 |
+
"🙂",
|
| 1139 |
+
),
|
| 1140 |
+
)
|
| 1141 |
+
def test_pure_label_network_fragment_rejects_adjacent_identifier_symbols(
|
| 1142 |
+
label,
|
| 1143 |
+
side,
|
| 1144 |
+
symbol,
|
| 1145 |
+
):
|
| 1146 |
+
target, proof = _pure_label_fragment_case(label)
|
| 1147 |
+
literal = f"{symbol}{label}" if side == "prefix" else f"{label}{symbol}"
|
| 1148 |
+
|
| 1149 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1150 |
+
f"請核對,{literal},完成。",
|
| 1151 |
+
target,
|
| 1152 |
+
(proof,),
|
| 1153 |
+
)
|
| 1154 |
+
|
| 1155 |
+
assert evidence.protected_range_count == 1
|
| 1156 |
+
assert evidence.passed is False
|
| 1157 |
+
|
| 1158 |
+
|
| 1159 |
+
@pytest.mark.parametrize("label", ("tour", "a"))
|
| 1160 |
+
def test_pure_label_network_fragment_accepts_ascii_sentence_punctuation(label):
|
| 1161 |
+
target, proof = _pure_label_fragment_case(label)
|
| 1162 |
+
|
| 1163 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1164 |
+
f"請核對,{label},完成.",
|
| 1165 |
+
target,
|
| 1166 |
+
(proof,),
|
| 1167 |
+
)
|
| 1168 |
+
|
| 1169 |
+
assert evidence.transcript_text == target
|
| 1170 |
+
assert evidence.passed is True
|
| 1171 |
+
|
| 1172 |
+
|
| 1173 |
+
def test_local_network_fragment_exact_gate_is_stricter_than_whole_chunk_cer():
|
| 1174 |
+
prefix = "甲" * 30 + ","
|
| 1175 |
+
suffix = "," + "乙" * 30 + "。"
|
| 1176 |
+
target, proof = _tour_help_fragment_case(prefix, suffix)
|
| 1177 |
+
wrong_spoken = target.replace("help", "herp", 1)
|
| 1178 |
+
|
| 1179 |
+
# One protected middle error is below the ordinary whole-chunk CER limit
|
| 1180 |
+
# and outside the prefix/suffix windows, but the range gate still rejects.
|
| 1181 |
+
assert compare_asr_text(target, wrong_spoken, max_cer=0.20).passed
|
| 1182 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1183 |
+
wrong_spoken,
|
| 1184 |
+
target,
|
| 1185 |
+
(proof,),
|
| 1186 |
+
)
|
| 1187 |
+
assert evidence.passed is False
|
| 1188 |
+
|
| 1189 |
+
|
| 1190 |
+
def test_local_network_fragment_provenance_is_range_bound_and_fails_closed():
|
| 1191 |
+
target, proof = _tour_help_fragment_case()
|
| 1192 |
+
forged = NetworkFragmentProof(
|
| 1193 |
+
span_index=proof.span_index,
|
| 1194 |
+
chunk_start=proof.chunk_start + 1,
|
| 1195 |
+
chunk_end=proof.chunk_end,
|
| 1196 |
+
parent_start=proof.parent_start,
|
| 1197 |
+
parent_end=proof.parent_end - 1,
|
| 1198 |
+
full_spoken_proof=proof.full_spoken_proof,
|
| 1199 |
+
)
|
| 1200 |
+
|
| 1201 |
+
with pytest.raises(ValueError, match="does not bind"):
|
| 1202 |
+
canonicalize_asr_network_fragments("tour.help", target, (forged,))
|
| 1203 |
+
with pytest.raises(ValueError, match="invalid type"):
|
| 1204 |
+
canonicalize_asr_network_fragments("tour.help", target, (object(),))
|
| 1205 |
+
|
| 1206 |
+
|
| 1207 |
+
def test_local_network_fragment_cannot_borrow_repeated_plain_spoken_text():
|
| 1208 |
+
fragment = "tour 點 help"
|
| 1209 |
+
prefix = f"先照字面念{fragment},再核對,"
|
| 1210 |
+
target, proof = _tour_help_fragment_case(prefix, "。")
|
| 1211 |
+
transcript = "先照字面念 tour.help,再核對,tour.herp。"
|
| 1212 |
+
|
| 1213 |
+
evidence = canonicalize_asr_network_fragments(transcript, target, (proof,))
|
| 1214 |
+
|
| 1215 |
+
assert evidence.passed is False
|
| 1216 |
+
|
| 1217 |
+
|
| 1218 |
+
def _hybrid_tour_help_fragment_case(prefix: str = "請核對,"):
|
| 1219 |
+
fragment = "tour 點 help"
|
| 1220 |
+
full_proof = normalize_spoken_forms("tour.help@islandmuseum.tw")
|
| 1221 |
+
parent_start = full_proof.index(fragment)
|
| 1222 |
+
target = f"{prefix}{fragment},完成。"
|
| 1223 |
+
chunk_start = len(prefix)
|
| 1224 |
+
return target, NetworkFragmentProof(
|
| 1225 |
+
span_index=0,
|
| 1226 |
+
chunk_start=chunk_start,
|
| 1227 |
+
chunk_end=chunk_start + len(fragment),
|
| 1228 |
+
parent_start=parent_start,
|
| 1229 |
+
parent_end=parent_start + len(fragment),
|
| 1230 |
+
full_spoken_proof=full_proof,
|
| 1231 |
+
)
|
| 1232 |
+
|
| 1233 |
+
|
| 1234 |
+
def test_local_network_fragment_accepts_exact_range_bound_hybrid_lexical_atoms():
|
| 1235 |
+
target, proof = _hybrid_tour_help_fragment_case()
|
| 1236 |
+
|
| 1237 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1238 |
+
"請核對,Tour.Help,完成。",
|
| 1239 |
+
target,
|
| 1240 |
+
(proof,),
|
| 1241 |
+
)
|
| 1242 |
+
|
| 1243 |
+
assert evidence.passed is True
|
| 1244 |
+
assert evidence.transcript_text == target
|
| 1245 |
+
|
| 1246 |
+
|
| 1247 |
+
@pytest.mark.parametrize(
|
| 1248 |
+
"literal",
|
| 1249 |
+
(
|
| 1250 |
+
"tours.help",
|
| 1251 |
+
"tour.helps",
|
| 1252 |
+
"our.help",
|
| 1253 |
+
"tour.hepl",
|
| 1254 |
+
"xtour.help",
|
| 1255 |
+
"tour.help/x",
|
| 1256 |
+
),
|
| 1257 |
+
)
|
| 1258 |
+
def test_hybrid_lexical_network_atoms_reject_substrings_and_mutations(literal):
|
| 1259 |
+
target, proof = _hybrid_tour_help_fragment_case()
|
| 1260 |
+
|
| 1261 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1262 |
+
f"請核對,{literal},完成。",
|
| 1263 |
+
target,
|
| 1264 |
+
(proof,),
|
| 1265 |
+
)
|
| 1266 |
+
|
| 1267 |
+
assert evidence.passed is False
|
| 1268 |
+
|
| 1269 |
+
|
| 1270 |
+
def _duplicate_network_fragment_case(*, protected_first: bool):
|
| 1271 |
+
fragment = "點 欸"
|
| 1272 |
+
full_proof = "H T T P S 冒號 斜線 斜線 tour 點 欸"
|
| 1273 |
+
prefix = "甲" * 20
|
| 1274 |
+
suffix = "乙" * 20 + "。"
|
| 1275 |
+
target = f"{prefix}{fragment},{fragment}{suffix}"
|
| 1276 |
+
if protected_first:
|
| 1277 |
+
chunk_start = len(prefix)
|
| 1278 |
+
else:
|
| 1279 |
+
chunk_start = len(prefix) + len(fragment) + 1
|
| 1280 |
+
parent_start = full_proof.index(fragment)
|
| 1281 |
+
return target, prefix, suffix, NetworkFragmentProof(
|
| 1282 |
+
span_index=0,
|
| 1283 |
+
chunk_start=chunk_start,
|
| 1284 |
+
chunk_end=chunk_start + len(fragment),
|
| 1285 |
+
parent_start=parent_start,
|
| 1286 |
+
parent_end=parent_start + len(fragment),
|
| 1287 |
+
full_spoken_proof=full_proof,
|
| 1288 |
+
)
|
| 1289 |
+
|
| 1290 |
+
|
| 1291 |
+
@pytest.mark.parametrize("protected_first", (False, True))
|
| 1292 |
+
def test_network_fragment_duplicate_cannot_borrow_any_optimal_alignment(
|
| 1293 |
+
protected_first,
|
| 1294 |
+
):
|
| 1295 |
+
target, prefix, suffix, proof = _duplicate_network_fragment_case(
|
| 1296 |
+
protected_first=protected_first
|
| 1297 |
+
)
|
| 1298 |
+
|
| 1299 |
+
# The lone literal can align equally well to either duplicate. The local
|
| 1300 |
+
# proof must fail because at least one optimal path deletes the protected
|
| 1301 |
+
# occurrence, regardless of deterministic backtrace tie-breaking.
|
| 1302 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1303 |
+
f"{prefix}.a{suffix}",
|
| 1304 |
+
target,
|
| 1305 |
+
(proof,),
|
| 1306 |
+
)
|
| 1307 |
+
|
| 1308 |
+
assert compare_asr_text(
|
| 1309 |
+
target,
|
| 1310 |
+
evidence.transcript_text,
|
| 1311 |
+
max_cer=0.20,
|
| 1312 |
+
).passed
|
| 1313 |
+
assert evidence.passed is False
|
| 1314 |
+
|
| 1315 |
+
|
| 1316 |
+
@pytest.mark.parametrize("protected_first", (False, True))
|
| 1317 |
+
def test_network_fragment_duplicate_exact_two_occurrences_remain_valid(
|
| 1318 |
+
protected_first,
|
| 1319 |
+
):
|
| 1320 |
+
target, prefix, suffix, proof = _duplicate_network_fragment_case(
|
| 1321 |
+
protected_first=protected_first
|
| 1322 |
+
)
|
| 1323 |
+
|
| 1324 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1325 |
+
f"{prefix}.a,.a{suffix}",
|
| 1326 |
+
target,
|
| 1327 |
+
(proof,),
|
| 1328 |
+
)
|
| 1329 |
+
|
| 1330 |
+
assert evidence.transcript_text == target
|
| 1331 |
+
assert evidence.passed is True
|
| 1332 |
+
|
| 1333 |
+
|
| 1334 |
+
@pytest.mark.parametrize("boundary", ("start", "end"))
|
| 1335 |
+
def test_network_fragment_rejects_insertions_on_closed_range_boundary(boundary):
|
| 1336 |
+
prefix = "甲" * 30 + ","
|
| 1337 |
+
suffix = "," + "乙" * 30 + "。"
|
| 1338 |
+
target, proof = _tour_help_fragment_case(prefix, suffix)
|
| 1339 |
+
offset = proof.chunk_start if boundary == "start" else proof.chunk_end
|
| 1340 |
+
transcript = target[:offset] + "錯" + target[offset:]
|
| 1341 |
+
|
| 1342 |
+
evidence = canonicalize_asr_network_fragments(
|
| 1343 |
+
transcript,
|
| 1344 |
+
target,
|
| 1345 |
+
(proof,),
|
| 1346 |
+
)
|
| 1347 |
+
|
| 1348 |
+
assert compare_asr_text(target, transcript, max_cer=0.20).passed
|
| 1349 |
+
assert evidence.passed is False
|
| 1350 |
|
| 1351 |
|
| 1352 |
def test_generation_network_planner_keeps_ordinary_semantic_chunks_unchanged():
|
|
|
|
| 1382 |
with pytest.raises(ValueError, match="non-ASCII IRI"):
|
| 1383 |
plan_generation_chunks(emoji_iri, normalize_spoken_forms(emoji_iri))
|
| 1384 |
|
| 1385 |
+
oversized = "請查詢 https://" + "a" * 150 + ".tw。"
|
| 1386 |
with pytest.raises(ValueError, match="indivisible network component"):
|
| 1387 |
plan_generation_chunks(oversized, normalize_spoken_forms(oversized))
|
| 1388 |
|
| 1389 |
+
with pytest.raises(ValueError, match="chunk limits are inconsistent"):
|
| 1390 |
+
plan_generation_chunks(
|
| 1391 |
+
"請查 https://example.tw。",
|
| 1392 |
+
normalize_spoken_forms("請查 https://example.tw。"),
|
| 1393 |
+
min_units=12,
|
| 1394 |
+
network_min_units=13,
|
| 1395 |
+
)
|
| 1396 |
|
| 1397 |
+
|
| 1398 |
+
@pytest.mark.parametrize("raw", ("a@b.co", "https://a.tw"))
|
| 1399 |
+
def test_short_network_identifier_relaxes_only_infeasible_preferred_cut(raw):
|
| 1400 |
+
normalized = normalize_spoken_forms(raw)
|
| 1401 |
+
|
| 1402 |
+
specs = plan_generation_chunks(raw, normalized)
|
| 1403 |
+
|
| 1404 |
+
assert len(specs) == 1
|
| 1405 |
+
spec = specs[0]
|
| 1406 |
+
assert spec.text == normalized
|
| 1407 |
+
assert spec.source_start == 0
|
| 1408 |
+
assert spec.source_end == len(normalized)
|
| 1409 |
+
assert spec.network_conditioned
|
| 1410 |
+
assert spec.boundary_after == "none"
|
| 1411 |
+
assert spec.network_span_indices == (0,)
|
| 1412 |
+
assert spec.network_full_spoken_proofs == (normalized,)
|
| 1413 |
+
assert len(spec.network_fragment_proofs) == 1
|
| 1414 |
+
proof = spec.network_fragment_proofs[0]
|
| 1415 |
+
assert (proof.chunk_start, proof.chunk_end) == (0, len(normalized))
|
| 1416 |
+
assert (proof.parent_start, proof.parent_end) == (0, len(normalized))
|
| 1417 |
+
assert proof.full_spoken_proof == normalized
|
| 1418 |
+
|
| 1419 |
+
|
| 1420 |
+
@pytest.mark.parametrize(
|
| 1421 |
+
("raw", "expected_texts", "expected_units"),
|
| 1422 |
+
(
|
| 1423 |
+
(
|
| 1424 |
+
"若要更換導覽場次,請寄信到 tour.help@islandmuseum.tw。",
|
| 1425 |
+
(
|
| 1426 |
+
"若要更換導覽場次,請寄信到,tour 點 help ",
|
| 1427 |
+
"小老鼠 islandmuseum 點 T W。",
|
| 1428 |
+
),
|
| 1429 |
+
(15, 9),
|
| 1430 |
+
),
|
| 1431 |
+
(
|
| 1432 |
+
"潮汐預報可查詢 https://coastwatch.example.tw/tide。",
|
| 1433 |
+
(
|
| 1434 |
+
"潮汐預報可查詢,H T T P S 冒號 斜線 斜線",
|
| 1435 |
+
" coastwatch 點 example 點 T W 斜線 tide。",
|
| 1436 |
+
),
|
| 1437 |
+
(18, 12),
|
| 1438 |
+
),
|
| 1439 |
+
),
|
| 1440 |
+
)
|
| 1441 |
+
def test_generation_planner_cannot_cross_mandatory_network_grammar_cuts(
|
| 1442 |
+
raw,
|
| 1443 |
+
expected_texts,
|
| 1444 |
+
expected_units,
|
| 1445 |
+
):
|
| 1446 |
+
normalized = normalize_spoken_forms(raw)
|
| 1447 |
+
specs = plan_generation_chunks(
|
| 1448 |
+
raw,
|
| 1449 |
+
normalized,
|
| 1450 |
+
target_units=80,
|
| 1451 |
+
network_max_units=80,
|
| 1452 |
)
|
| 1453 |
+
|
| 1454 |
+
assert tuple(spec.text for spec in specs) == expected_texts
|
| 1455 |
+
assert tuple(count_speech_units(spec.text) for spec in specs) == expected_units
|
| 1456 |
+
assert specs[0].boundary_after == "network_internal"
|
| 1457 |
+
assert specs[1].boundary_after == "none"
|
| 1458 |
+
mandatory_cut = specs[0].source_end
|
| 1459 |
+
assert all(
|
| 1460 |
+
not (spec.source_start < mandatory_cut < spec.source_end)
|
| 1461 |
+
for spec in specs
|
| 1462 |
+
)
|
| 1463 |
+
assert specs[0].network_fragment_proofs[0].parent_end == (
|
| 1464 |
+
specs[1].network_fragment_proofs[0].parent_start
|
| 1465 |
+
)
|
| 1466 |
+
|
| 1467 |
+
|
| 1468 |
+
def test_hybrid_network_frontend_keeps_labels_lexical_and_typed_atoms_explicit():
|
| 1469 |
+
assert normalize_spoken_forms(
|
| 1470 |
+
"若要更換導覽場次,請寄信到 tour.help@islandmuseum.tw。"
|
| 1471 |
+
) == (
|
| 1472 |
+
"若要更換導覽場次,請寄信到,tour 點 help 小老鼠 "
|
| 1473 |
+
"islandmuseum 點 T W。"
|
| 1474 |
+
)
|
| 1475 |
+
assert normalize_spoken_forms(
|
| 1476 |
+
"潮汐預報可查詢 https://coastwatch.example.tw/tide。"
|
| 1477 |
+
) == (
|
| 1478 |
+
"潮汐預報可查詢,H T T P S 冒號 斜線 斜線 coastwatch 點 "
|
| 1479 |
+
"example 點 T W 斜線 tide。"
|
| 1480 |
+
)
|
| 1481 |
+
assert normalize_spoken_forms(
|
| 1482 |
+
"HTTPS://WWW.Example.COM/API/v1?ID=RTX5090"
|
| 1483 |
+
) == (
|
| 1484 |
+
"H T T P S 冒號 斜線 斜線 W W W 點 Example 點 C O M 斜線 "
|
| 1485 |
+
"A P I 斜線 v 一 問號 I D 等於 R T X 五零九零"
|
| 1486 |
+
)
|
| 1487 |
+
|
| 1488 |
+
|
| 1489 |
+
def test_generation_network_planner_binds_duplicate_spoken_proof_to_raw_identifier():
|
| 1490 |
+
spoken = "H T T P S 冒號 斜線 斜線 example 點 T W"
|
| 1491 |
raw = f"先念 {spoken},再查 https://example.tw。"
|
| 1492 |
normalized = normalize_spoken_forms(raw)
|
| 1493 |
|
| 1494 |
specs = plan_generation_chunks(raw, normalized)
|
| 1495 |
|
| 1496 |
assert "".join(spec.text for spec in specs) == normalized
|
| 1497 |
+
assert [spec.network_conditioned for spec in specs] == [True]
|
| 1498 |
+
assert specs[0].text.count(spoken) == 2
|
| 1499 |
+
assert specs[0].network_full_spoken_proofs == (spoken,)
|
| 1500 |
+
proof = specs[0].network_fragment_proofs[0]
|
| 1501 |
+
assert proof.chunk_start == specs[0].text.rindex(spoken)
|
| 1502 |
+
assert specs[0].text[proof.chunk_start : proof.chunk_end] == spoken
|
| 1503 |
assert all(
|
| 1504 |
normalized[spec.source_start : spec.source_end] == spec.text
|
| 1505 |
for spec in specs
|
|
|
|
| 1751 |
assert select_generation_cps("AI TTS 測試", cjk_cps=5.2, ascii_cps=4.6) == 4.6
|
| 1752 |
|
| 1753 |
|
| 1754 |
+
def test_hybrid_network_text_keeps_the_ascii_generation_window():
|
| 1755 |
normalized = normalize_spoken_forms("https://example.tw/path")
|
| 1756 |
|
| 1757 |
+
assert any(character.isascii() and character.isalnum() for character in normalized)
|
| 1758 |
assert network_protected_spoken_spans(normalized)
|
| 1759 |
assert select_generation_cps(normalized, cjk_cps=5.2, ascii_cps=4.6) == 4.6
|
| 1760 |
|
|
|
|
| 1900 |
"無效 24點30分、15點60分。"
|
| 1901 |
)
|
| 1902 |
assert normalize_spoken_forms("IP 192.168.1.1,網址 https://example.test:30/path") == (
|
| 1903 |
+
"I P 192.168.1.1,網址,H T T P S 冒號 斜線 斜線 "
|
| 1904 |
+
"example 點 test 冒號 三零 斜線 path"
|
|
|
|
| 1905 |
)
|
| 1906 |
assert normalize_spoken_forms(
|
| 1907 |
"音量15點05分貝,區間15點30分鐘,比分15點05分。"
|
|
|
|
| 2337 |
assert normalize_spoken_forms(
|
| 2338 |
"網址 https://api.example.com/v1/items?q=RTX-5090&n=2。"
|
| 2339 |
) == (
|
| 2340 |
+
"網址,H T T P S 冒號 斜線 斜線 api 點 example 點 C O M "
|
| 2341 |
+
"斜線 v 一 斜線 items 問號 q 等於 R T X 橫線 五零九零 "
|
| 2342 |
+
"和 n 等於 二。"
|
|
|
|
| 2343 |
)
|
| 2344 |
assert normalize_spoken_forms("信箱 USER.name+tts@example.com。") == (
|
| 2345 |
+
"信箱,U S E R 點 name 加號 tts 小老鼠 example 點 C O M。"
|
|
|
|
| 2346 |
)
|
| 2347 |
assert normalize_spoken_forms(
|
| 2348 |
"版本 v1.2.3,候選 version 10.4.0-beta.1+build.5。"
|
|
|
|
| 2387 |
target = normalize_spoken_forms("https://example.tw/path")
|
| 2388 |
transcript = target.replace(
|
| 2389 |
_SPOKEN_PATH,
|
| 2390 |
+
f"pa{symbol}th",
|
| 2391 |
)
|
| 2392 |
comparison = compare_asr_text(
|
| 2393 |
target,
|
|
|
|
| 2406 |
def test_expanded_email_rejects_raw_symbol_inserted_inside_local_part(symbol):
|
| 2407 |
target = normalize_spoken_forms("museum@example.tw")
|
| 2408 |
transcript = target.replace(
|
| 2409 |
+
"museum",
|
| 2410 |
+
f"mu{symbol}seum",
|
| 2411 |
)
|
| 2412 |
comparison = compare_asr_text(
|
| 2413 |
target,
|
|
|
|
| 2564 |
)
|
| 2565 |
comparison = compare_asr_text(
|
| 2566 |
target,
|
| 2567 |
+
target.replace(_SPOKEN_PATH, "pat"),
|
| 2568 |
max_cer=0.20,
|
| 2569 |
max_prefix_cer=1.0,
|
| 2570 |
max_suffix_cer=1.0,
|
tests/test_quality_runtime.py
CHANGED
|
@@ -811,8 +811,8 @@ def test_quality_gate_rejects_inexact_network_span_below_whole_cer_limit():
|
|
| 811 |
"請先閱讀 https://museum.example.tw/path,確認展覽時間與集合位置後再回覆。"
|
| 812 |
)
|
| 813 |
transcript = target.replace(
|
| 814 |
-
"
|
| 815 |
-
"
|
| 816 |
)
|
| 817 |
result = verify_candidate(
|
| 818 |
CandidateObservation(
|
|
|
|
| 811 |
"請先閱讀 https://museum.example.tw/path,確認展覽時間與集合位置後再回覆。"
|
| 812 |
)
|
| 813 |
transcript = target.replace(
|
| 814 |
+
"museum",
|
| 815 |
+
"museuman",
|
| 816 |
)
|
| 817 |
result = verify_candidate(
|
| 818 |
CandidateObservation(
|
tests/test_release_pins.py
CHANGED
|
@@ -229,6 +229,11 @@ def test_readme_describes_coverage_refill_and_sequence_transition_scores():
|
|
| 229 |
assert "median-F0 軟成本" in readme
|
| 230 |
assert "5.2 CJK / 4.6 ASCII" in readme
|
| 231 |
assert "4.6 CJK / 4.0 ASCII" in readme
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 232 |
assert "boundary-only local rejects" in readme
|
| 233 |
assert "drop 不超過 0.15" in readme
|
| 234 |
assert "整段仍必須通過 similarity 0.105 與 boundary drop 0.095" in readme
|
|
@@ -352,6 +357,7 @@ def test_app_and_quality_runtime_pin_the_same_mixed_cfg_contract():
|
|
| 352 |
|
| 353 |
def test_app_rejects_silent_text_and_coalesces_before_runtime_budgeting():
|
| 354 |
source = (ROOT / "app.py").read_text(encoding="utf-8")
|
|
|
|
| 355 |
tree = ast.parse(source)
|
| 356 |
functions = {
|
| 357 |
node.name: node
|
|
@@ -372,8 +378,10 @@ def test_app_rejects_silent_text_and_coalesces_before_runtime_budgeting():
|
|
| 372 |
assert "chunks = tuple(spec.text for spec in chunk_specs)" in synthesize_source
|
| 373 |
assert "max_chunks=QUALITY_MAX_GENERATED_CHUNKS" in synthesize_source
|
| 374 |
assert "pre_faded_edges=True" in assemble_source
|
|
|
|
| 375 |
assert "NETWORK_GENERATION_TARGET_UNITS = 32" in source
|
| 376 |
assert "NETWORK_GENERATION_MAX_UNITS = 36" in source
|
|
|
|
| 377 |
assert "NETWORK_INTERNAL_FADE_MS = 5.0" in source
|
| 378 |
assert 'chunk_specs[index].boundary_after == "network_internal"' in (
|
| 379 |
assemble_source
|
|
@@ -381,6 +389,16 @@ def test_app_rejects_silent_text_and_coalesces_before_runtime_budgeting():
|
|
| 381 |
assert "join_audio_chunks_variable(" in assemble_source
|
| 382 |
assert "network_conditioned=network_flags" in synthesize_source
|
| 383 |
assert "network_conditioned=network_flag" in synthesize_source
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 384 |
|
| 385 |
|
| 386 |
def test_app_emits_one_canonical_content_free_evidence_line_per_terminal_outcome():
|
|
@@ -679,6 +697,69 @@ def test_space_hard_intersects_dual_asr_only_on_exact_whole_waveforms():
|
|
| 679 |
assert "except (RuntimeError, ValueError) as error:" in synthesize_source
|
| 680 |
|
| 681 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 682 |
def test_space_wires_bounded_k_best_paths_to_exact_assembled_whole_gate():
|
| 683 |
source = (ROOT / "app.py").read_text(encoding="utf-8")
|
| 684 |
|
|
|
|
| 229 |
assert "median-F0 軟成本" in readme
|
| 230 |
assert "5.2 CJK / 4.6 ASCII" in readme
|
| 231 |
assert "4.6 CJK / 4.0 ASCII" in readme
|
| 232 |
+
assert "URL/email-bearing chunks 8 units" in readme
|
| 233 |
+
assert "Email 第一輪必須在 `小老鼠`" in readme
|
| 234 |
+
assert "第一輪 DP 不得跨越這兩類" in readme
|
| 235 |
+
assert "grammar boundary" in readme
|
| 236 |
+
assert "所有最佳 edit alignment" in readme
|
| 237 |
assert "boundary-only local rejects" in readme
|
| 238 |
assert "drop 不超過 0.15" in readme
|
| 239 |
assert "整段仍必須通過 similarity 0.105 與 boundary drop 0.095" in readme
|
|
|
|
| 357 |
|
| 358 |
def test_app_rejects_silent_text_and_coalesces_before_runtime_budgeting():
|
| 359 |
source = (ROOT / "app.py").read_text(encoding="utf-8")
|
| 360 |
+
production_source = (ROOT / "production.py").read_text(encoding="utf-8")
|
| 361 |
tree = ast.parse(source)
|
| 362 |
functions = {
|
| 363 |
node.name: node
|
|
|
|
| 378 |
assert "chunks = tuple(spec.text for spec in chunk_specs)" in synthesize_source
|
| 379 |
assert "max_chunks=QUALITY_MAX_GENERATED_CHUNKS" in synthesize_source
|
| 380 |
assert "pre_faded_edges=True" in assemble_source
|
| 381 |
+
assert "NETWORK_GENERATION_MIN_UNITS = 8" in source
|
| 382 |
assert "NETWORK_GENERATION_TARGET_UNITS = 32" in source
|
| 383 |
assert "NETWORK_GENERATION_MAX_UNITS = 36" in source
|
| 384 |
+
assert "network_min_units=NETWORK_GENERATION_MIN_UNITS" in synthesize_source
|
| 385 |
assert "NETWORK_INTERNAL_FADE_MS = 5.0" in source
|
| 386 |
assert 'chunk_specs[index].boundary_after == "network_internal"' in (
|
| 387 |
assemble_source
|
|
|
|
| 389 |
assert "join_audio_chunks_variable(" in assemble_source
|
| 390 |
assert "network_conditioned=network_flags" in synthesize_source
|
| 391 |
assert "network_conditioned=network_flag" in synthesize_source
|
| 392 |
+
assert "mandatory_cut_offsets" in production_source
|
| 393 |
+
assert "any(start < cut < end for cut in active_mandatory_cuts)" in (
|
| 394 |
+
production_source
|
| 395 |
+
)
|
| 396 |
+
assert "plan = solve(mandatory_cuts - short_identifier_cuts)" in (
|
| 397 |
+
production_source
|
| 398 |
+
)
|
| 399 |
+
assert "_protected_ranges_are_exact_in_all_optimal_alignments(" in (
|
| 400 |
+
production_source
|
| 401 |
+
)
|
| 402 |
|
| 403 |
|
| 404 |
def test_app_emits_one_canonical_content_free_evidence_line_per_terminal_outcome():
|
|
|
|
| 697 |
assert "except (RuntimeError, ValueError) as error:" in synthesize_source
|
| 698 |
|
| 699 |
|
| 700 |
+
def test_network_fragment_relaxation_is_range_bound_and_local_only():
|
| 701 |
+
source = (ROOT / "app.py").read_text(encoding="utf-8")
|
| 702 |
+
tree = ast.parse(source)
|
| 703 |
+
functions = {
|
| 704 |
+
node.name: node
|
| 705 |
+
for node in tree.body
|
| 706 |
+
if isinstance(node, ast.FunctionDef)
|
| 707 |
+
}
|
| 708 |
+
verify_source = ast.get_source_segment(source, functions["_verify_trajectory_audio"])
|
| 709 |
+
proof_source = ast.get_source_segment(
|
| 710 |
+
source,
|
| 711 |
+
functions["_network_fragment_proof_rows"],
|
| 712 |
+
)
|
| 713 |
+
qualify_source = ast.get_source_segment(
|
| 714 |
+
source,
|
| 715 |
+
functions["_qualify_candidate_trajectory_audio"],
|
| 716 |
+
)
|
| 717 |
+
refill_source = ast.get_source_segment(
|
| 718 |
+
source,
|
| 719 |
+
functions["_verify_refill_candidate_trajectory_audio"],
|
| 720 |
+
)
|
| 721 |
+
independent_source = ast.get_source_segment(
|
| 722 |
+
source,
|
| 723 |
+
functions["_verify_independent_whole_audio"],
|
| 724 |
+
)
|
| 725 |
+
sequence_source = ast.get_source_segment(
|
| 726 |
+
source,
|
| 727 |
+
functions["_verify_sequence_trajectory_audio"],
|
| 728 |
+
)
|
| 729 |
+
synthesize_source = ast.get_source_segment(source, functions["_synthesize"])
|
| 730 |
+
|
| 731 |
+
assert all(
|
| 732 |
+
segment is not None
|
| 733 |
+
for segment in (
|
| 734 |
+
verify_source,
|
| 735 |
+
proof_source,
|
| 736 |
+
qualify_source,
|
| 737 |
+
refill_source,
|
| 738 |
+
independent_source,
|
| 739 |
+
sequence_source,
|
| 740 |
+
synthesize_source,
|
| 741 |
+
)
|
| 742 |
+
)
|
| 743 |
+
assert "canonicalize_asr_network_fragments(" in verify_source
|
| 744 |
+
assert "if fragment_evidence.passed" in verify_source
|
| 745 |
+
assert "network-conditioned chunk lacks exact fragment proof" in proof_source
|
| 746 |
+
assert "proof.span_index for proof in proofs" in proof_source
|
| 747 |
+
assert "proof.full_spoken_proof for proof in proofs" in proof_source
|
| 748 |
+
|
| 749 |
+
local_index = qualify_source.index("local_verification = _verify_trajectory_audio(")
|
| 750 |
+
joined_index = qualify_source.index("joined_verification = _verify_trajectory_audio(")
|
| 751 |
+
assert "network_fragment_proofs=" in qualify_source[local_index:joined_index]
|
| 752 |
+
assert "network_fragment_proofs=" not in qualify_source[joined_index:]
|
| 753 |
+
assert "network_fragment_proofs=" in refill_source
|
| 754 |
+
assert "network_fragment_proofs=" not in independent_source
|
| 755 |
+
assert "network_fragment_proofs=" not in sequence_source
|
| 756 |
+
|
| 757 |
+
final_index = synthesize_source.index("final_verification = _verify_trajectory_audio(")
|
| 758 |
+
assert "network_fragment_proofs=" not in synthesize_source[final_index:]
|
| 759 |
+
assert "generation_context_by_seed" in synthesize_source
|
| 760 |
+
assert "generation_chunk_specs(seed, candidate_chunks)" in synthesize_source
|
| 761 |
+
|
| 762 |
+
|
| 763 |
def test_space_wires_bounded_k_best_paths_to_exact_assembled_whole_gate():
|
| 764 |
source = (ROOT / "app.py").read_text(encoding="utf-8")
|
| 765 |
|