voidful commited on
Commit
1a6ae0b
·
1 Parent(s): 34ba6ed

Harden network generation and row-local recovery

Browse files
README.md CHANGED
@@ -45,7 +45,7 @@ Barbet 另固定在 `6fcd7ce4aa37f2250a3242995bef0fbc3b026ba8`,
45
 
46
  | 設定 | 值 |
47
  |---|---:|
48
- | CFG | fixed candidate schedule: offset 0/even = 3.0, positive odd = 2.0; UI primary CFG is locked |
49
  | NFE steps | fixed at the validated value 10 |
50
  | Target pace | 4.0 speech units/sec |
51
  | Initial whole-trajectory policy | base: 5.2 CJK / 4.6 ASCII-mixed units/sec + 1 latent step |
@@ -69,9 +69,9 @@ Barbet 另固定在 `6fcd7ce4aa37f2250a3242995bef0fbc3b026ba8`,
69
  | Sequence fallback | ragged DP with up to 3 culprit-diverse paths; boundary-only local rejects may enter through a frozen 0.15 cap, then every exact assembly must pass turbo + full large-v3 and the stricter 0.105/0.095 whole gate |
70
  | Final output gate | re-verify joined/faded/RMS-matched/speed-adjusted whole waveform; fail closed |
71
  | Runtime budget | at most 20 generated TTS chunks and 800 generated speech units per request; NFE fixed at 10 |
72
- | Maximum chunk | 80 speech units |
73
  | Minimum chunk | 12 speech units where feasible; genuine short requests/natural short sentence boundaries are preserved |
74
- | Crossfade / internal edge fade | 80 ms / 80 ms |
75
  | Chunk RMS adjustment | at most 4 dB |
76
  | Generated continuation context | 0 sec |
77
  | Generated-audio retry | disabled |
@@ -80,18 +80,19 @@ Barbet 另固定在 `6fcd7ce4aa37f2250a3242995bef0fbc3b026ba8`,
80
  每個請求會取得新的隨機 root seed。服務先且只先生成一次完整 trajectory;其中所有長文切段
81
  共用 root seed,並使用 base policy(5.2 CJK / 4.6 ASCII)。如果這條 exact whole path 未通過,
82
  服務會保留每一段的 local evidence,再依 `(coverage, refill attempts, chunk index)` 穩定順序,
83
- 只對 zero/low-coverage chunk 逐一生成 refill。大多數 offset 使用 safe-duration(4.6 CJK / 4.0 ASCII),
84
  每第四個 retry 使用 completion-headroom(4.2 CJK / 3.6 ASCII)。Refill 使用
85
  `root seed + global generation offset`,不會重新生成其他已覆蓋 chunks。三種 policy 都只改 native-duration endpoint estimate,
86
  margin 固定 +1 latent step、`min_len` 固定為 2,不會為了播放目標語速硬撐 generation loop。
87
  選中完整 trajectory 或 coverage-sequence-DP 混合路徑後,log 會記錄每個 chunk 的 candidate、seed、policy、實際 CFG,
88
- 並記錄每個已嘗試 offset 的 schedule CFG。Offset 0 與偶數候選使用 primary CFG 3.0,
89
- 正奇數候選使用 alternate CFG 2.0;含 URL/email 的 chunk 仍有 CFG 3.0 的內容安全下限。
90
  同一列 log 也記錄實際 generated chunks/text units 與各 row 的有效候選數。
91
  每次生成前另外記錄 seed、local chunk index、duration policy、scheduled CFG,
92
  以及 network/short-text floor 後的 effective CFG;因此即使最後 fail closed 也能重建嘗試軌跡。
93
- 終端 outcome 另輸出 content-free canonical evidence schema v2,包含 original chunk indices、
94
- CFG contract、20/800 預算用量與 selected-path 交叉檢查,不包含 target/transcript/audio/embedding。
 
95
  為了讓 release contract 與正式量測一致,UI 不提供其他 primary CFG。
96
  內部 hosted evaluator 可注入 `[0, 2^31)` 的固定 root seed 以重現結果;UI 不暴露這個參數,
97
  一般請求仍只在未注入 seed 時使用系統亂數。
@@ -139,6 +140,10 @@ zh-TW spoken form,例如 `2026/07/16`、`15:30`、`12.5%` 與 `10 km`。ASR
139
  Email 與 URL 會以可辨識的語義讀法展開:scheme、local/domain/path 的 opaque ASCII
140
  labels 逐字使用台灣華語字母名,數字逐位朗讀,分隔符明確朗讀(例如 `.tw`
141
  讀成「點、踢、達不溜」)。
 
 
 
 
142
  這能降低���型把不常見 TLD 自動補成 `.com` 的風險;一般英文句子不會套用這個規則。
143
  URL 與後續英文 prose 應以空白或中文標點分隔;未分隔的 RFC path punctuation 會視為 URL
144
  本身的一部分並納入 exact gate。Quoted email local-part 暫不支援,輸入時會直接 fail closed。
 
45
 
46
  | 設定 | 值 |
47
  |---|---:|
48
+ | CFG | fixed per-chunk schedule: row ordinal 0/even = 3.0, positive odd = 2.0; UI primary CFG is locked |
49
  | NFE steps | fixed at the validated value 10 |
50
  | Target pace | 4.0 speech units/sec |
51
  | Initial whole-trajectory policy | base: 5.2 CJK / 4.6 ASCII-mixed units/sec + 1 latent step |
 
69
  | Sequence fallback | ragged DP with up to 3 culprit-diverse paths; boundary-only local rejects may enter through a frozen 0.15 cap, then every exact assembly must pass turbo + full large-v3 and the stricter 0.105/0.095 whole gate |
70
  | Final output gate | re-verify joined/faded/RMS-matched/speed-adjusted whole waveform; fail closed |
71
  | Runtime budget | at most 20 generated TTS chunks and 800 generated speech units per request; NFE fixed at 10 |
72
+ | Maximum chunk | ordinary text 80 speech units; URL/email-bearing generation chunks 36 units |
73
  | Minimum chunk | 12 speech units where feasible; genuine short requests/natural short sentence boundaries are preserved |
74
+ | Crossfade / internal edge fade | semantic boundary 80 ms / 80 ms; proven URL/email internal boundary 5 ms / 5 ms with no inserted pause |
75
  | Chunk RMS adjustment | at most 4 dB |
76
  | Generated continuation context | 0 sec |
77
  | Generated-audio retry | disabled |
 
80
  每個請求會取得新的隨機 root seed。服務先且只先生成一次完整 trajectory;其中所有長文切段
81
  共用 root seed,並使用 base policy(5.2 CJK / 4.6 ASCII)。如果這條 exact whole path 未通過,
82
  服務會保留每一段的 local evidence,再依 `(coverage, refill attempts, chunk index)` 穩定順序,
83
+ 只對 zero/low-coverage chunk 逐一生成 refill。每個 chunk 依自己的 refill ordinal 輪替 policy;大多數 refill 使用 safe-duration(4.6 CJK / 4.0 ASCII),
84
  每第四個 retry 使用 completion-headroom(4.2 CJK / 3.6 ASCII)。Refill 使用
85
  `root seed + global generation offset`,不會重新生成其他已覆蓋 chunks。三種 policy 都只改 native-duration endpoint estimate,
86
  margin 固定 +1 latent step、`min_len` 固定為 2,不會為了播放目標語速硬撐 generation loop。
87
  選中完整 trajectory 或 coverage-sequence-DP 混合路徑後,log 會記錄每個 chunk 的 candidate、seed、policy、實際 CFG,
88
+ 並記錄每個已嘗試 row-local ordinal 的 schedule CFG。Ordinal 0 與正偶數使用 primary CFG 3.0,
89
+ 正奇數使用 alternate CFG 2.0;含 URL/email component 的 chunk 仍有 CFG 3.0 的內容安全下限。
90
  同一列 log 也記錄實際 generated chunks/text units 與各 row 的有效候選數。
91
  每次生成前另外記錄 seed、local chunk index、duration policy、scheduled CFG,
92
  以及 network/short-text floor 後的 effective CFG;因此即使最後 fail closed 也能重建嘗試軌跡。
93
+ 終端 outcome 另輸出 content-free canonical evidence schema v3,包含 original chunk indices、
94
+ row-local ordinal、network provenance、CFG contract、20/800 預算用量與 selected-path 交叉檢查,
95
+ 不包含 target/transcript/audio/embedding。
96
  為了讓 release contract 與正式量測一致,UI 不提供其他 primary CFG。
97
  內部 hosted evaluator 可注入 `[0, 2^31)` 的固定 root seed 以重現結果;UI 不暴露這個參數,
98
  一般請求仍只在未注入 seed 時使用系統亂數。
 
140
  Email 與 URL 會以可辨識的語義讀法展開:scheme、local/domain/path 的 opaque ASCII
141
  labels 逐字使用台灣華語字母名,數字逐位朗讀,分隔符明確朗讀(例如 `.tw`
142
  讀成「點、踢、達不溜」)。
143
+ 完整 identifier 仍保留為 joined、full-large-v3 與 final ASR 的 exact protected target;只有模型
144
+ generation 會在由原始 ASCII grammar 證明的 scheme/domain/path/query/email component 邊界切段,
145
+ 以 32 units 為目標、36 units 為硬上限。這些 identifier 內部邊界不插入 pause,只使用 5 ms
146
+ fade/crossfade;component proof 無法完整重建、非 ASCII IRI 或單一不可拆 component 超限時會 fail closed。
147
  這能降低���型把不常見 TLD 自動補成 `.com` 的風險;一般英文句子不會套用這個規則。
148
  URL 與後續英文 prose 應以空白或中文標點分隔;未分隔的 RFC path punctuation 會視為 URL
149
  本身的一部分並納入 exact gate。Quoted email local-part 暫不支援,輸入時會直接 fail closed。
app.py CHANGED
@@ -18,6 +18,7 @@ from transformers import PreTrainedTokenizerFast
18
  from bluemagpie import BlueMagpieModel
19
  from production import (
20
  StopHysteresisController,
 
21
  active_pace_correction_speed,
22
  apply_loudness_floor,
23
  coalesce_text_chunks,
@@ -27,14 +28,15 @@ from production import (
27
  endpoint_generation_plan,
28
  estimate_step_seconds,
29
  extract_windowed_speaker_embedding,
30
- fade_internal_edges,
31
  finish_audio,
32
- join_audio_chunks,
33
  match_chunk_rms,
34
  network_identifier_has_ambiguous_iri,
35
  network_protected_spoken_spans,
36
  normalize_spoken_forms,
37
  punctuation_pause_seconds,
 
38
  select_generation_cps,
39
  set_generation_seed,
40
  split_text_for_tts,
@@ -47,6 +49,7 @@ from quality_runtime import (
47
  VERIFICATION_WHISPER_REVISION,
48
  WHISPER_MODEL_ID,
49
  WHISPER_REVISION,
 
50
  CandidateGenerationEvidence,
51
  CandidateVerification,
52
  CandidateObservation,
@@ -101,7 +104,7 @@ TARGET_CPS = 4.0
101
  ACTIVE_PACE_TARGET_CPS = 4.00
102
  MIXED_CFG_PRIMARY = 3.0
103
  MIXED_CFG_ALTERNATE = 2.0
104
- MIXED_CFG_SCHEDULE = "offset_zero_and_even_primary_odd_alternate"
105
  MIN_ENDPOINT_CUE_UNITS = 6
106
  SHORT_TEXT_CFG_MIN = 3.0
107
  SHORT_TEXT_CFG_UNITS = 6
@@ -112,6 +115,9 @@ MIN_CHUNK_CHARS = 12
112
  CROSSFADE_MS = 80.0
113
  CHUNK_EDGE_FADE_MS = 80.0
114
  CHUNK_RMS_MATCH_DB = 4.0
 
 
 
115
  STOP_THRESHOLD = 0.50
116
  STOP_LATE_THRESHOLD = 0.05
117
  STOP_LATE_START_RATIO = 0.75
@@ -259,11 +265,16 @@ def _generate_chunk(
259
  steps: int,
260
  request_seed: int,
261
  policy: GenerationPolicy,
 
262
  ) -> np.ndarray:
263
- generation_cps = select_generation_cps(
264
- text,
265
- cjk_cps=policy.cjk_cps,
266
- ascii_cps=policy.ascii_cps,
 
 
 
 
267
  )
268
  model_text, expected_steps, hard_stop_steps = endpoint_generation_plan(
269
  text,
@@ -372,11 +383,18 @@ def _generate_trajectory(
372
  request_seed: int,
373
  policy: GenerationPolicy,
374
  network_cfg_min: float = NETWORK_TEXT_CFG_MIN,
 
375
  ) -> tuple[np.ndarray, ...]:
376
  scheduled_cfg = float(cfg)
377
  trajectory: list[np.ndarray] = []
 
 
378
  for local_chunk_index, chunk in enumerate(chunks):
379
- network_chunk = bool(network_protected_spoken_spans(chunk))
 
 
 
 
380
  network_floor_applied = bool(
381
  network_chunk and scheduled_cfg < float(network_cfg_min)
382
  )
@@ -408,6 +426,7 @@ def _generate_trajectory(
408
  steps=steps,
409
  request_seed=request_seed,
410
  policy=policy,
 
411
  )
412
  )
413
  return tuple(trajectory)
@@ -596,13 +615,49 @@ def _assemble_trajectory_audio(
596
  trajectory: tuple[np.ndarray, ...],
597
  chunks: tuple[str, ...],
598
  playback_speed: float,
 
599
  ) -> np.ndarray:
600
  """Assemble chunks exactly as they will be returned to the listener."""
601
 
602
  if not trajectory or len(trajectory) != len(chunks):
603
  raise ValueError("trajectory and text chunks must be non-empty and aligned")
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
604
  audio_chunks = [np.asarray(audio, dtype=np.float32).copy() for audio in trajectory]
605
  pauses: list[int] = []
 
 
606
  for index, chunk in enumerate(chunks):
607
  if index > 0:
608
  audio_chunks[index] = match_chunk_rms(
@@ -611,13 +666,34 @@ def _assemble_trajectory_audio(
611
  max_adjust_db=CHUNK_RMS_MATCH_DB,
612
  )
613
  if index + 1 < len(chunks):
614
- pauses.append(int(round(punctuation_pause_seconds(chunk) * SR)))
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
615
 
616
- audio_chunks = fade_internal_edges(audio_chunks, SR, fade_ms=CHUNK_EDGE_FADE_MS)
617
- waveform = join_audio_chunks(
618
  audio_chunks,
619
  pauses,
620
- crossfade_samples=int(round(CROSSFADE_MS * SR / 1000.0)),
 
 
 
621
  pre_faded_edges=True,
622
  )
623
  waveform = apply_loudness_floor(
@@ -657,6 +733,7 @@ def _qualify_candidate_trajectory_audio(
657
  independent_cache: WholeWaveformVerificationCache,
658
  *,
659
  candidate_seed: int,
 
660
  ):
661
  """Run whole-output qualification only after every local chunk passes."""
662
 
@@ -669,7 +746,12 @@ def _qualify_candidate_trajectory_audio(
669
  if not local_verification.passed:
670
  return CandidateVerification(local_verification)
671
 
672
- waveform = _assemble_trajectory_audio(trajectory, chunks, playback_speed)
 
 
 
 
 
673
  joined_verification = _verify_trajectory_audio(
674
  (waveform,),
675
  (whole_target_text,),
@@ -740,6 +822,7 @@ def _verify_sequence_trajectory_audio(
740
  anchor: np.ndarray,
741
  playback_speed: float,
742
  independent_cache: WholeWaveformVerificationCache,
 
743
  ):
744
  """Verify one ranked DP path after exact production assembly."""
745
 
@@ -747,6 +830,7 @@ def _verify_sequence_trajectory_audio(
747
  sequence_result.trajectory,
748
  chunks,
749
  playback_speed,
 
750
  )
751
  turbo_verification = _verify_trajectory_audio(
752
  (waveform,),
@@ -813,12 +897,13 @@ def _synthesize(
813
  raise gr.Error("CFG 必須介於 1.0 與 4.0。")
814
  if cfg_value != MIXED_CFG_PRIMARY:
815
  raise gr.Error(f"目前只支援已驗證的主 CFG {MIXED_CFG_PRIMARY:.1f}。")
816
- if network_identifier_has_ambiguous_iri(text):
 
817
  raise gr.Error("網址目前只支援 ASCII 字元;非 ASCII IRI 會與字母讀音混淆。")
818
- network_request = contains_network_identifier(text)
819
  request_cfg = MIXED_CFG_PRIMARY
820
  try:
821
- text = normalize_spoken_forms(text, locale="zh-TW")
822
  except ValueError as error:
823
  raise gr.Error(str(error)) from None
824
  if not text:
@@ -830,38 +915,66 @@ def _synthesize(
830
  if not np.isfinite(float(speed)) or not 0.85 <= float(speed) <= 1.05:
831
  raise gr.Error("後處理語速必須介於 0.85 與 1.05。")
832
 
833
- chunks = split_text_for_tts(text, max_chars=CHUNK_CHARS, min_chunk_chars=MIN_CHUNK_CHARS)
834
- chunks = coalesce_text_chunks(
835
- chunks,
836
- max_chunks=QUALITY_MAX_GENERATED_CHUNKS,
837
- max_units=CHUNK_CHARS,
838
- )
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
839
  request_seed = resolve_request_seed(request_seed, secrets.randbelow)
840
  anchor = _speaker_anchor_array(centroid)
841
  independent_cache = WholeWaveformVerificationCache()
842
 
843
- def candidate_cfg(seed: int) -> float:
844
  return generation_cfg_for_candidate_offset(
845
- seed - request_seed,
846
  primary_cfg=request_cfg,
847
  alternate_cfg=MIXED_CFG_ALTERNATE,
848
  )
849
 
850
  def chunk_cfg_evidence(
851
  chunk: str,
852
- candidate_offset: int,
 
 
853
  ) -> tuple[float, tuple[str, ...]]:
854
  scheduled = generation_cfg_for_candidate_offset(
855
- candidate_offset,
856
  primary_cfg=request_cfg,
857
  alternate_cfg=MIXED_CFG_ALTERNATE,
858
  )
859
  reasons: list[str] = []
860
  network_adjusted = scheduled
861
- if (
862
- network_protected_spoken_spans(chunk)
863
- and scheduled < NETWORK_TEXT_CFG_MIN
864
- ):
 
 
865
  network_adjusted = NETWORK_TEXT_CFG_MIN
866
  reasons.append("network")
867
  effective = effective_generation_cfg(
@@ -874,21 +987,105 @@ def _synthesize(
874
  reasons.append("short_text")
875
  return effective, tuple(reasons)
876
 
877
- def effective_chunk_cfg(chunk: str, candidate_offset: int) -> float:
878
- return chunk_cfg_evidence(chunk, candidate_offset)[0]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
879
 
880
  def candidate_generation_evidence(
881
  candidate_index: int,
882
  seed: int,
883
  chunk_indices: tuple[int, ...],
884
  candidate_chunks: tuple[str, ...],
 
 
885
  ) -> CandidateGenerationEvidence:
886
  if candidate_index != seed - request_seed:
887
  raise ValueError("candidate index does not match request seed offset")
888
- scheduled = candidate_cfg(seed)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
889
  rows = tuple(
890
- chunk_cfg_evidence(chunk, candidate_index)
891
- for chunk in candidate_chunks
 
 
 
 
 
 
 
 
892
  )
893
  return CandidateGenerationEvidence(
894
  chunk_indices=chunk_indices,
@@ -898,6 +1095,8 @@ def _synthesize(
898
  scheduled_cfg=scheduled,
899
  effective_cfgs=tuple(row[0] for row in rows),
900
  floor_reasons=tuple(row[1] for row in rows),
 
 
901
  )
902
 
903
  try:
@@ -905,14 +1104,7 @@ def _synthesize(
905
  cascade = run_coverage_adaptive_cascade(
906
  chunks,
907
  request_seed,
908
- lambda candidate_chunks, seed: _generate_trajectory(
909
- candidate_chunks,
910
- centroid,
911
- cfg=candidate_cfg(seed),
912
- steps=steps,
913
- request_seed=seed,
914
- policy=generation_policy_for_candidate_offset(seed - request_seed),
915
- ),
916
  lambda trajectory, candidate_chunks, seed: _qualify_candidate_trajectory_audio(
917
  trajectory,
918
  candidate_chunks,
@@ -921,6 +1113,7 @@ def _synthesize(
921
  speed,
922
  independent_cache,
923
  candidate_seed=seed,
 
924
  ),
925
  lambda trajectory, candidate_chunks, seed: (
926
  _qualify_candidate_trajectory_audio(
@@ -931,6 +1124,7 @@ def _synthesize(
931
  speed,
932
  independent_cache,
933
  candidate_seed=seed,
 
934
  )
935
  if len(chunks) == 1
936
  else _verify_refill_candidate_trajectory_audio(
@@ -948,6 +1142,7 @@ def _synthesize(
948
  anchor,
949
  speed,
950
  independent_cache,
 
951
  )
952
  ),
953
  generation_evidence_factory=candidate_generation_evidence,
@@ -971,24 +1166,49 @@ def _synthesize(
971
  except (RuntimeError, ValueError) as error:
972
  raise gr.Error("品質驗證暫時無法完成,未回傳未驗證的語音。") from error
973
 
 
 
 
 
 
 
 
 
 
 
 
 
 
974
  selected_policies = tuple(
975
- generation_policy_for_candidate_offset(index).name
976
- for index in cascade.chunk_candidate_indices
977
  )
978
  attempted_policies = tuple(
979
- generation_policy_for_candidate_offset(seed - request_seed).name
980
- for seed in cascade.attempted_seeds
 
 
 
 
 
 
981
  )
982
  selected_cfgs = tuple(
983
- effective_chunk_cfg(chunk, candidate_index)
984
- for chunk, candidate_index in zip(
 
 
 
 
985
  chunks,
986
- cascade.chunk_candidate_indices,
 
987
  strict=True,
988
  )
989
  )
990
  attempted_schedule_cfgs = tuple(
991
- candidate_cfg(seed) for seed in cascade.attempted_seeds
 
992
  )
993
  print(
994
  "[BlueMagpie] quality cascade "
@@ -1009,7 +1229,12 @@ def _synthesize(
1009
  f" cfg_schedule={MIXED_CFG_SCHEDULE}"
1010
  f" network_request={network_request}"
1011
  )
1012
- waveform = _assemble_trajectory_audio(cascade.trajectory, chunks, speed)
 
 
 
 
 
1013
  final_verification = _verify_trajectory_audio(
1014
  (waveform,),
1015
  (text,),
 
18
  from bluemagpie import BlueMagpieModel
19
  from production import (
20
  StopHysteresisController,
21
+ GenerationChunkSpec,
22
  active_pace_correction_speed,
23
  apply_loudness_floor,
24
  coalesce_text_chunks,
 
28
  endpoint_generation_plan,
29
  estimate_step_seconds,
30
  extract_windowed_speaker_embedding,
31
+ fade_variable_internal_edges,
32
  finish_audio,
33
+ join_audio_chunks_variable,
34
  match_chunk_rms,
35
  network_identifier_has_ambiguous_iri,
36
  network_protected_spoken_spans,
37
  normalize_spoken_forms,
38
  punctuation_pause_seconds,
39
+ plan_generation_chunks,
40
  select_generation_cps,
41
  set_generation_seed,
42
  split_text_for_tts,
 
49
  VERIFICATION_WHISPER_REVISION,
50
  WHISPER_MODEL_ID,
51
  WHISPER_REVISION,
52
+ CandidateGenerationContext,
53
  CandidateGenerationEvidence,
54
  CandidateVerification,
55
  CandidateObservation,
 
104
  ACTIVE_PACE_TARGET_CPS = 4.00
105
  MIXED_CFG_PRIMARY = 3.0
106
  MIXED_CFG_ALTERNATE = 2.0
107
+ MIXED_CFG_SCHEDULE = "row_ordinal_zero_and_even_primary_odd_alternate"
108
  MIN_ENDPOINT_CUE_UNITS = 6
109
  SHORT_TEXT_CFG_MIN = 3.0
110
  SHORT_TEXT_CFG_UNITS = 6
 
115
  CROSSFADE_MS = 80.0
116
  CHUNK_EDGE_FADE_MS = 80.0
117
  CHUNK_RMS_MATCH_DB = 4.0
118
+ NETWORK_GENERATION_TARGET_UNITS = 32
119
+ NETWORK_GENERATION_MAX_UNITS = 36
120
+ NETWORK_INTERNAL_FADE_MS = 5.0
121
  STOP_THRESHOLD = 0.50
122
  STOP_LATE_THRESHOLD = 0.05
123
  STOP_LATE_START_RATIO = 0.75
 
265
  steps: int,
266
  request_seed: int,
267
  policy: GenerationPolicy,
268
+ network_conditioned: bool = False,
269
  ) -> np.ndarray:
270
+ generation_cps = (
271
+ policy.ascii_cps
272
+ if network_conditioned
273
+ else select_generation_cps(
274
+ text,
275
+ cjk_cps=policy.cjk_cps,
276
+ ascii_cps=policy.ascii_cps,
277
+ )
278
  )
279
  model_text, expected_steps, hard_stop_steps = endpoint_generation_plan(
280
  text,
 
383
  request_seed: int,
384
  policy: GenerationPolicy,
385
  network_cfg_min: float = NETWORK_TEXT_CFG_MIN,
386
+ network_conditioned: tuple[bool, ...] | None = None,
387
  ) -> tuple[np.ndarray, ...]:
388
  scheduled_cfg = float(cfg)
389
  trajectory: list[np.ndarray] = []
390
+ if network_conditioned is not None and len(network_conditioned) != len(chunks):
391
+ raise ValueError("network provenance must align with generation chunks")
392
  for local_chunk_index, chunk in enumerate(chunks):
393
+ network_chunk = (
394
+ bool(network_conditioned[local_chunk_index])
395
+ if network_conditioned is not None
396
+ else bool(network_protected_spoken_spans(chunk))
397
+ )
398
  network_floor_applied = bool(
399
  network_chunk and scheduled_cfg < float(network_cfg_min)
400
  )
 
426
  steps=steps,
427
  request_seed=request_seed,
428
  policy=policy,
429
+ network_conditioned=network_chunk,
430
  )
431
  )
432
  return tuple(trajectory)
 
615
  trajectory: tuple[np.ndarray, ...],
616
  chunks: tuple[str, ...],
617
  playback_speed: float,
618
+ chunk_specs: tuple[GenerationChunkSpec, ...] | None = None,
619
  ) -> np.ndarray:
620
  """Assemble chunks exactly as they will be returned to the listener."""
621
 
622
  if not trajectory or len(trajectory) != len(chunks):
623
  raise ValueError("trajectory and text chunks must be non-empty and aligned")
624
+ if chunk_specs is not None and (
625
+ len(chunk_specs) != len(chunks)
626
+ or any(spec.text != chunk for spec, chunk in zip(chunk_specs, chunks, strict=True))
627
+ ):
628
+ raise ValueError("generation chunk provenance does not align with text")
629
+ if chunk_specs is not None:
630
+ for index, spec in enumerate(chunk_specs):
631
+ if (
632
+ spec.source_start < 0
633
+ or spec.source_end <= spec.source_start
634
+ or spec.boundary_after
635
+ not in {"semantic", "network_internal", "none"}
636
+ ):
637
+ raise ValueError("generation chunk provenance is invalid")
638
+ is_last = index + 1 == len(chunk_specs)
639
+ if is_last:
640
+ if spec.boundary_after != "none":
641
+ raise ValueError("final generation boundary must be none")
642
+ continue
643
+ following = chunk_specs[index + 1]
644
+ if (
645
+ spec.source_end != following.source_start
646
+ or spec.boundary_after == "none"
647
+ ):
648
+ raise ValueError("generation chunk boundaries are incomplete")
649
+ if spec.boundary_after == "network_internal" and (
650
+ not spec.network_conditioned
651
+ or not following.network_conditioned
652
+ or not set(spec.network_span_indices).intersection(
653
+ following.network_span_indices
654
+ )
655
+ ):
656
+ raise ValueError("network boundary lacks shared identifier proof")
657
  audio_chunks = [np.asarray(audio, dtype=np.float32).copy() for audio in trajectory]
658
  pauses: list[int] = []
659
+ fades_ms: list[float] = []
660
+ crossfades_ms: list[float] = []
661
  for index, chunk in enumerate(chunks):
662
  if index > 0:
663
  audio_chunks[index] = match_chunk_rms(
 
666
  max_adjust_db=CHUNK_RMS_MATCH_DB,
667
  )
668
  if index + 1 < len(chunks):
669
+ network_internal = bool(
670
+ chunk_specs is not None
671
+ and chunk_specs[index].boundary_after == "network_internal"
672
+ )
673
+ pauses.append(
674
+ 0
675
+ if network_internal
676
+ else int(round(punctuation_pause_seconds(chunk) * SR))
677
+ )
678
+ fades_ms.append(
679
+ NETWORK_INTERNAL_FADE_MS
680
+ if network_internal
681
+ else CHUNK_EDGE_FADE_MS
682
+ )
683
+ crossfades_ms.append(
684
+ NETWORK_INTERNAL_FADE_MS
685
+ if network_internal
686
+ else CROSSFADE_MS
687
+ )
688
 
689
+ audio_chunks = fade_variable_internal_edges(audio_chunks, SR, fades_ms)
690
+ waveform = join_audio_chunks_variable(
691
  audio_chunks,
692
  pauses,
693
+ crossfade_samples_by_boundary=[
694
+ int(round(crossfade_ms * SR / 1000.0))
695
+ for crossfade_ms in crossfades_ms
696
+ ],
697
  pre_faded_edges=True,
698
  )
699
  waveform = apply_loudness_floor(
 
733
  independent_cache: WholeWaveformVerificationCache,
734
  *,
735
  candidate_seed: int,
736
+ chunk_specs: tuple[GenerationChunkSpec, ...] | None = None,
737
  ):
738
  """Run whole-output qualification only after every local chunk passes."""
739
 
 
746
  if not local_verification.passed:
747
  return CandidateVerification(local_verification)
748
 
749
+ waveform = _assemble_trajectory_audio(
750
+ trajectory,
751
+ chunks,
752
+ playback_speed,
753
+ chunk_specs,
754
+ )
755
  joined_verification = _verify_trajectory_audio(
756
  (waveform,),
757
  (whole_target_text,),
 
822
  anchor: np.ndarray,
823
  playback_speed: float,
824
  independent_cache: WholeWaveformVerificationCache,
825
+ chunk_specs: tuple[GenerationChunkSpec, ...] | None = None,
826
  ):
827
  """Verify one ranked DP path after exact production assembly."""
828
 
 
830
  sequence_result.trajectory,
831
  chunks,
832
  playback_speed,
833
+ chunk_specs,
834
  )
835
  turbo_verification = _verify_trajectory_audio(
836
  (waveform,),
 
897
  raise gr.Error("CFG 必須介於 1.0 與 4.0。")
898
  if cfg_value != MIXED_CFG_PRIMARY:
899
  raise gr.Error(f"目前只支援已驗證的主 CFG {MIXED_CFG_PRIMARY:.1f}。")
900
+ raw_text = str(text)
901
+ if network_identifier_has_ambiguous_iri(raw_text):
902
  raise gr.Error("網址目前只支援 ASCII 字元;非 ASCII IRI 會與字母讀音混淆。")
903
+ network_request = contains_network_identifier(raw_text)
904
  request_cfg = MIXED_CFG_PRIMARY
905
  try:
906
+ text = normalize_spoken_forms(raw_text, locale="zh-TW")
907
  except ValueError as error:
908
  raise gr.Error(str(error)) from None
909
  if not text:
 
915
  if not np.isfinite(float(speed)) or not 0.85 <= float(speed) <= 1.05:
916
  raise gr.Error("後處理語速必須介於 0.85 與 1.05。")
917
 
918
+ chunk_specs: tuple[GenerationChunkSpec, ...] | None = None
919
+ try:
920
+ if network_request:
921
+ chunk_specs = plan_generation_chunks(
922
+ raw_text,
923
+ text,
924
+ min_units=MIN_CHUNK_CHARS,
925
+ target_units=NETWORK_GENERATION_TARGET_UNITS,
926
+ network_max_units=NETWORK_GENERATION_MAX_UNITS,
927
+ ordinary_max_units=CHUNK_CHARS,
928
+ )
929
+ chunks = tuple(spec.text for spec in chunk_specs)
930
+ if len(chunks) > QUALITY_MAX_GENERATED_CHUNKS:
931
+ raise ValueError(
932
+ "network request exceeds the generated-chunk budget"
933
+ )
934
+ else:
935
+ chunks = tuple(
936
+ coalesce_text_chunks(
937
+ split_text_for_tts(
938
+ text,
939
+ max_chars=CHUNK_CHARS,
940
+ min_chunk_chars=MIN_CHUNK_CHARS,
941
+ ),
942
+ max_chunks=QUALITY_MAX_GENERATED_CHUNKS,
943
+ max_units=CHUNK_CHARS,
944
+ )
945
+ )
946
+ except ValueError as error:
947
+ raise gr.Error("文字無法在已驗證的生成限制內安全分段。") from error
948
  request_seed = resolve_request_seed(request_seed, secrets.randbelow)
949
  anchor = _speaker_anchor_array(centroid)
950
  independent_cache = WholeWaveformVerificationCache()
951
 
952
+ def candidate_cfg(candidate_ordinal: int) -> float:
953
  return generation_cfg_for_candidate_offset(
954
+ candidate_ordinal,
955
  primary_cfg=request_cfg,
956
  alternate_cfg=MIXED_CFG_ALTERNATE,
957
  )
958
 
959
  def chunk_cfg_evidence(
960
  chunk: str,
961
+ candidate_ordinal: int,
962
+ *,
963
+ network_conditioned: bool | None = None,
964
  ) -> tuple[float, tuple[str, ...]]:
965
  scheduled = generation_cfg_for_candidate_offset(
966
+ candidate_ordinal,
967
  primary_cfg=request_cfg,
968
  alternate_cfg=MIXED_CFG_ALTERNATE,
969
  )
970
  reasons: list[str] = []
971
  network_adjusted = scheduled
972
+ is_network = (
973
+ bool(network_protected_spoken_spans(chunk))
974
+ if network_conditioned is None
975
+ else bool(network_conditioned)
976
+ )
977
+ if is_network and scheduled < NETWORK_TEXT_CFG_MIN:
978
  network_adjusted = NETWORK_TEXT_CFG_MIN
979
  reasons.append("network")
980
  effective = effective_generation_cfg(
 
987
  reasons.append("short_text")
988
  return effective, tuple(reasons)
989
 
990
+ def generation_network_flags(
991
+ chunk_indices: tuple[int, ...],
992
+ candidate_chunks: tuple[str, ...],
993
+ ) -> tuple[bool, ...]:
994
+ if len(chunk_indices) != len(candidate_chunks):
995
+ raise ValueError("generation provenance does not align with chunks")
996
+ if chunk_specs is None:
997
+ return tuple(
998
+ bool(network_protected_spoken_spans(chunk))
999
+ for chunk in candidate_chunks
1000
+ )
1001
+ flags: list[bool] = []
1002
+ for chunk_index, chunk in zip(
1003
+ chunk_indices,
1004
+ candidate_chunks,
1005
+ strict=True,
1006
+ ):
1007
+ if not 0 <= chunk_index < len(chunk_specs):
1008
+ raise ValueError("generation provenance index is out of range")
1009
+ spec = chunk_specs[chunk_index]
1010
+ if spec.text != chunk:
1011
+ raise ValueError("generation provenance text does not match")
1012
+ flags.append(spec.network_conditioned)
1013
+ return tuple(flags)
1014
+
1015
+ def effective_chunk_cfg(
1016
+ chunk: str,
1017
+ candidate_ordinal: int,
1018
+ *,
1019
+ network_conditioned: bool | None = None,
1020
+ ) -> float:
1021
+ return chunk_cfg_evidence(
1022
+ chunk,
1023
+ candidate_ordinal,
1024
+ network_conditioned=network_conditioned,
1025
+ )[0]
1026
+
1027
+ def candidate_generation(
1028
+ candidate_chunks: tuple[str, ...],
1029
+ seed: int,
1030
+ *,
1031
+ generation_context: CandidateGenerationContext,
1032
+ ) -> tuple[np.ndarray, ...]:
1033
+ if generation_context.seed != seed:
1034
+ raise ValueError("generation context seed does not match the request seed")
1035
+ ordinals = generation_context.chunk_candidate_ordinals
1036
+ if not ordinals or len(set(ordinals)) != 1:
1037
+ raise ValueError("one generation call must use one candidate ordinal")
1038
+ candidate_ordinal = ordinals[0]
1039
+ network_flags = generation_network_flags(
1040
+ generation_context.chunk_indices,
1041
+ candidate_chunks,
1042
+ )
1043
+ return _generate_trajectory(
1044
+ candidate_chunks,
1045
+ centroid,
1046
+ cfg=candidate_cfg(candidate_ordinal),
1047
+ steps=steps,
1048
+ request_seed=seed,
1049
+ policy=generation_policy_for_candidate_offset(candidate_ordinal),
1050
+ network_conditioned=network_flags,
1051
+ )
1052
 
1053
  def candidate_generation_evidence(
1054
  candidate_index: int,
1055
  seed: int,
1056
  chunk_indices: tuple[int, ...],
1057
  candidate_chunks: tuple[str, ...],
1058
+ *,
1059
+ generation_context: CandidateGenerationContext,
1060
  ) -> CandidateGenerationEvidence:
1061
  if candidate_index != seed - request_seed:
1062
  raise ValueError("candidate index does not match request seed offset")
1063
+ if (
1064
+ generation_context.candidate_index != candidate_index
1065
+ or generation_context.seed != seed
1066
+ or generation_context.chunk_indices != chunk_indices
1067
+ ):
1068
+ raise ValueError("candidate generation context does not match the attempt")
1069
+ candidate_ordinals = generation_context.chunk_candidate_ordinals
1070
+ if not candidate_ordinals or len(set(candidate_ordinals)) != 1:
1071
+ raise ValueError("one generation call must use one candidate ordinal")
1072
+ candidate_ordinal = candidate_ordinals[0]
1073
+ scheduled = candidate_cfg(candidate_ordinal)
1074
+ network_flags = generation_network_flags(
1075
+ chunk_indices,
1076
+ candidate_chunks,
1077
+ )
1078
  rows = tuple(
1079
+ chunk_cfg_evidence(
1080
+ chunk,
1081
+ candidate_ordinal,
1082
+ network_conditioned=network_flag,
1083
+ )
1084
+ for chunk, network_flag in zip(
1085
+ candidate_chunks,
1086
+ network_flags,
1087
+ strict=True,
1088
+ )
1089
  )
1090
  return CandidateGenerationEvidence(
1091
  chunk_indices=chunk_indices,
 
1095
  scheduled_cfg=scheduled,
1096
  effective_cfgs=tuple(row[0] for row in rows),
1097
  floor_reasons=tuple(row[1] for row in rows),
1098
+ chunk_candidate_ordinals=candidate_ordinals,
1099
+ network_conditioned=network_flags,
1100
  )
1101
 
1102
  try:
 
1104
  cascade = run_coverage_adaptive_cascade(
1105
  chunks,
1106
  request_seed,
1107
+ candidate_generation,
 
 
 
 
 
 
 
1108
  lambda trajectory, candidate_chunks, seed: _qualify_candidate_trajectory_audio(
1109
  trajectory,
1110
  candidate_chunks,
 
1113
  speed,
1114
  independent_cache,
1115
  candidate_seed=seed,
1116
+ chunk_specs=chunk_specs,
1117
  ),
1118
  lambda trajectory, candidate_chunks, seed: (
1119
  _qualify_candidate_trajectory_audio(
 
1124
  speed,
1125
  independent_cache,
1126
  candidate_seed=seed,
1127
+ chunk_specs=chunk_specs,
1128
  )
1129
  if len(chunks) == 1
1130
  else _verify_refill_candidate_trajectory_audio(
 
1142
  anchor,
1143
  speed,
1144
  independent_cache,
1145
+ chunk_specs,
1146
  )
1147
  ),
1148
  generation_evidence_factory=candidate_generation_evidence,
 
1166
  except (RuntimeError, ValueError) as error:
1167
  raise gr.Error("品質驗證暫時無法完成,未回傳未驗證的語音。") from error
1168
 
1169
+ attempts_by_index = {
1170
+ attempt.candidate_index: attempt for attempt in cascade.diagnostics.attempts
1171
+ }
1172
+ selected_ordinals: list[int] = []
1173
+ for chunk_index, candidate_index in enumerate(cascade.chunk_candidate_indices):
1174
+ attempt = attempts_by_index.get(candidate_index)
1175
+ if attempt is None or chunk_index not in attempt.chunk_indices:
1176
+ raise gr.Error("品質驗證紀錄不完整,未回傳未驗證的語音。")
1177
+ local_index = attempt.chunk_indices.index(chunk_index)
1178
+ try:
1179
+ selected_ordinals.append(attempt.chunk_candidate_ordinals[local_index])
1180
+ except IndexError as error:
1181
+ raise gr.Error("品質驗證紀錄不完整,未回傳未驗證的語音。") from error
1182
  selected_policies = tuple(
1183
+ generation_policy_for_candidate_offset(ordinal).name
1184
+ for ordinal in selected_ordinals
1185
  )
1186
  attempted_policies = tuple(
1187
+ generation_policy_for_candidate_offset(
1188
+ attempt.chunk_candidate_ordinals[0]
1189
+ ).name
1190
+ for attempt in cascade.diagnostics.attempts
1191
+ )
1192
+ selected_network_flags = generation_network_flags(
1193
+ tuple(range(len(chunks))),
1194
+ chunks,
1195
  )
1196
  selected_cfgs = tuple(
1197
+ effective_chunk_cfg(
1198
+ chunk,
1199
+ candidate_ordinal,
1200
+ network_conditioned=network_flag,
1201
+ )
1202
+ for chunk, candidate_ordinal, network_flag in zip(
1203
  chunks,
1204
+ selected_ordinals,
1205
+ selected_network_flags,
1206
  strict=True,
1207
  )
1208
  )
1209
  attempted_schedule_cfgs = tuple(
1210
+ candidate_cfg(attempt.chunk_candidate_ordinals[0])
1211
+ for attempt in cascade.diagnostics.attempts
1212
  )
1213
  print(
1214
  "[BlueMagpie] quality cascade "
 
1229
  f" cfg_schedule={MIXED_CFG_SCHEDULE}"
1230
  f" network_request={network_request}"
1231
  )
1232
+ waveform = _assemble_trajectory_audio(
1233
+ cascade.trajectory,
1234
+ chunks,
1235
+ speed,
1236
+ chunk_specs,
1237
+ )
1238
  final_verification = _verify_trajectory_audio(
1239
  (waveform,),
1240
  (text,),
production.py CHANGED
@@ -3,6 +3,7 @@
3
  from __future__ import annotations
4
 
5
  import math
 
6
  import random
7
  import re
8
  import unicodedata
@@ -1139,7 +1140,7 @@ def network_identifier_has_ambiguous_iri(text: str) -> bool:
1139
 
1140
  raw = unicodedata.normalize("NFC", str(text or ""))
1141
  return any(
1142
- any(character.isalnum() and not character.isascii() for character in match.group(0))
1143
  for match in _SPOKEN_URL_RE.finditer(raw)
1144
  )
1145
 
@@ -2332,6 +2333,10 @@ _PROTECTED_ASCII_SPAN_RE = re.compile(
2332
  )
2333
  _STRUCTURED_CLAUSE_BREAKS = frozenset(",,、::")
2334
  _STRUCTURED_CLAUSE_TARGET_UNITS = 32
 
 
 
 
2335
  _NORMALIZED_ZH_NUMBER_PATTERN = (
2336
  r"[正負]?[" + _ZH_DIGITS + r"十百千萬億兆]+"
2337
  r"(?:點[" + _ZH_DIGITS + r"]+)?"
@@ -2655,6 +2660,515 @@ def split_text_for_tts(text: str, max_chars: int = 80, min_chunk_chars: int = 12
2655
  return chunks or [text]
2656
 
2657
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2658
  def coalesce_text_chunks(
2659
  chunks: Sequence[str],
2660
  *,
@@ -2761,6 +3275,57 @@ def fade_internal_edges(chunks: list[np.ndarray], sample_rate: int, fade_ms: flo
2761
  return outputs
2762
 
2763
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2764
  def join_audio_chunks(
2765
  chunks: list[np.ndarray],
2766
  pauses: list[int],
@@ -2792,6 +3357,105 @@ def join_audio_chunks(
2792
  return output.astype(np.float32, copy=False)
2793
 
2794
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2795
  def apply_loudness_floor(
2796
  audio: np.ndarray,
2797
  min_rms: float = 0.07,
 
3
  from __future__ import annotations
4
 
5
  import math
6
+ import operator
7
  import random
8
  import re
9
  import unicodedata
 
1140
 
1141
  raw = unicodedata.normalize("NFC", str(text or ""))
1142
  return any(
1143
+ not match.group(0).isascii()
1144
  for match in _SPOKEN_URL_RE.finditer(raw)
1145
  )
1146
 
 
2333
  )
2334
  _STRUCTURED_CLAUSE_BREAKS = frozenset(",,、::")
2335
  _STRUCTURED_CLAUSE_TARGET_UNITS = 32
2336
+ _NETWORK_GENERATION_COMPONENT_AFTER = frozenset()
2337
+ _NETWORK_GENERATION_COMPONENT_BEFORE = frozenset(
2338
+ {"點", "小老鼠", "斜線", "問號", "井號", "冒號", "和"}
2339
+ )
2340
  _NORMALIZED_ZH_NUMBER_PATTERN = (
2341
  r"[正負]?[" + _ZH_DIGITS + r"十百千萬億兆]+"
2342
  r"(?:點[" + _ZH_DIGITS + r"]+)?"
 
2660
  return chunks or [text]
2661
 
2662
 
2663
+ @dataclass(frozen=True)
2664
+ class GenerationChunkSpec:
2665
+ """One generation-only slice with immutable full-identifier provenance."""
2666
+
2667
+ text: str
2668
+ source_start: int
2669
+ source_end: int
2670
+ network_span_indices: tuple[int, ...] = ()
2671
+ network_component_indices: tuple[tuple[int, int], ...] = ()
2672
+ network_full_spoken_proofs: tuple[str, ...] = ()
2673
+ boundary_after: str = "none"
2674
+
2675
+ @property
2676
+ def network_conditioned(self) -> bool:
2677
+ return bool(self.network_span_indices)
2678
+
2679
+
2680
+ @dataclass(frozen=True)
2681
+ class _NetworkGenerationSpan:
2682
+ start: int
2683
+ end: int
2684
+ spoken_proof: str
2685
+ component_ranges: tuple[tuple[int, int], ...]
2686
+
2687
+
2688
+ def _raw_network_generation_matches(
2689
+ text: str,
2690
+ ) -> tuple[tuple[int, int, str, str], ...]:
2691
+ """Return non-overlapping raw identifiers and their exact spoken proofs."""
2692
+
2693
+ raw = unicodedata.normalize("NFC", str(text or ""))
2694
+ matches: list[tuple[int, int, str, str]] = []
2695
+ for match in _SPOKEN_URL_RE.finditer(raw):
2696
+ value = match.group(0)
2697
+ if not value.isascii():
2698
+ raise ValueError("network generation does not support non-ASCII IRI")
2699
+ matches.append((match.start(), match.end(), _zh_url(value), "url"))
2700
+ for match in _SPOKEN_EMAIL_RE.finditer(raw):
2701
+ value = match.group(0)
2702
+ if not value.isascii():
2703
+ raise ValueError("network generation does not support non-ASCII email")
2704
+ matches.append((match.start(), match.end(), _zh_email(value), "email"))
2705
+
2706
+ selected: list[tuple[int, int, str, str]] = []
2707
+ for start, end, spoken, kind in sorted(
2708
+ matches,
2709
+ key=lambda item: (item[0], -item[1]),
2710
+ ):
2711
+ if any(
2712
+ start < kept_end and end > kept_start
2713
+ for kept_start, kept_end, _, _ in selected
2714
+ ):
2715
+ continue
2716
+ selected.append((start, end, normalize_tts_text(spoken), kind))
2717
+ return tuple(sorted(selected))
2718
+
2719
+
2720
+ def _network_component_ranges(
2721
+ spoken: str,
2722
+ *,
2723
+ absolute_start: int,
2724
+ hard_max_units: int,
2725
+ ) -> tuple[tuple[int, int], ...]:
2726
+ """Split one proven identifier only at audible grammar delimiters."""
2727
+
2728
+ token_matches = tuple(re.finditer(r"\S+", spoken))
2729
+ tokens = tuple(
2730
+ _SIMPLIFIED_NETWORK_SYMBOL_READINGS.get(match.group(0), match.group(0))
2731
+ for match in token_matches
2732
+ )
2733
+ cuts: set[int] = {0, len(spoken)}
2734
+ scheme_slashes: set[int] = set()
2735
+ for index in range(2, len(tokens)):
2736
+ if tokens[index - 2 : index + 1] == ("冒號", "斜線", "斜線"):
2737
+ scheme_slashes.update({index - 1, index})
2738
+ cuts.add(token_matches[index].end())
2739
+ for index, (token, match) in enumerate(zip(tokens, token_matches, strict=True)):
2740
+ if token in _NETWORK_GENERATION_COMPONENT_AFTER:
2741
+ cuts.add(match.end())
2742
+ if (
2743
+ token in _NETWORK_GENERATION_COMPONENT_BEFORE
2744
+ and index not in scheme_slashes
2745
+ and not (token == "冒號" and index + 2 < len(tokens) and tokens[index + 1 : index + 3] == ("斜線", "斜線"))
2746
+ ):
2747
+ cuts.add(match.start())
2748
+
2749
+ ordered = sorted(cuts)
2750
+ ranges: list[tuple[int, int]] = []
2751
+ for left, right in zip(ordered, ordered[1:]):
2752
+ component = normalize_tts_text(spoken[left:right])
2753
+ if not component:
2754
+ continue
2755
+ if count_speech_units(component) > hard_max_units:
2756
+ raise ValueError("an indivisible network component exceeds the generation limit")
2757
+ ranges.append((absolute_start + left, absolute_start + right))
2758
+ if not ranges:
2759
+ raise ValueError("network generation proof has no components")
2760
+ if ranges[0][0] != absolute_start or ranges[-1][1] != absolute_start + len(spoken):
2761
+ raise ValueError("network generation components do not cover their proof")
2762
+ return tuple(ranges)
2763
+
2764
+
2765
+ def _network_generation_spans(
2766
+ raw_text: str,
2767
+ normalized_text: str,
2768
+ *,
2769
+ hard_max_units: int,
2770
+ ) -> tuple[_NetworkGenerationSpan, ...]:
2771
+ """Bind raw ASCII grammar to exact offsets in the normalized model text."""
2772
+
2773
+ normalized = normalize_tts_text(normalized_text)
2774
+ expected = normalize_spoken_forms(raw_text)
2775
+ if normalized != expected:
2776
+ raise ValueError("normalized text does not match the raw network request")
2777
+ raw = unicodedata.normalize("NFC", str(raw_text or ""))
2778
+ raw_matches = _raw_network_generation_matches(raw)
2779
+ spans: list[_NetworkGenerationSpan] = []
2780
+ previous_raw_end: int | None = None
2781
+ previous_kind: str | None = None
2782
+ cursor = 0
2783
+ for raw_start, raw_end, spoken, kind in raw_matches:
2784
+ # Derive the offset from the exact raw prefix rather than searching for
2785
+ # a possibly duplicated spoken phrase. The frontend inserts a comma
2786
+ # on each unpunctuated side of a network identifier; a preceding
2787
+ # identifier at the end of the prefix needs its deferred suffix comma
2788
+ # accounted for as well.
2789
+ normalized_prefix = normalize_spoken_forms(raw[:raw_start])
2790
+ boundary_count = 0
2791
+ preceding = raw[:raw_start].rstrip()[-1:]
2792
+ if preceding and preceding not in _NETWORK_BOUNDARY_PUNCTUATION:
2793
+ boundary_count += 1
2794
+ if (
2795
+ previous_raw_end is not None
2796
+ and previous_kind == kind
2797
+ and not raw[previous_raw_end:raw_start].strip()
2798
+ ):
2799
+ following = raw[previous_raw_end:].lstrip()[:1]
2800
+ if following and following not in _NETWORK_BOUNDARY_PUNCTUATION:
2801
+ boundary_count += 1
2802
+ start = len(normalized_prefix) + boundary_count
2803
+ if start < cursor or normalized[start : start + len(spoken)] != spoken:
2804
+ raise ValueError(
2805
+ "network spoken proof offset does not match normalized text"
2806
+ )
2807
+ end = start + len(spoken)
2808
+ components = _network_component_ranges(
2809
+ spoken,
2810
+ absolute_start=start,
2811
+ hard_max_units=hard_max_units,
2812
+ )
2813
+ spans.append(
2814
+ _NetworkGenerationSpan(
2815
+ start=start,
2816
+ end=end,
2817
+ spoken_proof=spoken,
2818
+ component_ranges=components,
2819
+ )
2820
+ )
2821
+ cursor = end
2822
+ previous_raw_end = raw_end
2823
+ previous_kind = kind
2824
+ return tuple(spans)
2825
+
2826
+
2827
+ def _network_generation_group_plan(
2828
+ text: str,
2829
+ *,
2830
+ group_start: int,
2831
+ group_end: int,
2832
+ original_boundaries: Sequence[int],
2833
+ spans: Sequence[_NetworkGenerationSpan],
2834
+ minimum: int,
2835
+ target: int,
2836
+ network_maximum: int,
2837
+ ordinary_maximum: int,
2838
+ ) -> tuple[tuple[int, int], ...]:
2839
+ """Deterministically pack one network-bearing sentence near the target."""
2840
+
2841
+ group_spans = tuple(
2842
+ span for span in spans if span.start < group_end and span.end > group_start
2843
+ )
2844
+ if not group_spans:
2845
+ return tuple(
2846
+ zip(
2847
+ (group_start, *original_boundaries),
2848
+ (*original_boundaries, group_end),
2849
+ strict=True,
2850
+ )
2851
+ )
2852
+
2853
+ preferred_cuts: set[int] = {
2854
+ group_start,
2855
+ group_end,
2856
+ *original_boundaries,
2857
+ }
2858
+ component_cuts = {
2859
+ offset
2860
+ for span in group_spans
2861
+ for component in span.component_ranges
2862
+ for offset in component
2863
+ }
2864
+ preferred_cuts.update(component_cuts)
2865
+ network_interiors = {
2866
+ offset
2867
+ for span in group_spans
2868
+ for offset in range(span.start + 1, span.end)
2869
+ }
2870
+ for index in range(group_start, group_end):
2871
+ end = index + 1
2872
+ if text[index] in _STRUCTURED_CLAUSE_BREAKS and end not in network_interiors:
2873
+ preferred_cuts.add(end)
2874
+
2875
+ # Every source offset outside an identifier remains an emergency prose
2876
+ # split. Inside an identifier, only exact component boundaries are legal.
2877
+ cuts = {
2878
+ offset
2879
+ for offset in range(group_start, group_end + 1)
2880
+ if offset not in network_interiors or offset in component_cuts
2881
+ }
2882
+
2883
+ ordered = sorted(offset for offset in cuts if group_start <= offset <= group_end)
2884
+ best: dict[int, tuple[int, int, int, int, tuple[int, ...]]] = {
2885
+ group_end: (0, 0, 0, 0, ())
2886
+ }
2887
+ for start in reversed(ordered[:-1]):
2888
+ selected: tuple[int, int, int, int, tuple[int, ...]] | None = None
2889
+ for end in ordered:
2890
+ if end <= start:
2891
+ continue
2892
+ chunk = text[start:end]
2893
+ units = count_speech_units(chunk)
2894
+ network_conditioned = any(
2895
+ span.start < end and span.end > start for span in group_spans
2896
+ )
2897
+ maximum = network_maximum if network_conditioned else ordinary_maximum
2898
+ if units > maximum:
2899
+ continue
2900
+ if units < minimum and not (start == group_start and end == group_end):
2901
+ continue
2902
+ remainder = best.get(end)
2903
+ if remainder is None:
2904
+ continue
2905
+ candidate = (
2906
+ 1 + remainder[0],
2907
+ int(end != group_end and end not in preferred_cuts) + remainder[1],
2908
+ max(units, remainder[2]),
2909
+ abs(target - units) + remainder[3],
2910
+ (end,) + remainder[4],
2911
+ )
2912
+ if selected is None or candidate < selected:
2913
+ selected = candidate
2914
+ if selected is not None:
2915
+ best[start] = selected
2916
+ plan = best.get(group_start)
2917
+ if plan is None:
2918
+ raise ValueError("network-bearing sentence cannot satisfy generation limits")
2919
+ output: list[tuple[int, int]] = []
2920
+ start = group_start
2921
+ for end in plan[4]:
2922
+ output.append((start, end))
2923
+ start = end
2924
+ return tuple(output)
2925
+
2926
+
2927
+ def _semantic_sentence_ranges(text: str) -> tuple[tuple[int, int], ...]:
2928
+ """Return exact, contiguous strong-sentence source ranges."""
2929
+
2930
+ if not text:
2931
+ return ()
2932
+ ranges: list[tuple[int, int]] = []
2933
+ start = 0
2934
+ index = 0
2935
+ while index < len(text):
2936
+ if text[index] not in _SEMANTIC_STRONG_BREAKS:
2937
+ index += 1
2938
+ continue
2939
+ end = index + 1
2940
+ while end < len(text) and text[end] in _SEMANTIC_TRAILING_CLOSERS:
2941
+ end += 1
2942
+ while end < len(text) and text[end].isspace():
2943
+ end += 1
2944
+ if count_speech_units(text[start:end]) > 0:
2945
+ ranges.append((start, end))
2946
+ start = end
2947
+ index = end
2948
+ if start < len(text) and count_speech_units(text[start:]) > 0:
2949
+ ranges.append((start, len(text)))
2950
+ if not ranges or ranges[0][0] != 0 or ranges[-1][1] != len(text):
2951
+ raise ValueError("semantic source ranges do not cover normalized text")
2952
+ if any(
2953
+ left_end != right_start
2954
+ for (_, left_end), (right_start, _) in zip(ranges, ranges[1:])
2955
+ ):
2956
+ raise ValueError("semantic source ranges are not contiguous")
2957
+ return tuple(ranges)
2958
+
2959
+
2960
+ def _ordinary_generation_ranges(
2961
+ text: str,
2962
+ *,
2963
+ start: int,
2964
+ end: int,
2965
+ minimum: int,
2966
+ maximum: int,
2967
+ ) -> tuple[tuple[int, int], ...]:
2968
+ """Hard-split one exact non-network source range without re-normalizing it."""
2969
+
2970
+ source = text[start:end]
2971
+ semantic_chunks = split_text_for_tts(
2972
+ source,
2973
+ max_chars=maximum,
2974
+ min_chunk_chars=minimum,
2975
+ )
2976
+ if semantic_chunks and "".join(semantic_chunks) == source:
2977
+ ranges: list[tuple[int, int]] = []
2978
+ cursor = start
2979
+ for chunk in semantic_chunks:
2980
+ next_cursor = cursor + len(chunk)
2981
+ ranges.append((cursor, next_cursor))
2982
+ cursor = next_cursor
2983
+ if cursor == end:
2984
+ return tuple(ranges)
2985
+ if count_speech_units(source) <= maximum:
2986
+ return ((start, end),)
2987
+ output: list[tuple[int, int]] = []
2988
+ cursor = start
2989
+ while count_speech_units(text[cursor:end]) > maximum:
2990
+ local = text[cursor:end]
2991
+ forbidden = _protected_split_offsets(local, maximum)
2992
+ candidates: list[tuple[int, int, int]] = []
2993
+ for local_offset in range(1, len(local)):
2994
+ if local_offset in forbidden:
2995
+ continue
2996
+ head_units = count_speech_units(local[:local_offset])
2997
+ tail_units = count_speech_units(local[local_offset:])
2998
+ if not minimum <= head_units <= maximum:
2999
+ continue
3000
+ if 0 < tail_units < minimum:
3001
+ continue
3002
+ semantic = int(
3003
+ local[local_offset - 1] in _SEMANTIC_SOFT_BREAKS
3004
+ or local[local_offset].isspace()
3005
+ )
3006
+ candidates.append((semantic, head_units, -local_offset))
3007
+ if not candidates:
3008
+ raise ValueError("ordinary generation range cannot satisfy chunk limits")
3009
+ _, _, negative_offset = max(candidates)
3010
+ next_cursor = cursor - negative_offset
3011
+ output.append((cursor, next_cursor))
3012
+ cursor = next_cursor
3013
+ output.append((cursor, end))
3014
+ return tuple(output)
3015
+
3016
+
3017
+ def plan_generation_chunks(
3018
+ raw_text: str,
3019
+ normalized_text: str,
3020
+ *,
3021
+ min_units: int = 12,
3022
+ target_units: int = 32,
3023
+ network_max_units: int = 36,
3024
+ ordinary_max_units: int = 80,
3025
+ ) -> tuple[GenerationChunkSpec, ...]:
3026
+ """Plan model chunks while retaining exact whole-network verification proof.
3027
+
3028
+ Ordinary semantic chunking is unchanged. Only sentences that contain a raw
3029
+ ASCII URL/email are repacked at proven audible component delimiters. The
3030
+ caller must continue to verify the untouched ``normalized_text`` as one
3031
+ whole target before returning audio.
3032
+ """
3033
+
3034
+ if network_identifier_has_ambiguous_iri(raw_text):
3035
+ raise ValueError("network generation does not support non-ASCII IRI")
3036
+ minimum = max(1, int(min_units))
3037
+ target = max(minimum, int(target_units))
3038
+ network_maximum = max(minimum, int(network_max_units))
3039
+ ordinary_maximum = max(minimum, int(ordinary_max_units))
3040
+ if not minimum <= target <= network_maximum <= ordinary_maximum:
3041
+ raise ValueError("generation chunk limits are inconsistent")
3042
+
3043
+ normalized = normalize_tts_text(normalized_text)
3044
+ if not contains_network_identifier(raw_text):
3045
+ ordinary = split_text_for_tts(
3046
+ normalized,
3047
+ max_chars=ordinary_maximum,
3048
+ min_chunk_chars=minimum,
3049
+ )
3050
+ specs: list[GenerationChunkSpec] = []
3051
+ cursor = 0
3052
+ for index, chunk in enumerate(ordinary):
3053
+ start = normalized.find(chunk, cursor)
3054
+ if start < 0:
3055
+ raise ValueError("ordinary chunk is absent from normalized text")
3056
+ end = start + len(chunk)
3057
+ specs.append(
3058
+ GenerationChunkSpec(
3059
+ text=chunk,
3060
+ source_start=start,
3061
+ source_end=end,
3062
+ boundary_after=("semantic" if index + 1 < len(ordinary) else "none"),
3063
+ )
3064
+ )
3065
+ cursor = end
3066
+ return tuple(specs)
3067
+
3068
+ spans = _network_generation_spans(
3069
+ raw_text,
3070
+ normalized,
3071
+ hard_max_units=network_maximum,
3072
+ )
3073
+ if not spans:
3074
+ raise ValueError("network request has no bound spoken proof")
3075
+
3076
+ planned_ranges: list[tuple[int, int]] = []
3077
+ for sentence_start, sentence_end in _semantic_sentence_ranges(normalized):
3078
+ sentence_has_network = any(
3079
+ span.start < sentence_end and span.end > sentence_start
3080
+ for span in spans
3081
+ )
3082
+ if sentence_has_network:
3083
+ planned_ranges.extend(
3084
+ _network_generation_group_plan(
3085
+ normalized,
3086
+ group_start=sentence_start,
3087
+ group_end=sentence_end,
3088
+ original_boundaries=(),
3089
+ spans=spans,
3090
+ minimum=minimum,
3091
+ target=target,
3092
+ network_maximum=network_maximum,
3093
+ ordinary_maximum=ordinary_maximum,
3094
+ )
3095
+ )
3096
+ else:
3097
+ planned_ranges.extend(
3098
+ _ordinary_generation_ranges(
3099
+ normalized,
3100
+ start=sentence_start,
3101
+ end=sentence_end,
3102
+ minimum=minimum,
3103
+ maximum=ordinary_maximum,
3104
+ )
3105
+ )
3106
+
3107
+ if (
3108
+ not planned_ranges
3109
+ or planned_ranges[0][0] != 0
3110
+ or planned_ranges[-1][1] != len(normalized)
3111
+ or any(
3112
+ left_end != right_start
3113
+ for (_, left_end), (right_start, _) in zip(
3114
+ planned_ranges,
3115
+ planned_ranges[1:],
3116
+ )
3117
+ )
3118
+ ):
3119
+ raise ValueError("generation source ranges do not exactly cover the target")
3120
+
3121
+ specs = []
3122
+ internal_boundaries = {
3123
+ right
3124
+ for span in spans
3125
+ for _, right in span.component_ranges[:-1]
3126
+ }
3127
+ for start, end in planned_ranges:
3128
+ chunk = normalized[start:end]
3129
+ if not chunk or count_speech_units(chunk) <= 0:
3130
+ raise ValueError("generation planner produced an empty chunk")
3131
+ overlapping = tuple(
3132
+ span_index
3133
+ for span_index, span in enumerate(spans)
3134
+ if span.start < end and span.end > start
3135
+ )
3136
+ component_indices = tuple(
3137
+ (span_index, component_index)
3138
+ for span_index in overlapping
3139
+ for component_index, (left, right) in enumerate(
3140
+ spans[span_index].component_ranges
3141
+ )
3142
+ if left < end and right > start
3143
+ )
3144
+ specs.append(
3145
+ GenerationChunkSpec(
3146
+ text=chunk,
3147
+ source_start=start,
3148
+ source_end=end,
3149
+ network_span_indices=overlapping,
3150
+ network_component_indices=component_indices,
3151
+ network_full_spoken_proofs=tuple(
3152
+ spans[span_index].spoken_proof for span_index in overlapping
3153
+ ),
3154
+ boundary_after=(
3155
+ "network_internal"
3156
+ if end in internal_boundaries
3157
+ else ("semantic" if end < len(normalized) else "none")
3158
+ ),
3159
+ )
3160
+ )
3161
+ if "".join(spec.text for spec in specs) != normalized:
3162
+ raise ValueError("generation chunks do not reconstruct the normalized target")
3163
+ if any(
3164
+ count_speech_units(spec.text)
3165
+ > (network_maximum if spec.network_conditioned else ordinary_maximum)
3166
+ for spec in specs
3167
+ ):
3168
+ raise ValueError("generation chunk exceeds its provenance-specific limit")
3169
+ return tuple(specs)
3170
+
3171
+
3172
  def coalesce_text_chunks(
3173
  chunks: Sequence[str],
3174
  *,
 
3275
  return outputs
3276
 
3277
 
3278
+ def fade_variable_internal_edges(
3279
+ chunks: Sequence[np.ndarray],
3280
+ sample_rate: int,
3281
+ fade_ms_by_boundary: Sequence[float],
3282
+ ) -> list[np.ndarray]:
3283
+ """Apply independently bounded fades at each adjacent chunk boundary."""
3284
+
3285
+ outputs = [np.asarray(chunk, dtype=np.float32).reshape(-1).copy() for chunk in chunks]
3286
+ if len(fade_ms_by_boundary) != max(0, len(outputs) - 1):
3287
+ raise ValueError("fade boundary count must equal chunk count minus one")
3288
+ if sample_rate <= 0:
3289
+ raise ValueError("sample_rate must be positive")
3290
+ if any(output.size <= 0 for output in outputs):
3291
+ raise ValueError("audio chunks must be non-empty")
3292
+ widths: list[int] = []
3293
+ for index, raw_fade_ms in enumerate(fade_ms_by_boundary):
3294
+ try:
3295
+ fade_ms = float(raw_fade_ms)
3296
+ except (TypeError, ValueError, OverflowError) as error:
3297
+ raise ValueError("fade duration must be finite and non-negative") from error
3298
+ if not math.isfinite(fade_ms) or fade_ms < 0.0:
3299
+ raise ValueError("fade duration must be finite and non-negative")
3300
+ requested = int(round(fade_ms * sample_rate / 1000.0))
3301
+ widths.append(
3302
+ min(requested, outputs[index].size, outputs[index + 1].size)
3303
+ )
3304
+ if any(
3305
+ widths[index - 1] + widths[index] > outputs[index].size
3306
+ for index in range(1, len(outputs) - 1)
3307
+ ):
3308
+ raise ValueError("adjacent fades overlap inside an audio chunk")
3309
+ for index, count in enumerate(widths):
3310
+ if count <= 0:
3311
+ continue
3312
+ outputs[index][-count:] *= np.linspace(
3313
+ 1.0,
3314
+ 0.0,
3315
+ count,
3316
+ endpoint=True,
3317
+ dtype=np.float32,
3318
+ )
3319
+ outputs[index + 1][:count] *= np.linspace(
3320
+ 0.0,
3321
+ 1.0,
3322
+ count,
3323
+ endpoint=True,
3324
+ dtype=np.float32,
3325
+ )
3326
+ return outputs
3327
+
3328
+
3329
  def join_audio_chunks(
3330
  chunks: list[np.ndarray],
3331
  pauses: list[int],
 
3357
  return output.astype(np.float32, copy=False)
3358
 
3359
 
3360
+ def join_audio_chunks_variable(
3361
+ chunks: Sequence[np.ndarray],
3362
+ pauses: Sequence[int],
3363
+ crossfade_samples_by_boundary: Sequence[int],
3364
+ *,
3365
+ pre_faded_edges: bool = False,
3366
+ ) -> np.ndarray:
3367
+ """Join chunks with one validated crossfade width per boundary."""
3368
+
3369
+ if not chunks:
3370
+ if pauses or crossfade_samples_by_boundary:
3371
+ raise ValueError("empty chunks cannot carry boundary metadata")
3372
+ return np.zeros(0, dtype=np.float32)
3373
+ boundary_count = len(chunks) - 1
3374
+ if len(pauses) != boundary_count or len(crossfade_samples_by_boundary) != boundary_count:
3375
+ raise ValueError("pause/crossfade boundary counts must equal chunk count minus one")
3376
+ prepared = [
3377
+ np.asarray(chunk, dtype=np.float32).reshape(-1).copy()
3378
+ for chunk in chunks
3379
+ ]
3380
+ if any(chunk.size <= 0 for chunk in prepared):
3381
+ raise ValueError("audio chunks must be non-empty")
3382
+ validated_pauses: list[int] = []
3383
+ requested_crossfades: list[int] = []
3384
+ for raw_pause, raw_crossfade in zip(
3385
+ pauses,
3386
+ crossfade_samples_by_boundary,
3387
+ strict=True,
3388
+ ):
3389
+ if isinstance(raw_pause, (bool, np.bool_)) or isinstance(
3390
+ raw_crossfade,
3391
+ (bool, np.bool_),
3392
+ ):
3393
+ raise ValueError("pause and crossfade widths must be non-negative integers")
3394
+ try:
3395
+ pause = operator.index(raw_pause)
3396
+ crossfade = operator.index(raw_crossfade)
3397
+ except (TypeError, ValueError, OverflowError) as error:
3398
+ raise ValueError(
3399
+ "pause and crossfade widths must be non-negative integers"
3400
+ ) from error
3401
+ if pause < 0 or crossfade < 0:
3402
+ raise ValueError("pause and crossfade widths must be non-negative integers")
3403
+ validated_pauses.append(int(pause))
3404
+ requested_crossfades.append(int(crossfade))
3405
+ widths = [
3406
+ min(requested, prepared[index].size, prepared[index + 1].size)
3407
+ for index, requested in enumerate(requested_crossfades)
3408
+ ]
3409
+ if any(
3410
+ widths[index - 1] + widths[index] > prepared[index].size
3411
+ for index in range(1, len(prepared) - 1)
3412
+ ):
3413
+ raise ValueError("adjacent crossfades overlap inside an audio chunk")
3414
+
3415
+ output = prepared[0]
3416
+ for index, next_chunk in enumerate(prepared[1:]):
3417
+ pause = validated_pauses[index]
3418
+ crossfade = widths[index]
3419
+ if pause > 0:
3420
+ if crossfade > 0 and not pre_faded_edges:
3421
+ output[-crossfade:] *= np.linspace(
3422
+ 1.0,
3423
+ 0.0,
3424
+ crossfade,
3425
+ dtype=np.float32,
3426
+ )
3427
+ next_chunk[:crossfade] *= np.linspace(
3428
+ 0.0,
3429
+ 1.0,
3430
+ crossfade,
3431
+ dtype=np.float32,
3432
+ )
3433
+ output = np.concatenate(
3434
+ (output, np.zeros(pause, dtype=np.float32), next_chunk)
3435
+ )
3436
+ elif crossfade > 0:
3437
+ if pre_faded_edges:
3438
+ overlap = output[-crossfade:] + next_chunk[:crossfade]
3439
+ else:
3440
+ fade_out = np.linspace(
3441
+ 1.0,
3442
+ 0.0,
3443
+ crossfade,
3444
+ endpoint=False,
3445
+ dtype=np.float32,
3446
+ )
3447
+ overlap = (
3448
+ output[-crossfade:] * fade_out
3449
+ + next_chunk[:crossfade] * (1.0 - fade_out)
3450
+ )
3451
+ output = np.concatenate(
3452
+ (output[:-crossfade], overlap, next_chunk[crossfade:])
3453
+ )
3454
+ else:
3455
+ output = np.concatenate((output, next_chunk))
3456
+ return output.astype(np.float32, copy=False)
3457
+
3458
+
3459
  def apply_loudness_floor(
3460
  audio: np.ndarray,
3461
  min_rms: float = 0.07,
quality_runtime.py CHANGED
@@ -8,6 +8,7 @@ load the real models lazily while unit tests remain deterministic and offline.
8
  from __future__ import annotations
9
 
10
  import hashlib
 
11
  import json
12
  import math
13
  import operator
@@ -46,7 +47,7 @@ RELEASE_SPEAKER_TRIGGER_SECONDS = 1.48
46
  SEQUENCE_FALLBACK_MAX_LOCAL_BOUNDARY_SPEAKER_DROP = 0.15
47
  SEQUENCE_FALLBACK_SPEAKER_WEIGHT = 0.05
48
  SEQUENCE_FALLBACK_BOUNDARY_WEIGHT = 0.10
49
- CASCADE_EVIDENCE_SCHEMA_VERSION = 2
50
  CASCADE_EVIDENCE_LOG_PREFIX = "[BlueMagpie] cascade evidence "
51
  CASCADE_EVIDENCE_MAX_ATTEMPTS = ADAPTIVE_CASCADE_STAGE_LIMITS[-1]
52
  CASCADE_EVIDENCE_MAX_LOCAL_RESULTS = 20
@@ -93,6 +94,23 @@ class GenerationPolicy:
93
  hard_stop_margin_steps: int
94
 
95
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
96
  BASE_GENERATION_POLICY = GenerationPolicy(
97
  name="base",
98
  cjk_cps=5.2,
@@ -111,7 +129,7 @@ COMPLETION_HEADROOM_GENERATION_POLICY = GenerationPolicy(
111
  ascii_cps=3.6,
112
  hard_stop_margin_steps=1,
113
  )
114
- MIXED_CFG_SCHEDULE = "offset_zero_and_even_primary_odd_alternate"
115
  MIXED_CFG_PRIMARY = 3.0
116
  MIXED_CFG_ALTERNATE = 2.0
117
  MIXED_CFG_SHORT_TEXT_MAX_UNITS = 6
@@ -1181,6 +1199,8 @@ class CandidateGenerationEvidence:
1181
  scheduled_cfg: float
1182
  effective_cfgs: tuple[float, ...]
1183
  floor_reasons: tuple[tuple[str, ...], ...]
 
 
1184
 
1185
 
1186
  @dataclass(frozen=True)
@@ -1196,6 +1216,8 @@ class CandidateAttemptEvidence:
1196
  local_results: tuple[CandidateGateEvidence, ...]
1197
  chunk_indices: tuple[int, ...]
1198
  chunk_text_units: tuple[int, ...]
 
 
1199
  scheduled_cfg: float | None
1200
  effective_cfgs: tuple[float, ...]
1201
  floor_reasons: tuple[tuple[str, ...], ...]
@@ -1804,12 +1826,20 @@ def _trajectory_gate_evidence_payload(
1804
  def _candidate_attempt_evidence_payload(
1805
  evidence: CandidateAttemptEvidence,
1806
  ) -> dict[str, Any]:
 
 
 
 
 
 
 
 
 
 
1807
  return {
1808
  "candidate_index": evidence.candidate_index,
1809
  "seed": evidence.seed,
1810
- "policy": generation_policy_for_candidate_offset(
1811
- evidence.candidate_index
1812
- ).name,
1813
  "trajectory_passed": evidence.trajectory_passed,
1814
  "trajectory_score": evidence.trajectory_score,
1815
  "trajectory_reasons": list(
@@ -1817,6 +1847,9 @@ def _candidate_attempt_evidence_payload(
1817
  ),
1818
  "chunk_indices": list(evidence.chunk_indices),
1819
  "chunk_text_units": list(evidence.chunk_text_units),
 
 
 
1820
  "scheduled_cfg": evidence.scheduled_cfg,
1821
  "effective_cfgs": list(evidence.effective_cfgs),
1822
  "floor_reasons": [list(reasons) for reasons in evidence.floor_reasons],
@@ -1891,7 +1924,9 @@ def _selected_generation_evidence_payload(
1891
  scheduled_cfgs: list[float | None] = []
1892
  effective_cfgs: list[float | None] = []
1893
  floor_reasons: list[list[str]] = []
 
1894
  policies: list[str | None] = []
 
1895
  complete = True
1896
  for chunk_index, candidate_index in enumerate(
1897
  selection.chunk_candidate_indices
@@ -1902,7 +1937,9 @@ def _selected_generation_evidence_payload(
1902
  scheduled_cfgs.append(None)
1903
  effective_cfgs.append(None)
1904
  floor_reasons.append([])
 
1905
  policies.append(None)
 
1906
  continue
1907
  try:
1908
  local_index = attempt.chunk_indices.index(chunk_index)
@@ -1913,20 +1950,37 @@ def _selected_generation_evidence_payload(
1913
  scheduled_cfgs.append(attempt.scheduled_cfg)
1914
  effective_cfgs.append(None)
1915
  floor_reasons.append([])
 
1916
  policies.append(None)
 
1917
  continue
 
 
 
 
 
1918
  scheduled_cfgs.append(attempt.scheduled_cfg)
1919
  effective_cfgs.append(effective)
1920
  floor_reasons.append(list(reasons))
 
1921
  policies.append(
1922
- generation_policy_for_candidate_offset(candidate_index).name
 
 
1923
  )
 
 
 
 
 
1924
  return {
1925
  "complete": complete,
1926
  "chunk_scheduled_cfgs": scheduled_cfgs,
1927
  "chunk_effective_cfgs": effective_cfgs,
1928
  "chunk_floor_reasons": floor_reasons,
 
1929
  "chunk_policies": policies,
 
1930
  }
1931
 
1932
 
@@ -2008,6 +2062,8 @@ def format_cascade_evidence_log(
2008
  generation_evidence_complete = bool(attempts) and all(
2009
  attempt.scheduled_cfg is not None
2010
  and len(attempt.chunk_indices) == len(attempt.chunk_text_units)
 
 
2011
  == len(attempt.effective_cfgs)
2012
  == len(attempt.floor_reasons)
2013
  and bool(attempt.chunk_indices)
@@ -2469,6 +2525,8 @@ def _validated_candidate_generation_evidence(
2469
  *,
2470
  chunk_indices: tuple[int, ...],
2471
  chunks: tuple[str, ...],
 
 
2472
  ) -> CandidateGenerationEvidence:
2473
  """Validate deterministic CFG evidence supplied by the hosted app."""
2474
 
@@ -2483,14 +2541,60 @@ def _validated_candidate_generation_evidence(
2483
  raise ValueError("generation evidence text units do not match the attempt")
2484
  if len(value.effective_cfgs) != len(chunks) or len(value.floor_reasons) != len(chunks):
2485
  raise ValueError("generation evidence CFG rows do not match the attempt")
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2486
  scheduled = _finite_float(value.scheduled_cfg, minimum=1.0, maximum=4.0)
2487
  if scheduled is None:
2488
  raise ValueError("generation evidence scheduled CFG is invalid")
 
 
 
2489
  effective_cfgs: list[float] = []
2490
  floor_reasons: list[tuple[str, ...]] = []
2491
- for effective_value, raw_reasons in zip(
2492
  value.effective_cfgs,
2493
  value.floor_reasons,
 
2494
  strict=True,
2495
  ):
2496
  effective = _finite_float(effective_value, minimum=1.0, maximum=4.0)
@@ -2507,6 +2611,17 @@ def _validated_candidate_generation_evidence(
2507
  reasons and effective <= scheduled
2508
  ):
2509
  raise ValueError("generation evidence floor reasons disagree with CFG")
 
 
 
 
 
 
 
 
 
 
 
2510
  effective_cfgs.append(effective)
2511
  floor_reasons.append(reasons)
2512
  return CandidateGenerationEvidence(
@@ -2515,6 +2630,8 @@ def _validated_candidate_generation_evidence(
2515
  scheduled_cfg=scheduled,
2516
  effective_cfgs=tuple(effective_cfgs),
2517
  floor_reasons=tuple(floor_reasons),
 
 
2518
  )
2519
 
2520
 
@@ -2547,6 +2664,12 @@ def _candidate_attempt_evidence(
2547
  chunk_text_units=(
2548
  generation.chunk_text_units if generation is not None else ()
2549
  ),
 
 
 
 
 
 
2550
  scheduled_cfg=(generation.scheduled_cfg if generation is not None else None),
2551
  effective_cfgs=(generation.effective_cfgs if generation is not None else ()),
2552
  floor_reasons=(generation.floor_reasons if generation is not None else ()),
@@ -3093,7 +3216,7 @@ def _coverage_final_passed(verification: Any) -> bool:
3093
  def run_coverage_adaptive_cascade(
3094
  chunks: Sequence[str],
3095
  root_seed: int,
3096
- candidate_generator: Callable[[tuple[str, ...], int], Any],
3097
  whole_candidate_verifier: (
3098
  Callable[
3099
  [Any, tuple[str, ...], int],
@@ -3111,10 +3234,7 @@ def run_coverage_adaptive_cascade(
3111
  [CascadeResult, tuple[str, ...]],
3112
  TrajectoryGateResult,
3113
  ],
3114
- generation_evidence_factory: Callable[
3115
- [int, int, tuple[int, ...], tuple[str, ...]],
3116
- CandidateGenerationEvidence,
3117
- ] | None = None,
3118
  max_generated_chunks: int = 20,
3119
  max_generated_text_units: int = 800,
3120
  max_sequence_paths: int = 3,
@@ -3127,6 +3247,11 @@ def run_coverage_adaptive_cascade(
3127
  retained. Later seeds generate exactly one low-coverage chunk under hard
3128
  generated-chunk and generated-text-unit budgets. Ragged DP paths are
3129
  never returned without the supplied exact whole-waveform verifier.
 
 
 
 
 
3130
  """
3131
 
3132
  if not all(
@@ -3143,6 +3268,34 @@ def run_coverage_adaptive_cascade(
3143
  generation_evidence_factory
3144
  ):
3145
  raise ValueError("generation_evidence_factory must be callable")
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3146
  try:
3147
  chunk_tuple = tuple(str(chunk) for chunk in chunks)
3148
  except TypeError as error:
@@ -3211,30 +3364,56 @@ def run_coverage_adaptive_cascade(
3211
  generated_units = initial_units
3212
 
3213
  def generation_evidence(
3214
- candidate_index: int,
3215
- seed: int,
3216
- source_chunk_indices: tuple[int, ...],
3217
  candidate_chunks: tuple[str, ...],
3218
  ) -> CandidateGenerationEvidence | None:
3219
  if generation_evidence_factory is None:
3220
  return None
3221
  try:
3222
- raw_evidence = generation_evidence_factory(
3223
- candidate_index,
3224
- seed,
3225
- source_chunk_indices,
3226
- candidate_chunks,
3227
- )
 
 
 
 
 
 
 
 
 
 
 
 
 
3228
  return _validated_candidate_generation_evidence(
3229
  raw_evidence,
3230
- chunk_indices=source_chunk_indices,
3231
  chunks=candidate_chunks,
 
 
3232
  )
3233
  except Exception as error:
3234
  raise RuntimeError("candidate generation evidence is invalid") from error
3235
 
 
 
 
 
 
 
3236
  try:
3237
- raw_initial_trajectory = candidate_generator(chunk_tuple, base_seed)
 
 
 
 
 
 
 
3238
  except Exception as error:
3239
  raise RuntimeError("initial trajectory generation failed") from error
3240
  initial_trajectory = _coverage_trajectory_tuple(
@@ -3242,9 +3421,7 @@ def run_coverage_adaptive_cascade(
3242
  chunk_count,
3243
  )
3244
  initial_generation_evidence = generation_evidence(
3245
- 0,
3246
- base_seed,
3247
- tuple(range(chunk_count)),
3248
  chunk_tuple,
3249
  )
3250
  attempted_seeds.append(base_seed)
@@ -3366,15 +3543,26 @@ def run_coverage_adaptive_cascade(
3366
  )
3367
  seed = base_seed + next_candidate_index
3368
  refill_chunks = (chunk_tuple[chunk_index],)
 
 
 
 
 
 
3369
  try:
3370
- raw_refill_trajectory = candidate_generator(refill_chunks, seed)
 
 
 
 
 
 
 
3371
  except Exception as error:
3372
  raise RuntimeError("chunk refill generation failed") from error
3373
  refill_trajectory = _coverage_trajectory_tuple(raw_refill_trajectory, 1)
3374
  refill_generation_evidence = generation_evidence(
3375
- next_candidate_index,
3376
- seed,
3377
- (chunk_index,),
3378
  refill_chunks,
3379
  )
3380
  attempted_seeds.append(seed)
 
8
  from __future__ import annotations
9
 
10
  import hashlib
11
+ import inspect
12
  import json
13
  import math
14
  import operator
 
47
  SEQUENCE_FALLBACK_MAX_LOCAL_BOUNDARY_SPEAKER_DROP = 0.15
48
  SEQUENCE_FALLBACK_SPEAKER_WEIGHT = 0.05
49
  SEQUENCE_FALLBACK_BOUNDARY_WEIGHT = 0.10
50
+ CASCADE_EVIDENCE_SCHEMA_VERSION = 3
51
  CASCADE_EVIDENCE_LOG_PREFIX = "[BlueMagpie] cascade evidence "
52
  CASCADE_EVIDENCE_MAX_ATTEMPTS = ADAPTIVE_CASCADE_STAGE_LIMITS[-1]
53
  CASCADE_EVIDENCE_MAX_LOCAL_RESULTS = 20
 
94
  hard_stop_margin_steps: int
95
 
96
 
97
+ @dataclass(frozen=True)
98
+ class CandidateGenerationContext:
99
+ """Immutable global identity and row-local schedule for one generation.
100
+
101
+ ``candidate_index`` and ``seed`` retain the request-global attempt identity.
102
+ ``chunk_candidate_ordinals`` is independent: zero denotes the initial
103
+ trajectory and positive values count refill attempts within each source
104
+ chunk. Context-aware callbacks can therefore rotate generation policies
105
+ per chunk without changing the canonical seed schedule.
106
+ """
107
+
108
+ candidate_index: int
109
+ seed: int
110
+ chunk_indices: tuple[int, ...]
111
+ chunk_candidate_ordinals: tuple[int, ...]
112
+
113
+
114
  BASE_GENERATION_POLICY = GenerationPolicy(
115
  name="base",
116
  cjk_cps=5.2,
 
129
  ascii_cps=3.6,
130
  hard_stop_margin_steps=1,
131
  )
132
+ MIXED_CFG_SCHEDULE = "row_ordinal_zero_and_even_primary_odd_alternate"
133
  MIXED_CFG_PRIMARY = 3.0
134
  MIXED_CFG_ALTERNATE = 2.0
135
  MIXED_CFG_SHORT_TEXT_MAX_UNITS = 6
 
1199
  scheduled_cfg: float
1200
  effective_cfgs: tuple[float, ...]
1201
  floor_reasons: tuple[tuple[str, ...], ...]
1202
+ chunk_candidate_ordinals: tuple[int, ...] = ()
1203
+ network_conditioned: tuple[bool, ...] = ()
1204
 
1205
 
1206
  @dataclass(frozen=True)
 
1216
  local_results: tuple[CandidateGateEvidence, ...]
1217
  chunk_indices: tuple[int, ...]
1218
  chunk_text_units: tuple[int, ...]
1219
+ chunk_candidate_ordinals: tuple[int, ...]
1220
+ network_conditioned: tuple[bool, ...]
1221
  scheduled_cfg: float | None
1222
  effective_cfgs: tuple[float, ...]
1223
  floor_reasons: tuple[tuple[str, ...], ...]
 
1826
  def _candidate_attempt_evidence_payload(
1827
  evidence: CandidateAttemptEvidence,
1828
  ) -> dict[str, Any]:
1829
+ candidate_ordinals = evidence.chunk_candidate_ordinals
1830
+ chunk_policies = [
1831
+ generation_policy_for_candidate_offset(ordinal).name
1832
+ for ordinal in candidate_ordinals
1833
+ ]
1834
+ policy = (
1835
+ chunk_policies[0]
1836
+ if chunk_policies and len(set(chunk_policies)) == 1
1837
+ else generation_policy_for_candidate_offset(evidence.candidate_index).name
1838
+ )
1839
  return {
1840
  "candidate_index": evidence.candidate_index,
1841
  "seed": evidence.seed,
1842
+ "policy": policy,
 
 
1843
  "trajectory_passed": evidence.trajectory_passed,
1844
  "trajectory_score": evidence.trajectory_score,
1845
  "trajectory_reasons": list(
 
1847
  ),
1848
  "chunk_indices": list(evidence.chunk_indices),
1849
  "chunk_text_units": list(evidence.chunk_text_units),
1850
+ "chunk_candidate_ordinals": list(candidate_ordinals),
1851
+ "chunk_policies": chunk_policies,
1852
+ "network_conditioned": list(evidence.network_conditioned),
1853
  "scheduled_cfg": evidence.scheduled_cfg,
1854
  "effective_cfgs": list(evidence.effective_cfgs),
1855
  "floor_reasons": [list(reasons) for reasons in evidence.floor_reasons],
 
1924
  scheduled_cfgs: list[float | None] = []
1925
  effective_cfgs: list[float | None] = []
1926
  floor_reasons: list[list[str]] = []
1927
+ candidate_ordinals: list[int | None] = []
1928
  policies: list[str | None] = []
1929
+ network_conditioned: list[bool | None] = []
1930
  complete = True
1931
  for chunk_index, candidate_index in enumerate(
1932
  selection.chunk_candidate_indices
 
1937
  scheduled_cfgs.append(None)
1938
  effective_cfgs.append(None)
1939
  floor_reasons.append([])
1940
+ candidate_ordinals.append(None)
1941
  policies.append(None)
1942
+ network_conditioned.append(None)
1943
  continue
1944
  try:
1945
  local_index = attempt.chunk_indices.index(chunk_index)
 
1950
  scheduled_cfgs.append(attempt.scheduled_cfg)
1951
  effective_cfgs.append(None)
1952
  floor_reasons.append([])
1953
+ candidate_ordinals.append(None)
1954
  policies.append(None)
1955
+ network_conditioned.append(None)
1956
  continue
1957
+ try:
1958
+ ordinal = attempt.chunk_candidate_ordinals[local_index]
1959
+ except IndexError:
1960
+ complete = False
1961
+ ordinal = None
1962
  scheduled_cfgs.append(attempt.scheduled_cfg)
1963
  effective_cfgs.append(effective)
1964
  floor_reasons.append(list(reasons))
1965
+ candidate_ordinals.append(ordinal)
1966
  policies.append(
1967
+ None
1968
+ if ordinal is None
1969
+ else generation_policy_for_candidate_offset(ordinal).name
1970
  )
1971
+ try:
1972
+ network_conditioned.append(attempt.network_conditioned[local_index])
1973
+ except IndexError:
1974
+ complete = False
1975
+ network_conditioned.append(None)
1976
  return {
1977
  "complete": complete,
1978
  "chunk_scheduled_cfgs": scheduled_cfgs,
1979
  "chunk_effective_cfgs": effective_cfgs,
1980
  "chunk_floor_reasons": floor_reasons,
1981
+ "chunk_candidate_ordinals": candidate_ordinals,
1982
  "chunk_policies": policies,
1983
+ "network_conditioned": network_conditioned,
1984
  }
1985
 
1986
 
 
2062
  generation_evidence_complete = bool(attempts) and all(
2063
  attempt.scheduled_cfg is not None
2064
  and len(attempt.chunk_indices) == len(attempt.chunk_text_units)
2065
+ == len(attempt.chunk_candidate_ordinals)
2066
+ == len(attempt.network_conditioned)
2067
  == len(attempt.effective_cfgs)
2068
  == len(attempt.floor_reasons)
2069
  and bool(attempt.chunk_indices)
 
2525
  *,
2526
  chunk_indices: tuple[int, ...],
2527
  chunks: tuple[str, ...],
2528
+ expected_candidate_ordinals: tuple[int, ...],
2529
+ require_explicit_candidate_ordinals: bool,
2530
  ) -> CandidateGenerationEvidence:
2531
  """Validate deterministic CFG evidence supplied by the hosted app."""
2532
 
 
2541
  raise ValueError("generation evidence text units do not match the attempt")
2542
  if len(value.effective_cfgs) != len(chunks) or len(value.floor_reasons) != len(chunks):
2543
  raise ValueError("generation evidence CFG rows do not match the attempt")
2544
+ if value.network_conditioned:
2545
+ if (
2546
+ not isinstance(value.network_conditioned, tuple)
2547
+ or len(value.network_conditioned) != len(chunks)
2548
+ or any(type(flag) is not bool for flag in value.network_conditioned)
2549
+ ):
2550
+ raise ValueError(
2551
+ "generation evidence network provenance does not match the attempt"
2552
+ )
2553
+ network_conditioned = value.network_conditioned
2554
+ else:
2555
+ network_conditioned = (False,) * len(chunks)
2556
+ raw_ordinals = value.chunk_candidate_ordinals
2557
+ if not raw_ordinals:
2558
+ if require_explicit_candidate_ordinals:
2559
+ raise ValueError(
2560
+ "context-aware generation evidence must provide chunk candidate ordinals"
2561
+ )
2562
+ candidate_ordinals = expected_candidate_ordinals
2563
+ else:
2564
+ if not isinstance(raw_ordinals, tuple) or len(raw_ordinals) != len(chunks):
2565
+ raise ValueError("generation evidence candidate ordinals do not match the attempt")
2566
+ candidate_ordinals_list: list[int] = []
2567
+ for raw_ordinal in raw_ordinals:
2568
+ if isinstance(raw_ordinal, (bool, np.bool_)):
2569
+ raise ValueError("generation evidence candidate ordinal is invalid")
2570
+ try:
2571
+ ordinal = operator.index(raw_ordinal)
2572
+ except (TypeError, ValueError, OverflowError) as error:
2573
+ raise ValueError(
2574
+ "generation evidence candidate ordinal is invalid"
2575
+ ) from error
2576
+ if ordinal < 0:
2577
+ raise ValueError("generation evidence candidate ordinal is invalid")
2578
+ candidate_ordinals_list.append(int(ordinal))
2579
+ candidate_ordinals = tuple(candidate_ordinals_list)
2580
+ if candidate_ordinals != expected_candidate_ordinals:
2581
+ raise ValueError(
2582
+ "generation evidence candidate ordinals do not match the schedule"
2583
+ )
2584
+ if len(set(candidate_ordinals)) != 1:
2585
+ raise ValueError("one generation call must use one candidate schedule ordinal")
2586
  scheduled = _finite_float(value.scheduled_cfg, minimum=1.0, maximum=4.0)
2587
  if scheduled is None:
2588
  raise ValueError("generation evidence scheduled CFG is invalid")
2589
+ expected_scheduled = generation_cfg_for_candidate_offset(candidate_ordinals[0])
2590
+ if scheduled != expected_scheduled:
2591
+ raise ValueError("generation evidence scheduled CFG disagrees with the schedule")
2592
  effective_cfgs: list[float] = []
2593
  floor_reasons: list[tuple[str, ...]] = []
2594
+ for effective_value, raw_reasons, is_network in zip(
2595
  value.effective_cfgs,
2596
  value.floor_reasons,
2597
+ network_conditioned,
2598
  strict=True,
2599
  ):
2600
  effective = _finite_float(effective_value, minimum=1.0, maximum=4.0)
 
2611
  reasons and effective <= scheduled
2612
  ):
2613
  raise ValueError("generation evidence floor reasons disagree with CFG")
2614
+ network_floor_required = bool(
2615
+ is_network and scheduled < MIXED_CFG_NETWORK_MIN
2616
+ )
2617
+ if ("network" in reasons) != network_floor_required:
2618
+ raise ValueError(
2619
+ "generation evidence network floor disagrees with provenance"
2620
+ )
2621
+ if is_network and effective < MIXED_CFG_NETWORK_MIN:
2622
+ raise ValueError(
2623
+ "generation evidence network CFG is below the frozen minimum"
2624
+ )
2625
  effective_cfgs.append(effective)
2626
  floor_reasons.append(reasons)
2627
  return CandidateGenerationEvidence(
 
2630
  scheduled_cfg=scheduled,
2631
  effective_cfgs=tuple(effective_cfgs),
2632
  floor_reasons=tuple(floor_reasons),
2633
+ chunk_candidate_ordinals=candidate_ordinals,
2634
+ network_conditioned=network_conditioned,
2635
  )
2636
 
2637
 
 
2664
  chunk_text_units=(
2665
  generation.chunk_text_units if generation is not None else ()
2666
  ),
2667
+ chunk_candidate_ordinals=(
2668
+ generation.chunk_candidate_ordinals if generation is not None else ()
2669
+ ),
2670
+ network_conditioned=(
2671
+ generation.network_conditioned if generation is not None else ()
2672
+ ),
2673
  scheduled_cfg=(generation.scheduled_cfg if generation is not None else None),
2674
  effective_cfgs=(generation.effective_cfgs if generation is not None else ()),
2675
  floor_reasons=(generation.floor_reasons if generation is not None else ()),
 
3216
  def run_coverage_adaptive_cascade(
3217
  chunks: Sequence[str],
3218
  root_seed: int,
3219
+ candidate_generator: Callable[..., Any],
3220
  whole_candidate_verifier: (
3221
  Callable[
3222
  [Any, tuple[str, ...], int],
 
3234
  [CascadeResult, tuple[str, ...]],
3235
  TrajectoryGateResult,
3236
  ],
3237
+ generation_evidence_factory: Callable[..., CandidateGenerationEvidence] | None = None,
 
 
 
3238
  max_generated_chunks: int = 20,
3239
  max_generated_text_units: int = 800,
3240
  max_sequence_paths: int = 3,
 
3247
  retained. Later seeds generate exactly one low-coverage chunk under hard
3248
  generated-chunk and generated-text-unit budgets. Ragged DP paths are
3249
  never returned without the supplied exact whole-waveform verifier.
3250
+
3251
+ Existing two-argument generation callbacks remain unchanged. A callback
3252
+ that explicitly accepts the keyword-only ``generation_context`` opts into
3253
+ per-chunk refill scheduling. Its evidence factory must accept the same
3254
+ keyword and report the supplied row-local candidate ordinals.
3255
  """
3256
 
3257
  if not all(
 
3268
  generation_evidence_factory
3269
  ):
3270
  raise ValueError("generation_evidence_factory must be callable")
3271
+
3272
+ def accepts_generation_context(callback: Callable[..., Any]) -> bool:
3273
+ try:
3274
+ parameters = inspect.signature(callback).parameters.values()
3275
+ except (TypeError, ValueError):
3276
+ return False
3277
+ return any(
3278
+ parameter.kind is inspect.Parameter.VAR_KEYWORD
3279
+ or (
3280
+ parameter.name == "generation_context"
3281
+ and parameter.kind
3282
+ in (
3283
+ inspect.Parameter.POSITIONAL_OR_KEYWORD,
3284
+ inspect.Parameter.KEYWORD_ONLY,
3285
+ )
3286
+ )
3287
+ for parameter in parameters
3288
+ )
3289
+
3290
+ context_aware_generator = accepts_generation_context(candidate_generator)
3291
+ context_aware_evidence = (
3292
+ generation_evidence_factory is not None
3293
+ and accepts_generation_context(generation_evidence_factory)
3294
+ )
3295
+ if context_aware_generator != context_aware_evidence:
3296
+ raise ValueError(
3297
+ "context-aware generation and evidence callbacks must opt in together"
3298
+ )
3299
  try:
3300
  chunk_tuple = tuple(str(chunk) for chunk in chunks)
3301
  except TypeError as error:
 
3364
  generated_units = initial_units
3365
 
3366
  def generation_evidence(
3367
+ context: CandidateGenerationContext,
 
 
3368
  candidate_chunks: tuple[str, ...],
3369
  ) -> CandidateGenerationEvidence | None:
3370
  if generation_evidence_factory is None:
3371
  return None
3372
  try:
3373
+ if context_aware_evidence:
3374
+ raw_evidence = generation_evidence_factory(
3375
+ context.candidate_index,
3376
+ context.seed,
3377
+ context.chunk_indices,
3378
+ candidate_chunks,
3379
+ generation_context=context,
3380
+ )
3381
+ expected_ordinals = context.chunk_candidate_ordinals
3382
+ else:
3383
+ raw_evidence = generation_evidence_factory(
3384
+ context.candidate_index,
3385
+ context.seed,
3386
+ context.chunk_indices,
3387
+ candidate_chunks,
3388
+ )
3389
+ expected_ordinals = (context.candidate_index,) * len(
3390
+ candidate_chunks
3391
+ )
3392
  return _validated_candidate_generation_evidence(
3393
  raw_evidence,
3394
+ chunk_indices=context.chunk_indices,
3395
  chunks=candidate_chunks,
3396
+ expected_candidate_ordinals=expected_ordinals,
3397
+ require_explicit_candidate_ordinals=context_aware_evidence,
3398
  )
3399
  except Exception as error:
3400
  raise RuntimeError("candidate generation evidence is invalid") from error
3401
 
3402
+ initial_context = CandidateGenerationContext(
3403
+ candidate_index=0,
3404
+ seed=base_seed,
3405
+ chunk_indices=tuple(range(chunk_count)),
3406
+ chunk_candidate_ordinals=(0,) * chunk_count,
3407
+ )
3408
  try:
3409
+ if context_aware_generator:
3410
+ raw_initial_trajectory = candidate_generator(
3411
+ chunk_tuple,
3412
+ base_seed,
3413
+ generation_context=initial_context,
3414
+ )
3415
+ else:
3416
+ raw_initial_trajectory = candidate_generator(chunk_tuple, base_seed)
3417
  except Exception as error:
3418
  raise RuntimeError("initial trajectory generation failed") from error
3419
  initial_trajectory = _coverage_trajectory_tuple(
 
3421
  chunk_count,
3422
  )
3423
  initial_generation_evidence = generation_evidence(
3424
+ initial_context,
 
 
3425
  chunk_tuple,
3426
  )
3427
  attempted_seeds.append(base_seed)
 
3543
  )
3544
  seed = base_seed + next_candidate_index
3545
  refill_chunks = (chunk_tuple[chunk_index],)
3546
+ refill_context = CandidateGenerationContext(
3547
+ candidate_index=next_candidate_index,
3548
+ seed=seed,
3549
+ chunk_indices=(chunk_index,),
3550
+ chunk_candidate_ordinals=(refill_attempts[chunk_index] + 1,),
3551
+ )
3552
  try:
3553
+ if context_aware_generator:
3554
+ raw_refill_trajectory = candidate_generator(
3555
+ refill_chunks,
3556
+ seed,
3557
+ generation_context=refill_context,
3558
+ )
3559
+ else:
3560
+ raw_refill_trajectory = candidate_generator(refill_chunks, seed)
3561
  except Exception as error:
3562
  raise RuntimeError("chunk refill generation failed") from error
3563
  refill_trajectory = _coverage_trajectory_tuple(raw_refill_trajectory, 1)
3564
  refill_generation_evidence = generation_evidence(
3565
+ refill_context,
 
 
3566
  refill_chunks,
3567
  )
3568
  attempted_seeds.append(seed)
tests/test_coverage_adaptive.py CHANGED
@@ -6,6 +6,7 @@ import pytest
6
 
7
  from quality_runtime import (
8
  CASCADE_EVIDENCE_LOG_PREFIX,
 
9
  CandidateGenerationEvidence,
10
  CandidateObservation,
11
  CandidateVerification,
@@ -14,6 +15,7 @@ from quality_runtime import (
14
  SEQUENCE_FALLBACK_MAX_LOCAL_BOUNDARY_SPEAKER_DROP,
15
  TrajectoryGateResult,
16
  format_cascade_evidence_log,
 
17
  run_coverage_adaptive_cascade,
18
  select_culprit_diverse_candidate_sequences,
19
  trajectory_gate_evidence,
@@ -180,7 +182,7 @@ def test_refill_budget_accounts_exact_generated_chunks_and_text_units():
180
  selection=result,
181
  )
182
  payload = json.loads(line.removeprefix(CASCADE_EVIDENCE_LOG_PREFIX))
183
- assert payload["schema_version"] == 2
184
  assert payload["generation_evidence_complete"] is True
185
  assert payload["request_chunk_count"] == 2
186
  assert payload["generated_chunk_count"] == 3
@@ -190,6 +192,7 @@ def test_refill_budget_accounts_exact_generated_chunks_and_text_units():
190
  "generated_text_units": 6,
191
  }
192
  assert payload["attempts"][1]["chunk_indices"] == [0]
 
193
  assert payload["attempts"][1]["scheduled_cfg"] == 2.0
194
  assert payload["attempts"][1]["effective_cfgs"] == [3.0]
195
  assert payload["attempts"][1]["floor_reasons"] == [["short_text"]]
@@ -198,7 +201,9 @@ def test_refill_budget_accounts_exact_generated_chunks_and_text_units():
198
  "chunk_scheduled_cfgs": [2.0, 3.0],
199
  "chunk_effective_cfgs": [3.0, 3.0],
200
  "chunk_floor_reasons": [["short_text"], []],
 
201
  "chunk_policies": ["safe_duration", "base"],
 
202
  }
203
 
204
 
@@ -361,6 +366,240 @@ def test_refill_schedule_and_selected_path_are_deterministic():
361
  assert first.chunk_candidate_counts == second.chunk_candidate_counts == (2, 2, 2)
362
 
363
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
364
  def test_zero_coverage_rows_are_refilled_before_low_coverage_rows_round_robin():
365
  chunks = ("第一段", "第二段", "第三段")
366
  generated = []
 
6
 
7
  from quality_runtime import (
8
  CASCADE_EVIDENCE_LOG_PREFIX,
9
+ CandidateGenerationContext,
10
  CandidateGenerationEvidence,
11
  CandidateObservation,
12
  CandidateVerification,
 
15
  SEQUENCE_FALLBACK_MAX_LOCAL_BOUNDARY_SPEAKER_DROP,
16
  TrajectoryGateResult,
17
  format_cascade_evidence_log,
18
+ generation_cfg_for_candidate_offset,
19
  run_coverage_adaptive_cascade,
20
  select_culprit_diverse_candidate_sequences,
21
  trajectory_gate_evidence,
 
182
  selection=result,
183
  )
184
  payload = json.loads(line.removeprefix(CASCADE_EVIDENCE_LOG_PREFIX))
185
+ assert payload["schema_version"] == 3
186
  assert payload["generation_evidence_complete"] is True
187
  assert payload["request_chunk_count"] == 2
188
  assert payload["generated_chunk_count"] == 3
 
192
  "generated_text_units": 6,
193
  }
194
  assert payload["attempts"][1]["chunk_indices"] == [0]
195
+ assert payload["attempts"][1]["chunk_candidate_ordinals"] == [1]
196
  assert payload["attempts"][1]["scheduled_cfg"] == 2.0
197
  assert payload["attempts"][1]["effective_cfgs"] == [3.0]
198
  assert payload["attempts"][1]["floor_reasons"] == [["short_text"]]
 
201
  "chunk_scheduled_cfgs": [2.0, 3.0],
202
  "chunk_effective_cfgs": [3.0, 3.0],
203
  "chunk_floor_reasons": [["short_text"], []],
204
+ "chunk_candidate_ordinals": [1, 0],
205
  "chunk_policies": ["safe_duration", "base"],
206
+ "network_conditioned": [False, False],
207
  }
208
 
209
 
 
366
  assert first.chunk_candidate_counts == second.chunk_candidate_counts == (2, 2, 2)
367
 
368
 
369
+ def test_context_aware_refills_rotate_cfg_and_endpoint_policy_per_chunk():
370
+ chunks = ("第一段", "第二段")
371
+ contexts = []
372
+
373
+ def generator(candidate_chunks, seed, *, generation_context):
374
+ assert isinstance(generation_context, CandidateGenerationContext)
375
+ contexts.append(generation_context)
376
+ return generation_context.chunk_candidate_ordinals
377
+
378
+ def generation_evidence(
379
+ candidate_index,
380
+ seed,
381
+ chunk_indices,
382
+ candidate_chunks,
383
+ *,
384
+ generation_context,
385
+ ):
386
+ assert candidate_index == generation_context.candidate_index
387
+ assert seed == generation_context.seed
388
+ assert chunk_indices == generation_context.chunk_indices
389
+ ordinal = generation_context.chunk_candidate_ordinals[0]
390
+ assert all(
391
+ candidate_ordinal == ordinal
392
+ for candidate_ordinal in generation_context.chunk_candidate_ordinals
393
+ )
394
+ scheduled = generation_cfg_for_candidate_offset(ordinal)
395
+ return CandidateGenerationEvidence(
396
+ chunk_indices=chunk_indices,
397
+ chunk_text_units=tuple(3 for _ in candidate_chunks),
398
+ scheduled_cfg=scheduled,
399
+ effective_cfgs=tuple(scheduled for _ in candidate_chunks),
400
+ floor_reasons=tuple(() for _ in candidate_chunks),
401
+ chunk_candidate_ordinals=(ordinal,) * len(candidate_chunks),
402
+ )
403
+
404
+ result = run_coverage_adaptive_cascade(
405
+ chunks,
406
+ 100,
407
+ generator,
408
+ lambda trajectory, candidate_chunks, seed: _joined_rejected_local(
409
+ candidate_chunks,
410
+ passing=(False, False),
411
+ ),
412
+ lambda trajectory, candidate_chunks, seed: _local_verification(
413
+ candidate_chunks,
414
+ passing=(trajectory[0] >= 4,),
415
+ ),
416
+ sequence_final_verifier=lambda result, candidate_chunks: _exact_final(
417
+ candidate_chunks
418
+ ),
419
+ generation_evidence_factory=generation_evidence,
420
+ max_generated_chunks=10,
421
+ max_generated_text_units=100,
422
+ max_sequence_paths=1,
423
+ )
424
+
425
+ assert [context.candidate_index for context in contexts] == list(range(9))
426
+ assert [context.seed for context in contexts] == list(range(100, 109))
427
+ assert [context.chunk_indices for context in contexts] == [
428
+ (0, 1),
429
+ (0,),
430
+ (1,),
431
+ (0,),
432
+ (1,),
433
+ (0,),
434
+ (1,),
435
+ (0,),
436
+ (1,),
437
+ ]
438
+ assert [context.chunk_candidate_ordinals for context in contexts] == [
439
+ (0, 0),
440
+ (1,),
441
+ (1,),
442
+ (2,),
443
+ (2,),
444
+ (3,),
445
+ (3,),
446
+ (4,),
447
+ (4,),
448
+ ]
449
+ assert result.attempted_seeds == tuple(range(100, 109))
450
+ assert result.chunk_candidate_indices == (7, 8)
451
+ assert result.chunk_seeds == (107, 108)
452
+
453
+ line = format_cascade_evidence_log(
454
+ result.diagnostics,
455
+ outcome="returned",
456
+ generated_chunk_limit=10,
457
+ generated_text_unit_limit=100,
458
+ selection=result,
459
+ )
460
+ payload = json.loads(line.removeprefix(CASCADE_EVIDENCE_LOG_PREFIX))
461
+ assert payload["generation_evidence_complete"] is True
462
+ assert payload["attempts"][7]["candidate_index"] == 7
463
+ assert payload["attempts"][7]["seed"] == 107
464
+ assert payload["attempts"][7]["chunk_candidate_ordinals"] == [4]
465
+ assert payload["attempts"][7]["scheduled_cfg"] == 3.0
466
+ assert payload["attempts"][7]["policy"] == "completion_headroom"
467
+ assert payload["attempts"][7]["chunk_policies"] == [
468
+ "completion_headroom"
469
+ ]
470
+ assert payload["selection"]["generation"]["chunk_candidate_ordinals"] == [
471
+ 4,
472
+ 4,
473
+ ]
474
+ assert payload["selection"]["generation"]["chunk_policies"] == [
475
+ "completion_headroom",
476
+ "completion_headroom",
477
+ ]
478
+
479
+
480
+ def test_context_generation_and_evidence_must_opt_in_together_and_match():
481
+ def context_generator(chunks, seed, *, generation_context):
482
+ return tuple(chunks)
483
+
484
+ with pytest.raises(ValueError, match="must opt in together"):
485
+ run_coverage_adaptive_cascade(
486
+ ("完整內容",),
487
+ 10,
488
+ context_generator,
489
+ lambda trajectory, chunks, seed: _local_verification(chunks),
490
+ lambda trajectory, chunks, seed: _local_verification(chunks),
491
+ sequence_final_verifier=lambda result, chunks: _exact_final(chunks),
492
+ )
493
+
494
+ def mismatched_evidence(
495
+ candidate_index,
496
+ seed,
497
+ chunk_indices,
498
+ candidate_chunks,
499
+ *,
500
+ generation_context,
501
+ ):
502
+ return CandidateGenerationEvidence(
503
+ chunk_indices=chunk_indices,
504
+ chunk_text_units=(4,),
505
+ scheduled_cfg=2.0,
506
+ effective_cfgs=(2.0,),
507
+ floor_reasons=((),),
508
+ chunk_candidate_ordinals=(1,),
509
+ )
510
+
511
+ with pytest.raises(RuntimeError, match="generation evidence is invalid"):
512
+ run_coverage_adaptive_cascade(
513
+ ("完整內容",),
514
+ 10,
515
+ context_generator,
516
+ lambda trajectory, chunks, seed: _local_verification(chunks),
517
+ lambda trajectory, chunks, seed: _local_verification(chunks),
518
+ sequence_final_verifier=lambda result, chunks: _exact_final(chunks),
519
+ generation_evidence_factory=mismatched_evidence,
520
+ )
521
+
522
+
523
+ def test_context_generation_records_and_enforces_network_cfg_provenance():
524
+ def generator(chunks, seed, *, generation_context):
525
+ return tuple(chunks)
526
+
527
+ def network_evidence(
528
+ candidate_index,
529
+ seed,
530
+ chunk_indices,
531
+ candidate_chunks,
532
+ *,
533
+ generation_context,
534
+ ):
535
+ ordinal = generation_context.chunk_candidate_ordinals[0]
536
+ scheduled = generation_cfg_for_candidate_offset(ordinal)
537
+ return CandidateGenerationEvidence(
538
+ chunk_indices=chunk_indices,
539
+ chunk_text_units=(4,),
540
+ scheduled_cfg=scheduled,
541
+ effective_cfgs=(3.0,),
542
+ floor_reasons=(("network",),) if scheduled < 3.0 else ((),),
543
+ chunk_candidate_ordinals=(ordinal,),
544
+ network_conditioned=(True,),
545
+ )
546
+
547
+ result = run_coverage_adaptive_cascade(
548
+ ("完整內容",),
549
+ 10,
550
+ generator,
551
+ lambda trajectory, chunks, seed: _joined_rejected_local(chunks),
552
+ lambda trajectory, chunks, seed: _local_verification(chunks),
553
+ sequence_final_verifier=lambda result, chunks: _exact_final(chunks),
554
+ generation_evidence_factory=network_evidence,
555
+ max_generated_chunks=2,
556
+ max_generated_text_units=8,
557
+ max_sequence_paths=1,
558
+ )
559
+
560
+ assert result.diagnostics.attempts[1].chunk_candidate_ordinals == (1,)
561
+ assert result.diagnostics.attempts[1].network_conditioned == (True,)
562
+ assert result.diagnostics.attempts[1].floor_reasons == (("network",),)
563
+ payload = json.loads(
564
+ format_cascade_evidence_log(
565
+ result.diagnostics,
566
+ outcome="returned",
567
+ generated_chunk_limit=2,
568
+ generated_text_unit_limit=8,
569
+ selection=result,
570
+ ).removeprefix(CASCADE_EVIDENCE_LOG_PREFIX)
571
+ )
572
+ assert payload["attempts"][1]["network_conditioned"] == [True]
573
+ assert payload["selection"]["generation"]["network_conditioned"] == [True]
574
+
575
+ def forged_network_evidence(*args, generation_context, **kwargs):
576
+ ordinal = generation_context.chunk_candidate_ordinals[0]
577
+ scheduled = generation_cfg_for_candidate_offset(ordinal)
578
+ return CandidateGenerationEvidence(
579
+ chunk_indices=generation_context.chunk_indices,
580
+ chunk_text_units=(4,),
581
+ scheduled_cfg=scheduled,
582
+ effective_cfgs=(3.0,),
583
+ floor_reasons=((),),
584
+ chunk_candidate_ordinals=(ordinal,),
585
+ network_conditioned=(True,),
586
+ )
587
+
588
+ with pytest.raises(RuntimeError, match="generation evidence is invalid"):
589
+ run_coverage_adaptive_cascade(
590
+ ("完整內容",),
591
+ 10,
592
+ generator,
593
+ lambda trajectory, chunks, seed: _joined_rejected_local(chunks),
594
+ lambda trajectory, chunks, seed: _local_verification(chunks),
595
+ sequence_final_verifier=lambda result, chunks: _exact_final(chunks),
596
+ generation_evidence_factory=forged_network_evidence,
597
+ max_generated_chunks=2,
598
+ max_generated_text_units=8,
599
+ max_sequence_paths=1,
600
+ )
601
+
602
+
603
  def test_zero_coverage_rows_are_refilled_before_low_coverage_rows_round_robin():
604
  chunks = ("第一段", "第二段", "第三段")
605
  generated = []
tests/test_production.py CHANGED
@@ -20,13 +20,16 @@ from production import (
20
  ensure_terminal_punctuation,
21
  finish_audio,
22
  fade_internal_edges,
 
23
  join_audio_chunks,
 
24
  normalize_asr_spoken_forms,
25
  mandarin_acoustic_units,
26
  network_protected_spoken_spans,
27
  normalize_spoken_forms,
28
  normalize_tts_eval_text,
29
  normalize_tts_text,
 
30
  select_candidate_sequence,
31
  select_generation_cps,
32
  split_leading_clause,
@@ -927,6 +930,168 @@ def test_frontend_clause_split_matrix_for_structured_and_network_holdouts():
927
  ) == 4.6
928
 
929
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
930
  def test_split_preserves_long_strong_sentences_and_is_deterministic():
931
  sentences = (
932
  "甲" * 59 + "。",
@@ -987,6 +1152,60 @@ def test_onset_split_requires_a_natural_boundary():
987
  ]
988
 
989
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
990
  def test_prefaded_join_does_not_apply_a_second_edge_ramp():
991
  sample_rate = 1_000
992
  chunks = fade_internal_edges(
 
20
  ensure_terminal_punctuation,
21
  finish_audio,
22
  fade_internal_edges,
23
+ fade_variable_internal_edges,
24
  join_audio_chunks,
25
+ join_audio_chunks_variable,
26
  normalize_asr_spoken_forms,
27
  mandarin_acoustic_units,
28
  network_protected_spoken_spans,
29
  normalize_spoken_forms,
30
  normalize_tts_eval_text,
31
  normalize_tts_text,
32
+ plan_generation_chunks,
33
  select_candidate_sequence,
34
  select_generation_cps,
35
  split_leading_clause,
 
930
  ) == 4.6
931
 
932
 
933
+ def test_generation_only_network_component_planner_holdout_matrix():
934
+ raw_h11 = (
935
+ "山區步道的志工預計在 2028/10/21 上午 07:15 集合,"
936
+ "先用定位器 AX-520 核對座標,再分組檢查木棧道、里程牌與飲水站。"
937
+ "若氣象網站 https://trailweather.example.tw 顯示降雨機率超過 65%,"
938
+ "領隊就取消高海拔路線,改走較短的林間環線。"
939
+ "途中若發現落石或樹枝阻斷通行,請拍照並寄到 "
940
+ "patrol@forestmail.tw,不要自行搬動大型障礙物。"
941
+ "所有隊員回到登山口後,還要清點無線電與急救包,"
942
+ "確認沒有任何人落單,才結束當天的巡查。"
943
+ )
944
+ matrix = {
945
+ "H06": (
946
+ "若要更換導覽場次,請寄信到 tour.help@islandmuseum.tw。",
947
+ (24, 25),
948
+ ),
949
+ "H07": (
950
+ "潮汐預報可查詢 https://coastwatch.example.tw/tide。",
951
+ (34, 23),
952
+ ),
953
+ "H11": (raw_h11, (30, 29, 18, 35, 33, 28, 34, 38)),
954
+ }
955
+
956
+ for raw, expected_units in matrix.values():
957
+ normalized = normalize_spoken_forms(raw)
958
+ specs = plan_generation_chunks(raw, normalized)
959
+
960
+ assert tuple(count_speech_units(spec.text) for spec in specs) == expected_units
961
+ assert normalize_tts_text(" ".join(spec.text for spec in specs)) == normalized
962
+ assert all(
963
+ spec.source_end == following.source_start
964
+ for spec, following in zip(specs, specs[1:])
965
+ )
966
+ assert all(
967
+ 12 <= count_speech_units(spec.text) <= (36 if spec.network_conditioned else 80)
968
+ for spec in specs
969
+ )
970
+ assert any(spec.boundary_after == "network_internal" for spec in specs)
971
+ for spec in specs:
972
+ assert len(spec.network_span_indices) == len(
973
+ spec.network_full_spoken_proofs
974
+ )
975
+
976
+
977
+ def test_generation_network_planner_keeps_ordinary_semantic_chunks_unchanged():
978
+ raw = (
979
+ "清晨開館以前,水族館人員會先量測各池的水溫與鹽度,"
980
+ "再觀察魚群是否正常進食。確認設備都正常後才開門。"
981
+ )
982
+ normalized = normalize_spoken_forms(raw)
983
+
984
+ specs = plan_generation_chunks(raw, normalized)
985
+
986
+ assert [spec.text for spec in specs] == split_text_for_tts(
987
+ normalized,
988
+ max_chars=80,
989
+ min_chunk_chars=12,
990
+ )
991
+ assert all(not spec.network_conditioned for spec in specs)
992
+
993
+
994
+ def test_generation_network_planner_fails_closed_on_invalid_provenance():
995
+ raw = "請寄到 tour.help@islandmuseum.tw。"
996
+ different = normalize_spoken_forms(
997
+ "請寄到 tour.help@differentmuseum.tw。"
998
+ )
999
+ with pytest.raises(ValueError, match="does not match"):
1000
+ plan_generation_chunks(raw, different)
1001
+
1002
+ iri = "請查詢 https://example.tw/愛。"
1003
+ with pytest.raises(ValueError, match="non-ASCII IRI"):
1004
+ plan_generation_chunks(iri, normalize_spoken_forms(iri))
1005
+
1006
+ emoji_iri = "請查詢 https://example.tw/🐦。"
1007
+ with pytest.raises(ValueError, match="non-ASCII IRI"):
1008
+ plan_generation_chunks(emoji_iri, normalize_spoken_forms(emoji_iri))
1009
+
1010
+ oversized = "請查詢 https://" + "a" * 50 + ".tw。"
1011
+ with pytest.raises(ValueError, match="indivisible network component"):
1012
+ plan_generation_chunks(oversized, normalize_spoken_forms(oversized))
1013
+
1014
+
1015
+ def test_generation_network_planner_binds_duplicate_spoken_proof_to_raw_identifier():
1016
+ spoken = (
1017
+ "艾取 踢 踢 批 艾斯 冒號 斜線 斜線 "
1018
+ "伊 艾克斯 欸 艾姆 批 艾爾 伊 點 踢 達不溜"
1019
+ )
1020
+ raw = f"先念 {spoken},再查 https://example.tw。"
1021
+ normalized = normalize_spoken_forms(raw)
1022
+
1023
+ specs = plan_generation_chunks(raw, normalized)
1024
+
1025
+ assert "".join(spec.text for spec in specs) == normalized
1026
+ assert [spec.network_conditioned for spec in specs] == [False, True]
1027
+ assert spoken in specs[0].text
1028
+ assert specs[1].network_full_spoken_proofs == (spoken,)
1029
+ assert all(
1030
+ normalized[spec.source_start : spec.source_end] == spec.text
1031
+ for spec in specs
1032
+ )
1033
+
1034
+
1035
+ def test_generation_network_planner_is_range_first_for_english_and_long_prose():
1036
+ cases = (
1037
+ "Please visit https://example.tw。",
1038
+ "甲" * 35 + "請查 https://c.tw。",
1039
+ )
1040
+ for raw in cases:
1041
+ normalized = normalize_spoken_forms(raw)
1042
+ specs = plan_generation_chunks(raw, normalized)
1043
+
1044
+ assert "".join(spec.text for spec in specs) == normalized
1045
+ assert all(
1046
+ count_speech_units(spec.text)
1047
+ <= (36 if spec.network_conditioned else 80)
1048
+ for spec in specs
1049
+ )
1050
+ prose_specs = plan_generation_chunks(cases[1], normalize_spoken_forms(cases[1]))
1051
+ assert any(
1052
+ not spec.network_conditioned and count_speech_units(spec.text) > 36
1053
+ for spec in prose_specs
1054
+ )
1055
+
1056
+
1057
+ def test_generation_network_planner_splits_long_and_repeated_identifiers_by_proof():
1058
+ cases = (
1059
+ "請查 https://example.tw/a/b/c/d/e/f/g/h/i/j/k/l/m/n/o/p。",
1060
+ "https://a.tw https://a.tw",
1061
+ )
1062
+ for raw in cases:
1063
+ normalized = normalize_spoken_forms(raw)
1064
+ specs = plan_generation_chunks(raw, normalized)
1065
+
1066
+ assert "".join(spec.text for spec in specs) == normalized
1067
+ assert all(count_speech_units(spec.text) <= 36 for spec in specs)
1068
+ assert all(spec.network_conditioned for spec in specs)
1069
+ repeated = plan_generation_chunks(cases[1], normalize_spoken_forms(cases[1]))
1070
+ assert {index for spec in repeated for index in spec.network_span_indices} == {0, 1}
1071
+
1072
+
1073
+ @pytest.mark.parametrize(
1074
+ "raw",
1075
+ (
1076
+ "a@b.co https://c.de",
1077
+ "https://a.co b@c.de",
1078
+ "a@b.co b@c.de",
1079
+ "https://a.co https://b.de",
1080
+ ),
1081
+ )
1082
+ def test_generation_network_planner_maps_whitespace_adjacent_identifiers(raw):
1083
+ normalized = normalize_spoken_forms(raw)
1084
+
1085
+ specs = plan_generation_chunks(raw, normalized)
1086
+
1087
+ assert "".join(spec.text for spec in specs) == normalized
1088
+ assert {index for spec in specs for index in spec.network_span_indices} == {
1089
+ 0,
1090
+ 1,
1091
+ }
1092
+ assert all(spec.network_conditioned for spec in specs)
1093
+
1094
+
1095
  def test_split_preserves_long_strong_sentences_and_is_deterministic():
1096
  sentences = (
1097
  "甲" * 59 + "。",
 
1152
  ]
1153
 
1154
 
1155
+ def test_variable_network_boundary_uses_short_fade_without_losing_chunks():
1156
+ sample_rate = 1_000
1157
+ chunks = [np.ones(1_000, dtype=np.float32) for _ in range(3)]
1158
+
1159
+ faded = fade_variable_internal_edges(
1160
+ chunks,
1161
+ sample_rate,
1162
+ fade_ms_by_boundary=[5.0, 80.0],
1163
+ )
1164
+ joined = join_audio_chunks_variable(
1165
+ faded,
1166
+ pauses=[0, 10],
1167
+ crossfade_samples_by_boundary=[5, 80],
1168
+ pre_faded_edges=True,
1169
+ )
1170
+
1171
+ assert joined.shape == (3_005,)
1172
+ assert np.isfinite(joined).all()
1173
+ assert np.max(np.abs(joined)) <= 1.0
1174
+ assert np.count_nonzero(faded[0]) >= 995
1175
+ assert np.count_nonzero(faded[1]) >= 915
1176
+ with pytest.raises(ValueError, match="boundary count"):
1177
+ fade_variable_internal_edges(chunks, sample_rate, [5.0])
1178
+ with pytest.raises(ValueError, match="boundary counts"):
1179
+ join_audio_chunks_variable(chunks, [0, 0], [5])
1180
+
1181
+
1182
+ def test_variable_audio_boundaries_reject_empty_overlap_and_invalid_widths():
1183
+ sample_rate = 1_000
1184
+ empty_middle = [
1185
+ np.ones(10, dtype=np.float32),
1186
+ np.zeros(0, dtype=np.float32),
1187
+ np.ones(10, dtype=np.float32),
1188
+ ]
1189
+ with pytest.raises(ValueError, match="non-empty"):
1190
+ fade_variable_internal_edges(empty_middle, sample_rate, [5.0, 5.0])
1191
+ with pytest.raises(ValueError, match="non-empty"):
1192
+ join_audio_chunks_variable(empty_middle, [0, 0], [5, 5])
1193
+
1194
+ short_middle = [
1195
+ np.ones(10, dtype=np.float32),
1196
+ np.ones(6, dtype=np.float32),
1197
+ np.ones(10, dtype=np.float32),
1198
+ ]
1199
+ with pytest.raises(ValueError, match="overlap"):
1200
+ fade_variable_internal_edges(short_middle, sample_rate, [4.0, 4.0])
1201
+ with pytest.raises(ValueError, match="overlap"):
1202
+ join_audio_chunks_variable(short_middle, [0, 0], [4, 4])
1203
+ with pytest.raises(ValueError, match="non-negative integers"):
1204
+ join_audio_chunks_variable(short_middle, [-1, 0], [0, 0])
1205
+ with pytest.raises(ValueError, match="non-negative integers"):
1206
+ join_audio_chunks_variable(short_middle, [0, 0], [1.5, 0])
1207
+
1208
+
1209
  def test_prefaded_join_does_not_apply_a_second_edge_ramp():
1210
  sample_rate = 1_000
1211
  chunks = fade_internal_edges(
tests/test_release_pins.py CHANGED
@@ -237,11 +237,13 @@ def test_readme_describes_coverage_refill_and_sequence_transition_scores():
237
  assert "1→5→10→15→20" not in app_source
238
 
239
 
240
- def test_app_wires_candidate_offset_to_explicit_generation_policy_and_logs_it():
241
  source = (ROOT / "app.py").read_text(encoding="utf-8")
242
 
243
  assert source.count("policy: GenerationPolicy") == 2
244
- assert "policy=generation_policy_for_candidate_offset(seed - request_seed)" in source
 
 
245
  assert 'f"name={policy.name}' in source
246
  assert "chunk_policies={selected_policies}" in source
247
  assert '"min_len": min_len' in source
@@ -255,16 +257,17 @@ def test_app_wires_candidate_offset_to_explicit_generation_policy_and_logs_it():
255
  def test_app_rejects_ambiguous_iri_before_frontend_normalization():
256
  source = (ROOT / "app.py").read_text(encoding="utf-8")
257
  synthesize_start = source.index("def _synthesize(")
 
258
  iri_guard = source.index(
259
- "if network_identifier_has_ambiguous_iri(text):",
260
  synthesize_start,
261
  )
262
  normalization = source.index(
263
- 'text = normalize_spoken_forms(text, locale="zh-TW")',
264
  synthesize_start,
265
  )
266
 
267
- assert synthesize_start < iri_guard < normalization
268
  assert "非 ASCII IRI 必須先轉成 ASCII/percent-encoded" in (
269
  ROOT / "README.md"
270
  ).read_text(encoding="utf-8")
@@ -302,13 +305,17 @@ def test_app_applies_fixed_mixed_cfg_schedule_after_global_quality_floor():
302
  assert "MIXED_CFG_PRIMARY = 3.0" in source
303
  assert "MIXED_CFG_ALTERNATE = 2.0" in source
304
  assert (
305
- 'MIXED_CFG_SCHEDULE = "offset_zero_and_even_primary_odd_alternate"'
306
  in source
307
  )
308
- assert "network_request = contains_network_identifier(text)" in synthesize_source
309
  assert "cfg_value != MIXED_CFG_PRIMARY" in synthesize_source
310
  assert "request_cfg = MIXED_CFG_PRIMARY" in synthesize_source
311
- assert "cfg=candidate_cfg(seed)" in synthesize_source
 
 
 
 
312
  assert "generation_cfg_for_candidate_offset(" in synthesize_source
313
  assert "chunk_cfgs={selected_cfgs}" in synthesize_source
314
  assert "attempted_schedule_cfgs={attempted_schedule_cfgs}" in synthesize_source
@@ -360,9 +367,20 @@ def test_app_rejects_silent_text_and_coalesces_before_runtime_budgeting():
360
  assert synthesize_source is not None
361
  assert assemble_source is not None
362
  assert "if count_speech_units(text) <= 0" in synthesize_source
363
- assert "chunks = coalesce_text_chunks(" in synthesize_source
 
 
364
  assert "max_chunks=QUALITY_MAX_GENERATED_CHUNKS" in synthesize_source
365
  assert "pre_faded_edges=True" in assemble_source
 
 
 
 
 
 
 
 
 
366
 
367
 
368
  def test_app_emits_one_canonical_content_free_evidence_line_per_terminal_outcome():
@@ -468,7 +486,7 @@ def test_app_reverifies_the_post_join_speed_adjusted_whole_waveform():
468
  production_source = (ROOT / "production.py").read_text(encoding="utf-8")
469
  assert "trailing_silence_ms: float = 180.0" in production_source
470
  assemble_index = synthesize_source.index(
471
- "waveform = _assemble_trajectory_audio(cascade.trajectory, chunks, speed)"
472
  )
473
  verify_index = synthesize_source.index("final_verification = _verify_trajectory_audio(")
474
  require_index = synthesize_source.index(
@@ -549,10 +567,9 @@ def test_whole_candidate_qualification_uses_the_exact_return_assembler_after_loc
549
  assert "QUALITY_FINAL_ASR_MAX_NEW_TOKENS" in qualify_source[joined_index:]
550
  assert "release_speaker_gate=True" in qualify_source[joined_index:]
551
  assert "qualify_trajectory_with_joined_output(" in qualify_source[joined_index:]
552
- assert (
553
- "waveform = _assemble_trajectory_audio(cascade.trajectory, chunks, speed)"
554
- in synthesize_source
555
- )
556
  assert "require_verified_final_output(final_verification)" in synthesize_source
557
 
558
 
@@ -635,7 +652,7 @@ def test_space_hard_intersects_dual_asr_only_on_exact_whole_waveforms():
635
  "cascade = run_coverage_adaptive_cascade("
636
  )
637
  final_assemble = synthesize_source.index(
638
- "waveform = _assemble_trajectory_audio(cascade.trajectory, chunks, speed)"
639
  )
640
  final_turbo = synthesize_source.index("final_verification = _verify_trajectory_audio(")
641
  final_turbo_require = synthesize_source.index(
 
237
  assert "1→5→10→15→20" not in app_source
238
 
239
 
240
+ def test_app_wires_row_local_candidate_ordinal_to_generation_policy_and_logs_it():
241
  source = (ROOT / "app.py").read_text(encoding="utf-8")
242
 
243
  assert source.count("policy: GenerationPolicy") == 2
244
+ assert "generation_context.chunk_candidate_ordinals" in source
245
+ assert "policy=generation_policy_for_candidate_offset(candidate_ordinal)" in source
246
+ assert "generation_context.seed != seed" in source
247
  assert 'f"name={policy.name}' in source
248
  assert "chunk_policies={selected_policies}" in source
249
  assert '"min_len": min_len' in source
 
257
  def test_app_rejects_ambiguous_iri_before_frontend_normalization():
258
  source = (ROOT / "app.py").read_text(encoding="utf-8")
259
  synthesize_start = source.index("def _synthesize(")
260
+ raw_text = source.index("raw_text = str(text)", synthesize_start)
261
  iri_guard = source.index(
262
+ "if network_identifier_has_ambiguous_iri(raw_text):",
263
  synthesize_start,
264
  )
265
  normalization = source.index(
266
+ 'text = normalize_spoken_forms(raw_text, locale="zh-TW")',
267
  synthesize_start,
268
  )
269
 
270
+ assert synthesize_start < raw_text < iri_guard < normalization
271
  assert "非 ASCII IRI 必須先轉成 ASCII/percent-encoded" in (
272
  ROOT / "README.md"
273
  ).read_text(encoding="utf-8")
 
305
  assert "MIXED_CFG_PRIMARY = 3.0" in source
306
  assert "MIXED_CFG_ALTERNATE = 2.0" in source
307
  assert (
308
+ 'MIXED_CFG_SCHEDULE = "row_ordinal_zero_and_even_primary_odd_alternate"'
309
  in source
310
  )
311
+ assert "network_request = contains_network_identifier(raw_text)" in synthesize_source
312
  assert "cfg_value != MIXED_CFG_PRIMARY" in synthesize_source
313
  assert "request_cfg = MIXED_CFG_PRIMARY" in synthesize_source
314
+ assert "cfg=candidate_cfg(candidate_ordinal)" in synthesize_source
315
+ assert "candidate_ordinals = generation_context.chunk_candidate_ordinals" in (
316
+ synthesize_source
317
+ )
318
+ assert "chunk_candidate_ordinals=candidate_ordinals" in synthesize_source
319
  assert "generation_cfg_for_candidate_offset(" in synthesize_source
320
  assert "chunk_cfgs={selected_cfgs}" in synthesize_source
321
  assert "attempted_schedule_cfgs={attempted_schedule_cfgs}" in synthesize_source
 
367
  assert synthesize_source is not None
368
  assert assemble_source is not None
369
  assert "if count_speech_units(text) <= 0" in synthesize_source
370
+ assert "coalesce_text_chunks(" in synthesize_source
371
+ assert "chunk_specs = plan_generation_chunks(" in synthesize_source
372
+ assert "chunks = tuple(spec.text for spec in chunk_specs)" in synthesize_source
373
  assert "max_chunks=QUALITY_MAX_GENERATED_CHUNKS" in synthesize_source
374
  assert "pre_faded_edges=True" in assemble_source
375
+ assert "NETWORK_GENERATION_TARGET_UNITS = 32" in source
376
+ assert "NETWORK_GENERATION_MAX_UNITS = 36" in source
377
+ assert "NETWORK_INTERNAL_FADE_MS = 5.0" in source
378
+ assert 'chunk_specs[index].boundary_after == "network_internal"' in (
379
+ assemble_source
380
+ )
381
+ assert "join_audio_chunks_variable(" in assemble_source
382
+ assert "network_conditioned=network_flags" in synthesize_source
383
+ assert "network_conditioned=network_flag" in synthesize_source
384
 
385
 
386
  def test_app_emits_one_canonical_content_free_evidence_line_per_terminal_outcome():
 
486
  production_source = (ROOT / "production.py").read_text(encoding="utf-8")
487
  assert "trailing_silence_ms: float = 180.0" in production_source
488
  assemble_index = synthesize_source.index(
489
+ "waveform = _assemble_trajectory_audio("
490
  )
491
  verify_index = synthesize_source.index("final_verification = _verify_trajectory_audio(")
492
  require_index = synthesize_source.index(
 
567
  assert "QUALITY_FINAL_ASR_MAX_NEW_TOKENS" in qualify_source[joined_index:]
568
  assert "release_speaker_gate=True" in qualify_source[joined_index:]
569
  assert "qualify_trajectory_with_joined_output(" in qualify_source[joined_index:]
570
+ assert "waveform = _assemble_trajectory_audio(" in synthesize_source
571
+ assert "cascade.trajectory" in synthesize_source
572
+ assert "chunk_specs" in synthesize_source
 
573
  assert "require_verified_final_output(final_verification)" in synthesize_source
574
 
575
 
 
652
  "cascade = run_coverage_adaptive_cascade("
653
  )
654
  final_assemble = synthesize_source.index(
655
+ "waveform = _assemble_trajectory_audio("
656
  )
657
  final_turbo = synthesize_source.index("final_verification = _verify_trajectory_audio(")
658
  final_turbo_require = synthesize_source.index(