docs(evals): F38 numeric refinements from the repricing verification writeup
Browse filesThe parallel verification (gpt56_repricing_counterfactual.md) showed the
priced >272k count is 176 calls by provider-reported input (191 counted by
the approximate context estimate), and the ~$0.55/M cached-input break-even
is a mix-specific aggregate (22-turn subset ~$2.91/M, 36-turn ~$0.39/M).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
evals.md
CHANGED
|
@@ -197,7 +197,7 @@ Numbered findings; each states what it was tested on. Convention: entries are ne
|
|
| 197 |
|
| 198 |
- **F37 (2026-07-15) β A chunk of "compaction's cost" was an implementation artifact, not physics: LangChain's stock summarizer breaks the provider cache on the summary-generation call itself, and a prefix-preserving compaction request removes that β summarizer cache hit 0%β94%, summarization cost β87%, total β14%/trajectory vs the stock arm, with equal-or-better memory. But the *post-compaction* prefix break is structural, so compaction still costs +14% vs capped full history: F35's ranking stands, with a smaller gap. (Same regime as F35/F36 β DeepSeek-V4-Flash first-party, absolute numbers not cross-comparable with the Gemini fleets.)** *Experiment: **DeepSeek compaction study, stage 1 follow-up β prefix-preserving ("structured") compaction.** The stock `SummarizationMiddleware` serializes the selected messages into a fresh XML prompt to ask for the summary, so even summary *generation* pays 100% cache-miss prices on ~150k-token inputs; `PrefixPreservingCompactionMiddleware` (`app/chat_service.py`, `summarization_strategy="structured_prefix"`) instead sends the *unchanged current request prefix* plus one checkpoint instruction, binding the same model/tools/settings, so the summarizer call rides the same cache as the agent. Single arm `exp_c200_cap10k_structured` β identical to `exp_c200_cap10k` except the strategy β run `runs/deepseek_structured_only_stage1_20260715` on the same `battery_sessions_v2` Γ 3 trials (414 turns, 0 errors), compared pairwise per sessionΓtrial against the four stage-1 arms; graded through the same blinded-subagent pipeline (429 judgments, 0 missing). Writeup `runs/deepseek_structured_only_stage1_20260715/report/structured_findings.md`; raw on HF `eval_runs/deepseek_compaction_stage1/`. A 2026-07-15 source-level check of the open-source Codex CLI confirmed this strategy is a faithful analog of Codex's local compaction β unchanged history + appended instruction, explicitly cache-preserving β while Codex's OpenAI-default *remote v2* path is the provider-native opaque-compaction primitive a client middleware cannot replicate; comparison + adoption ideas in `runs/deepseek_structured_only_stage1_20260715/report/codex_comparison.md`.* Cost: vs the stock XML arm, **β$0.0217/trajectory (β14.0%, cheaper in 11/15 pairs) at identical compaction dose** (1.27 events/trajectory both): the summarizer call's cache-hit ratio goes **0.0% β 94.2%**, so summarization cost falls $0.0276 β $0.0036/trajectory (β87%) β even though the structured arm *bills more* raw input (10.6M vs 9.9M tokens/trajectory; the checkpoint request re-reads the whole prefix at the ~50Γ cache-read discount). Tokens β dollars, again (F2/F25). vs capped full history it remains **+13.7% (cheaper in only 4/15)**: what's left is the structural boundary β installing the summary necessarily rewrites the *next* agent call's prefix, which no client-side middleware can avoid (that would take a provider-native compaction/continuation primitive) β plus the extra agent work (3.5 vs 3.1 LLM calls/turn). Memory: equal-best of the compaction arms β on audit-clean probes (excluding the two battery-compromised sessions, see F36) it scores **9/9**, tied with both full-history arms and capped-XML compaction, while raw compaction scores 7/9; the pre-audit "6/6 vs 5/6" long-horizon edge over the XML arm rode entirely on the compromised fullstack session and should not be cited β whether the intact-prefix summary *also* preserves facts better than the XML rewrite is unmeasured at this n. Holistic flat at 96%. Its single overall probe miss is a contradiction hedge in `python_colab_to_local_22t` β the confounded session β a battery artifact, not an arm signal. Trigger gate: compaction fired in 12/15 trajectories (19 events, min pre-compaction observation 200,053 tokens; the 3 dormant trajectories all Tier-1, and 100% of Tier-2 probes ran under compaction). Caveats: the memory edge is 1β2 probes at n=6 long-horizon / n=15 total β directional; the arm ran solo, not turn-level-interleaved with the stage-1 arms (token/cost metrics are run-time-independent, but treat its latency rows as confounded, F27); holistic judge not human-validated (F34); DeepSeek's ~50Γ cache discount is load-bearing (F25's provider caveat).
|
| 199 |
|
| 200 |
-
- **F38 (2026-07-15) β The keep-everything-vs-compact crossover is set by one number: the cached-input price. Repricing our stage-1 traces call-by-call shows "full history wins" is a fact about DeepSeek's pricing, not a law: at GPT-5.6 Sol's public rates the *same recorded workload* already flips to favor compaction at 36 turns β quantitatively consistent with why OpenAI's Codex auto-compacts at ~272k while our tutor, on DeepSeek, should not compact at all. (Analysis-only finding β no new run; adds the missing fourth boundary to F32's list.)** *Experiment: same-trace repricing over the stage-1 arms (runs `runs/deepseek_compaction_stage1_triggerfix_20260715` + `runs/deepseek_structured_only_stage1_20260715`, F35/F37): every recorded model call's (fresh, cached, output) token triple re-costed at GPT-5.6 Sol's public API rates ($5 / $0.50 / $30 per M β a 10Γ cache discount vs DeepSeek's ~50Γ), with and without the API's >272k long-context cliff (2Γ input / 1.5Γ output on the whole request; per OpenAI's Tibo, the cliff is NOT charged on Codex subscriptions β the stated real driver is accumulated cache-read cost scaling with the live window re-read on every tool call).* Result: `exp_fh_cap10k` vs `exp_c200_cap10k_structured` goes from **$0.117 vs $0.134/trajectory (full history β13%) at DeepSeek prices** to **$9.85 vs $9.94 (βtied) at GPT rates** β and the tie hides a horizon split: full history still wins the 22-turn tier ($6.29 vs $6.76) but **already loses the 36-turn tier ($15.18 vs $14.72), no pricing cliff involved** (directional: 3/6 tier-2 pairs, mean β$0.46; the API cliff widens it to $12.56 vs $9.94 overall
|
| 201 |
|
| 202 |
**Screen winners β promotion candidates:** `incontext_history_retrieval` (needs the long-session test), `retrieval_budget_30k` and `clear_retrieval_kb` (cheap Axis-B wins; clear_retrieval_kb is cheaper than prod with similar memory + better recall/key-points), anchored by `full_history`. **Drop:** `context_reset`, `observation_truncation`, `selective_retention`.
|
| 203 |
|
|
|
|
| 197 |
|
| 198 |
- **F37 (2026-07-15) β A chunk of "compaction's cost" was an implementation artifact, not physics: LangChain's stock summarizer breaks the provider cache on the summary-generation call itself, and a prefix-preserving compaction request removes that β summarizer cache hit 0%β94%, summarization cost β87%, total β14%/trajectory vs the stock arm, with equal-or-better memory. But the *post-compaction* prefix break is structural, so compaction still costs +14% vs capped full history: F35's ranking stands, with a smaller gap. (Same regime as F35/F36 β DeepSeek-V4-Flash first-party, absolute numbers not cross-comparable with the Gemini fleets.)** *Experiment: **DeepSeek compaction study, stage 1 follow-up β prefix-preserving ("structured") compaction.** The stock `SummarizationMiddleware` serializes the selected messages into a fresh XML prompt to ask for the summary, so even summary *generation* pays 100% cache-miss prices on ~150k-token inputs; `PrefixPreservingCompactionMiddleware` (`app/chat_service.py`, `summarization_strategy="structured_prefix"`) instead sends the *unchanged current request prefix* plus one checkpoint instruction, binding the same model/tools/settings, so the summarizer call rides the same cache as the agent. Single arm `exp_c200_cap10k_structured` β identical to `exp_c200_cap10k` except the strategy β run `runs/deepseek_structured_only_stage1_20260715` on the same `battery_sessions_v2` Γ 3 trials (414 turns, 0 errors), compared pairwise per sessionΓtrial against the four stage-1 arms; graded through the same blinded-subagent pipeline (429 judgments, 0 missing). Writeup `runs/deepseek_structured_only_stage1_20260715/report/structured_findings.md`; raw on HF `eval_runs/deepseek_compaction_stage1/`. A 2026-07-15 source-level check of the open-source Codex CLI confirmed this strategy is a faithful analog of Codex's local compaction β unchanged history + appended instruction, explicitly cache-preserving β while Codex's OpenAI-default *remote v2* path is the provider-native opaque-compaction primitive a client middleware cannot replicate; comparison + adoption ideas in `runs/deepseek_structured_only_stage1_20260715/report/codex_comparison.md`.* Cost: vs the stock XML arm, **β$0.0217/trajectory (β14.0%, cheaper in 11/15 pairs) at identical compaction dose** (1.27 events/trajectory both): the summarizer call's cache-hit ratio goes **0.0% β 94.2%**, so summarization cost falls $0.0276 β $0.0036/trajectory (β87%) β even though the structured arm *bills more* raw input (10.6M vs 9.9M tokens/trajectory; the checkpoint request re-reads the whole prefix at the ~50Γ cache-read discount). Tokens β dollars, again (F2/F25). vs capped full history it remains **+13.7% (cheaper in only 4/15)**: what's left is the structural boundary β installing the summary necessarily rewrites the *next* agent call's prefix, which no client-side middleware can avoid (that would take a provider-native compaction/continuation primitive) β plus the extra agent work (3.5 vs 3.1 LLM calls/turn). Memory: equal-best of the compaction arms β on audit-clean probes (excluding the two battery-compromised sessions, see F36) it scores **9/9**, tied with both full-history arms and capped-XML compaction, while raw compaction scores 7/9; the pre-audit "6/6 vs 5/6" long-horizon edge over the XML arm rode entirely on the compromised fullstack session and should not be cited β whether the intact-prefix summary *also* preserves facts better than the XML rewrite is unmeasured at this n. Holistic flat at 96%. Its single overall probe miss is a contradiction hedge in `python_colab_to_local_22t` β the confounded session β a battery artifact, not an arm signal. Trigger gate: compaction fired in 12/15 trajectories (19 events, min pre-compaction observation 200,053 tokens; the 3 dormant trajectories all Tier-1, and 100% of Tier-2 probes ran under compaction). Caveats: the memory edge is 1β2 probes at n=6 long-horizon / n=15 total β directional; the arm ran solo, not turn-level-interleaved with the stage-1 arms (token/cost metrics are run-time-independent, but treat its latency rows as confounded, F27); holistic judge not human-validated (F34); DeepSeek's ~50Γ cache discount is load-bearing (F25's provider caveat).
|
| 199 |
|
| 200 |
+
- **F38 (2026-07-15) β The keep-everything-vs-compact crossover is set by one number: the cached-input price. Repricing our stage-1 traces call-by-call shows "full history wins" is a fact about DeepSeek's pricing, not a law: at GPT-5.6 Sol's public rates the *same recorded workload* already flips to favor compaction at 36 turns β quantitatively consistent with why OpenAI's Codex auto-compacts at ~272k while our tutor, on DeepSeek, should not compact at all. (Analysis-only finding β no new run; adds the missing fourth boundary to F32's list.)** *Experiment: same-trace repricing over the stage-1 arms (runs `runs/deepseek_compaction_stage1_triggerfix_20260715` + `runs/deepseek_structured_only_stage1_20260715`, F35/F37): every recorded model call's (fresh, cached, output) token triple re-costed at GPT-5.6 Sol's public API rates ($5 / $0.50 / $30 per M β a 10Γ cache discount vs DeepSeek's ~50Γ), with and without the API's >272k long-context cliff (2Γ input / 1.5Γ output on the whole request; per OpenAI's Tibo, the cliff is NOT charged on Codex subscriptions β the stated real driver is accumulated cache-read cost scaling with the live window re-read on every tool call).* Result: `exp_fh_cap10k` vs `exp_c200_cap10k_structured` goes from **$0.117 vs $0.134/trajectory (full history β13%) at DeepSeek prices** to **$9.85 vs $9.94 (βtied) at GPT rates** β and the tie hides a horizon split: full history still wins the 22-turn tier ($6.29 vs $6.76) but **already loses the 36-turn tier ($15.18 vs $14.72), no pricing cliff involved** (directional: 3/6 tier-2 pairs, mean β$0.46; the API cliff widens it to $12.56 vs $9.94 overall β 176/1,247 full-history calls exceed 272k by provider-reported input tokens, the count the dollar figure uses; an earlier draft said 191, which counted by the app's approximate context estimate). Cache reads are **60%** of full history's repriced spend (vs ~28% at DeepSeek prices) β exactly OpenAI's named mechanism, visible in our own traces. Closed form at this 22/36-turn horizon: full history wins iff cached input costs < **~$0.55/M** β DeepSeek charges $0.0028 (200Γ under the line), GPT-5.6 Sol $0.50 (right at it), so the crossover turn falls from "beyond everything we tested" to "inside ordinary sessions" purely by the price constant. (2026-07-16 refinement: $0.55/M is a mix-specific aggregate over the 9-short/6-long trajectory mix β the 22-turn subset alone ties at ~$2.91/M and the 36-turn subset at ~$0.39/M; full sensitivity analysis, token receipts, and the GPT-5.6 cache-write caveat in `runs/deepseek_structured_only_stage1_20260715/report/gpt56_repricing_counterfactual.md`.) Reading: F9/F15/F25/F35's "keep everything wins" carries an implicit precondition F32's list omitted β **a deep-enough cache discount** β the fourth boundary alongside F32's three (static doc, small window, no caching). It rationalizes Codex's design (compact once cache-read accumulation crosses the compaction tax; not earlier than quality forces, since compaction damage is dose-responsive β F36 β matching Codex's own repeated-compaction accuracy warning; and don't buy bigger windows for quality β flat above 272k per OpenAI, consistent with F6/F26/F35 where 880k-token full-history contexts stayed perfectly accurate) but does **not** derive the specific 272k, which is OpenAI-specific (GPT-5.5-lineage standard-context serving boundary; the API cliff sits exactly there; a Codex 372k trial was reverted for usage cost). Caveats: same-trajectory repricing assumes GPT-5.6 would reproduce DeepSeek's call pattern and ~96% cache-hit rate (imperfect caching hurts full history *more* β F25's routing note β so the real crossover is likely earlier, not later); excludes cache-write billing on both sides; subscription "usage" is OpenAI-internal accounting, not API dollars; trigger sensitivity is unrun (`exp_c400_cap10k`/`exp_c800_cap10k` built, pending), so the optimal-trigger curve is unknown even on DeepSeek.
|
| 201 |
|
| 202 |
**Screen winners β promotion candidates:** `incontext_history_retrieval` (needs the long-session test), `retrieval_budget_30k` and `clear_retrieval_kb` (cheap Axis-B wins; clear_retrieval_kb is cheaper than prod with similar memory + better recall/key-points), anchored by `full_history`. **Drop:** `context_reset`, `observation_truncation`, `selective_retention`.
|
| 203 |
|