Download evals.md from towardsai-tutors/ai-tutor-chatbot: direct link, hf CLI and curl.
- Browser
- Download file 99.2 kB
-
https://huggingface.co/spaces/towardsai-tutors/ai-tutor-chatbot/resolve/1a2b802acfecc2ab6a1e56d795540fa60acfdd8a/evals.md
- Command line
-
hf download hf://spaces/towardsai-tutors/ai-tutor-chatbot@1a2b802acfecc2ab6a1e56d795540fa60acfdd8a/evals.md
-
curl -L -o evals.md https://huggingface.co/spaces/towardsai-tutors/ai-tutor-chatbot/resolve/1a2b802acfecc2ab6a1e56d795540fa60acfdd8a/evals.md
Evaluating the AI tutor
This document explains how we test the AI tutor with repeatable conversations and measured results. The main question is: when the tutor has a long chat history, which memory strategy gives the best answers without wasting too many tokens, dollars, or seconds?
The eval setup compares answer quality, memory retention, retrieval accuracy, tokens, cost, and latency for the June 2026 workshop. The same harness also becomes the ongoing quality program afterwards. Research sources and design rationale live in evals/background.md.
Running a new experiment? See evals/contributing.md for the step-by-step (run β upload data to HF β record an F<N> finding here β merge) and the definition-of-done checklist, so the next person inherits both your result and the data to reproduce it.
Plain-English map:
- A battery is a dataset of test conversations or test questions.
- A run is one battery executed with one model and one memory setting.
- A trace bundle is the saved JSON record for one tutor turn: the question, answer, tool calls, sources, token usage, timing, and any error.
- A grade is the pass/fail or metric row computed later from saved trace bundles.
- A memory preset is a named setting that controls how much chat history the tutor keeps, summarizes, or stores as a student profile.
- An axis is one independent dimension the experiment varies. Part B varied a single axis (the memory preset). Part C uses two: A β memory & context management (how conversation history is kept or compacted) and B β retrieval & tool outputs (how retrieved docs and tool results are sized, cleared, or whether the KB tool exists). A single-axis run changes one variant on one axis and holds everything else at the production baseline, so any difference is attributable to that one change; a cross-product tests every Axis-A Γ Axis-B combination at once (far more runs, far bigger bill).
What we evaluate, and what varies
The system under test is the production tutor (app/): the same agent, retrieval tools, prompts, and telemetry used by the app. The tutor can answer from the course/docs corpus through retrieve_tutor_context, browse the local knowledge base through run_kb_command, and keep per-thread conversation memory.
The single variable is how the tutor manages context β 13 arms across two axes in the Part B/C screen (A = memory/history, B = retrieval/tool-outputs), with later experiments adding more arms (F29's summarization family, F35/F37's exp_* factorial + structured-prefix arms; exp_c400/c800 built, unrun). Within each experiment everything else is held constant β system prompt, retrieval configuration, selected sources, web tools off, temperature, and the model. The Part B/C screen held the model at gemini-3.5-flash; later experiments varied it (DeepSeek-V4-Flash β now the production default β in F25/F26/F35βF38, Gemini 2.5 Flash in F29, three Ollama SLMs in F30/F31), with per-finding comparability caveats: absolute numbers never cross fleets. Most arms are a MemoryConfig preset (app/memory_presets.py, part of the agent cache key); a few vary a retrieval knob or a tool toggle instead. In this document, compaction means reducing the prompt by summarizing older messages or clearing old tool outputs.
The arms β every strategy we tested:
| Arm | Axis | Technique | What it tests | Phase |
|---|---|---|---|---|
full_history |
β | keep everything (no compaction) | quality/memory upper bound; token worst case | B + C anchor |
prod |
A+B | summarization + tool-result clearing | the live production baseline | B + C anchor, D |
aggressive |
A+B | early summarization + clearing | high-pressure compaction | B |
profile_memory |
A (store) | semantic profile store + prod compaction | long-term personalization | B, D rerun |
summarization_only / editing_only |
A / B | one compaction technique alone | isolate each half of prod's bundle | defined, folded/dropped (F3) |
sliding_window |
A | trimming / keep last N | recency-only memory | C |
prompt_compression |
A | rewrite history to fewer tokens | shrink text, drop no facts | C |
selective_retention |
A | constraint-preserving summarization | quality-preserving summary | C |
context_reset |
A | summary-seeded fresh state | aggressive prefix rewrite | C |
incontext_history_retrieval |
A | retrieve relevant old turns | retrieve vs carry/summarize all history | C |
clear_retrieval_kb |
B | clear tool outputs incl. retrieval/KB | the F3 fix (prod exempts them) | C |
observation_truncation |
B | head/tail-truncate tool outputs | trim the dominant token source | C |
retrieval_budget_30k |
B | retrieval token-budget knob (100kβ30k) | does a tighter budget drop recall (F1) | C |
kb_off |
B | disable the KB browse tool | agentic browse vs top-k RAG | C |
The datasets β what we ran the arms on (full schemas in data/eval/README.md; detail in "The dataset" below):
| Battery | Size | Tests | What actually ran |
|---|---|---|---|
battery_singleturn_v1 |
60 one-shot Qs | answer quality + response type | B (all 60), C (24-case subset) |
battery_sessions_v1 |
32 sessions (113 probes) | memory under compaction | only 3 sessions ever (s01/s02/s08), in B and C |
battery_personas_v1 |
10 personas Γ 4 Qs | long-term profile memory | B (all) |
replay_n1_v1 |
30 "next staff reply" cases | multi-turn behavior on real threads | built, never run |
battery_sessions_v2 |
6 sessions, 3 tiers (contradiction / long-horizon / entity) | where full_history might lose |
Gemini partial (prod all 6, full_history 2 Tier-1), then in full on DeepSeek: the 12-arm v2 matrix (F25/F26) and the Γ3-trial stage-1 factorial (F35βF37). Deprecated 2026-07-15 (filler confound, see harness corrections) |
battery_sessions_v2_1 |
v2 with 7 recycled filler turns repaired (probes/plants identical) | same tiers, confound-free | built 2026-07-15; the required v2 successor for new runs |
What actually ran, by phase: Part B β the 4-arm bake-off (full_history, prod, profile_memory, aggressive) over the full single-turn + personas batteries + 3 sessions, Γ2 trials (β F1βF13). Part C β the 11-arm screen (the 2 anchors + 9 variants) over a 24-case single-turn subset + 3 sessions, Γ1 trial (β F14βF20, F23). D β re-ran profile_memory with its store actually engaged (β F22). v2 β a partial long-context / contradiction run (β the open v2 items in "Part C β what remains" below). Later phases moved off Gemini: the DeepSeek v2 matrix (F25/F26), GraphRAG head-to-head (F28), the Gemini-2.5 compaction study (F29), the SLM studies (F30/F31), and the DeepSeek stage-1 factorial + structured-compaction follow-up (F35βF38, Γ3 trials on the v2 sessions battery). Nothing has been run at promotion grade (full batteries Γ 3 trials); the earlier phases are 1β2-trial screens.
Why the BβC order: Part B's findings drove it β tokens live in tool outputs not history (F1/F3) and compaction can cost more than it saves (F2/F9), so the cheap retrieval/tool-output variants ran first and the history-compaction variants ran to confirm that inversion. That narrative lives in the Results sections below; this section is the complete catalog.
What we measure
| Layer | Metrics | Source |
|---|---|---|
| Runtime and cost | Time to first answer text, total turn time, input/output tokens, cached tokens, estimated dollars, number of model calls, and whether compaction fired. | The context_stats event emitted at the end of each turn (app/telemetry.py). This does not depend on LangSmith. |
| Tool behavior | Which tools the tutor called, how often, and whether it re-searched for information it had already seen. | Tool-call events saved in each trace bundle. |
| Retrieval | Whether the retrieved chunks included the correct course or lesson. recall@shown means "was the right source among the results shown to the model?" MRR rewards putting the right lesson earlier in the shown results. |
Ground truth stored on each real discussion case: source_key and lesson_url. |
| Response behavior | Whether the tutor did the correct kind of thing: answer from course content, answer generally, redirect a platform/support issue, or acknowledge feedback. | Cheap code checks for early signals; final behavior accuracy comes from human grades. |
| Answer quality | Whether the answer covered the required key points, passed graded memory probes, and (holistic) whether course staff would approve sending it as a genuinely helpful, learning-oriented reply. A separate faithfulness check (every claim grounded in retrieved evidence) is built but parked β see the cap correction below. | Human pass/fail grading in a blinded workbook, or the validated blind subagent judge (same per-item rubric). LLM judges are used only after they are validated against human labels. Holistic and probe quality are complementary β holistic is graded blind to conversation history, so it cannot see a memory failure (F34); read them together. |
| Memory | Whether the tutor remembered or updated facts from earlier in the session, followed preferences, resolved references like "that thing from earlier," and used stored student profiles. | Session probes and persona checks. |
Cost is reported both as raw tokens and as estimated dollars after cached-token discounts. These can rank presets differently (finding F2).
The dataset
Source: data/academy_discussion_eval.jsonl β 151 real student posts from the academy discussion boards with real staff answers, annotated by gemini-3.5-flash. All eval inputs are real or hand-authored; we never generate reference answers β ground truth is staff replies (distilled to key_points) or facts we wrote ourselves.
Review. All 62 gold + 30 not-time-bound usable annotations (92) were re-reviewed against the live KB with file/line evidence. 60 kept, 32 excluded. Main corrections: the annotator badly under-flagged staleness (course content was updated after many posts β ~14 exclusions had premises no longer true in today's corpus), 5 near-duplicates, 1 fabricated URL, 1 misaligned question/answer pair, and key points rewritten to claims that are atomic, still true, and binary-checkable. Colab-notebook dependency was checked per case: 1 of 92 truly required the notebook (excluded); lessons embed the cells everywhere else. Full audit trail: data/eval/review_log_v1.md (11 judgment-call cases flagged for human review).
Batteries (data/eval/, schemas and glossary in its README.md; files are gitignored because they contain real student text):
| Battery | Contents | Tests |
|---|---|---|
battery_singleturn_v1 |
60 reviewed real questions asked as one-off chats. They include course-content questions, support/platform issues, general AI/programming questions, and course feedback. | Basic answer quality and correct response type. Memory presets should score about the same here; big differences suggest a bug. |
battery_sessions_v1 |
32 multi-turn study sessions. Early turns give the tutor facts about the student, middle turns make the chat long, and later turns check whether important facts survived. | Memory under pressure. A probe is one of the later turns that gets graded. Probe examples: remember a fact, respect a preference, understand "that thing from earlier," or use a newer fact after the student changed their mind. |
battery_personas_v1 |
10 stored student profiles with 4 questions each. The questions are written so the best answer needs the profile. | Long-term profile memory: does a stored profile actually improve personalization across fresh chats? |
replay_n1_v1 |
30 tests made from real multi-turn discussion threads. For each test, we give the tutor the conversation up to just before a real staff reply, ask it to write the next reply, and compare it with what staff actually wrote. | Multi-turn behavior on fully real data. This is secondary because it needs human or validated judge grading. |
How it runs
uv run -m evals.run_battery --battery data/eval/battery_singleturn_v1.jsonl --preset prod --trials 2 --out runs/exp
uv run -m evals.grade --run runs/exp # code checks + handgrade_sheet.csv
uv run -m evals.check_triggers --runs runs/exp # gate: compaction fired where probes assume
uv run -m evals.report --runs runs/expA runs/expB # side-by-side tables + token curves
uv run -m evals.handgrade_workbook build|merge ... # blinded human-grading workbook
Command roles:
run_batterytalks to the tutor and saves one JSON trace bundle per turn.gradereads saved bundles and writes automatic grades plus a CSV for human judgments.check_triggersverifies that session probes really happened after compaction. A memory test is not useful if memory trimming never occurred.reportbuilds side-by-side tables and token curves from graded runs.handgrade_workbookbuilds and merges a blinded grading workbook, so the grader does not know which preset produced an answer.
Every turn persists a JSON trace bundle, so grading and reporting can re-run offline without touching the API. Runs are resume-safe: re-run the same command after an interruption and completed units are skipped. Each turn has a 10-minute timeout, and LangSmith is off by default. Running the tutor needs COHERE_API_KEY plus the key for the model under test β DEEPSEEK_API_KEY for today's default (with only a Gemini key the app substitutes gemini-2.5-flash outright and the harness flags every bundle); the Part B/C runs used GEMINI_API_KEY; the full 4-preset comparison run below cost β$323 at corrected pricing (the pre-2026-06-13 price table under-priced gemini-3.5-flash ~4.4Γ; the often-quoted "$73" is a pre-correction figure β see Harness corrections).
Setup for collaborators. Code and docs are in git; the datasets and run results contain real student text and live only in the private HF dataset (towardsai-tutors/ai-tutor-data β git is force-pushed to the public prod Space on deploys, so student data never enters git). With an HF_TOKEN that can read it:
export HF_TOKEN=hf_... # token with READ access; the inline cmd does NOT load .env (or: uv run huggingface-cli login)
mkdir -p data/eval runs/_grading # runs/ is gitignored, so a fresh clone has no runs/ dir yet
uv run python -c "from huggingface_hub import snapshot_download as d; d(repo_id='towardsai-tutors/ai-tutor-data', repo_type='dataset', allow_patterns=['eval/**','eval_runs/**'], ignore_patterns=['eval/README.md'], local_dir='.')"
mv eval/* data/eval/ && mv eval_runs/part_*/* runs/ # restore working paths (part_b, part_c, ...)
mv eval_runs/slm_compaction/axis_a/* runs/ 2>/dev/null # SLM Axis-A runs -> runs/axisa_* (F31); Axis-B (F30) stays under eval_runs/slm_compaction/axis_b/, see evals/slm_compaction.md
mv eval_runs/deepseek_compaction_stage1/deepseek_* runs/ 2>/dev/null && mv eval_runs/deepseek_compaction_stage1/_grading/* runs/_grading/ 2>/dev/null # DeepSeek stage-1 compaction runs + judge verdicts (F35-F38)
Results β Part B comparison run, 2026-06-12
4 presets Γ 2 trials over the full single-turn battery, full persona battery, and 3 sessions (s01 13-turn, s02 agentic, s08 fact-update). 1,232 turns, zero API errors. Tables: runs/b_report/report.md; token curves: tokens_by_turn.csv. Auto-graded only means the current table includes metrics computed by code. Answer quality and session probe accuracy appear after the human grading workbook is filled and merged.
| full_history | prod | profile_memory | aggressive | |
|---|---|---|---|---|
| personalization pass (personas, n=70) | 63% | 67% | 94% | 56% |
| sessions: est. cost/turn (corrected pricing) | $0.112 | $0.237 | $0.218 | $0.296 |
| sessions: median time to first answer text | 17s | 21s | 22s | 43s |
| sessions: tool calls/turn | 2.8 | 3.5 | 3.9 | 8.3 |
| sessions: cumulative input tokens | 2.9M | 1.7M | 1.6M | 1.75M |
| single-turn: behavior proxy from code checks (n=106) | 88% | 86% | 86% | 78% |
| single-turn: retrieval recall@shown source | 51% | 50% | 49% | 47% |
Row definitions:
- personalization pass β % of persona question-runs whose answer passed every authored check (expected regex matched, e.g.
condafor the conda persona) with no anti-pattern hit (e.g. bashexportfor a Windows persona). n=70: 40 questions Γ 2 trials, minus 10 runs whose checks need human judgment. - est. cost/turn β mean estimated $ per conversation turn: token counts Γ
MODEL_PRICING, cached input billed at the cache-read discount (our table, not the invoice). - median time to first answer text β median seconds from user message to first visible answer text. It includes tool calls and internal model rounds before the visible answer starts, so it approximates the user's perceived wait.
- tool calls/turn β mean retrieval + KB-command invocations per turn; here it's the re-work signal (compaction presets re-search for evidence their compressed history lost).
- cumulative input tokens β all input tokens billed across an entire session (every turn, every internal call), mean over the 6 session-runs per preset.
- behavior proxy from code checks β % of single-turn case-runs where a cheap code check confirms the right kind of response (course question β called retrieval/KB; support issue β points to support; feedback β acknowledges). n=106 of 120:
answer_generalhas no proxy check. This is only an early signal; authoritative behavior accuracy comes from human grades. - retrieval recall@shown source β % of case-runs where the reranked retrieval results the agent saw contained at least one chunk from the correct course. Stricter
recall@lesson(right lesson page, 30β36%) and the right-lesson ranking score (MRR) are in the full report.
Observations
Numbered findings; each states what it was tested on. Convention: entries are never edited, only superseded.
- F1 β Retrieval payloads dominate input tokens, not conversation history. Each retrieval call may return up to
DEFAULT_CONTEXT_TOKEN_BUDGET = 100_000tokens; turns average ~200k input. (All runs; high confidence.) β retrieval budget became a Part C variant dimension. - F2 β Compaction saves tokens but not necessarily dollars. Summarization rewrites the prompt prefix and invalidates Gemini's implicit cache: on the same session,
full_historyhad 86.8% of input billed at the ~4x cache discount;aggressiveused 44% fewer tokens yet cost 68% more. (s03 Γ 3 presets, then confirmed n=24 session-runs; provider-specific β Anthropic caching is explicit.) - F3 β Clearing old tool-output messages never fires on this workload.
ClearToolUsesEditexcludes retrieval results β where the tokens are; 0 clears across all session runs. βediting_onlydropped from Part B; "clear retrieval results too" variant queued for Part C. - F4 β Aggressive compaction degrades even single-turn behavior. 18.0 vs 9.6 LLM calls/turn, 57s vs 39s median, behavior proxy β10pts vs full_history (60 cases Γ 2 trials). Mid-turn compaction churn changes agent behavior, it doesn't just trim.
- F5 β Gemini reasoning tokens are ~90%+ of output even with reasoning display off. Billed as output; dominates latency; recorded per model in
usage_by_model. - F6 β Long context costs latency, not correctness. Zero API errors in 1,018+ turns at any size (max 6.06M tokens across one turn's calls; largest single context 274k). Median time to first answer text scales 22s β 76s from <100k to >800k input tokens/turn. Compaction's value here is responsiveness and spend, not keeping the model functional.
- F7 β ~0.8% of turns produce no answer text despite tool calls and billed reasoning tokens; preset- and size-independent. Candidate golden-case assertion ("answer non-empty").
- F8 β Profile memory wins on quality AND cost. 94% personalization vs 56β67% without, while the cheapest persona preset ($0.049 vs $0.081β0.095/turn, 7.6 vs 10β11.2 tool calls): the stored profile saves the agent from re-searching for user context. (40 questions Γ 2 trials Γ 4 presets, auto checks; LLM-check rows pending human grades.)
- F9 β Full history is cheapest AND fastest up to 13 turns; compaction causes re-work. Sessions: $0.034/turn, time to first answer text 17s, 2.8 tool calls for
full_historyvs $0.051β0.066, 21β43s, 3.5β8.3 elsewhere β despite ~2x the tokens (F2's cache mechanism plus a second one: raw history lets the agent re-use earlier retrieval evidence; summaries force re-retrieval). (3 sessions Γ 4 presets Γ 2 trials, n=24, zero errors.) The conventional pitch inverts: under modern prompt caching, the naive baseline wins short-to-medium sessions; compaction must justify itself on quality (pending human grades) and long horizons. Where the crossover actually sits is a Part C question. - F10 β Compaction degrades within-session memory, and what it drops is old facts. Session-probe accuracy:
full_history92% vsprod38% /profile_memory38% /aggressive42% (n=24/preset). The collapse is entirely turn-0 material βfact_recall100% β 17β25%,preference_compliance100% β 0% β whilefact_updatestays 100% across all presets (the mid-session update sits inside the kept-recent window; summarization evicts the early planted facts, not the recent ones). With F9 (full_history also cheapest and fastest here), the naive baseline wins cost, speed, AND memory on β€13-turn sessions β the workshop headline. (3 sessions Γ 4 presets Γ 2 trials; provisional β LLM-graded under 4 reviewer rubric policies, pending human review of the 96 probes.) - F11 β Profile memory helps personalization, not working memory.
profile_memoryreached 92% persona personalization (vs 54β66% without) yet scored 38% on session probes β identical toprod: the long-term student-profile store improves fresh-chat answers while the live thread is still summarized exactly as inprod. Long-term store β in-session retention; they are independent subsystems and a preset can win one while tying the other. (Personas n=80 auto; sessions n=24, provisional.) Refines F8. - F12 β Aggressive compaction is dominated on every axis. Beyond F4's cost/latency blowup (18 LLM calls/turn, 57s median), it is also worst on quality: key-point coverage 36% vs 72% (
full_history), single-turn behavior 67% vs 83β92%, session probes 42% vs 92%. No measured metric favors it. (Single-turn quality n thin at 12β25; sessions n=24; provisional.) - F13 β Session-probe grades are human-confirmed (zero overrides), so F10βF12's memory numbers are no longer provisional. 2026-06-15 Omar reviewed all 96 session probes β the 83 high-confidence LLM verdicts by sign-off and the 13 low-confidence ones individually β and agreed with every grade. The provisional LLM grading therefore stands as human ground truth, confirming the 92%-vs-38% probe-accuracy result (F10) and the probe components of F11/F12. Caveat: only the session-probe tier was human-reviewed; the non-probe quality rows (persona, key_point, behavior β feeding key-point coverage and single-turn behavior in F8/F12) remain LLM-graded. The judge has not yet been validated against these labels (next step), so judge-graded Part C numbers stay gated.
Part C screen (2026-06-15)
11 arms (prod + full_history anchors + 9 variants), each on a fixed subset (24 stratified single-turn + 3 sessions s01/s02/s08), 1 trial, ~660 turns, 0 errors, β$163 at correct pricing. Graded by the validated subagent judge. Full table: runs/c_report/report.md. Probe-accuracy n=12/arm (1 trial) β treat percentages as coarse rankings, not precise rates; promotion (Γ3 trials, full batteries) would firm them up.
F14 β The LLM judge is validated; Part C quality columns are reportable. A blind subagent re-grade of the 96 human-confirmed probes reproduced them at 98% agreement / TPR 100% / TNR 96% (clears the >90% gate; this measures grader reproducibility, since the labels were sign-offs, not blind-from-scratch). The validated grader is the subagent workflow (same per-item-type rubric as
evals/judge.py), run on subscription at no API cost; all Part C arms are graded that way. Supersedes F13's "judge not yet validated" caveat. (runs/b_report/judge_val/.)F15 β The screen reproduces F9/F10 on independent arms.
full_historyis cheapest on sessions ($0.10/turn) AND best memory (100% probe); every compaction arm is both pricier and weaker.prompt_compressionalso reaches 100% memory β it shrinks message text but drops no facts β yet saves little (highest tokens), so "don't drop content" preserves memory but is not a cost win. (Screen; n thin.)F16 β
incontext_history_retrievalworks but is short-session-neutral. 83% probe accuracy (best after full_history): retrieving the relevant old turns restores what summarization loses. But on β€13-turn sessions it can't drop much, so it costs β full_history (~$0.20/turn) plus embed overhead; its cost payoff needs LONG sessions (a_v2battery). The principled cost answer to F9, pending a long-horizon test; separate from the higher-priority v2 contradiction-precision question.F17 β Turning the KB off inverts retrieval recall (Axis B tradeoff).
kb_offforcesretrieve_tutor_contextinstead of browsing β single-turn recall@shown jumps 50%β96%, but it is the priciest session arm ($0.31/turn) and worst memory (33%). KB browsing is cheaper and better for memory yet actively lowers top-k retrieval recall (the agent browses instead of retrieving the labeled source). Not a clear win either way.F18 β The 100k retrieval budget is over-provisioned.
retrieval_budget_30kmatches prod's recall@shown (50%) at a third of the budget β tokens can be cut with no recall loss on this subset (direct confirmation of F1). The 10k rung is untested.F19 β
observation_truncationbackfires (Axis B negative). Head/tail-truncating tool outputs in the model's view makes the agent re-call tools to recover what was cut: 22.8 LLM calls/turn, 430s p95 latency (single-turn), memory only 33% β the same churn pathology asaggressive(F4).F20 β Axis-A summary/trim arms do not rescue memory.
context_reset(17% probe, 56% key-points β dominated on every axis),selective_retention(25%), andsliding_window(42%) all stay well belowfull_history/prompt_compressionand none beatsprod(58%): dropping or summarizing early turns loses the planted facts the probes test (consistent with F10).F21 (2026-06-16) β Supersedes F11 in part:
profile_memorywas dormant on the sessions battery, so its 38% isprod, not a test of the profile.evals/run_battery.run_sessionpasses nostudent_id, andStudentProfileMiddlewareinjection plus the post-turn write-back both no-op without one β so on sessionsprofile_memoryreduces exactly toprodand could not differ. The 38% is real but comes from the live conversation thread + prod's lossy summary (the student states the facts in-thread; summarization keeps the recent ones and a compressed summary of the rest), not from any profile β which is why it is ~38%, not 0. F11's reading that this shows "long-term store β in-session retention, independent subsystems" is therefore unsupported: the store was never engaged in-session. Whether a store that captures in-session facts verbatim and re-injects them recovers recall under compaction is open β thecollection_memoryv1 test (sessions get astudent_id+ an append-atomic-facts write-back; goalposts: compaction 17β25% vsfull_history100%). F11's personas number (94%) stands; personas do set astudent_id.F22 (2026-06-16) β Activating
profile_memoryon sessions rescues much of the old-fact loss, but does not replace full history. Experiment A reran the three Part-C sessions withrun_sessionpassing a per-session/per-trialstudent_id, so the profile store was actually read/written and injected through the system prompt on every turn. With compaction active at 100% of probes, probe accuracy rose to 75% (9/12) vs same-screenprod58% (7/12), whilefull_historystayed 100% (12/12). The gain was concentrated exactly where F10/F21 predicted:fact_recallimproved to 83% (5/6) vsprod33% (2/6), andpreference_complianceto 100% (1/1) vsprod0%. It still missed two reference/consistency probes, and was not a cost win on this short screen ($0.265/turn vsprod$0.229 andfull_history$0.103). Interpretation: injecting stored facts outside summarized history can recover many facts lost by compaction, but the current "5 durable lines" profile write-back is an incomplete working-memory store. Next test:collection_memory/ verbatim atomic facts to determine whether the remaining gap is extraction/storage loss vs answer synthesis. Caveat (thin n): the only robust signal isfact_recall(5/6 vs 2/6) and the 9/12-vs-7/12 aggregate β thepreference_compliance/fact_updatecells are n=1 anecdotes; andprofile_memoryactually regressed on anaphora / anaphora_consistency (50% vs prod's 100%, n=2 each), possibly profile injection distracting from reference resolution, possibly noise. (Codex handgrade at Omar's request; one embeddings row was borderline-pass, strict fail would make this 67% (8/12) without changing the direction.)F23 (2026-06-17) β F17's recall inversion was a measurement artifact:
recall@shownis blind to KB grounding, and a KB-fair metric erases the gap.recall@shown/recall@lesson/MRR count onlyretrieve_tutor_contextmatches, so whenprodgrounds by browsing the KB instead of retrieving, the labeled source is invisible to the metric β which is exactlykb_off's only structural difference, so the 50%β96% "inversion" was partly definitional. Re-grading the same Part C single-turn bundles (no re-run;runs/kbfair_report/) with two tool-agnostic measures: recall source (any tool: retrieval+KB) =prod100% vskb_off96% (n=24), and cited-correct source/lesson (does the answer cite the labeled source/lesson, resolved from both tools viakb_manifest) = 100%/100% for both (n=14 corpus). Verified the any-tool credit is real, not a regex hit: in all 12 prod cases where retrieval missed, the agent had browsed that exact gold source viarun_kb_command. So turning the KB off does not improve grounding β both configs find and cite the right source.kb_off's genuine, non-artifact wins remain latency (19s vs 38s p50 TTFT), efficiency (3.6 vs 10.1 LLM calls/turn), and a slight single-turn quality edge (behavior 88% vs 79%, key-points 82% vs 78%) β all from doing one retrieval call instead of many browse rounds, not from better recall. Supersedes F17's recall reading; F17's cost/memory points stand. Caveats: n=24 screen subset, 1 trial;cited_correctsaturates at 100% here (coarse at this n), so any-tool recall is the discriminating fair metric; bundle KB output is capped at 6000 chars (could undercount KB hits, which only biases against prod β the true gap can be smaller, not larger). New auto metrics inevals/grade.py:recall_anytool_source,cited_correct_source/lesson(free to recompute on any saved bundle). Caveat on the efficiency read:kb_off's low tool-calls/turn (2.7 vs 8.7) is partly structural β an entire tool was removed β so it is not on the same "re-work" axis as the compaction arms' tool-call counts.F24 (2026-06-19) β F3's "0 clears" was a blind-signal artifact:
cleared_tool_outputscannot observeClearToolUsesEdit, which fires but invisibly.ContextEditingMiddleware.wrap_model_callapplies the clear to adeepcopyof the per-call message view and returns it viarequest.override(messages=...)(langchain/agents/middleware/context_editing.py:251-255) β it never writes the placeholder back to the checkpoint, buttelemetry.context_window_statscounts placeholders in the checkpointedstate_messages(chat_service.py:1876). So the counter reads 0 for every preset includingaggressiveacross allruns/, regardless of whether clearing fired β "0 clears" is structural, not evidence of absence. The real effect lands in input tokens: on sessionsclear_retrieval_kbcut mean input 145.6kβ120.5k (~17%, $0.229β$0.203/turn) vsprodby clearing the retrieval payloadsprodexempts, while on single-turn it backfired to 308k/$0.477 (re-retrieval churn to recover what it cut β theaggressive/F19 pathology). So clearing is real and mixed (session win, single-turn loss), not absent. Supersedes F3's "clearing never fires" reading; F3's structural point stands βprod'sexclude_tools=("retrieve_tutor_context",)provably leaves the dominant retrieval payloads in place by construction, so prod-clearing still can't touch the big tokens. Caveats: the session token saving is 1-trial screen n; library clearing stays telemetry-invisible (the probe gate survives only because summarization, which does persist, co-fires in these arms β a clearing-only arm would false-failcheck_triggers). Detect clearing via input-token deltas, notcleared_tool_outputs.F25 (2026-06-19) β DeepSeek-V4-Flash replicates F9/F10 and extends them to long horizon:
full_historyis cheapest AND best-memory on every tier, so the cost crossover never appears. A full Part-C-style v2 matrix re-run on DeepSeek-V4-Flash (first-party API, not Gemini) β 12 arms, 840 turns, 0 errors; Tier 1 contradiction (3Γ22-turn) + Tier 2 long-horizon (2Γ36-turn); reportsruns/ds_v2_t{1,2}_report/, raw on HFeval_runs/part_f_deepseek_v2.full_historyis the cheapest arm on both tiers ($0.0069/turn at 22t, $0.0101/turn at 36t), undercutting every compaction arm β while billing the most tokens (1.78M input tok/turn at long horizon) because ~97% are cache-hit. DeepSeek's ~50Γ cache discount ($0.14 cache-miss vs $0.0028 cache-hit input) makes the enormous cached prefix cheaper than any compaction's cache-breaking rewrite, so the summarization arms are priciest (prod$0.0336,selective_retention$0.0357). F9's inversion therefore holds even at 36 turns, and harder than on Gemini (10Γ discount β crossover would appear sooner; DeepSeek's 50Γ pushes it past every length tested). Per-turn cost is ~10-15Γ below Gemini on the comparable tier. Provider caveat:incontext_history_retrievaldoes not beatfull_historyon cost here ($0.0147 vs $0.0101) β the opposite of a teammate's OpenRouter run, where fallback routing fragments the prefix cache; cost comparisons must hold the provider/caching path constant (first-party single-endpoint caching verified live,cache_readnon-zero). Caveats: thin n (Tier 1 n=3, Tier 2 n=2 per arm, 1 trial) β coarse rankings.F26 (2026-06-19) β The v2 contradiction question (Q2) is answered on DeepSeek:
full_historyresolves contradictions, compaction loses them, so no temporal memory is needed. Same matrix, judge-graded via the validated subagent path (28 probes, blind,judge.pyrubric β F14).full_historyscores 3/3 contradiction and 2/2 long-horizon recall β perfect memory on both tiers, including the regime where it could in theory lose (it carries both A and Aβ² raw and must pick the current Aβ²).prodsummarization fails contradictions 0/3 (evicts the updated fact β F10 confirmed and sharpened) and 1/2 long-horizon.incontext_history_retrievalmatchesfull_history(3/3, 2/2) by retrieving the old turns, but costs more (F25);profile_memorypartially rescues contradictions (2/3) via its store (the F22 pattern). Among long-horizon compaction arms,prompt_compressionis least-bad (2/2 β shrinks text, drops no facts) andcontext_resetis dominated (0/2). Validity: compaction fired at 100% of probes for every compaction arm, 0% forfull_history(gate passed). Resolves item 4's conditional βtemporal_graph_memoryis not triggered, sincefull_historydid not fail contradictions. Caveats: thin n (3/2 per arm, 1 trial); the judge is reproducibility-validated (F14), not independent-human.F27 (2026-06-20) β On single-turn,
full_history's latency edge is a tail effect fromprodsummarizing mid-turn, not a median win; andtotal β ttftcannot de-confound latency. Re-read of the Part B/C single-turn bundles (no re-run). Paired per question, theprodβfull_historymedian TTFT delta is +0.43s (Part B, n=120, 2 trials) / +3.5s (Part C, n=24, 1 trial) andprodis slower in only 62/120 cases β a near-tie at the median, not the sessions gap (F9's 17s-vs-21β43s is multi-turn and cannot apply with no prior history). The gap lives in the tail:SummarizationMiddlewarefires on 29/120 single-turnprodturns (24%) because the agent's own retrieval/browse output crosses the 30k trigger within one turn; those turns run 52s TTFT / 16 LLM calls vs 29s / 8 when it doesn't, and on the same questionsproddoes 2β3Γfull_history's calls (e.g. 9β27, 12β32) β the F9/F4 re-work pathology (summary drops raw evidence β re-browse) inside a single turn.full_historynever summarizes, so it skips the tail; that lifts its mean (Part C 35s vs 51s) while the median stays tied. Methodological corollary (extends the time-of-run caveat below):total β ttftdoes not strip API-load noise β it is the answer-streaming tail (~6.0s for both presets; +1.5s even when summarization fires), because the entire agentic loop incl. the summarization tax lands inttft, so subtracting it deletes the signal. Latency β work Γ time-per-work, and the confound smears across ~20 calls, so no algebra on one sequential run removes it; the unconfounded read is the work metrics β LLM calls/turn and tokens/turn (run-time-independent like cost/recall, already in every bundle) β which are also the upstream cause. Live LangSmith spot-check (3 interleaved pairs, one browse-heavy single-turn Q) reproduced both directions:prodfaster once (summary collapsed 268kβ23k ctx, 65s) and slower twice (re-work: 23/32 calls, 116s/124s, $2.27/$1.04 vsfull_history$0.74/$0.33), and LangSmith's independent latency matched telemetry to the ms (prod65.07s vs 65.16s;full_history103.64s vs 104.28s), trees confirmingfull_historycarries noSummarizationMiddleware. Caveats: thin n and time-of-run confounded for absolute seconds; the live pairs are n=3 on one non-representative question. Same class as F24 (single-turn clearing backfire) on the other compaction half.F28 (2026-06-19) β GraphRAG does not beat classical hybrid RAG on single-turn course Q&A: it ties on grounding accuracy and costs +44% $/turn, empirically confirming the earlier decision to drop it. A scoped, fair head-to-head where the only variable is the retrieval backend behind
retrieve_tutor_context: production hybridLocalChromaRetriever(dense Cohereembed-v4+ BM25 β RRF β Cohere rerank β token budget) vs a true Microsoft GraphRAG index (app/graph_rag.pyGraphRAGRetrieverβ local-search-style context assembly over entities/text-units/community-reports, mapped back to realsource/urlviacorpus_manifest, then the same Cohere rerank + token budget as classical; context-provider only, so the agent's Gemini 3.5 Flash is the sole generation model in both arms). 41full_stack_ai_engineeringsingle-turn cases (27answer_from_corpus), Gemini 3.5 Flash,--disable-kbso the retriever is the only variable,--scope-sources, 1 trial, 0 errors/arm. Both arms surface and cite the right source 100% of the time and tie on lesson recall (76%) and cited-correct lesson (85%); classical is slightly better at ranking the right lesson (MRR 0.70 vs 0.65). GraphRAG pulls community-report context, so it spends +61% input tokens (178k vs 110k) and +44% $/turn ($0.212 vs $0.147) and is a touch slower, with no accuracy payoff (the +3pt behavior proxy is a coarse n=37 code check, not a quality signal). So GraphRAG's community structure does not help typical single-hop course Q&A β this empirically supports the prior "dropped on faith" decision (see "Deliberately not pursued" below) rather than asserting it. Reusable artifact: the scoped index (90 docs β 15,059 entities / 32,748 relationships / 3,371 community reports, built with Gemini 2.5 Flash for $44.96) is published to the private HF dataset (graphrag/output/, pulled on demand, not by prod cold-start βensure_local_vector_dbignoresgraphrag/**), so the eval re-runs without the ~$45 rebuild. Full methodology, build/run commands, and the indexing gotchas (route Gemini 2.5 via its OpenAI-compatible endpoint;run_graphrag_index.shloopsgraphrag indexuntilentities.parquetexists since one exhausted retry aborts the stage) live inevals/graphrag.md. Connects to F17/F23 (thekb_offgrounding discussion β classical RAG already grounds well). Caveats: scoped to one source, single-turn, n=41 (27 corpus-answer), 1 trial, auto metrics only (key-point/behavior human grades pending);recall@sourceis saturated by--scope-sources, so lesson recall / MRR / cited-correct-lesson are the discriminating numbers. GraphRAG's theoretical edge is multi-hop / cross-document synthesis, which this single-hop battery barely exercises β so this says it doesn't help on single-hop course Q&A, not that it never helps; a fair test of that strength needs a multi-hop_v2probe set (out of scope here). (Branchexperiment/graphrag-vs-rag, PR #2.)F29 (2026-06-20) β Keeping a large static document in context is a different regime from conversational
full_history, and it is one boundary where "keep everything" stops winning: for a single long lesson queried repeatedly, retrieving the relevant slice matches keep-it-all at ~1/13 the cost. (Boundary on F9; standalone Gemini-2.5-Flash fleet β absolute numbers not cross-comparable.) A keep-vs-compact-vs-retrieve study on the corpus's largest lesson (~37.5k tok) loaded once in turn 0, then 15 question-turns with no tools (the only context is what each method keeps or builds), each answer judged against the full lesson; 11 arms, Gemini 2.5 Flash, 1 trial.rag(fetch the relevant chunk per question) is cheapest and tied-best (60% / 3.2k tok / $0.020);full_history(keep the whole lesson in the prompt) is the priciest in-context arm and not the best (53% / 43k tok / $0.134); the summarization family cuts ~35% of tokens for a quality drop (40β53%);incontext_history_retrievaltiesragon quality but at ~13Γ the tokens;hierarchical_summarizationis the worst trade β priciest ($0.605) and slowest (41.8s p50) for the lowest quality (33%).graphrag(53%, $0.045) again does not beat plainrag(consistent with F28). Read it as a lesson, not a head-to-head: this is not "RAG beats thefull_historymemory strategy" β herefull_historyonly means "the whole lesson sits in the prompt," not its multi-turn-memory role. It says F9's "keep everything wins" is about conversation history; for a big static document you could retrieve from, hoarding it in context is the priciest option and buys nothing, and in the production tutor retrieval (Axis B) and a memory policy (Axis A) compose, they don't compete. Companion boundary: F30 (small context window). New reusable presets landed inapp/memory_presets.py:delta_summarization,hierarchical_summarization. Full matrix, the F-C1βF-C6 detail, and reproduce steps inevals/compaction.md. Caveats: n=15, 1 session, 1 trial (coarse rankings, not precise rates); judge is Gemini 2.5 Flash (same family, reads the full lesson as ground truth); separate fleet from the Part B/C 3.5-Flash runs β quality is not cross-comparable across models. (Branchexperiment/context-compaction, PR #5.)F30 (2026-06-20) β A small context window is a second boundary where "keep everything" stops winning, and a harder one than F29: when the document physically does not fit, "shove it all" (
full_context) is forced into truncation and is strictly dominated, so retrieving the relevant slice (rag) is not just cheaper but necessary. (Small-window companion to F29, Axis B; local-SLM fleet β absolute numbers not cross-comparable.) "Fit one37.7k-token lesson into context to answer 15 questions," 7 methods Γ 3 Ollama SLMs (llama3.1:8b, qwen2.5:7b-instruct, qwen3:8b no-think) on an M1 Pro/16 GB at a 32k window (12β15Γ the latency** (266β346s vs ~20s) β silent information loss plus a huge latency tax (e.g. qwen2.5--num-ctx 32768), no prompt caching; judge stays on Gemini 2.5 Flash (reads the full lesson as ground truth β the SLM cannot hold it, and a model must never grade itself), 1 trial.full_context(the whole document placed in the prompt,ctx['full_context'] = lesson, stateless per question β this is document-stuffing, not the conversationalfull_historymemory policy) is truncated on 15/15 turns (37.7k β 32,767) and dominated: equal-or-worse quality (67β80%) at **full_contextanswered "no model was mentioned" for a fact the truncated intro had named).rag(chunk + Cohere-embed + retrieve top-k per question) is top or tied-top on all three (qwen2.5 100%, qwen3 100%, llama 80%) at ~2.9k context tokens and ~20β26s.graphragtiesragon quality but at ~2.8Γ tokens and ~2.5Γ latency β no payoff, consistent with F28/F29. Summarization is the floor (llamasummary0%);trim/selectiveland mid (47β73%). This marks where F9/F15's precondition β a window big enough to hold everything, cheaply, under caching β simply fails on the hardware a workshop attendee runs. Distinct from the conversational-memory inversion (F31):full_contexthere is document-stuffing (Axis B), notfull_history's multi-turn role, exactly the distinction F29 drew. Methodology fix: the first pass scored ~64% of judge verdicts as fails because a long JSONreasonhitmax_tokensand truncated the verdict; hardened the judge (25-word reason cap + leading-boolean recovery) and re-graded the saved answers offline (evals/rejudge_compaction.py, 0 unparseable after) β the directly-measured axes (tokens, latency, overflow) were never affected. Caveats: n=15/method, 1 trial, one lesson, one domain β read as a ranking, not rates; separate fleet (local Ollama SLMs, no caching, Gemini-2.5-Flash judge) β quality/cost not cross-comparable with the Part B/C 3.5-Flash, the F29 2.5-Flash, or the DeepSeek runs. Full writeup:evals/slm_compaction.md(Experiment 1). (Branchexperiment/slm-compaction, PR #3.)F31 (2026-06-20) β On an SLM, compacting a growing conversation history has no universal best method, and "keep everything" (
full_history) never wins β and model capability dominates the choice of method. (SLM counterpart to F9/F15 on Axis A; same local fleet β absolute numbers not cross-comparable.) Same lesson loaded turn 0 of a session, 15 question-turns, retrieval off, run through the real app middlewares (the production memory presets) on the 3 SLMs atnum_ctx=32768, judged on Gemini against the full lesson, 1 trial. The best preset is model-dependent βprompt_compressionon qwen2.5 (60%),hierarchical_summarizationon llama3.1 (40%),summarization_only/delta_summarizationon qwen3 (73%) β so there is no single "best compaction" to recommend.full_history(the genuine conversational memory policy β keep every prior turn, no summarization or context-editing;summarization=False, context_editing=False) tops no model (27% / 7% / 67%) and is near-bottom on the two weaker ones;sliding_windowis reliably among the worst (13% / 13% / 40%). Model capability dominates the method: qwen3 (40β73% across every preset) >> qwen2.5 (13β60%) >> llama3.1 (0β40%) β choosing a stronger small model buys more than choosing the best strategy. Mechanism (why keep-all can't win here): on an SLM "keep everything" (full_history) can't even be truncated-and-kept β Ollama's server-side context fitting evicts the whole oversized turn-0 lesson from turn 1 on (the checkpoint still holds it; the model receives only ~0.3β1.4k tok of accumulated Q&A), so a method that summarizes the lesson into a small injected message is the only way to retain its gist; on Gemini's large window this never happens (F9'sfull_historyheld ~43k/turn).hierarchical_summarizationis by far the priciest (β30-call map-reduce, p50 233β333s, ~25β29k retained tok) for no quality lead except on the weakest model. This inverts F9/F15's conversational-memory result (keep-all cheapest and best) in the regime where the runtime physically cannot keep it. Caveats: n=15, 1 trial, LLM-judge variance moves individual cells Β±1β3/15 (read only the big gaps); thefull_historynumbers are Ollama-eviction-specific (a different serving layer could keep a truncated lesson instead); same separate-fleet comparability caveat as F30. Writeup:evals/slm_compaction.md(Experiment 2). (Branchexperiment/slm-compaction, PR #3.)F32 (2026-06-22) β Synthesis: "keep everything in context" wins only under a specific precondition β a large window that holds the whole history cheaply under prompt caching β and we have now mapped three boundaries where that precondition breaks and keep-all stops winning. Two distinct "keep-all" arms β never conflate them:
full_history(Axis A β F9/F15/F31) is the memory policy that keeps every prior conversation turn with no compaction (summarization=False, context_editing=Falseinapp/memory_presets.py);full_context(Axis B β F30) is document-stuffing β the whole source document placed in the prompt (ctx['full_context'] = lesson). Same slogan ("keep everything"), different mechanisms. F29 is the trap to watch: its keep-all arm ran thefull_historypreset but over a single static lesson parked in turn 0 (turns = [lesson] + questions), so what it actually measured is thefull_contextidea (a document held in the prompt), notfull_history's multi-turn-memory role β which is why F29 sits on the document boundary (1) below, alongside F30, and not the conversational boundary (3). F9/F15 established the default (rawfull_historyis cheapest and best memory on conversational sessions), and F2/F25/F26 explain why it holds and even strengthens with horizon (the big cached prefix is cheaper than any compaction's cache-breaking rewrite; DeepSeek's ~50Γ cache pushes the crossover past 36 turns). Remove a piece of that precondition and keep-all loses, along three boundaries: (1) a large static document you could retrieve from (F29, Axis B, large cached window) β retrieving the slice matches keep-all at ~1/13 the cost, so hoarding the doc is the priciest option and buys nothing; (2) a context window too small to hold the document (F30, Axis B, SLM) β keep-all (full_context) is forced into truncation and is strictly dominated, so retrieval is necessary, not merely cheaper; (3) a growing history on a small window (F31, Axis A, SLM) β the runtime evicts the oversized turn, so keep-all (full_history) never tops the table and some compaction is mandatory, though no single method wins. Two cross-cutting lessons fall out: model capability dominates the compaction method on weak local models (F30/F31 β a better small model beats a better strategy), and the axes compose, they don't compete (F29) β retrieval (Axis B, how to fit a document) and a memory policy (Axis A, how to compact history) are orthogonal, so "RAG vsfull_history" is a category error. Net guidance: keep-everything stays the right default in the production regime (large cached cloud window + conversational history; F9/F15/F25/F26); reach for retrieval or compaction exactly when one precondition breaks β a big static doc, a small window, or no caching. Comparability: this ties the lessons across fleets (3.5-Flash screen, 2.5-Flash F29, DeepSeek F25/F26, local SLMs F30/F31), not their absolute numbers, which are not cross-comparable.F33 (2026-06-23) β New quality axis (holistic "would staff send this?"): on single-turn, compaction degrades answer quality itself, and the degradation dose-responds to how often summarization fires within a single turn. (Independent quality-axis confirmation of F4/F19/F27.) Added two judge rubrics in
evals/judge.py: holistic (whole-answer staff-approval / does-it-help-the-student-learn gate, emitted for every answered turn) and faithfulness (groundedness vs retrieved evidence β parked, see the correction below). Graded blind by the validated subagent path (F14) across all single-turn runs; holistic rows added toevals/report.py, tagged[judge]. The single-turn premise is that memory presets should barely differ ("Memory presets should score about the same here; big differences suggest a bug", Β§The dataset) β yet holistic spans 57%β100%, and it is not a bug: it tracks within-turn summarization. Dose-response: arms that never summarize mid-turn score 92β100% (full_history92% Part C / 92% Part B n=120;kb_off100% at just 3.6 LLM calls;sliding_window/incontext_history_retrieval/selective_retention/retrieval_budget_30k96%), while arms whose own retrieval/KB output trips the trigger mid-turn sink in trigger order βobservation_truncation71% (22.8 LLM calls/turn, 12/24 turns summarized),context_reset62% (18/24),aggressive57% (Part B n=120, 110/120 turns summarized mid-turn). So the F4/F19/F27 within-turn re-work pathology is not merely slower/pricier β it makes the answers worse, at a product-relevant magnitude (aggressive: ~half its single-turn answers would not be sent). Caveats: Part C n=24 (directional); the firm cell is Part Baggressive57% vsfull_history92% at n=120; judge reproducibility-validated on probes (F14), not independently on holistic.F34 (2026-06-23) β Across a session, standard compaction's damage is surgical: it collapses memory while answer quality stays ~100% β but that "quality" is partly an artifact of a history-blind judge, so "looks fine" β "served the student". Only the most aggressive rewrite (
context_reset) degrades the answers themselves, progressively over the session. (Refines F10/F12 on the holistic axis.) Holistic graded on every answered turn ofb_se+c_se(683 turns, 15 arms, blind subagent path). The decoupling:prodprobe 58%(c)/38%(b) yet holistic 97/99%;kb_offprobe 33% / holistic 100%;selective_retentionprobe 25% / holistic 100%;profile_memoryprobe 38% / holistic 100% β memory and answer-quality are independent failure axes. The lone exception iscontext_reset(probe 17%, holistic 69%, holistic dropping β24% earlyβlate turn;sliding_windowa milder β16% late): broad, progressive quality degradation appears only under the most destructive history rewriting, otherwise quality holds flat (Ξ lateβearlyβ 0 for all standard arms). But "high holistic" is misleading on exactly the turns that matter β holistic is graded blind to the session history (sessions carry no per-turn gold answer, and the judge sees only question+answer), so on the memory-failed probe turns it still rated the (generic, context-ignoring) answer "good" 100% of the time for every standard arm (c_se_prod5/5,b_se_prod15/15,kb_off8/8,selective_retention9/9;forgot+looks-bad= 0) β onlycontext_resetforgot so badly the blind judge caught it (6/10 forgotten-fact turns also looked bad). The failure mode is therefore confidently generic: a well-formed answer that silently ignores what the student stated earlier, which only the targeted probe detects. Lesson: holistic (context-free quality) and probe (context-aware memory) measure different things and must be read together β "answer quality looks fine" does not mean the student was served. This refines F10 (memory dies but general answer quality survives) and F12 (only the aggressive rewrites also degrade quality). Caveats:c_seis 3 sessions Γ 1 trial (~18 turns/early-late bucket),b_seΓ2 trials; session holistic graded without a staff reference (the within-arm early/late comparison controls for that); judge reproducibility-validated on probes (F14), not on holistic β the blind-spot above is itself why a held-out human validation of holistic is the next gate.F35 (2026-07-15) β 2Γ2 factorial on DeepSeek: the cheap win is a stable tool-output cap, not compaction β capping cuts cost β38% at no measured quality/memory loss while preserving the cache, and summarization adds cost on BOTH sides of the factorial, so it still doesn't pay for itself even after capping. (Decomposes F25; Axis A Γ Axis B; DeepSeek-V4-Flash first-party β absolute numbers not cross-comparable with the Gemini fleets.) Experiment: DeepSeek long-context compaction study, stage 1. The study asks whether summarization-based history compaction ever pays for itself on the production tutor running its default model (DeepSeek-V4-Flash, first-party API) over long tutoring sessions, once prompt caching and tool-output size are accounted for. F25 answered "no,
full_historystays cheapest through 36 turns" β but its compaction arms confounded two mechanisms, the history policy and the giant raw tool outputs that dominate input tokens (F1), so it couldn't say where the money actually goes. Stage 1 is the clean decomposition: a 2Γ2 factorial crossing the history policy (keep everything vs summarize at a 200k trigger, retaining a 50k recent tail) with tool-output handling (raw vs a stable 10k-token cap applied once when the output enters history), giving four arms βexp_fh_raw/exp_fh_cap10k/exp_c200_raw/exp_c200_cap10k(app/memory_presets.py) β whose paired contrasts isolate each mechanism's cost/memory effect and their interaction. Stage 2 (trigger-threshold sensitivity:exp_c400_cap10k/exp_c800_cap10k) is built but not yet run. Method: batterybattery_sessions_v2(22-/36-turn scripted tutoring sessions that plant student facts early and probe them late), 3 trials, run via theevals.run_compaction_experimentpaired lockstep runner (all four arms advance turn-by-turn together in randomized within-turn order, each arm/session/trial under a distinct DeepSeek cacheuser_idso no arm warms another's KV cache; immutable run fingerprint;evals/contributing.mdΒ§2). "triggerfix" in the run name marks this as the corrected rerun after the earlier stage-1 attempt's summarization-trigger bug. Runruns/deepseek_compaction_stage1_triggerfix_20260715; full writeupruns/deepseek_compaction_stage1_triggerfix_20260715/report/stage1_findings.md; raw on HFeval_runs/deepseek_compaction_stage1/. Four arms {full_history, summarize@200k keep-50k} Γ {raw tool outputs, stable 10k-token cap (40k bytes) applied once when the output enters history} on sessions-v2 Tier-1 contradiction (3Γ22t) + Tier-2 long-horizon (2Γ36t), Γ3 trials, run turn-level-lockstep with randomized within-turn order and a distinct DeepSeek cacheuser_idper arm/session/trial (no cross-arm KV warming); 414 turns/arm, 0 errors. The earlier stage-1 trigger bug is fixed and evidenced: every compaction fired at β₯200,056 pre-compaction tokens with 92kβ198k-token summary inputs, and thefull_historyarms compacted 0 times. Cost (mean $/trajectory): capping alone $0.189 β $0.117 (β38%, cheaper in 14/15 paired trajectories) with the cache-hit ratio intact (96.0%β95.9%) β it removes ~1MB of repeated tool payload per trajectory without rewriting the prefix, so it composes with caching instead of fighting it (the F1βF24/F25 thread: the tokens live in tool outputs, and cutting them stably is nearly free). Summarization is a net cost add on both sides: +50% vs raw full-history (14/15) and +32% vs capped full-history (12/15); capping shrinks the compaction penalty (interaction β$0.057) but does not reverse it. Quality holds where it should: holistic 96β98% for all four arms, and the cap costs no memory (probe 87%, identical to raw full-history). Contrast with F19: per-call-view head/tail truncation churned (re-calls, 33% memory); this checkpoint-level insertion-time cap shows no churn at all (3.0 vs 3.1 LLM calls/turn, 87% memory) β though the arms differ in retained size and model too, the stability of what the model re-reads looks like the operative difference. Caveats: probes n=15/arm and only 5 session templates (the writeup's cluster-bootstrap intervals are descriptive, not population estimates); holistic is judge-graded, not human-validated (F34); DeepSeek's ~50Γ cache discount is load-bearing for the ranking (F25's provider caveat).F36 (2026-07-15) β Where compaction hurts memory, the damage tracks the number of lossy rewrites, not the size of the live context β and raw tool outputs double that number, so Axis-B bloat amplifies Axis-A damage. "Too many tokens in the answering call" is not the failure mode:
full_historyrecalls perfectly from up to 879k-token contexts. (Refines F10/F26's mechanism.) Experiment: same DeepSeek stage-1 compaction factorial as F35 (runruns/deepseek_compaction_stage1_triggerfix_20260715, batterybattery_sessions_v2; probe-level evidence in the samestage1_findings.mdwriteup) β this is the memory half of that run's results. The battery's two tiers each target one memory failure mode: Tier 2 long-horizon (2 sessions Γ 36 turns Γ 3 trials) plants student facts/constraints early and probes them ~30 turns later, after compaction has had every chance to evict them; Tier 1 contradiction (3 Γ 22t Γ 3) plants a fact, updates it mid-session, and probes whether the agent uses the current value. Probes judge-graded via the F14-validated blinded subagent path. Long-horizon recall (plant facts early in a 36-turn session, probe at the end): the extended 2026-07-15 audit found one of the two Tier-2 sessions (v2_t2_fullstack_long_persona_36t) carries recycled "project pivot" filler that outright contradicts the probed PHI constraint, so its cells cannot separate eviction-by-compression from eviction-by-believed-pivot and are excluded (repaired inbattery_sessions_v2_1.jsonl). On the clean Tier-2 session (v2_t2_agentic_long_persona_36t, whose off-persona filler is orthogonal to its probed governance facts; n=3/arm): 0 summarization events β 3/3 and 3/3 (both full-history arms), ~2β3 events β 3/3 and 3/3 (exp_c200_cap10kand F37's structured arm), ~5β6 events β 1/3 (exp_c200_raw). The dose mechanism survives in coarser form β only the high-dose raw arm measurably loses far-back facts, and low-dose capped compaction is indistinguishable from full history at this n β but the fine "6/6 β 5/6 β 3/6" gradient quoted before the audit mixed in the compromised session and should not be cited. The raw arm compacts twice as often because uncapped tool payloads regrow the history to the 200k trigger faster; each event then compresses150k tokens into ~1.7k (90Γ), and on the clean session the judge's fail reason is genuine eviction ("explicitly denies knowing the student's constraints and offers only a generic list of common governance options"). Read-time context size points the other way: at probe time the compaction arms answered from 95β195k-token contexts and failed, whilefull_historyanswered from 363β879k and never missed β so the "context rot" here lives in the summarization pipeline (a planted fact that is a tiny fraction of a payload-dominated span must survive every ~90Γ rewrite), not in the answering model's attention. Separately, the contradiction tier: all 11 contradiction misses across all arms sit in one session (python_colab_to_local_22t), and the miss mode is hedging (recap both environments, ask "which machine are you on?") rather than stale-fact use β but the 2026-07-15 battery audit found that session confounded, so treat its cells as unmeasured rather than as a finding: its recycled filler injects a first-person third environment ("my iMac", turn 16) after the turn-7 update and before the turn-21 probe, which makes the hedge a defensible reading of genuinely contradictory history, not a demonstrable memory failure. The other two contradiction sessions are confound-free (audited: zero post-update first-person environment claims) and had zero misses in any arm β so this run shows no contradiction gap between arms and leaves F26 standing; whetherfull_historyhedges when it truly holds A and Aβ² needs the repaired session rerun (fix as a new battery version β v1/v2 files are frozen). Caveats: the clean Tier-2 evidence is n=3/arm β read "high dose fails, low dose doesn't", nothing finer; Tier-1 does not guarantee eviction before its probe (all five ofexp_c200_cap10k's compaction-dormant trajectories are Tier-1, so in capped arms Tier-1 tests update-adherence, not contradiction-under-eviction); Tier-2's probed constraints are partially guessable from world knowledge (a governance rule for a procurement agent), which inflates every arm's absolute recall equally and biases against the dose signal. Aggregate probe accuracy restricted to audit-clean probes (9/arm): full-history arms 9/9 and 9/9, capped compaction 9/9 (XML) / 9/9 (structured), raw compaction 7/9.F37 (2026-07-15) β A chunk of "compaction's cost" was an implementation artifact, not physics: LangChain's stock summarizer breaks the provider cache on the summary-generation call itself, and a prefix-preserving compaction request removes that β summarizer cache hit 0%β94%, summarization cost β87%, total β14%/trajectory vs the stock arm, with equal-or-better memory. But the post-compaction prefix break is structural, so compaction still costs +14% vs capped full history: F35's ranking stands, with a smaller gap. (Same regime as F35/F36 β DeepSeek-V4-Flash first-party, absolute numbers not cross-comparable with the Gemini fleets.) Experiment: DeepSeek compaction study, stage 1 follow-up β prefix-preserving ("structured") compaction. The stock
SummarizationMiddlewareserializes the selected messages into a fresh XML prompt to ask for the summary, so even summary generation pays 100% cache-miss prices on ~150k-token inputs;PrefixPreservingCompactionMiddleware(app/chat_service.py,summarization_strategy="structured_prefix") instead sends the unchanged current request prefix plus one checkpoint instruction, binding the same model/tools/settings, so the summarizer call rides the same cache as the agent. Single armexp_c200_cap10k_structuredβ identical toexp_c200_cap10kexcept the strategy β runruns/deepseek_structured_only_stage1_20260715on the samebattery_sessions_v2Γ 3 trials (414 turns, 0 errors), compared pairwise per sessionΓtrial against the four stage-1 arms; graded through the same blinded-subagent pipeline (429 judgments, 0 missing). Writeupruns/deepseek_structured_only_stage1_20260715/report/structured_findings.md; raw on HFeval_runs/deepseek_compaction_stage1/. A 2026-07-15 source-level check of the open-source Codex CLI confirmed this strategy is a faithful analog of Codex's local compaction β unchanged history + appended instruction, explicitly cache-preserving β while Codex's OpenAI-default remote v2 path is the provider-native opaque-compaction primitive a client middleware cannot replicate; comparison + adoption ideas inruns/deepseek_structured_only_stage1_20260715/report/codex_comparison.md. Cost: vs the stock XML arm, β$0.0217/trajectory (β14.0%, cheaper in 11/15 pairs) at identical compaction dose (1.27 events/trajectory both): the summarizer call's cache-hit ratio goes 0.0% β 94.2%, so summarization cost falls $0.0276 β $0.0036/trajectory (β87%) β even though the structured arm bills more raw input (10.6M vs 9.9M tokens/trajectory; the checkpoint request re-reads the whole prefix at the ~50Γ cache-read discount). Tokens β dollars, again (F2/F25). vs capped full history it remains +13.7% (cheaper in only 4/15): what's left is the structural boundary β installing the summary necessarily rewrites the next agent call's prefix, which no client-side middleware can avoid (that would take a provider-native compaction/continuation primitive) β plus the extra agent work (3.5 vs 3.1 LLM calls/turn). Memory: equal-best of the compaction arms β on audit-clean probes (excluding the two battery-compromised sessions, see F36) it scores 9/9, tied with both full-history arms and capped-XML compaction, while raw compaction scores 7/9; the pre-audit "6/6 vs 5/6" long-horizon edge over the XML arm rode entirely on the compromised fullstack session and should not be cited β whether the intact-prefix summary also preserves facts better than the XML rewrite is unmeasured at this n. Holistic flat at 96%. Its single overall probe miss is a contradiction hedge inpython_colab_to_local_22tβ the confounded session β a battery artifact, not an arm signal. Trigger gate: compaction fired in 12/15 trajectories (19 events, min pre-compaction observation 200,053 tokens; the 3 dormant trajectories all Tier-1, and 100% of Tier-2 probes ran under compaction). Caveats: the memory edge is 1β2 probes at n=6 long-horizon / n=15 total β directional; the arm ran solo, not turn-level-interleaved with the stage-1 arms (token/cost metrics are run-time-independent, but treat its latency rows as confounded, F27); holistic judge not human-validated (F34); DeepSeek's ~50Γ cache discount is load-bearing (F25's provider caveat).F38 (2026-07-15) β The keep-everything-vs-compact crossover is set by one number: the cached-input price. Repricing our stage-1 traces call-by-call shows "full history wins" is a fact about DeepSeek's pricing, not a law: at GPT-5.6 Sol's public rates the same recorded workload already flips to favor compaction at 36 turns β quantitatively consistent with why OpenAI's Codex auto-compacts at ~272k while our tutor, on DeepSeek, should not compact at all. (Analysis-only finding β no new run; adds the missing fourth boundary to F32's list.) Experiment: same-trace repricing over the stage-1 arms (runs
runs/deepseek_compaction_stage1_triggerfix_20260715+runs/deepseek_structured_only_stage1_20260715, F35/F37): every recorded model call's (fresh, cached, output) token triple re-costed at GPT-5.6 Sol's public API rates ($5 / $0.50 / $30 per M β a 10Γ cache discount vs DeepSeek's ~50Γ), with and without the API's >272k long-context cliff (2Γ input / 1.5Γ output on the whole request; per OpenAI's Tibo, the cliff is NOT charged on Codex subscriptions β the stated real driver is accumulated cache-read cost scaling with the live window re-read on every tool call). Result:exp_fh_cap10kvsexp_c200_cap10k_structuredgoes from $0.117 vs $0.134/trajectory (full history β13%) at DeepSeek prices to $9.85 vs $9.94 (βtied) at GPT rates β and the tie hides a horizon split: full history still wins the 22-turn tier ($6.29 vs $6.76) but already loses the 36-turn tier ($15.18 vs $14.72), no pricing cliff involved (directional: 3/6 tier-2 pairs, mean β$0.46; the API cliff widens it to $12.56 vs $9.94 overall β 176/1,247 full-history calls exceed 272k by provider-reported input tokens, the count the dollar figure uses; an earlier draft said 191, which counted by the app's approximate context estimate). Cache reads are 60% of full history's repriced spend (vs28% at DeepSeek prices) β exactly OpenAI's named mechanism, visible in our own traces. Closed form at this 22/36-turn horizon: full history wins iff cached input costs < **$0.55/M** β DeepSeek charges $0.0028 (200Γ under the line), GPT-5.6 Sol $0.50 (right at it), so the crossover turn falls from "beyond everything we tested" to "inside ordinary sessions" purely by the price constant. (2026-07-16 refinement: $0.55/M is a mix-specific aggregate over the 9-short/6-long trajectory mix β the 22-turn subset alone ties at ~$2.91/M and the 36-turn subset at ~$0.39/M; full sensitivity analysis, token receipts, and the GPT-5.6 cache-write caveat inruns/deepseek_structured_only_stage1_20260715/report/gpt56_repricing_counterfactual.md.) Reading: F9/F15/F25/F35's "keep everything wins" carries an implicit precondition F32's list omitted β a deep-enough cache discount β the fourth boundary alongside F32's three (static doc, small window, no caching). It rationalizes Codex's design (compact once cache-read accumulation crosses the compaction tax; not earlier than quality forces, since compaction damage is dose-responsive β F36 β matching Codex's own repeated-compaction accuracy warning; and don't buy bigger windows for quality β flat above 272k per OpenAI, consistent with F6/F26/F35 where 880k-token full-history contexts stayed perfectly accurate) but does not derive the specific 272k, which is OpenAI-specific (GPT-5.5-lineage standard-context serving boundary; the API cliff sits exactly there; a Codex 372k trial was reverted for usage cost). Caveats: same-trajectory repricing assumes GPT-5.6 would reproduce DeepSeek's call pattern and ~96% cache-hit rate (imperfect caching hurts full history more β F25's routing note β so the real crossover is likely earlier, not later); excludes cache-write billing on both sides; subscription "usage" is OpenAI-internal accounting, not API dollars; trigger sensitivity is unrun (exp_c400_cap10k/exp_c800_cap10kbuilt, pending), so the optimal-trigger curve is unknown even on DeepSeek.
Screen winners β promotion candidates: incontext_history_retrieval (needs the long-session test), retrieval_budget_30k and clear_retrieval_kb (cheap Axis-B wins; clear_retrieval_kb is cheaper than prod with similar memory + better recall/key-points), anchored by full_history. Drop: context_reset, observation_truncation, selective_retention.
Part C β what remains (beyond the workshop). The screen answered the core question; everything below hardens or extends it and is gated on a spend or data decision from Omar β none is required for the workshop.
- Promotion (rigor, not new findings). Re-run the 2-3 winners (+ 1-2 combos) on the full batteries (60 single-turn + 32 sessions + 30 replay) Γ 3 trials, with paired statistical tests and per-variant failure-taxonomy diffs, to turn the screen's coarse n=12 rankings into confident numbers; optionally an Anthropic (Haiku) re-run of the prefix-rewriting arms (
prompt_compression,context_reset,aggressive) to test whether explicit caching flips F2's cost ranking. ~$1,300-1,800 at corrected pricing (the old "$300-400" was at pre-correction prices; Γ the 4.4 fix β estimate precisely before running, since the per-turn cost is now ~$0.25). - Finish the v2 contradiction tier β the highest-value new dataset (already built and partially run; see the V2 execution note below). Moderate sessions: plant A, update to Aβ², then probe after prod compaction has evicted A from the kept-recent window. This is the one correctness regime where
full_historyitself might lose (it sees both A and Aβ² and must choose the current one) β yet the partial run so far shows the opposite (full_historycontradiction 2/2 vsprod0/3), so that regime has not actually appeared. What remains is finishing Tier 1 (the missingfull_historysession, then Tier-1profile_memoryand optionallyincontext_history_retrieval) and adding trials, not building the dataset. (2026-06-19 β now run in full on DeepSeek-V4-Flash rather than Gemini: F26 confirmsfull_history3/3 contradiction vsprod0/3, so the regime did not appear on a second model either; the Gemini Tier-1 finish is now optional cross-model confirmation.) - The
incontext_history_retrievallong-session cost test β principled but less product-critical. The screen showedincontextworks (83%, F16) but couldn't show a cost benefit on β€13-turn sessions. Whether retrieve-old-turns beatsfull_historyonce sessions get genuinely long needs a token-calibrated long-session v2 tier; useful for the workshop argument, but gated by whether that regime matters in real tutor telemetry. (2026-06-19 β run on DeepSeek (F25) at 36 turns:full_historystays cheapest andincontextworks but does not beat it on cost under first-party caching; whether the answer flips on a 10Γ cache provider like Gemini/Anthropic is still open.) - Conditional builds. Build
temporal_graph_memoryonly iffull_historyfails contradictions (F26: it did not on DeepSeek β 3/3 β so this stays unbuilt),delta_summarizationonly if the long-horizon cost tier shows room to beat full history (F25:full_historystill cheapest at 36t on DeepSeek, so no room there β though a 10Γ cache provider is untested), andentity_memoryonly if multi-project probes show fact bleed.collection_memoryis a separate cheap v1/product follow-up for verbatim fact storage (F22 tested the current profile write-back, not that).hierarchical_summarizationandsleeptime_consolidationare deferred. The step-by-step v2 plan is inevals/part_c_plan.mdβ "Battery v2".
Deliberately not pursued. Stretch (heavier, likely later): sub-agent isolation, reframed as a tool-output-token play β keep noisy run_kb_command output out of the main context (ties to F1); temporal-graph memory is already the conditional build in item 4 above. Dropped as low-information given the findings: the skills / lazy-prompt-loading family (only a ~458-token win β see the talk-outline correction below), procedural memory, GraphRAG (now empirically tested rather than dropped on faith β F28: ties classical RAG on grounding, costs +44% $/turn, no payoff on single-hop Q&A), and multi-agent / parallel-research agents. Context for the kb_off arm (F17/F23): the KB is the dominant grounding tool β β89% of Part B turns, ~7.7 KB vs ~0.9 retrieval calls/turn over 1,088 turns β so disabling it forces the 100k-token retrieval fallback, which is why it raised cost/latency.
V2 execution note (2026-06-16). A private/gitignored 6-session battery_sessions_v2.jsonl was built and partially run. The full v2 screen is too slow under the lower-tier Gemini key: full_history at concurrency 3 hit the 3M input-tokens/minute quota, and concurrency 1 works but makes the full 4-arm screen a multi-hour job. Current local artifacts (runs/e2_v2_partial_report/report.md) are directional, not final: prod completed all 6 sessions with 0 errors, all probes under compaction, 14% probe accuracy (1/7) and contradiction 0/3; full_history completed 2 Tier-1 contradiction sessions with 0 errors, no compaction, contradiction 2/2. Priority is now Tier 1 only: finish the missing full_history contradiction session, then run Tier-1 profile_memory and optionally Tier-1 incontext_history_retrieval. (2026-07-15: superseded β the stage-1 factorial ran the full battery Γ3 trials on DeepSeek (F35βF37), and the contradiction cells here sit on the session the audit later found confounded; new runs must use battery_sessions_v2_1.jsonl.)
Harness corrections (bugs in our measurement, not findings): the overnight 06-12 stall was machine sleep hanging API streams (all four pipelines stopped the same minute; fixed with the per-turn timeout); battery lesson-URLs carried a /discussions/ suffix that silently zeroed recall@lesson until normalized (caught by the A3 smoke, re-graded from bundles without re-running).
Pricing correction (2026-06-13).
MODEL_PRICINGunder-pricedgemini-3.5-flashbefore this date ($0.30/$2.50 per MTok vs the correct $1.50 input / $9.00 output / $0.15 cache-read, verified against Google's price sheet). Every dollar figure generated earlier β the Part B table and the inline costs in findings F2/F8/F9/F12 β is **4.4Γ too low**; token counts are unaffected. Part B bundles were re-costed from their saved token counts (2026-06-15) andruns/b_report/report.mdregenerated at correct pricing: relative rankings are unchanged (F9 still hasfull_historycheapest, $0.11/turn sessions), only absolute dollars move. Real Part A+B spend β $338, not ~$73; the comprehensive program β $590 across all Gemini 3.5 Flash runs (Part B ~$323 Β· Part C ~$163 Β· rest ~$103). Use the regenerated report for absolute costs; finding dollars above are pre-correction.Stable
sheet_row_id(2026-06-15).evals.gradebuiltsheet_row_idwith builtinhash(), which is salted per process, so regenerating ahandgrade_sheet.csvproduced ids that no longer matched the frozen workbook keymap β silently emptyinghandgrade_workbook merge. Switched to ahashlib.md5hash; the Part B re-merge was rebuilt via the deterministicrun_idto recover the human grades.cleared_tool_outputsis blind to library clearing (2026-06-19).ContextEditingMiddlewareedits a per-calldeepcopyand never persists the placeholder to the checkpoint thatcontext_window_statsreads, so the signal is 0 for all presets whether or notClearToolUsesEditfired (full detail + token evidence in F24). Detect clearing via input-token deltas; the custom Part C view-mechanisms avoid this by reporting through theapp.telemetryturn-signal registry.Faithfulness parked; bundle tool-output cap raised 6kβ40k (2026-06-23). The new faithfulness rubric (groundedness vs retrieved evidence, F33) is not reportable on existing runs:
evals/run_battery.pycapped each saved tool output atTOOL_OUTPUT_MAX_CHARS = 6_000while the KB shell returns up to 40k (app/kb_shell.pyDEFAULT_MAX_OUTPUT_CHARS), so on KB-heavy arms the judge saw a fraction of what the agent actually grounded on and scored capture-completeness, not grounding β faithfulness tracked the KB:retrieval ratio almost monotonically (b_st_prodkb:ret 10.3 β 4%;c_st_prod7.7 β 25%;kb_off/graphragkb:ret 0 β 50β76%), the same blind-signal class as F23/F24. Raised the cap to 40k (captures what the agent saw;output_charsstill records the true length so any overflow stays visible) so future re-records make faithfulness measurable; existing bundles cannot be backfilled (content past 6k is gone β needs a re-run). Until thenevals/grade.pyemits faithfulness only when evidence was captured in full (evidence_is_complete, self-re-enabling after the cap raise), and the column is dropped downstream (grading_merge.DROP_ITEM_TYPES;report.pyomits it directly). Holistic (F33/F34) is unaffected β it grades the answer, not the evidence.Report mislabeled judge grades as human (2026-06-19, fixed).
evals.reporthardcoded "(human grade)" on the behavior / key-point / probe-accuracy rows, but Part C runs were filled by the LLM judge (judge_filled.csv), not a person β soruns/c_report/report.mdasserted a human label over judge numbers. Fix:evals.gradenow records agrade_sourceper merged grade (the judge prefixes its note[judgeβ¦]; humans do not), andevals.reporttags each column[judge]/[human]/[mixed]and renames the rows to "(graded)". Numbers are unchanged (and judgeβhuman per F14); only the provenance label is now honest. Same label-vs-measurement class as F23/F24. All Part C arms were re-graded andc_reportregenerated.input tok/turnis billed-across-calls, not context size (2026-06-19, clarified). Telemetry sumsusage_metadataover every internal model call in a turn, so a high-loop arm (observation_truncation, 22.8 calls) bills lots of input from small contexts while a big-payload arm (kb_off) bills similar totals from few large ones. Correct for cost, but it is not "how big is the context window." The report now shows both: input tok/turn (billed, all calls) and a new context tokens/turn (window size) row (fromcontext_tokens_approx, always computed but previously unsurfaced). On single-turn the two even rank arms differently (e.g.kb_offhas the largest window at 73k from its 100k retrieval payloads, but mid-pack billed input). Read F1/F6's "~200k input" as billed volume, not window size.Latency rows are confounded by time-of-run (caveat, not yet fixed).
run_part_c_screen.shruns arms sequentially, and TTFT/total-ms depend on API load at that moment, so cross-arm latency deltas (e.g. F19's 430s p95, F23's 19s-vs-38s) partly reflect when each arm ran. Token/cost/recall are run-time-independent and unaffected. Any latency claim that drives a decision should come from arms run interleaved (or repeated trials), not a single sequential screen.total β ttftdoes not remove this β it is only the answer-streaming tail (~6s, preset-insensitive); the agentic latency incl. the summarization tax is all inttft, so subtracting it deletes the signal. Report the unconfounded work metrics (LLM calls/turn, tokens/turn) as the headline and treat seconds as supporting. See F27.Latency de-confounding tooling (2026-06-22). Acting on F27: (1)
runs/run_part_c_screen_interleaved.shruns the screen question-outer / arm-inner (writesruns/ci_*, notruns/c_*) so every question's arms share one API-load window β same total runs as the sequential screen, so interleaving is free. (2)context_statsnow also emitstime_to_first_token_ms(first reasoning/tool/text token, always β€ttft_ms), separating raw model responsiveness from the tool-call loop thatttft_msincludes; it flows throughevals/grade.pyto the report and showsβon runs recorded before it existed. (3)evals/report.pynow leads with the run-time-independent work metrics (LLM/tool calls/turn, tokens/turn) and tags every latency row[confounded], with a runtime note. Allruns/*report*were regenerated (layout/labels only; numbers unchanged). The scripts live in gitignoredruns/like the rest of the run tooling.Mixed-judge grading confound (2026-07-15, fixed). While grading the DeepSeek stage-1 compaction factorial (F35/F36,
runs/deepseek_compaction_stage1_triggerfix_20260715), the first quality merge silently mixed two judges: an interrupted Codex (gpt-5.6-sol) subscription pass (evals.run_subscription_grading) had graded chunks 000β011 of the grading plan (runs/_grading/deepseek_compaction_stage1_full_20260715) β almost exactly theexp_c200_cap10krows β before the blinded subagent workflow (evals/grade_workflow.js) graded the rest. The GPT judge is one-directionally harsher on holistic: on the 462 double-graded rows, 67% agreement, with GPT failing 151 rows the subagent judge passed and zero the reverse (fail rates 35% vs 3%; its extra fails are mostly technical-accuracy critiques the blind rubric can't adjudicate without evidence). That madeexp_c200_cap10k's holistic read 63% vs 96β98% for the other arms β a grader artifact, not an arm effect (the same rows re-graded by the subagent judge score 98%; the tell was the straddling chunk, where the same arm scored 76% under GPT and 97% under the subagent judge). Fix: chunks 000β011 re-graded so all 44 chunks share the F14-validated subagent judge; the Codex verdicts are preserved for comparison (runs/_grading/deepseek_compaction_stage1_full_20260715/codex_partial/). Lessons: one judge per battery β grade provenance is part of the measurement (same class as the judge-as-human label bug above); and two frontier judges disagreeing this much, one-directionally, on holistic is further evidence that holistic needs the held-out human validation F34 called for.battery_sessions_v2filler recycling is systematically confounded (2026-07-15, two audit passes; repaired asbattery_sessions_v2_1.jsonl). v2's filler was recycled verbatim from v1 turns, and the recycling blindly included v1 persona/fact-update/project-pivot turns β each a first-person state claim β dropped into a different persona: 7 such turns across 4 of the 6 sessions. Two sessions are unusable as authored:v2_t1_python_colab_to_local_22t(a "my iMac" turn injects a third environment between the update and the open-ended contradiction probe β every stage-1 contradiction miss in every arm sits here) andv2_t2_fullstack_long_persona_36t(two "project pivot" turns relabel the student's project to hosted architectures that contradict the probed no-PHI constraint, so its long-horizon cells measure the injected contradiction, not memory); two more are muddied (v2_t2_agentic_long_persona_36t,v2_t3_fullstack_two_projects_20tβ injected claims orthogonal to the probed facts). F36/F37 were corrected to audit-clean probe subsets. The same audit cleared everything else:battery_sessions_v1(all 32 sessions β no confounds; the update/pivot turns there are the intended probed facts),battery_singleturn_v1, andreplay_n1_v1all clean and matching their README counts. Separate defect: 5 ofbattery_personas_v1's authored anti-pattern checks are negation-insensitive substrings that false-fail correct answers (e.g. antiimport openaimatches the official Node SDK'simport OpenAI from 'openai'; antiinstall Pythonmatches "you don't need to install Python") β an unquantified downward bias on Part B's persona auto-check rows (F8/F11); repair queued. Fixes shipped:battery_sessions_v2_1.jsonlreplaces the 7 recycled turns with neutral same-course corpus questions (probes/plants byte-identical), and the runners now refusebattery_sessions_v2.jsonlfor new runs (--allow-deprecated-batteryreproduces historical runs only). Standing lesson: recycled real-student filler must be scanned for first-person state claims against the planted facts before a battery freezes.Pattern worth naming β most of our measurement bugs are telemetry/label blind-spots, not wrong math. Six now share one shape: a number or label asserted a property the instrumentation could not observe for that config β F21 (mechanism dormant:
profile_memorynever engaged on sessions), F23 (recall blind to KB browsing), F24 (clearing invisible to the checkpoint signal), the judge-as-human label above, faithfulness graded on truncated evidence (the 6k cap above β the judge scored what it could see, not what the agent saw), and F34's holistic-on-session blind spot (the quality judge sees no conversation history, so it credits answers that silently ignored the student's earlier-stated facts β "looks good" on 100% of memory-failed turns).check_triggersenforces "did the mechanism fire?" for compaction only; nothing yet does for retrieval/profile/KB engagement, grade provenance, or whether the grader saw what it needed. Standing guard for any new arm or metric: before trusting a per-arm number, confirm the signal/label/grader reflects what the config actually did.
Talk-outline correction: the planned "skills / just-in-time KB instructions" change (Change 4) was never built, and its premise is off β the KB instructions block measures 458 tokens (not ~4k) inside a 1,655-token system prompt; max saving ~2% of a turn, mostly cache-discounted. Recommended reframe: demo progressive disclosure on tool outputs (F1), where our numbers actually are.
Remaining work
Omar
- Grade the session probes (the critical tier). 2026-06-15: all 96 reviewed β the 83 high-confidence LLM verdicts by sign-off, the 13 low-confidence ones individually β with zero overrides, so the provisional LLM grades are now human-confirmed ground truth (see F13). The non-probe rows (persona / key_point / behavior in
review_other.md) were not part of this pass and remain LLM-graded; bless them too if those columns need a human label. - Audit
data/eval/review_log_v1.md(the 11 flagged judgment calls first). - Verify and update the
gemini-3.5-flashentry inapp/telemetry.py:MODEL_PRICINGagainst the current price sheet.
Workshop
- Replace placeholder slide numbers with measured ones; pick 2β3 failure traces from the graded probes for the failureβfixβnumber beats (a
fact_updatefailure under summarization is the target demo). - Token-vs-turn plot: CSV is ready; add matplotlib (or plot elsewhere) for the PNG.
- Next.js meter component over the
data-context-statsSSE part, if the live demo needs it. - Skills section: reframe per the 458-token measurement above.
Beyond the workshop. Part C itself is done and reported above (findings F14βF24); what remains only hardens or extends it and is consolidated in "Part C β what remains" above β promotion, finishing the v2 contradiction tier, the incontext_history_retrieval long-session test, the conditional subsystem builds, and the deliberately-not-pursued list. The full execution detail (variant wiring, telemetry signals, the subset run matrix, the cost model, and the v2 battery plan) lives in evals/part_c_plan.md.
Product quality track (post-workshop)
Golden cases as a CI gate (5β10 critical-path cases with deterministic assertions, including the F7 non-empty check) β error-analysis cycles (an expert reads 50β100 traces in a small viewer, records pass/fail plus the first failure reason, and stops when new failure types stop appearing) β failure taxonomy β code assertions for recurring modes β LLM judges built from the hand-grade critiques and validated to >90% true-positive and true-negative rates on held-out human labels before any judge-graded number is reported β nightly battery + weekly trace sampling. Full methodology behind each step: evals/background.md.
Files
evals.mdβ this file: the what, the data, the results, the queue.evals/contributing.mdβ contributor playbook: how to run an experiment and land it (run β upload data to HF β record anF<N>finding β merge), with a definition-of-done gate.evals/part_c_plan.mdβ Part C execution plan: orchestration (workflow vs direct), the variant catalog with exact wiring + telemetry signals, the subset run matrix, and cost.evals/background.mdβ research sources (Hamel, howtoeval, OpenAI macro-evals) and design rationale.evals/graphrag.mdβ GraphRAG vs classical RAG experiment (branchexperiment/graphrag-vs-rag): a scoped, true-GraphRAG head-to-head on the single-turn battery. Revisits the dropped GraphRAG idea with a fair test.evals/compaction.mdβ compaction-vs-keep-everything study (branchexperiment/context-compaction): on the largest lesson, keep-all vs every compaction method vs retrieve-per-question. Standalone fleet on Gemini 2.5 Flash β not comparable to the 3.5-Flash results above.evals/slm_compaction.mdβ knowledge-compaction on small local models (branchexperiment/slm-compaction): two experiments on 3 Ollama SLMs (llama3.1:8b, qwen2.5:7b, qwen3:8b) at a 32k window where the lesson overflows. Axis B (fit a document: RAG vs stuff vs summarize, F30) and Axis A (compact growing history via the real memory presets, F31). The regime where compaction is forced; RAG wins Axis B, no universal winner on Axis A, and "keep everything" never wins on an SLM (synthesis in F32).data/eval/README.mdβ battery schemas + glossary of terms;review_log_v1.mdβ dataset audit trail.evals/β the harness code;app/memory_presets.py,app/telemetry.pyβ the app-side hooks.runs/b_report/,runs/c_report/β generated tables, token curves, blinded grading workbook (Part B and the Part C screen). Re-grade / follow-up reports:runs/kbfair_report/(F23 KB-fair recall),runs/d_report_profile_memory_active/(F22 profile rerun),runs/e2_v2_partial_report/(v2 partial).runs/ds_v2_t1_report/,runs/ds_v2_t2_report/β DeepSeek-V4-Flash v2 matrix (F25/F26: cost + contradiction + long-horizon), graded via the subagent path (runs/ds_judge/). Shared on HF aseval_runs/part_f_deepseek_v2/(the standardmv eval_runs/part_*/* runs/download restores them). Provider added inapp/chat_service.py(deepseek/openrouter), pricing inapp/telemetry.py.runs/deepseek_compaction_stage1_triggerfix_20260715/report/,runs/deepseek_structured_only_stage1_20260715/report/β DeepSeek long-context compaction study, stage 1 (F35/F36): the 2Γ2 {full_history, summarize@200k} Γ {raw, stable 10k tool-output cap} factorial, plus the follow-up prefix-preserving compaction arm (F37:exp_c200_cap10k_structured,summarization_strategy="structured_prefix"inapp/chat_service.pyβ the summarizer call rides the provider cache instead of LangChain's cache-missing XML re-serialization; β14%/trajectory vs the XML arm). Writeups:stage1_findings.md/structured_findings.mdin each report dir (incl.audit_20260715.md, the validity audit); runnerevals.run_compaction_experiment; graded via the blinded subagent path intoruns/_grading/deepseek_*_20260715/. Shared on HF aseval_runs/deepseek_compaction_stage1/(runs +_grading/judge verdicts; restore line in the download snippet above).battery_sessions_v2.jsonlis deprecated for new runs (python_colab confound β see harness corrections); the runners refuse it, use the repairedbattery_sessions_v2_1.jsonl.runs/axisa_*_report/β SLM compaction study (F30/F31/F32, branchexperiment/slm-compaction): Axis B fit-a-document and Axis A compact-history on 3 Ollama SLMs at a 32k window. Shared on HF undereval_runs/slm_compaction/(axis_b/<model>/,axis_a/axisa_<model>_<preset>/) + batterieseval/compaction/; the snippet above restores Axis-A toruns/axisa_*, Axis-B stays undereval_runs/slm_compaction/axis_b/β full restore command inevals/slm_compaction.md. Providerollamaadded inapp/chat_service.py.