|
Download evals.md from towardsai-tutors/ai-tutor-chatbot: direct link, hf CLI and curl.
- Browser
- Download file 99.2 kB
-
https://huggingface.co/spaces/towardsai-tutors/ai-tutor-chatbot/resolve/1a2b802acfecc2ab6a1e56d795540fa60acfdd8a/evals.md
- Command line
-
hf download hf://spaces/towardsai-tutors/ai-tutor-chatbot@1a2b802acfecc2ab6a1e56d795540fa60acfdd8a/evals.md
-
curl -L -o evals.md https://huggingface.co/spaces/towardsai-tutors/ai-tutor-chatbot/resolve/1a2b802acfecc2ab6a1e56d795540fa60acfdd8a/evals.md
99.2 kB
| # Evaluating the AI tutor | |
| This document explains how we test the AI tutor with repeatable conversations and measured results. The main question is: when the tutor has a long chat history, which memory strategy gives the best answers without wasting too many tokens, dollars, or seconds? | |
| The eval setup compares answer quality, memory retention, retrieval accuracy, tokens, cost, and latency for the June 2026 workshop. The same harness also becomes the ongoing quality program afterwards. Research sources and design rationale live in `evals/background.md`. | |
| **Running a new experiment?** See `evals/contributing.md` for the step-by-step (run β upload data to HF β record an `F<N>` finding here β merge) and the definition-of-done checklist, so the next person inherits both your result and the data to reproduce it. | |
| Plain-English map: | |
| - A **battery** is a dataset of test conversations or test questions. | |
| - A **run** is one battery executed with one model and one memory setting. | |
| - A **trace bundle** is the saved JSON record for one tutor turn: the question, answer, tool calls, sources, token usage, timing, and any error. | |
| - A **grade** is the pass/fail or metric row computed later from saved trace bundles. | |
| - A **memory preset** is a named setting that controls how much chat history the tutor keeps, summarizes, or stores as a student profile. | |
| - An **axis** is one independent dimension the experiment varies. Part B varied a single axis (the memory preset). Part C uses two: **A β memory & context management** (how conversation history is kept or compacted) and **B β retrieval & tool outputs** (how retrieved docs and tool results are sized, cleared, or whether the KB tool exists). A **single-axis** run changes one variant on one axis and holds everything else at the production baseline, so any difference is attributable to that one change; a **cross-product** tests every Axis-A Γ Axis-B combination at once (far more runs, far bigger bill). | |
| ## What we evaluate, and what varies | |
| The system under test is the production tutor (`app/`): the same agent, retrieval tools, prompts, and telemetry used by the app. The tutor can answer from the course/docs corpus through `retrieve_tutor_context`, browse the local knowledge base through `run_kb_command`, and keep per-thread conversation memory. | |
| **The single variable is how the tutor manages context** β 13 arms across two axes in the Part B/C screen (**A** = memory/history, **B** = retrieval/tool-outputs), with later experiments adding more arms (F29's summarization family, F35/F37's `exp_*` factorial + structured-prefix arms; `exp_c400/c800` built, unrun). Within each experiment everything else is held constant β system prompt, retrieval configuration, selected sources, web tools off, temperature, and the model. **The Part B/C screen held the model at `gemini-3.5-flash`; later experiments varied it** (DeepSeek-V4-Flash β now the production default β in F25/F26/F35βF38, Gemini 2.5 Flash in F29, three Ollama SLMs in F30/F31), with per-finding comparability caveats: absolute numbers never cross fleets. Most arms are a `MemoryConfig` preset (`app/memory_presets.py`, part of the agent cache key); a few vary a retrieval knob or a tool toggle instead. In this document, **compaction** means reducing the prompt by summarizing older messages or clearing old tool outputs. | |
| **The arms β every strategy we tested:** | |
| | Arm | Axis | Technique | What it tests | Phase | | |
| |---|---|---|---|---| | |
| | `full_history` | β | keep everything (no compaction) | quality/memory upper bound; token worst case | B + C anchor | | |
| | `prod` | A+B | summarization + tool-result clearing | the live production baseline | B + C anchor, D | | |
| | `aggressive` | A+B | early summarization + clearing | high-pressure compaction | B | | |
| | `profile_memory` | A (store) | semantic profile store + prod compaction | long-term personalization | B, D rerun | | |
| | `summarization_only` / `editing_only` | A / B | one compaction technique alone | isolate each half of prod's bundle | defined, folded/dropped (F3) | | |
| | `sliding_window` | A | trimming / keep last N | recency-only memory | C | | |
| | `prompt_compression` | A | rewrite history to fewer tokens | shrink text, drop no facts | C | | |
| | `selective_retention` | A | constraint-preserving summarization | quality-preserving summary | C | | |
| | `context_reset` | A | summary-seeded fresh state | aggressive prefix rewrite | C | | |
| | `incontext_history_retrieval` | A | retrieve relevant old turns | retrieve vs carry/summarize all history | C | | |
| | `clear_retrieval_kb` | B | clear tool outputs incl. retrieval/KB | the F3 fix (prod exempts them) | C | | |
| | `observation_truncation` | B | head/tail-truncate tool outputs | trim the dominant token source | C | | |
| | `retrieval_budget_30k` | B | retrieval token-budget knob (100kβ30k) | does a tighter budget drop recall (F1) | C | | |
| | `kb_off` | B | disable the KB browse tool | agentic browse vs top-k RAG | C | | |
| **The datasets β what we ran the arms on** (full schemas in `data/eval/README.md`; detail in "The dataset" below): | |
| | Battery | Size | Tests | What actually ran | | |
| |---|---|---|---| | |
| | `battery_singleturn_v1` | 60 one-shot Qs | answer quality + response type | B (all 60), C (24-case subset) | | |
| | `battery_sessions_v1` | 32 sessions (113 probes) | memory under compaction | **only 3 sessions ever** (s01/s02/s08), in B and C | | |
| | `battery_personas_v1` | 10 personas Γ 4 Qs | long-term profile memory | B (all) | | |
| | `replay_n1_v1` | 30 "next staff reply" cases | multi-turn behavior on real threads | **built, never run** | | |
| | `battery_sessions_v2` | 6 sessions, 3 tiers (contradiction / long-horizon / entity) | where `full_history` might lose | Gemini partial (`prod` all 6, `full_history` 2 Tier-1), then in full on DeepSeek: the 12-arm v2 matrix (F25/F26) and the Γ3-trial stage-1 factorial (F35βF37). **Deprecated 2026-07-15** (filler confound, see harness corrections) | | |
| | `battery_sessions_v2_1` | v2 with 7 recycled filler turns repaired (probes/plants identical) | same tiers, confound-free | built 2026-07-15; the required v2 successor for new runs | | |
| **What actually ran, by phase:** **Part B** β the 4-arm bake-off (`full_history`, `prod`, `profile_memory`, `aggressive`) over the full single-turn + personas batteries + 3 sessions, Γ2 trials (β F1βF13). **Part C** β the 11-arm screen (the 2 anchors + 9 variants) over a 24-case single-turn subset + 3 sessions, Γ1 trial (β F14βF20, F23). **D** β re-ran `profile_memory` with its store actually engaged (β F22). **v2** β a partial long-context / contradiction run (β the open v2 items in "Part C β what remains" below). Later phases moved off Gemini: the **DeepSeek v2 matrix** (F25/F26), **GraphRAG head-to-head** (F28), the **Gemini-2.5 compaction study** (F29), the **SLM studies** (F30/F31), and the **DeepSeek stage-1 factorial + structured-compaction follow-up** (F35βF38, Γ3 trials on the v2 sessions battery). Nothing has been run at promotion grade (full batteries Γ 3 trials); the earlier phases are 1β2-trial screens. | |
| *Why the BβC order:* Part B's findings drove it β tokens live in tool outputs not history (F1/F3) and compaction can cost more than it saves (F2/F9), so the cheap retrieval/tool-output variants ran first and the history-compaction variants ran to confirm that inversion. That narrative lives in the Results sections below; this section is the complete catalog. | |
| ## What we measure | |
| | Layer | Metrics | Source | | |
| |---|---|---| | |
| | Runtime and cost | Time to first answer text, total turn time, input/output tokens, cached tokens, estimated dollars, number of model calls, and whether compaction fired. | The `context_stats` event emitted at the end of each turn (`app/telemetry.py`). This does not depend on LangSmith. | | |
| | Tool behavior | Which tools the tutor called, how often, and whether it re-searched for information it had already seen. | Tool-call events saved in each trace bundle. | | |
| | Retrieval | Whether the retrieved chunks included the correct course or lesson. `recall@shown` means "was the right source among the results shown to the model?" `MRR` rewards putting the right lesson earlier in the shown results. | Ground truth stored on each real discussion case: `source_key` and `lesson_url`. | | |
| | Response behavior | Whether the tutor did the correct kind of thing: answer from course content, answer generally, redirect a platform/support issue, or acknowledge feedback. | Cheap code checks for early signals; final behavior accuracy comes from human grades. | | |
| | Answer quality | Whether the answer covered the required key points, passed graded memory probes, and (holistic) whether course staff would approve sending it as a genuinely helpful, learning-oriented reply. A separate **faithfulness** check (every claim grounded in retrieved evidence) is built but parked β see the cap correction below. | Human pass/fail grading in a blinded workbook, or the validated blind subagent judge (same per-item rubric). LLM judges are used only after they are validated against human labels. Holistic and probe quality are complementary β holistic is graded blind to conversation history, so it cannot see a memory failure (F34); read them together. | | |
| | Memory | Whether the tutor remembered or updated facts from earlier in the session, followed preferences, resolved references like "that thing from earlier," and used stored student profiles. | Session probes and persona checks. | | |
| Cost is reported both as raw tokens and as estimated dollars after cached-token discounts. These can rank presets differently (finding F2). | |
| ## The dataset | |
| Source: `data/academy_discussion_eval.jsonl` β 151 real student posts from the academy discussion boards with real staff answers, annotated by gemini-3.5-flash. All eval inputs are real or hand-authored; **we never generate reference answers** β ground truth is staff replies (distilled to `key_points`) or facts we wrote ourselves. | |
| **Review.** All 62 gold + 30 not-time-bound usable annotations (92) were re-reviewed against the live KB with file/line evidence. 60 kept, 32 excluded. Main corrections: the annotator badly under-flagged staleness (course content was updated after many posts β ~14 exclusions had premises no longer true in today's corpus), 5 near-duplicates, 1 fabricated URL, 1 misaligned question/answer pair, and key points rewritten to claims that are atomic, still true, and binary-checkable. Colab-notebook dependency was checked per case: **1 of 92** truly required the notebook (excluded); lessons embed the cells everywhere else. Full audit trail: `data/eval/review_log_v1.md` (11 judgment-call cases flagged for human review). | |
| **Batteries** (`data/eval/`, schemas and glossary in its `README.md`; files are gitignored because they contain real student text): | |
| | Battery | Contents | Tests | | |
| |---|---|---| | |
| | `battery_singleturn_v1` | 60 reviewed real questions asked as one-off chats. They include course-content questions, support/platform issues, general AI/programming questions, and course feedback. | Basic answer quality and correct response type. Memory presets should score about the same here; big differences suggest a bug. | | |
| | `battery_sessions_v1` | 32 multi-turn study sessions. Early turns give the tutor facts about the student, middle turns make the chat long, and later turns check whether important facts survived. | Memory under pressure. A **probe** is one of the later turns that gets graded. Probe examples: remember a fact, respect a preference, understand "that thing from earlier," or use a newer fact after the student changed their mind. | | |
| | `battery_personas_v1` | 10 stored student profiles with 4 questions each. The questions are written so the best answer needs the profile. | Long-term profile memory: does a stored profile actually improve personalization across fresh chats? | | |
| | `replay_n1_v1` | 30 tests made from real multi-turn discussion threads. For each test, we give the tutor the conversation up to just before a real staff reply, ask it to write the next reply, and compare it with what staff actually wrote. | Multi-turn behavior on fully real data. This is secondary because it needs human or validated judge grading. | | |
| ## How it runs | |
| ```bash | |
| uv run -m evals.run_battery --battery data/eval/battery_singleturn_v1.jsonl --preset prod --trials 2 --out runs/exp | |
| uv run -m evals.grade --run runs/exp # code checks + handgrade_sheet.csv | |
| uv run -m evals.check_triggers --runs runs/exp # gate: compaction fired where probes assume | |
| uv run -m evals.report --runs runs/expA runs/expB # side-by-side tables + token curves | |
| uv run -m evals.handgrade_workbook build|merge ... # blinded human-grading workbook | |
| ``` | |
| Command roles: | |
| - `run_battery` talks to the tutor and saves one JSON trace bundle per turn. | |
| - `grade` reads saved bundles and writes automatic grades plus a CSV for human judgments. | |
| - `check_triggers` verifies that session probes really happened after compaction. A memory test is not useful if memory trimming never occurred. | |
| - `report` builds side-by-side tables and token curves from graded runs. | |
| - `handgrade_workbook` builds and merges a blinded grading workbook, so the grader does not know which preset produced an answer. | |
| Every turn persists a JSON trace bundle, so grading and reporting can re-run offline without touching the API. Runs are resume-safe: re-run the same command after an interruption and completed units are skipped. Each turn has a 10-minute timeout, and LangSmith is off by default. Running the tutor needs `COHERE_API_KEY` plus the key for the model under test β `DEEPSEEK_API_KEY` for today's default (with only a Gemini key the app substitutes gemini-2.5-flash outright and the harness flags every bundle); the Part B/C runs used `GEMINI_API_KEY`; the full 4-preset comparison run below cost **β$323** at corrected pricing (the pre-2026-06-13 price table under-priced `gemini-3.5-flash` ~4.4Γ; the often-quoted "$73" is a pre-correction figure β see Harness corrections). | |
| **Setup for collaborators.** Code and docs are in git; the datasets and run results contain real student text and live only in the private HF dataset (`towardsai-tutors/ai-tutor-data` β git is force-pushed to the public prod Space on deploys, so student data never enters git). With an `HF_TOKEN` that can read it: | |
| ```bash | |
| export HF_TOKEN=hf_... # token with READ access; the inline cmd does NOT load .env (or: uv run huggingface-cli login) | |
| mkdir -p data/eval runs/_grading # runs/ is gitignored, so a fresh clone has no runs/ dir yet | |
| uv run python -c "from huggingface_hub import snapshot_download as d; d(repo_id='towardsai-tutors/ai-tutor-data', repo_type='dataset', allow_patterns=['eval/**','eval_runs/**'], ignore_patterns=['eval/README.md'], local_dir='.')" | |
| mv eval/* data/eval/ && mv eval_runs/part_*/* runs/ # restore working paths (part_b, part_c, ...) | |
| mv eval_runs/slm_compaction/axis_a/* runs/ 2>/dev/null # SLM Axis-A runs -> runs/axisa_* (F31); Axis-B (F30) stays under eval_runs/slm_compaction/axis_b/, see evals/slm_compaction.md | |
| mv eval_runs/deepseek_compaction_stage1/deepseek_* runs/ 2>/dev/null && mv eval_runs/deepseek_compaction_stage1/_grading/* runs/_grading/ 2>/dev/null # DeepSeek stage-1 compaction runs + judge verdicts (F35-F38) | |
| ``` | |
| ## Results β Part B comparison run, 2026-06-12 | |
| 4 presets Γ 2 trials over the full single-turn battery, full persona battery, and 3 sessions (s01 13-turn, s02 agentic, s08 fact-update). 1,232 turns, zero API errors. Tables: `runs/b_report/report.md`; token curves: `tokens_by_turn.csv`. **Auto-graded only** means the current table includes metrics computed by code. Answer quality and session probe accuracy appear after the human grading workbook is filled and merged. | |
| | | full_history | prod | profile_memory | aggressive | | |
| |---|---|---|---|---| | |
| | personalization pass (personas, n=70) | 63% | 67% | **94%** | 56% | | |
| | sessions: est. cost/turn (corrected pricing) | **$0.112** | $0.237 | $0.218 | $0.296 | | |
| | sessions: median time to first answer text | **17s** | 21s | 22s | 43s | | |
| | sessions: tool calls/turn | **2.8** | 3.5 | 3.9 | 8.3 | | |
| | sessions: cumulative input tokens | 2.9M | 1.7M | 1.6M | 1.75M | | |
| | single-turn: behavior proxy from code checks (n=106) | **88%** | 86% | 86% | 78% | | |
| | single-turn: retrieval recall@shown source | 51% | 50% | 49% | 47% | | |
| Row definitions: | |
| - **personalization pass** β % of persona question-runs whose answer passed every authored check (expected regex matched, e.g. `conda` for the conda persona) with no anti-pattern hit (e.g. bash `export` for a Windows persona). n=70: 40 questions Γ 2 trials, minus 10 runs whose checks need human judgment. | |
| - **est. cost/turn** β mean estimated $ per conversation turn: token counts Γ `MODEL_PRICING`, cached input billed at the cache-read discount (our table, not the invoice). | |
| - **median time to first answer text** β median seconds from user message to first visible answer text. It includes tool calls and internal model rounds before the visible answer starts, so it approximates the user's perceived wait. | |
| - **tool calls/turn** β mean retrieval + KB-command invocations per turn; here it's the re-work signal (compaction presets re-search for evidence their compressed history lost). | |
| - **cumulative input tokens** β all input tokens billed across an entire session (every turn, every internal call), mean over the 6 session-runs per preset. | |
| - **behavior proxy from code checks** β % of single-turn case-runs where a cheap code check confirms the right *kind* of response (course question β called retrieval/KB; support issue β points to support; feedback β acknowledges). n=106 of 120: `answer_general` has no proxy check. This is only an early signal; authoritative behavior accuracy comes from human grades. | |
| - **retrieval recall@shown source** β % of case-runs where the reranked retrieval results the agent saw contained at least one chunk from the correct course. Stricter `recall@lesson` (right lesson page, 30β36%) and the right-lesson ranking score (MRR) are in the full report. | |
| ### Observations | |
| Numbered findings; each states what it was tested on. Convention: entries are never edited, only superseded. | |
| - **F1 β Retrieval payloads dominate input tokens, not conversation history.** Each retrieval call may return up to `DEFAULT_CONTEXT_TOKEN_BUDGET = 100_000` tokens; turns average ~200k input. (All runs; high confidence.) β retrieval budget became a Part C variant dimension. | |
| - **F2 β Compaction saves tokens but not necessarily dollars.** Summarization rewrites the prompt prefix and invalidates Gemini's implicit cache: on the same session, `full_history` had 86.8% of input billed at the ~4x cache discount; `aggressive` used 44% fewer tokens yet cost 68% more. (s03 Γ 3 presets, then confirmed n=24 session-runs; provider-specific β Anthropic caching is explicit.) | |
| - **F3 β Clearing old tool-output messages never fires on this workload.** `ClearToolUsesEdit` excludes retrieval results β where the tokens are; 0 clears across all session runs. β `editing_only` dropped from Part B; "clear retrieval results too" variant queued for Part C. | |
| - **F4 β Aggressive compaction degrades even single-turn behavior.** 18.0 vs 9.6 LLM calls/turn, 57s vs 39s median, behavior proxy β10pts vs full_history (60 cases Γ 2 trials). Mid-turn compaction churn changes agent behavior, it doesn't just trim. | |
| - **F5 β Gemini reasoning tokens are ~90%+ of output even with reasoning display off.** Billed as output; dominates latency; recorded per model in `usage_by_model`. | |
| - **F6 β Long context costs latency, not correctness.** Zero API errors in 1,018+ turns at any size (max 6.06M tokens across one turn's calls; largest single context 274k). Median time to first answer text scales 22s β 76s from <100k to >800k input tokens/turn. Compaction's value here is responsiveness and spend, not keeping the model functional. | |
| - **F7 β ~0.8% of turns produce no answer text** despite tool calls and billed reasoning tokens; preset- and size-independent. Candidate golden-case assertion ("answer non-empty"). | |
| - **F8 β Profile memory wins on quality AND cost.** 94% personalization vs 56β67% without, while the cheapest persona preset ($0.049 vs $0.081β0.095/turn, 7.6 vs 10β11.2 tool calls): the stored profile saves the agent from re-searching for user context. (40 questions Γ 2 trials Γ 4 presets, auto checks; LLM-check rows pending human grades.) | |
| - **F9 β Full history is cheapest AND fastest up to 13 turns; compaction causes re-work.** Sessions: $0.034/turn, time to first answer text 17s, 2.8 tool calls for `full_history` vs $0.051β0.066, 21β43s, 3.5β8.3 elsewhere β despite ~2x the tokens (F2's cache mechanism plus a second one: raw history lets the agent re-use earlier retrieval evidence; summaries force re-retrieval). (3 sessions Γ 4 presets Γ 2 trials, n=24, zero errors.) The conventional pitch inverts: under modern prompt caching, the naive baseline wins short-to-medium sessions; compaction must justify itself on quality (pending human grades) and long horizons. Where the crossover actually sits is a Part C question. | |
| - **F10 β Compaction degrades within-session memory, and what it drops is *old* facts.** Session-probe accuracy: `full_history` **92%** vs `prod` 38% / `profile_memory` 38% / `aggressive` 42% (n=24/preset). The collapse is entirely turn-0 material β `fact_recall` 100% β 17β25%, `preference_compliance` 100% β 0% β while `fact_update` stays **100% across all presets** (the mid-session update sits inside the kept-recent window; summarization evicts the early planted facts, not the recent ones). With F9 (full_history also cheapest and fastest here), the naive baseline wins cost, speed, AND memory on β€13-turn sessions β the workshop headline. (3 sessions Γ 4 presets Γ 2 trials; **provisional β LLM-graded under 4 reviewer rubric policies, pending human review of the 96 probes**.) | |
| - **F11 β Profile memory helps personalization, not working memory.** `profile_memory` reached 92% persona personalization (vs 54β66% without) yet scored **38% on session probes β identical to `prod`**: the long-term student-profile store improves fresh-chat answers while the live thread is still summarized exactly as in `prod`. Long-term store β in-session retention; they are independent subsystems and a preset can win one while tying the other. (Personas n=80 auto; sessions n=24, provisional.) Refines F8. | |
| - **F12 β Aggressive compaction is dominated on every axis.** Beyond F4's cost/latency blowup (18 LLM calls/turn, 57s median), it is also worst on quality: key-point coverage 36% vs 72% (`full_history`), single-turn behavior 67% vs 83β92%, session probes 42% vs 92%. No measured metric favors it. (Single-turn quality n thin at 12β25; sessions n=24; provisional.) | |
| - **F13 β Session-probe grades are human-confirmed (zero overrides), so F10βF12's memory numbers are no longer provisional.** 2026-06-15 Omar reviewed all 96 session probes β the 83 high-confidence LLM verdicts by sign-off and the 13 low-confidence ones individually β and agreed with every grade. The provisional LLM grading therefore stands as human ground truth, confirming the 92%-vs-38% probe-accuracy result (F10) and the probe components of F11/F12. Caveat: only the session-probe tier was human-reviewed; the non-probe quality rows (persona, key_point, behavior β feeding key-point coverage and single-turn behavior in F8/F12) remain LLM-graded. The judge has **not** yet been validated against these labels (next step), so judge-graded Part C numbers stay gated. | |
| ### Part C screen (2026-06-15) | |
| 11 arms (prod + full_history anchors + 9 variants), each on a fixed subset (24 stratified single-turn + 3 sessions s01/s02/s08), 1 trial, ~660 turns, 0 errors, β$163 at correct pricing. Graded by the validated subagent judge. Full table: `runs/c_report/report.md`. Probe-accuracy n=12/arm (1 trial) β treat percentages as coarse rankings, not precise rates; promotion (Γ3 trials, full batteries) would firm them up. | |
| - **F14 β The LLM judge is validated; Part C quality columns are reportable.** A blind subagent re-grade of the 96 human-confirmed probes reproduced them at 98% agreement / TPR 100% / TNR 96% (clears the >90% gate; this measures grader *reproducibility*, since the labels were sign-offs, not blind-from-scratch). The validated grader is the subagent workflow (same per-item-type rubric as `evals/judge.py`), run on subscription at no API cost; all Part C arms are graded that way. Supersedes F13's "judge not yet validated" caveat. (`runs/b_report/judge_val/`.) | |
| - **F15 β The screen reproduces F9/F10 on independent arms.** `full_history` is cheapest on sessions ($0.10/turn) AND best memory (100% probe); every compaction arm is both pricier and weaker. `prompt_compression` also reaches 100% memory β it shrinks message text but drops no facts β yet saves little (highest tokens), so "don't drop content" preserves memory but is not a cost win. (Screen; n thin.) | |
| - **F16 β `incontext_history_retrieval` works but is short-session-neutral.** 83% probe accuracy (best after full_history): retrieving the relevant old turns restores what summarization loses. But on β€13-turn sessions it can't drop much, so it costs β full_history (~$0.20/turn) plus embed overhead; its cost payoff needs LONG sessions (a `_v2` battery). The principled cost answer to F9, pending a long-horizon test; separate from the higher-priority v2 contradiction-precision question. | |
| - **F17 β Turning the KB off inverts retrieval recall (Axis B tradeoff).** `kb_off` forces `retrieve_tutor_context` instead of browsing β single-turn recall@shown jumps 50%β96%, but it is the priciest session arm ($0.31/turn) and worst memory (33%). KB browsing is cheaper and better for memory yet actively *lowers* top-k retrieval recall (the agent browses instead of retrieving the labeled source). Not a clear win either way. | |
| - **F18 β The 100k retrieval budget is over-provisioned.** `retrieval_budget_30k` matches prod's recall@shown (50%) at a third of the budget β tokens can be cut with no recall loss on this subset (direct confirmation of F1). The 10k rung is untested. | |
| - **F19 β `observation_truncation` backfires (Axis B negative).** Head/tail-truncating tool outputs in the model's view makes the agent re-call tools to recover what was cut: 22.8 LLM calls/turn, 430s p95 latency (single-turn), memory only 33% β the same churn pathology as `aggressive` (F4). | |
| - **F20 β Axis-A summary/trim arms do not rescue memory.** `context_reset` (17% probe, 56% key-points β dominated on every axis), `selective_retention` (25%), and `sliding_window` (42%) all stay well below `full_history`/`prompt_compression` and none beats `prod` (58%): dropping or summarizing early turns loses the planted facts the probes test (consistent with F10). | |
| - **F21 (2026-06-16) β Supersedes F11 in part: `profile_memory` was dormant on the sessions battery, so its 38% is `prod`, not a test of the profile.** `evals/run_battery.run_session` passes no `student_id`, and `StudentProfileMiddleware` injection plus the post-turn write-back both no-op without one β so on sessions `profile_memory` reduces exactly to `prod` and could not differ. The 38% is real but comes from the live conversation thread + prod's lossy summary (the student states the facts in-thread; summarization keeps the recent ones and a compressed summary of the rest), **not** from any profile β which is why it is ~38%, not 0. F11's reading that this shows "long-term store β in-session retention, independent subsystems" is therefore unsupported: the store was never engaged in-session. Whether a store that captures in-session facts verbatim and re-injects them recovers recall under compaction is **open** β the `collection_memory` v1 test (sessions get a `student_id` + an append-atomic-facts write-back; goalposts: compaction 17β25% vs `full_history` 100%). F11's personas number (94%) stands; personas do set a `student_id`. | |
| - **F22 (2026-06-16) β Activating `profile_memory` on sessions rescues much of the old-fact loss, but does not replace full history.** Experiment A reran the three Part-C sessions with `run_session` passing a per-session/per-trial `student_id`, so the profile store was actually read/written and injected through the system prompt on every turn. With compaction active at 100% of probes, probe accuracy rose to **75% (9/12)** vs same-screen `prod` **58% (7/12)**, while `full_history` stayed **100% (12/12)**. The gain was concentrated exactly where F10/F21 predicted: `fact_recall` improved to **83% (5/6)** vs `prod` **33% (2/6)**, and `preference_compliance` to **100% (1/1)** vs `prod` 0%. It still missed two reference/consistency probes, and was not a cost win on this short screen ($0.265/turn vs `prod` $0.229 and `full_history` $0.103). Interpretation: injecting stored facts outside summarized history can recover many facts lost by compaction, but the current "5 durable lines" profile write-back is an incomplete working-memory store. Next test: `collection_memory` / verbatim atomic facts to determine whether the remaining gap is extraction/storage loss vs answer synthesis. **Caveat (thin n):** the only robust signal is `fact_recall` (5/6 vs 2/6) and the 9/12-vs-7/12 aggregate β the `preference_compliance`/`fact_update` cells are n=1 anecdotes; and `profile_memory` actually *regressed* on anaphora / anaphora_consistency (50% vs prod's 100%, n=2 each), possibly profile injection distracting from reference resolution, possibly noise. (Codex handgrade at Omar's request; one embeddings row was borderline-pass, strict fail would make this **67% (8/12)** without changing the direction.) | |
| - **F23 (2026-06-17) β F17's recall inversion was a measurement artifact: `recall@shown` is blind to KB grounding, and a KB-fair metric erases the gap.** `recall@shown`/`recall@lesson`/MRR count only `retrieve_tutor_context` matches, so when `prod` grounds by *browsing* the KB instead of retrieving, the labeled source is invisible to the metric β which is exactly `kb_off`'s only structural difference, so the 50%β96% "inversion" was partly definitional. Re-grading the same Part C single-turn bundles (no re-run; `runs/kbfair_report/`) with two tool-agnostic measures: **recall source (any tool: retrieval+KB)** = `prod` **100%** vs `kb_off` 96% (n=24), and **cited-correct source/lesson** (does the *answer* cite the labeled source/lesson, resolved from both tools via `kb_manifest`) = **100%/100% for both** (n=14 corpus). Verified the any-tool credit is real, not a regex hit: in all **12** prod cases where retrieval missed, the agent had browsed that exact gold source via `run_kb_command`. So turning the KB off does **not** improve grounding β both configs find and cite the right source. `kb_off`'s genuine, non-artifact wins remain latency (19s vs 38s p50 TTFT), efficiency (3.6 vs 10.1 LLM calls/turn), and a slight single-turn quality edge (behavior 88% vs 79%, key-points 82% vs 78%) β all from doing one retrieval call instead of many browse rounds, not from better recall. **Supersedes F17's recall reading; F17's cost/memory points stand.** Caveats: n=24 screen subset, 1 trial; `cited_correct` saturates at 100% here (coarse at this n), so any-tool recall is the discriminating fair metric; bundle KB output is capped at 6000 chars (could undercount KB hits, which only biases *against* prod β the true gap can be smaller, not larger). New auto metrics in `evals/grade.py`: `recall_anytool_source`, `cited_correct_source/lesson` (free to recompute on any saved bundle). Caveat on the efficiency read: `kb_off`'s low tool-calls/turn (2.7 vs 8.7) is partly *structural* β an entire tool was removed β so it is not on the same "re-work" axis as the compaction arms' tool-call counts. | |
| - **F24 (2026-06-19) β F3's "0 clears" was a blind-signal artifact: `cleared_tool_outputs` cannot observe `ClearToolUsesEdit`, which fires but invisibly.** `ContextEditingMiddleware.wrap_model_call` applies the clear to a `deepcopy` of the per-call message view and returns it via `request.override(messages=...)` (`langchain/agents/middleware/context_editing.py:251-255`) β it never writes the placeholder back to the checkpoint, but `telemetry.context_window_stats` counts placeholders in the *checkpointed* `state_messages` (`chat_service.py:1876`). So the counter reads **0 for every preset including `aggressive`** across all `runs/`, regardless of whether clearing fired β "0 clears" is structural, not evidence of absence. The real effect lands in input tokens: on sessions `clear_retrieval_kb` cut mean input **145.6kβ120.5k (~17%, $0.229β$0.203/turn)** vs `prod` by clearing the retrieval payloads `prod` exempts, while on single-turn it **backfired to 308k/$0.477** (re-retrieval churn to recover what it cut β the `aggressive`/F19 pathology). So clearing is real and *mixed* (session win, single-turn loss), not absent. **Supersedes F3's "clearing never fires" reading; F3's structural point stands** β `prod`'s `exclude_tools=("retrieve_tutor_context",)` provably leaves the dominant retrieval payloads in place by construction, so prod-clearing still can't touch the big tokens. Caveats: the session token saving is 1-trial screen n; library clearing stays telemetry-invisible (the probe gate survives only because summarization, which *does* persist, co-fires in these arms β a clearing-only arm would false-fail `check_triggers`). Detect clearing via input-token deltas, not `cleared_tool_outputs`. | |
| - **F25 (2026-06-19) β DeepSeek-V4-Flash replicates F9/F10 and extends them to long horizon: `full_history` is cheapest AND best-memory on every tier, so the cost crossover never appears.** A full Part-C-style v2 matrix re-run on **DeepSeek-V4-Flash (first-party API, not Gemini)** β 12 arms, 840 turns, 0 errors; Tier 1 contradiction (3Γ22-turn) + Tier 2 long-horizon (2Γ36-turn); reports `runs/ds_v2_t{1,2}_report/`, raw on HF `eval_runs/part_f_deepseek_v2`. `full_history` is the **cheapest** arm on both tiers (**$0.0069/turn** at 22t, **$0.0101/turn** at 36t), undercutting every compaction arm β while billing the *most* tokens (1.78M input tok/turn at long horizon) because ~**97% are cache-hit**. DeepSeek's ~50Γ cache discount ($0.14 cache-miss vs $0.0028 cache-hit input) makes the enormous *cached* prefix cheaper than any compaction's cache-breaking rewrite, so the summarization arms are priciest (`prod` $0.0336, `selective_retention` $0.0357). F9's inversion therefore holds **even at 36 turns**, and harder than on Gemini (10Γ discount β crossover would appear sooner; DeepSeek's 50Γ pushes it past every length tested). Per-turn cost is ~10-15Γ below Gemini on the comparable tier. **Provider caveat:** `incontext_history_retrieval` does *not* beat `full_history` on cost here ($0.0147 vs $0.0101) β the opposite of a teammate's OpenRouter run, where fallback routing fragments the prefix cache; cost comparisons must hold the provider/caching path constant (first-party single-endpoint caching verified live, `cache_read` non-zero). Caveats: thin n (Tier 1 n=3, Tier 2 n=2 per arm, 1 trial) β coarse rankings. | |
| - **F26 (2026-06-19) β The v2 contradiction question (Q2) is answered on DeepSeek: `full_history` resolves contradictions, compaction loses them, so no temporal memory is needed.** Same matrix, judge-graded via the validated **subagent path** (28 probes, blind, `judge.py` rubric β F14). `full_history` scores **3/3 contradiction** and **2/2 long-horizon recall** β perfect memory on both tiers, including the regime where it could in theory lose (it carries both A and Aβ² raw and must pick the current Aβ²). `prod` summarization **fails contradictions 0/3** (evicts the updated fact β F10 confirmed and sharpened) and 1/2 long-horizon. `incontext_history_retrieval` matches `full_history` (3/3, 2/2) by retrieving the old turns, but costs more (F25); `profile_memory` partially rescues contradictions (2/3) via its store (the F22 pattern). Among long-horizon compaction arms, `prompt_compression` is least-bad (2/2 β shrinks text, drops no facts) and `context_reset` is dominated (0/2). Validity: compaction fired at **100% of probes** for every compaction arm, **0%** for `full_history` (gate passed). **Resolves item 4's conditional β `temporal_graph_memory` is not triggered**, since `full_history` did not fail contradictions. Caveats: thin n (3/2 per arm, 1 trial); the judge is reproducibility-validated (F14), not independent-human. | |
| - **F27 (2026-06-20) β On single-turn, `full_history`'s latency edge is a *tail* effect from `prod` summarizing mid-turn, not a median win; and `total β ttft` cannot de-confound latency.** Re-read of the Part B/C single-turn bundles (no re-run). Paired per question, the `prod` β `full_history` median **TTFT** delta is **+0.43s** (Part B, n=120, 2 trials) / +3.5s (Part C, n=24, 1 trial) and `prod` is slower in only **62/120** cases β a near-tie at the median, *not* the sessions gap (F9's 17s-vs-21β43s is multi-turn and cannot apply with no prior history). The gap lives in the **tail**: `SummarizationMiddleware` fires on **29/120** single-turn `prod` turns (24%) because the agent's *own* retrieval/browse output crosses the 30k trigger **within one turn**; those turns run **52s TTFT / 16 LLM calls** vs **29s / 8** when it doesn't, and on the *same* questions `prod` does 2β3Γ `full_history`'s calls (e.g. 9β27, 12β32) β the F9/F4 re-work pathology (summary drops raw evidence β re-browse) inside a single turn. `full_history` never summarizes, so it skips the tail; that lifts its *mean* (Part C 35s vs 51s) while the *median* stays tied. **Methodological corollary (extends the time-of-run caveat below):** `total β ttft` does **not** strip API-load noise β it is the answer-streaming tail (**~6.0s for both presets**; +1.5s even when summarization fires), because the entire agentic loop incl. the summarization tax lands in `ttft`, so subtracting it deletes the signal. Latency β work Γ time-per-work, and the confound smears across ~20 calls, so no algebra on one sequential run removes it; the **unconfounded** read is the work metrics β **LLM calls/turn and tokens/turn** (run-time-independent like cost/recall, already in every bundle) β which are also the upstream cause. Live LangSmith spot-check (3 interleaved pairs, one browse-heavy single-turn Q) reproduced **both** directions: `prod` faster once (summary collapsed 268kβ23k ctx, 65s) and slower twice (re-work: 23/32 calls, 116s/124s, $2.27/$1.04 vs `full_history` $0.74/$0.33), and LangSmith's independent latency matched telemetry to the ms (`prod` 65.07s vs 65.16s; `full_history` 103.64s vs 104.28s), trees confirming `full_history` carries no `SummarizationMiddleware`. Caveats: thin n and time-of-run confounded for absolute seconds; the live pairs are n=3 on one non-representative question. Same class as F24 (single-turn *clearing* backfire) on the other compaction half. | |
| - **F28 (2026-06-19) β GraphRAG does not beat classical hybrid RAG on single-turn course Q&A: it ties on grounding accuracy and costs +44% $/turn, empirically confirming the earlier decision to drop it.** A scoped, fair head-to-head where the *only* variable is the retrieval backend behind `retrieve_tutor_context`: production hybrid `LocalChromaRetriever` (dense Cohere `embed-v4` + BM25 β RRF β Cohere rerank β token budget) vs a true Microsoft GraphRAG index (`app/graph_rag.py` `GraphRAGRetriever` β local-search-style context assembly over entities/text-units/community-reports, mapped back to real `source`/`url` via `corpus_manifest`, then the **same Cohere rerank + token budget** as classical; context-provider only, so the agent's **Gemini 3.5 Flash is the sole generation model in both arms**). 41 `full_stack_ai_engineering` single-turn cases (27 `answer_from_corpus`), Gemini 3.5 Flash, `--disable-kb` so the retriever is the only variable, `--scope-sources`, 1 trial, **0 errors/arm**. Both arms **surface and cite the right source 100%** of the time and tie on lesson recall (**76%**) and cited-correct lesson (**85%**); classical is slightly *better* at ranking the right lesson (**MRR 0.70 vs 0.65**). GraphRAG pulls community-report context, so it spends **+61% input tokens (178k vs 110k) and +44% $/turn ($0.212 vs $0.147)** and is a touch slower, with no accuracy payoff (the +3pt behavior proxy is a coarse n=37 code check, not a quality signal). So GraphRAG's community structure does not help typical single-hop course Q&A β this **empirically supports** the prior "dropped on faith" decision (see "Deliberately not pursued" below) rather than asserting it. **Reusable artifact:** the scoped index (90 docs β 15,059 entities / 32,748 relationships / 3,371 community reports, built with Gemini 2.5 Flash for **$44.96**) is published to the private HF dataset (`graphrag/output/`, pulled on demand, *not* by prod cold-start β `ensure_local_vector_db` ignores `graphrag/**`), so the eval re-runs without the ~$45 rebuild. Full methodology, build/run commands, and the indexing gotchas (route Gemini 2.5 via its OpenAI-compatible endpoint; `run_graphrag_index.sh` loops `graphrag index` until `entities.parquet` exists since one exhausted retry aborts the stage) live in **`evals/graphrag.md`**. Connects to F17/F23 (the `kb_off` grounding discussion β classical RAG already grounds well). Caveats: scoped to one source, single-turn, n=41 (27 corpus-answer), 1 trial, auto metrics only (key-point/behavior human grades pending); `recall@source` is saturated by `--scope-sources`, so lesson recall / MRR / cited-correct-lesson are the discriminating numbers. GraphRAG's theoretical edge is **multi-hop / cross-document synthesis**, which this single-hop battery barely exercises β so this says it doesn't help on single-hop course Q&A, *not* that it never helps; a fair test of that strength needs a multi-hop `_v2` probe set (out of scope here). (Branch `experiment/graphrag-vs-rag`, PR #2.) | |
| - **F29 (2026-06-20) β Keeping a large *static document* in context is a different regime from conversational `full_history`, and it is one boundary where "keep everything" stops winning: for a single long lesson queried repeatedly, retrieving the relevant slice matches keep-it-all at ~1/13 the cost. (Boundary on F9; standalone Gemini-2.5-Flash fleet β absolute numbers not cross-comparable.)** A keep-vs-compact-vs-retrieve study on the corpus's largest lesson (~37.5k tok) loaded once in turn 0, then **15 question-turns with no tools** (the only context is what each method keeps or builds), each answer judged against the full lesson; 11 arms, Gemini 2.5 Flash, 1 trial. `rag` (fetch the relevant chunk per question) is **cheapest *and* tied-best** (60% / 3.2k tok / $0.020); `full_history` (keep the whole lesson in the prompt) is the **priciest in-context arm and not the best** (53% / 43k tok / $0.134); the summarization family cuts ~35% of tokens for a quality drop (40β53%); `incontext_history_retrieval` ties `rag` on quality but at ~13Γ the tokens; **`hierarchical_summarization` is the worst trade** β priciest ($0.605) and slowest (41.8s p50) for the lowest quality (33%). `graphrag` (53%, $0.045) again does not beat plain `rag` (consistent with **F28**). **Read it as a lesson, not a head-to-head:** this is *not* "RAG beats the `full_history` memory strategy" β here `full_history` only means "the whole lesson sits in the prompt," not its multi-turn-memory role. It says **F9's "keep everything wins" is about *conversation history*; for a big *static document* you could retrieve from, hoarding it in context is the priciest option and buys nothing**, and in the production tutor retrieval (Axis B) and a memory policy (Axis A) **compose**, they don't compete. Companion boundary: F30 (small context window). New reusable presets landed in `app/memory_presets.py`: `delta_summarization`, `hierarchical_summarization`. Full matrix, the F-C1βF-C6 detail, and reproduce steps in **`evals/compaction.md`**. Caveats: n=15, 1 session, 1 trial (coarse rankings, not precise rates); judge is Gemini 2.5 Flash (same family, reads the full lesson as ground truth); **separate fleet from the Part B/C 3.5-Flash runs β quality is not cross-comparable across models**. (Branch `experiment/context-compaction`, PR #5.) | |
| - **F30 (2026-06-20) β A small context window is a second boundary where "keep everything" stops winning, and a harder one than F29: when the document physically does not fit, "shove it all" (`full_context`) is forced into truncation and is strictly *dominated*, so retrieving the relevant slice (`rag`) is not just cheaper but necessary. (Small-window companion to F29, Axis B; local-SLM fleet β absolute numbers not cross-comparable.)** "Fit one ~37.7k-token lesson into context to answer 15 questions," 7 methods Γ 3 Ollama SLMs (llama3.1:8b, qwen2.5:7b-instruct, qwen3:8b no-think) on an M1 Pro/16 GB at a 32k window (`--num-ctx 32768`), no prompt caching; judge stays on Gemini 2.5 Flash (reads the full lesson as ground truth β the SLM cannot hold it, and a model must never grade itself), 1 trial. `full_context` (the whole *document* placed in the prompt, `ctx['full_context'] = lesson`, stateless per question β this is **document-stuffing**, *not* the conversational `full_history` memory policy) is **truncated on 15/15 turns** (37.7k β 32,767) and **dominated**: equal-or-worse quality (67β80%) at **~12β15Γ the latency** (266β346s vs ~20s) β silent information loss *plus* a huge latency tax (e.g. qwen2.5 `full_context` answered "no model was mentioned" for a fact the truncated intro had named). `rag` (chunk + Cohere-embed + retrieve top-k per question) is **top or tied-top on all three** (qwen2.5 100%, qwen3 100%, llama 80%) at ~2.9k context tokens and ~20β26s. `graphrag` ties `rag` on quality but at ~2.8Γ tokens and ~2.5Γ latency β no payoff, **consistent with F28/F29**. Summarization is the floor (llama `summary` **0%**); `trim`/`selective` land mid (47β73%). This marks where F9/F15's precondition β a window big enough to hold everything, cheaply, under caching β simply fails on the hardware a workshop attendee runs. **Distinct from the conversational-memory inversion (F31):** `full_context` here is *document-stuffing* (Axis B), not `full_history`'s multi-turn role, exactly the distinction F29 drew. **Methodology fix:** the first pass scored ~64% of judge verdicts as fails because a long JSON `reason` hit `max_tokens` and truncated the verdict; hardened the judge (25-word reason cap + leading-boolean recovery) and re-graded the saved answers offline (`evals/rejudge_compaction.py`, **0 unparseable** after) β the directly-measured axes (tokens, latency, overflow) were never affected. Caveats: n=15/method, 1 trial, one lesson, one domain β read as a ranking, not rates; **separate fleet (local Ollama SLMs, no caching, Gemini-2.5-Flash judge) β quality/cost not cross-comparable** with the Part B/C 3.5-Flash, the F29 2.5-Flash, or the DeepSeek runs. Full writeup: **`evals/slm_compaction.md`** (Experiment 1). (Branch `experiment/slm-compaction`, PR #3.) | |
| - **F31 (2026-06-20) β On an SLM, compacting a *growing conversation history* has no universal best method, and "keep everything" (`full_history`) never wins β and model capability dominates the choice of method. (SLM counterpart to F9/F15 on Axis A; same local fleet β absolute numbers not cross-comparable.)** Same lesson loaded turn 0 of a session, 15 question-turns, **retrieval off**, run through the real app middlewares (the production memory presets) on the 3 SLMs at `num_ctx=32768`, judged on Gemini against the full lesson, 1 trial. The best preset is **model-dependent** β `prompt_compression` on qwen2.5 (60%), `hierarchical_summarization` on llama3.1 (40%), `summarization_only`/`delta_summarization` on qwen3 (73%) β so there is **no single "best compaction" to recommend**. `full_history` (the genuine conversational memory policy β keep *every prior turn*, no summarization or context-editing; `summarization=False, context_editing=False`) **tops no model** (27% / 7% / 67%) and is near-bottom on the two weaker ones; `sliding_window` is reliably among the worst (13% / 13% / 40%). **Model capability dominates the method:** qwen3 (40β73% across *every* preset) >> qwen2.5 (13β60%) >> llama3.1 (0β40%) β choosing a stronger small model buys more than choosing the best strategy. **Mechanism (why keep-all can't win here):** on an SLM "keep everything" (`full_history`) can't even be truncated-and-kept β Ollama's server-side context fitting **evicts the whole oversized turn-0 lesson** from turn 1 on (the checkpoint still holds it; the model receives only ~0.3β1.4k tok of accumulated Q&A), so a method that *summarizes the lesson into a small injected message* is the only way to retain its gist; on Gemini's large window this never happens (F9's `full_history` held ~43k/turn). `hierarchical_summarization` is by far the priciest (β30-call map-reduce, p50 233β333s, ~25β29k retained tok) for no quality lead except on the weakest model. This **inverts F9/F15's conversational-memory result** (keep-all cheapest *and* best) in the regime where the runtime physically cannot keep it. Caveats: n=15, 1 trial, LLM-judge variance moves individual cells Β±1β3/15 (read only the big gaps); the `full_history` numbers are **Ollama-eviction-specific** (a different serving layer could keep a truncated lesson instead); **same separate-fleet comparability caveat as F30**. Writeup: **`evals/slm_compaction.md`** (Experiment 2). (Branch `experiment/slm-compaction`, PR #3.) | |
| - **F32 (2026-06-22) β Synthesis: "keep everything in context" wins only under a specific precondition β a large window that holds the whole history cheaply under prompt caching β and we have now mapped three boundaries where that precondition breaks and keep-all stops winning.** **Two distinct "keep-all" arms β never conflate them:** `full_history` (Axis A β F9/F15/F31) is the *memory policy* that keeps every prior **conversation turn** with no compaction (`summarization=False, context_editing=False` in `app/memory_presets.py`); `full_context` (Axis B β F30) is **document-stuffing** β the whole *source document* placed in the prompt (`ctx['full_context'] = lesson`). Same slogan ("keep everything"), different mechanisms. **F29 is the trap to watch:** its keep-all arm ran the `full_history` *preset* but over a single static lesson parked in turn 0 (`turns = [lesson] + questions`), so what it actually measured is the `full_context` idea (a document held in the prompt), **not** `full_history`'s multi-turn-memory role β which is why F29 sits on the *document* boundary (1) below, alongside F30, and **not** the conversational boundary (3). F9/F15 established the default (raw `full_history` is cheapest *and* best memory on conversational sessions), and F2/F25/F26 explain *why* it holds and even strengthens with horizon (the big **cached** prefix is cheaper than any compaction's cache-breaking rewrite; DeepSeek's ~50Γ cache pushes the crossover past 36 turns). Remove a piece of that precondition and keep-all loses, along three boundaries: **(1) a large static document you could retrieve from** (F29, Axis B, large cached window) β retrieving the slice *matches* keep-all at ~1/13 the cost, so hoarding the doc is the priciest option and buys nothing; **(2) a context window too small to hold the document** (F30, Axis B, SLM) β keep-all (`full_context`) is forced into truncation and is strictly dominated, so retrieval is *necessary*, not merely cheaper; **(3) a growing history on a small window** (F31, Axis A, SLM) β the runtime *evicts* the oversized turn, so keep-all (`full_history`) never tops the table and some compaction is mandatory, though no single method wins. Two cross-cutting lessons fall out: **model capability dominates the compaction method** on weak local models (F30/F31 β a better small model beats a better strategy), and the **axes compose, they don't compete** (F29) β retrieval (Axis B, *how to fit a document*) and a memory policy (Axis A, *how to compact history*) are orthogonal, so "RAG vs `full_history`" is a category error. **Net guidance:** keep-everything stays the right default in the production regime (large cached cloud window + conversational history; F9/F15/F25/F26); reach for retrieval or compaction exactly when one precondition breaks β a big static doc, a small window, or no caching. **Comparability:** this ties the *lessons* across fleets (3.5-Flash screen, 2.5-Flash F29, DeepSeek F25/F26, local SLMs F30/F31), **not** their absolute numbers, which are not cross-comparable. | |
| - **F33 (2026-06-23) β New quality axis (holistic "would staff send this?"): on single-turn, compaction degrades *answer quality itself*, and the degradation dose-responds to how often summarization fires *within a single turn*. (Independent quality-axis confirmation of F4/F19/F27.)** Added two judge rubrics in `evals/judge.py`: **holistic** (whole-answer staff-approval / does-it-help-the-student-learn gate, emitted for every answered turn) and **faithfulness** (groundedness vs retrieved evidence β *parked*, see the correction below). Graded blind by the validated subagent path (F14) across all single-turn runs; holistic rows added to `evals/report.py`, tagged `[judge]`. The single-turn premise is that memory presets should barely differ ("Memory presets should score about the same here; big differences suggest a bug", Β§The dataset) β yet holistic spans **57%β100%**, and it is *not* a bug: it tracks within-turn summarization. Dose-response: arms that never summarize mid-turn score 92β100% (`full_history` 92% Part C / 92% Part B n=120; `kb_off` **100%** at just 3.6 LLM calls; `sliding_window`/`incontext_history_retrieval`/`selective_retention`/`retrieval_budget_30k` 96%), while arms whose *own* retrieval/KB output trips the trigger mid-turn sink in trigger order β `observation_truncation` **71%** (22.8 LLM calls/turn, 12/24 turns summarized), `context_reset` **62%** (18/24), `aggressive` **57%** (Part B n=120, **110/120** turns summarized mid-turn). So the F4/F19/F27 within-turn re-work pathology is not merely slower/pricier β it makes the *answers worse*, at a product-relevant magnitude (`aggressive`: ~half its single-turn answers would not be sent). Caveats: Part C n=24 (directional); the firm cell is Part B `aggressive` 57% vs `full_history` 92% at n=120; judge reproducibility-validated on probes (F14), **not** independently on holistic. | |
| - **F34 (2026-06-23) β Across a session, standard compaction's damage is *surgical*: it collapses memory while answer quality stays ~100% β but that "quality" is partly an artifact of a history-blind judge, so "looks fine" β "served the student". Only the most aggressive rewrite (`context_reset`) degrades the answers themselves, progressively over the session. (Refines F10/F12 on the holistic axis.)** Holistic graded on every answered turn of `b_se`+`c_se` (683 turns, 15 arms, blind subagent path). **The decoupling:** `prod` probe **58%**(c)/**38%**(b) yet holistic **97/99%**; `kb_off` probe 33% / holistic 100%; `selective_retention` probe 25% / holistic 100%; `profile_memory` probe 38% / holistic 100% β memory and answer-quality are *independent failure axes*. The lone exception is `context_reset` (probe 17%, holistic **69%**, holistic dropping **β24% earlyβlate turn**; `sliding_window` a milder β16% late): broad, progressive quality degradation appears only under the most destructive history rewriting, otherwise quality holds flat (`Ξ lateβearly` β 0 for all standard arms). **But "high holistic" is misleading on exactly the turns that matter** β holistic is graded *blind to the session history* (sessions carry no per-turn gold answer, and the judge sees only question+answer), so on the **memory-failed probe turns** it still rated the (generic, context-ignoring) answer "good" **100% of the time** for every standard arm (`c_se_prod` 5/5, `b_se_prod` 15/15, `kb_off` 8/8, `selective_retention` 9/9; `forgot+looks-bad` = **0**) β only `context_reset` forgot so badly the blind judge caught it (6/10 forgotten-fact turns also looked bad). The failure mode is therefore *confidently generic*: a well-formed answer that silently ignores what the student stated earlier, which only the targeted probe detects. **Lesson:** holistic (context-free quality) and probe (context-aware memory) measure different things and must be read together β "answer quality looks fine" does **not** mean the student was served. This refines F10 (memory dies but general answer quality survives) and F12 (only the *aggressive* rewrites *also* degrade quality). Caveats: `c_se` is 3 sessions Γ 1 trial (~18 turns/early-late bucket), `b_se` Γ2 trials; session holistic graded without a staff reference (the within-arm early/late comparison controls for that); judge reproducibility-validated on probes (F14), not on holistic β the blind-spot above is itself why a held-out human validation of holistic is the next gate. | |
| - **F35 (2026-07-15) β 2Γ2 factorial on DeepSeek: the cheap win is a *stable* tool-output cap, not compaction β capping cuts cost β38% at no measured quality/memory loss while preserving the cache, and summarization adds cost on BOTH sides of the factorial, so it still doesn't pay for itself even after capping. (Decomposes F25; Axis A Γ Axis B; DeepSeek-V4-Flash first-party β absolute numbers not cross-comparable with the Gemini fleets.)** *Experiment: **DeepSeek long-context compaction study, stage 1.** The study asks whether summarization-based history compaction ever pays for itself on the production tutor running its default model (DeepSeek-V4-Flash, first-party API) over long tutoring sessions, once prompt caching and tool-output size are accounted for. F25 answered "no, `full_history` stays cheapest through 36 turns" β but its compaction arms confounded two mechanisms, the history policy and the giant raw tool outputs that dominate input tokens (F1), so it couldn't say* where *the money actually goes. Stage 1 is the clean decomposition: a 2Γ2 factorial crossing the history policy (keep everything vs summarize at a 200k trigger, retaining a 50k recent tail) with tool-output handling (raw vs a stable 10k-token cap applied once when the output enters history), giving four arms β `exp_fh_raw` / `exp_fh_cap10k` / `exp_c200_raw` / `exp_c200_cap10k` (`app/memory_presets.py`) β whose paired contrasts isolate each mechanism's cost/memory effect and their interaction. Stage 2 (trigger-threshold sensitivity: `exp_c400_cap10k` / `exp_c800_cap10k`) is built but not yet run. Method: battery `battery_sessions_v2` (22-/36-turn scripted tutoring sessions that plant student facts early and probe them late), 3 trials, run via the `evals.run_compaction_experiment` paired lockstep runner (all four arms advance turn-by-turn together in randomized within-turn order, each arm/session/trial under a distinct DeepSeek cache `user_id` so no arm warms another's KV cache; immutable run fingerprint; `evals/contributing.md` Β§2). "triggerfix" in the run name marks this as the corrected rerun after the earlier stage-1 attempt's summarization-trigger bug. Run `runs/deepseek_compaction_stage1_triggerfix_20260715`; full writeup `runs/deepseek_compaction_stage1_triggerfix_20260715/report/stage1_findings.md`; raw on HF `eval_runs/deepseek_compaction_stage1/`.* Four arms {`full_history`, summarize@200k keep-50k} Γ {raw tool outputs, stable 10k-token cap (40k bytes) applied once when the output enters history} on sessions-v2 Tier-1 contradiction (3Γ22t) + Tier-2 long-horizon (2Γ36t), Γ3 trials, run turn-level-lockstep with randomized within-turn order and a distinct DeepSeek cache `user_id` per arm/session/trial (no cross-arm KV warming); 414 turns/arm, 0 errors. The earlier stage-1 trigger bug is fixed and evidenced: every compaction fired at β₯200,056 pre-compaction tokens with 92kβ198k-token summary inputs, and the `full_history` arms compacted 0 times. Cost (mean $/trajectory): capping alone **$0.189 β $0.117 (β38%, cheaper in 14/15 paired trajectories)** with the cache-hit ratio intact (96.0%β95.9%) β it removes ~1MB of repeated tool payload per trajectory *without* rewriting the prefix, so it composes with caching instead of fighting it (the F1βF24/F25 thread: the tokens live in tool outputs, and cutting them *stably* is nearly free). Summarization is a net cost **add** on both sides: +50% vs raw full-history (14/15) and **+32% vs capped full-history (12/15)**; capping shrinks the compaction penalty (interaction β$0.057) but does not reverse it. Quality holds where it should: holistic 96β98% for all four arms, and the cap costs no memory (probe 87%, identical to raw full-history). Contrast with F19: per-call-view head/tail truncation churned (re-calls, 33% memory); this checkpoint-level insertion-time cap shows no churn at all (3.0 vs 3.1 LLM calls/turn, 87% memory) β though the arms differ in retained size and model too, the stability of what the model re-reads looks like the operative difference. Caveats: probes n=15/arm and only 5 session templates (the writeup's cluster-bootstrap intervals are descriptive, not population estimates); holistic is judge-graded, not human-validated (F34); DeepSeek's ~50Γ cache discount is load-bearing for the ranking (F25's provider caveat). | |
| - **F36 (2026-07-15) β Where compaction hurts memory, the damage tracks the *number of lossy rewrites*, not the size of the live context β and raw tool outputs double that number, so Axis-B bloat *amplifies* Axis-A damage. "Too many tokens in the answering call" is not the failure mode: `full_history` recalls perfectly from up to 879k-token contexts. (Refines F10/F26's mechanism.)** *Experiment: same DeepSeek stage-1 compaction factorial as F35 (run `runs/deepseek_compaction_stage1_triggerfix_20260715`, battery `battery_sessions_v2`; probe-level evidence in the same `stage1_findings.md` writeup) β this is the memory half of that run's results. The battery's two tiers each target one memory failure mode: Tier 2 long-horizon (2 sessions Γ 36 turns Γ 3 trials) plants student facts/constraints early and probes them ~30 turns later, after compaction has had every chance to evict them; Tier 1 contradiction (3 Γ 22t Γ 3) plants a fact, updates it mid-session, and probes whether the agent uses the current value. Probes judge-graded via the F14-validated blinded subagent path.* Long-horizon recall (plant facts early in a 36-turn session, probe at the end): the extended 2026-07-15 audit found one of the two Tier-2 sessions (`v2_t2_fullstack_long_persona_36t`) carries recycled "project pivot" filler that outright contradicts the probed PHI constraint, so its cells cannot separate eviction-by-compression from eviction-by-believed-pivot and are excluded (repaired in `battery_sessions_v2_1.jsonl`). On the clean Tier-2 session (`v2_t2_agentic_long_persona_36t`, whose off-persona filler is orthogonal to its probed governance facts; n=3/arm): **0 summarization events β 3/3 and 3/3** (both full-history arms), **~2β3 events β 3/3 and 3/3** (`exp_c200_cap10k` and F37's structured arm), **~5β6 events β 1/3** (`exp_c200_raw`). The dose mechanism survives in coarser form β only the high-dose raw arm measurably loses far-back facts, and low-dose capped compaction is indistinguishable from full history at this n β but the fine "6/6 β 5/6 β 3/6" gradient quoted before the audit mixed in the compromised session and should not be cited. The raw arm compacts twice as often because uncapped tool payloads regrow the history to the 200k trigger faster; each event then compresses ~150k tokens into ~1.7k (~90Γ), and on the clean session the judge's fail reason is genuine eviction ("explicitly denies knowing the student's constraints and offers only a generic list of common governance options"). Read-time context size points the *other* way: at probe time the compaction arms answered from 95β195k-token contexts and failed, while `full_history` answered from 363β879k and never missed β so the "context rot" here lives in the summarization pipeline (a planted fact that is a tiny fraction of a payload-dominated span must survive every ~90Γ rewrite), not in the answering model's attention. Separately, the contradiction tier: all 11 contradiction misses across all arms sit in **one session** (`python_colab_to_local_22t`), and the miss mode is *hedging* (recap both environments, ask "which machine are you on?") rather than stale-fact use β but the 2026-07-15 battery audit found that session **confounded**, so treat its cells as unmeasured rather than as a finding: its recycled filler injects a first-person *third* environment ("my iMac", turn 16) *after* the turn-7 update and *before* the turn-21 probe, which makes the hedge a defensible reading of genuinely contradictory history, not a demonstrable memory failure. The other two contradiction sessions are confound-free (audited: zero post-update first-person environment claims) and had **zero misses in any arm** β so this run shows no contradiction gap between arms and leaves F26 standing; whether `full_history` hedges when it truly holds A and Aβ² needs the repaired session rerun (fix as a new battery version β v1/v2 files are frozen). Caveats: the clean Tier-2 evidence is n=3/arm β read "high dose fails, low dose doesn't", nothing finer; Tier-1 does not guarantee eviction before its probe (all five of `exp_c200_cap10k`'s compaction-dormant trajectories are Tier-1, so in capped arms Tier-1 tests update-adherence, not contradiction-under-eviction); Tier-2's probed constraints are partially guessable from world knowledge (a governance rule for a procurement agent), which inflates every arm's absolute recall equally and biases *against* the dose signal. Aggregate probe accuracy restricted to audit-clean probes (9/arm): full-history arms **9/9 and 9/9**, capped compaction **9/9** (XML) / **9/9** (structured), raw compaction **7/9**. | |
| - **F37 (2026-07-15) β A chunk of "compaction's cost" was an implementation artifact, not physics: LangChain's stock summarizer breaks the provider cache on the summary-generation call itself, and a prefix-preserving compaction request removes that β summarizer cache hit 0%β94%, summarization cost β87%, total β14%/trajectory vs the stock arm, with equal-or-better memory. But the *post-compaction* prefix break is structural, so compaction still costs +14% vs capped full history: F35's ranking stands, with a smaller gap. (Same regime as F35/F36 β DeepSeek-V4-Flash first-party, absolute numbers not cross-comparable with the Gemini fleets.)** *Experiment: **DeepSeek compaction study, stage 1 follow-up β prefix-preserving ("structured") compaction.** The stock `SummarizationMiddleware` serializes the selected messages into a fresh XML prompt to ask for the summary, so even summary *generation* pays 100% cache-miss prices on ~150k-token inputs; `PrefixPreservingCompactionMiddleware` (`app/chat_service.py`, `summarization_strategy="structured_prefix"`) instead sends the *unchanged current request prefix* plus one checkpoint instruction, binding the same model/tools/settings, so the summarizer call rides the same cache as the agent. Single arm `exp_c200_cap10k_structured` β identical to `exp_c200_cap10k` except the strategy β run `runs/deepseek_structured_only_stage1_20260715` on the same `battery_sessions_v2` Γ 3 trials (414 turns, 0 errors), compared pairwise per sessionΓtrial against the four stage-1 arms; graded through the same blinded-subagent pipeline (429 judgments, 0 missing). Writeup `runs/deepseek_structured_only_stage1_20260715/report/structured_findings.md`; raw on HF `eval_runs/deepseek_compaction_stage1/`. A 2026-07-15 source-level check of the open-source Codex CLI confirmed this strategy is a faithful analog of Codex's local compaction β unchanged history + appended instruction, explicitly cache-preserving β while Codex's OpenAI-default *remote v2* path is the provider-native opaque-compaction primitive a client middleware cannot replicate; comparison + adoption ideas in `runs/deepseek_structured_only_stage1_20260715/report/codex_comparison.md`.* Cost: vs the stock XML arm, **β$0.0217/trajectory (β14.0%, cheaper in 11/15 pairs) at identical compaction dose** (1.27 events/trajectory both): the summarizer call's cache-hit ratio goes **0.0% β 94.2%**, so summarization cost falls $0.0276 β $0.0036/trajectory (β87%) β even though the structured arm *bills more* raw input (10.6M vs 9.9M tokens/trajectory; the checkpoint request re-reads the whole prefix at the ~50Γ cache-read discount). Tokens β dollars, again (F2/F25). vs capped full history it remains **+13.7% (cheaper in only 4/15)**: what's left is the structural boundary β installing the summary necessarily rewrites the *next* agent call's prefix, which no client-side middleware can avoid (that would take a provider-native compaction/continuation primitive) β plus the extra agent work (3.5 vs 3.1 LLM calls/turn). Memory: equal-best of the compaction arms β on audit-clean probes (excluding the two battery-compromised sessions, see F36) it scores **9/9**, tied with both full-history arms and capped-XML compaction, while raw compaction scores 7/9; the pre-audit "6/6 vs 5/6" long-horizon edge over the XML arm rode entirely on the compromised fullstack session and should not be cited β whether the intact-prefix summary *also* preserves facts better than the XML rewrite is unmeasured at this n. Holistic flat at 96%. Its single overall probe miss is a contradiction hedge in `python_colab_to_local_22t` β the confounded session β a battery artifact, not an arm signal. Trigger gate: compaction fired in 12/15 trajectories (19 events, min pre-compaction observation 200,053 tokens; the 3 dormant trajectories all Tier-1, and 100% of Tier-2 probes ran under compaction). Caveats: the memory edge is 1β2 probes at n=6 long-horizon / n=15 total β directional; the arm ran solo, not turn-level-interleaved with the stage-1 arms (token/cost metrics are run-time-independent, but treat its latency rows as confounded, F27); holistic judge not human-validated (F34); DeepSeek's ~50Γ cache discount is load-bearing (F25's provider caveat). | |
| - **F38 (2026-07-15) β The keep-everything-vs-compact crossover is set by one number: the cached-input price. Repricing our stage-1 traces call-by-call shows "full history wins" is a fact about DeepSeek's pricing, not a law: at GPT-5.6 Sol's public rates the *same recorded workload* already flips to favor compaction at 36 turns β quantitatively consistent with why OpenAI's Codex auto-compacts at ~272k while our tutor, on DeepSeek, should not compact at all. (Analysis-only finding β no new run; adds the missing fourth boundary to F32's list.)** *Experiment: same-trace repricing over the stage-1 arms (runs `runs/deepseek_compaction_stage1_triggerfix_20260715` + `runs/deepseek_structured_only_stage1_20260715`, F35/F37): every recorded model call's (fresh, cached, output) token triple re-costed at GPT-5.6 Sol's public API rates ($5 / $0.50 / $30 per M β a 10Γ cache discount vs DeepSeek's ~50Γ), with and without the API's >272k long-context cliff (2Γ input / 1.5Γ output on the whole request; per OpenAI's Tibo, the cliff is NOT charged on Codex subscriptions β the stated real driver is accumulated cache-read cost scaling with the live window re-read on every tool call).* Result: `exp_fh_cap10k` vs `exp_c200_cap10k_structured` goes from **$0.117 vs $0.134/trajectory (full history β13%) at DeepSeek prices** to **$9.85 vs $9.94 (βtied) at GPT rates** β and the tie hides a horizon split: full history still wins the 22-turn tier ($6.29 vs $6.76) but **already loses the 36-turn tier ($15.18 vs $14.72), no pricing cliff involved** (directional: 3/6 tier-2 pairs, mean β$0.46; the API cliff widens it to $12.56 vs $9.94 overall β 176/1,247 full-history calls exceed 272k by provider-reported input tokens, the count the dollar figure uses; an earlier draft said 191, which counted by the app's approximate context estimate). Cache reads are **60%** of full history's repriced spend (vs ~28% at DeepSeek prices) β exactly OpenAI's named mechanism, visible in our own traces. Closed form at this 22/36-turn horizon: full history wins iff cached input costs < **~$0.55/M** β DeepSeek charges $0.0028 (200Γ under the line), GPT-5.6 Sol $0.50 (right at it), so the crossover turn falls from "beyond everything we tested" to "inside ordinary sessions" purely by the price constant. (2026-07-16 refinement: $0.55/M is a mix-specific aggregate over the 9-short/6-long trajectory mix β the 22-turn subset alone ties at ~$2.91/M and the 36-turn subset at ~$0.39/M; full sensitivity analysis, token receipts, and the GPT-5.6 cache-write caveat in `runs/deepseek_structured_only_stage1_20260715/report/gpt56_repricing_counterfactual.md`.) Reading: F9/F15/F25/F35's "keep everything wins" carries an implicit precondition F32's list omitted β **a deep-enough cache discount** β the fourth boundary alongside F32's three (static doc, small window, no caching). It rationalizes Codex's design (compact once cache-read accumulation crosses the compaction tax; not earlier than quality forces, since compaction damage is dose-responsive β F36 β matching Codex's own repeated-compaction accuracy warning; and don't buy bigger windows for quality β flat above 272k per OpenAI, consistent with F6/F26/F35 where 880k-token full-history contexts stayed perfectly accurate) but does **not** derive the specific 272k, which is OpenAI-specific (GPT-5.5-lineage standard-context serving boundary; the API cliff sits exactly there; a Codex 372k trial was reverted for usage cost). Caveats: same-trajectory repricing assumes GPT-5.6 would reproduce DeepSeek's call pattern and ~96% cache-hit rate (imperfect caching hurts full history *more* β F25's routing note β so the real crossover is likely earlier, not later); excludes cache-write billing on both sides; subscription "usage" is OpenAI-internal accounting, not API dollars; trigger sensitivity is unrun (`exp_c400_cap10k`/`exp_c800_cap10k` built, pending), so the optimal-trigger curve is unknown even on DeepSeek. | |
| **Screen winners β promotion candidates:** `incontext_history_retrieval` (needs the long-session test), `retrieval_budget_30k` and `clear_retrieval_kb` (cheap Axis-B wins; clear_retrieval_kb is cheaper than prod with similar memory + better recall/key-points), anchored by `full_history`. **Drop:** `context_reset`, `observation_truncation`, `selective_retention`. | |
| **Part C β what remains (beyond the workshop).** The screen answered the core question; everything below *hardens* or *extends* it and is **gated on a spend or data decision from Omar** β none is required for the workshop. | |
| 1. **Promotion (rigor, not new findings).** Re-run the 2-3 winners (+ 1-2 combos) on the full batteries (60 single-turn + 32 sessions + 30 replay) Γ 3 trials, with paired statistical tests and per-variant failure-taxonomy diffs, to turn the screen's coarse n=12 rankings into confident numbers; optionally an Anthropic (Haiku) re-run of the prefix-rewriting arms (`prompt_compression`, `context_reset`, `aggressive`) to test whether explicit caching flips F2's cost ranking. **~$1,300-1,800 at corrected pricing** (the old "$300-400" was at pre-correction prices; Γ the 4.4 fix β estimate precisely before running, since the per-turn cost is now ~$0.25). | |
| 2. **Finish the v2 contradiction tier β the highest-value new dataset (already built and partially run; see the V2 execution note below).** Moderate sessions: plant A, update to Aβ², then probe after prod compaction has evicted A from the kept-recent window. This is the one correctness regime where `full_history` itself might lose (it sees both A and Aβ² and must choose the current one) β yet the partial run so far shows the opposite (`full_history` contradiction **2/2** vs `prod` **0/3**), so that regime has not actually appeared. What remains is finishing Tier 1 (the missing `full_history` session, then Tier-1 `profile_memory` and optionally `incontext_history_retrieval`) and adding trials, not building the dataset. **(2026-06-19 β now run in full on DeepSeek-V4-Flash rather than Gemini: F26 confirms `full_history` 3/3 contradiction vs `prod` 0/3, so the regime did not appear on a second model either; the Gemini Tier-1 finish is now optional cross-model confirmation.)** | |
| 3. **The `incontext_history_retrieval` long-session cost test β principled but less product-critical.** The screen showed `incontext` works (83%, F16) but couldn't show a cost benefit on β€13-turn sessions. Whether retrieve-old-turns beats `full_history` *once sessions get genuinely long* needs a token-calibrated long-session v2 tier; useful for the workshop argument, but gated by whether that regime matters in real tutor telemetry. **(2026-06-19 β run on DeepSeek (F25) at 36 turns: `full_history` stays cheapest and `incontext` works but does not beat it on cost under first-party caching; whether the answer flips on a 10Γ cache provider like Gemini/Anthropic is still open.)** | |
| 4. **Conditional builds.** Build `temporal_graph_memory` only if `full_history` fails contradictions (F26: it did **not** on DeepSeek β 3/3 β so this stays unbuilt), `delta_summarization` only if the long-horizon cost tier shows room to beat full history (F25: `full_history` still cheapest at 36t on DeepSeek, so no room there β though a 10Γ cache provider is untested), and `entity_memory` only if multi-project probes show fact bleed. `collection_memory` is a separate cheap v1/product follow-up for verbatim fact storage (F22 tested the current profile write-back, not that). `hierarchical_summarization` and `sleeptime_consolidation` are deferred. The step-by-step v2 plan is in `evals/part_c_plan.md` β "Battery v2". | |
| **Deliberately not pursued.** *Stretch (heavier, likely later):* sub-agent isolation, reframed as a tool-output-token play β keep noisy `run_kb_command` output out of the main context (ties to F1); temporal-graph memory is already the conditional build in item 4 above. *Dropped as low-information given the findings:* the skills / lazy-prompt-loading family (only a ~458-token win β see the talk-outline correction below), procedural memory, GraphRAG (now **empirically tested** rather than dropped on faith β **F28**: ties classical RAG on grounding, costs +44% $/turn, no payoff on single-hop Q&A), and multi-agent / parallel-research agents. Context for the `kb_off` arm (F17/F23): the KB is the dominant grounding tool β β89% of Part B turns, ~7.7 KB vs ~0.9 retrieval calls/turn over 1,088 turns β so disabling it forces the 100k-token retrieval fallback, which is why it raised cost/latency. | |
| **V2 execution note (2026-06-16).** A private/gitignored 6-session `battery_sessions_v2.jsonl` was built and partially run. The full v2 screen is too slow under the lower-tier Gemini key: `full_history` at concurrency 3 hit the 3M input-tokens/minute quota, and concurrency 1 works but makes the full 4-arm screen a multi-hour job. Current local artifacts (`runs/e2_v2_partial_report/report.md`) are directional, not final: `prod` completed all 6 sessions with 0 errors, all probes under compaction, **14% probe accuracy (1/7)** and contradiction **0/3**; `full_history` completed 2 Tier-1 contradiction sessions with 0 errors, no compaction, contradiction **2/2**. Priority is now **Tier 1 only**: finish the missing `full_history` contradiction session, then run Tier-1 `profile_memory` and optionally Tier-1 `incontext_history_retrieval`. **(2026-07-15: superseded β the stage-1 factorial ran the full battery Γ3 trials on DeepSeek (F35βF37), and the contradiction cells here sit on the session the audit later found confounded; new runs must use `battery_sessions_v2_1.jsonl`.)** | |
| Harness corrections (bugs in our measurement, not findings): the overnight 06-12 stall was machine sleep hanging API streams (all four pipelines stopped the same minute; fixed with the per-turn timeout); battery lesson-URLs carried a `/discussions/` suffix that silently zeroed recall@lesson until normalized (caught by the A3 smoke, re-graded from bundles without re-running). | |
| - **Pricing correction (2026-06-13).** `MODEL_PRICING` under-priced `gemini-3.5-flash` before this date (~$0.30/$2.50 per MTok vs the correct **$1.50 input / $9.00 output / $0.15 cache-read**, verified against Google's price sheet). Every dollar figure generated earlier β the Part B table and the inline costs in findings **F2/F8/F9/F12** β is **~4.4Γ too low**; token counts are unaffected. Part B bundles were re-costed from their saved token counts (2026-06-15) and `runs/b_report/report.md` regenerated at correct pricing: **relative rankings are unchanged** (F9 still has `full_history` cheapest, $0.11/turn sessions), only absolute dollars move. Real Part A+B spend β **$338**, not ~$73; the comprehensive program β **$590** across all Gemini 3.5 Flash runs (Part B ~$323 Β· Part C ~$163 Β· rest ~$103). Use the regenerated report for absolute costs; finding dollars above are pre-correction. | |
| - **Stable `sheet_row_id` (2026-06-15).** `evals.grade` built `sheet_row_id` with builtin `hash()`, which is salted per process, so regenerating a `handgrade_sheet.csv` produced ids that no longer matched the frozen workbook keymap β silently emptying `handgrade_workbook merge`. Switched to a `hashlib.md5` hash; the Part B re-merge was rebuilt via the deterministic `run_id` to recover the human grades. | |
| - **`cleared_tool_outputs` is blind to library clearing (2026-06-19).** `ContextEditingMiddleware` edits a per-call `deepcopy` and never persists the placeholder to the checkpoint that `context_window_stats` reads, so the signal is 0 for all presets whether or not `ClearToolUsesEdit` fired (full detail + token evidence in **F24**). Detect clearing via input-token deltas; the custom Part C view-mechanisms avoid this by reporting through the `app.telemetry` turn-signal registry. | |
| - **Faithfulness parked; bundle tool-output cap raised 6kβ40k (2026-06-23).** The new **faithfulness** rubric (groundedness vs retrieved evidence, F33) is *not reportable on existing runs*: `evals/run_battery.py` capped each saved tool output at `TOOL_OUTPUT_MAX_CHARS = 6_000` while the KB shell returns up to **40k** (`app/kb_shell.py` `DEFAULT_MAX_OUTPUT_CHARS`), so on KB-heavy arms the judge saw a fraction of what the agent actually grounded on and scored *capture-completeness, not grounding* β faithfulness tracked the KB:retrieval ratio almost monotonically (`b_st_prod` kb:ret 10.3 β **4%**; `c_st_prod` 7.7 β 25%; `kb_off`/`graphrag` kb:ret 0 β 50β76%), the same blind-signal class as **F23/F24**. Raised the cap to **40k** (captures what the agent saw; `output_chars` still records the true length so any overflow stays visible) so *future* re-records make faithfulness measurable; existing bundles cannot be backfilled (content past 6k is gone β needs a re-run). Until then `evals/grade.py` emits faithfulness **only when evidence was captured in full** (`evidence_is_complete`, self-re-enabling after the cap raise), and the column is dropped downstream (`grading_merge.DROP_ITEM_TYPES`; `report.py` omits it directly). Holistic (F33/F34) is unaffected β it grades the answer, not the evidence. | |
| - **Report mislabeled judge grades as human (2026-06-19, fixed).** `evals.report` hardcoded "(human grade)" on the behavior / key-point / probe-accuracy rows, but Part C runs were filled by the LLM judge (`judge_filled.csv`), not a person β so `runs/c_report/report.md` asserted a human label over judge numbers. Fix: `evals.grade` now records a `grade_source` per merged grade (the judge prefixes its note `[judgeβ¦]`; humans do not), and `evals.report` tags each column `[judge]`/`[human]`/`[mixed]` and renames the rows to "(graded)". Numbers are unchanged (and judgeβhuman per F14); only the provenance label is now honest. Same label-vs-measurement class as F23/F24. All Part C arms were re-graded and `c_report` regenerated. | |
| - **`input tok/turn` is billed-across-calls, not context size (2026-06-19, clarified).** Telemetry sums `usage_metadata` over every internal model call in a turn, so a high-loop arm (`observation_truncation`, 22.8 calls) bills lots of input from small contexts while a big-payload arm (`kb_off`) bills similar totals from few large ones. Correct for *cost*, but it is not "how big is the context window." The report now shows both: **input tok/turn (billed, all calls)** and a new **context tokens/turn (window size)** row (from `context_tokens_approx`, always computed but previously unsurfaced). On single-turn the two even rank arms differently (e.g. `kb_off` has the *largest* window at 73k from its 100k retrieval payloads, but mid-pack billed input). Read F1/F6's "~200k input" as billed volume, not window size. | |
| - **Latency rows are confounded by time-of-run (caveat, not yet fixed).** `run_part_c_screen.sh` runs arms sequentially, and TTFT/total-ms depend on API load at that moment, so cross-arm latency deltas (e.g. F19's 430s p95, F23's 19s-vs-38s) partly reflect *when* each arm ran. Token/cost/recall are run-time-independent and unaffected. Any latency claim that drives a decision should come from arms run interleaved (or repeated trials), not a single sequential screen. **`total β ttft` does *not* remove this** β it is only the answer-streaming tail (~6s, preset-insensitive); the agentic latency incl. the summarization tax is all in `ttft`, so subtracting it deletes the signal. Report the unconfounded work metrics (LLM calls/turn, tokens/turn) as the headline and treat seconds as supporting. See **F27**. | |
| - **Latency de-confounding tooling (2026-06-22).** Acting on F27: (1) `runs/run_part_c_screen_interleaved.sh` runs the screen question-outer / arm-inner (writes `runs/ci_*`, not `runs/c_*`) so every question's arms share one API-load window β same total runs as the sequential screen, so interleaving is free. (2) `context_stats` now also emits `time_to_first_token_ms` (first reasoning/tool/text token, always β€ `ttft_ms`), separating raw model responsiveness from the tool-call loop that `ttft_ms` includes; it flows through `evals/grade.py` to the report and shows `β` on runs recorded before it existed. (3) `evals/report.py` now leads with the run-time-independent work metrics (LLM/tool calls/turn, tokens/turn) and tags every latency row `[confounded]`, with a runtime note. All `runs/*report*` were regenerated (layout/labels only; numbers unchanged). The scripts live in gitignored `runs/` like the rest of the run tooling. | |
| - **Mixed-judge grading confound (2026-07-15, fixed).** While grading the DeepSeek stage-1 compaction factorial (F35/F36, `runs/deepseek_compaction_stage1_triggerfix_20260715`), the first quality merge silently mixed two judges: an interrupted Codex (`gpt-5.6-sol`) subscription pass (`evals.run_subscription_grading`) had graded chunks 000β011 of the grading plan (`runs/_grading/deepseek_compaction_stage1_full_20260715`) β almost exactly the `exp_c200_cap10k` rows β before the blinded subagent workflow (`evals/grade_workflow.js`) graded the rest. The GPT judge is one-directionally harsher on holistic: on the 462 double-graded rows, 67% agreement, with GPT failing 151 rows the subagent judge passed and zero the reverse (fail rates 35% vs 3%; its extra fails are mostly technical-accuracy critiques the blind rubric can't adjudicate without evidence). That made `exp_c200_cap10k`'s holistic read 63% vs 96β98% for the other arms β a grader artifact, not an arm effect (the same rows re-graded by the subagent judge score 98%; the tell was the straddling chunk, where the *same arm* scored 76% under GPT and 97% under the subagent judge). Fix: chunks 000β011 re-graded so all 44 chunks share the F14-validated subagent judge; the Codex verdicts are preserved for comparison (`runs/_grading/deepseek_compaction_stage1_full_20260715/codex_partial/`). Lessons: **one judge per battery** β grade provenance is part of the measurement (same class as the judge-as-human label bug above); and two frontier judges disagreeing this much, one-directionally, on holistic is further evidence that holistic needs the held-out human validation F34 called for. | |
| - **`battery_sessions_v2` filler recycling is systematically confounded (2026-07-15, two audit passes; repaired as `battery_sessions_v2_1.jsonl`).** v2's filler was recycled verbatim from v1 turns, and the recycling blindly included v1 *persona/fact-update/project-pivot* turns β each a first-person state claim β dropped into a different persona: 7 such turns across 4 of the 6 sessions. Two sessions are unusable as authored: `v2_t1_python_colab_to_local_22t` (a "my iMac" turn injects a third environment between the update and the open-ended contradiction probe β every stage-1 contradiction miss in every arm sits here) and `v2_t2_fullstack_long_persona_36t` (two "project pivot" turns relabel the student's project to hosted architectures that contradict the probed no-PHI constraint, so its long-horizon cells measure the injected contradiction, not memory); two more are muddied (`v2_t2_agentic_long_persona_36t`, `v2_t3_fullstack_two_projects_20t` β injected claims orthogonal to the probed facts). F36/F37 were corrected to audit-clean probe subsets. The same audit cleared everything else: `battery_sessions_v1` (all 32 sessions β no confounds; the update/pivot turns there are the *intended* probed facts), `battery_singleturn_v1`, and `replay_n1_v1` all clean and matching their README counts. Separate defect: 5 of `battery_personas_v1`'s authored anti-pattern checks are negation-insensitive substrings that false-fail correct answers (e.g. anti `import openai` matches the official Node SDK's `import OpenAI from 'openai'`; anti `install Python` matches "you don't need to install Python") β an unquantified downward bias on Part B's persona auto-check rows (F8/F11); repair queued. Fixes shipped: `battery_sessions_v2_1.jsonl` replaces the 7 recycled turns with neutral same-course corpus questions (probes/plants byte-identical), and the runners now **refuse `battery_sessions_v2.jsonl`** for new runs (`--allow-deprecated-battery` reproduces historical runs only). Standing lesson: recycled real-student filler must be scanned for first-person state claims against the planted facts before a battery freezes. | |
| - **Pattern worth naming β most of our measurement bugs are telemetry/label blind-spots, not wrong math.** Six now share one shape: a number or label asserted a property the instrumentation could not observe for that config β **F21** (mechanism dormant: `profile_memory` never engaged on sessions), **F23** (recall blind to KB browsing), **F24** (clearing invisible to the checkpoint signal), the judge-as-human label above, **faithfulness graded on truncated evidence** (the 6k cap above β the judge scored what it could *see*, not what the agent *saw*), and **F34's holistic-on-session blind spot** (the quality judge sees no conversation history, so it credits answers that silently ignored the student's earlier-stated facts β "looks good" on 100% of memory-failed turns). `check_triggers` enforces "did the mechanism fire?" for *compaction* only; nothing yet does for retrieval/profile/KB engagement, grade provenance, or whether the *grader* saw what it needed. Standing guard for any new arm or metric: before trusting a per-arm number, confirm the signal/label/grader reflects what the config actually did. | |
| Talk-outline correction: the planned "skills / just-in-time KB instructions" change (Change 4) was never built, and its premise is off β the KB instructions block measures **458 tokens** (not ~4k) inside a 1,655-token system prompt; max saving ~2% of a turn, mostly cache-discounted. Recommended reframe: demo progressive disclosure on tool *outputs* (F1), where our numbers actually are. | |
| ## Remaining work | |
| **Omar** | |
| - [x] Grade the session probes (the critical tier). 2026-06-15: all 96 reviewed β the 83 high-confidence LLM verdicts by sign-off, the 13 low-confidence ones individually β with **zero overrides**, so the provisional LLM grades are now human-confirmed ground truth (see F13). The non-probe rows (persona / key_point / behavior in `review_other.md`) were not part of this pass and remain LLM-graded; bless them too if those columns need a human label. | |
| - [ ] Audit `data/eval/review_log_v1.md` (the 11 flagged judgment calls first). | |
| - [x] Verify and update the `gemini-3.5-flash` entry in `app/telemetry.py:MODEL_PRICING` against the current price sheet. | |
| **Workshop** | |
| - [ ] Replace placeholder slide numbers with measured ones; pick 2β3 failure traces from the graded probes for the failureβfixβnumber beats (a `fact_update` failure under summarization is the target demo). | |
| - [ ] Token-vs-turn plot: CSV is ready; add matplotlib (or plot elsewhere) for the PNG. | |
| - [ ] Next.js meter component over the `data-context-stats` SSE part, if the live demo needs it. | |
| - [ ] Skills section: reframe per the 458-token measurement above. | |
| **Beyond the workshop.** Part C itself is done and reported above (findings F14βF24); what remains only *hardens* or *extends* it and is consolidated in "Part C β what remains" above β promotion, finishing the v2 contradiction tier, the `incontext_history_retrieval` long-session test, the conditional subsystem builds, and the deliberately-not-pursued list. The full execution detail (variant wiring, telemetry signals, the subset run matrix, the cost model, and the v2 battery plan) lives in `evals/part_c_plan.md`. | |
| **Product quality track (post-workshop)** | |
| Golden cases as a CI gate (5β10 critical-path cases with deterministic assertions, including the F7 non-empty check) β error-analysis cycles (an expert reads 50β100 traces in a small viewer, records pass/fail plus the first failure reason, and stops when new failure types stop appearing) β failure taxonomy β code assertions for recurring modes β LLM judges built from the hand-grade critiques and validated to >90% true-positive and true-negative rates on held-out human labels before any judge-graded number is reported β nightly battery + weekly trace sampling. Full methodology behind each step: `evals/background.md`. | |
| ## Files | |
| - `evals.md` β this file: the what, the data, the results, the queue. | |
| - `evals/contributing.md` β contributor playbook: how to run an experiment and land it (run β upload data to HF β record an `F<N>` finding β merge), with a definition-of-done gate. | |
| - `evals/part_c_plan.md` β Part C execution plan: orchestration (workflow vs direct), the variant catalog with exact wiring + telemetry signals, the subset run matrix, and cost. | |
| - `evals/background.md` β research sources (Hamel, howtoeval, OpenAI macro-evals) and design rationale. | |
| - `evals/graphrag.md` β GraphRAG vs classical RAG experiment (branch `experiment/graphrag-vs-rag`): a scoped, true-GraphRAG head-to-head on the single-turn battery. Revisits the dropped GraphRAG idea with a fair test. | |
| - `evals/compaction.md` β compaction-vs-keep-everything study (branch `experiment/context-compaction`): on the largest lesson, keep-all vs every compaction method vs retrieve-per-question. **Standalone fleet on Gemini 2.5 Flash β not comparable to the 3.5-Flash results above.** | |
| - `evals/slm_compaction.md` β knowledge-compaction on small local models (branch `experiment/slm-compaction`): two experiments on 3 Ollama SLMs (llama3.1:8b, qwen2.5:7b, qwen3:8b) at a 32k window where the lesson overflows. **Axis B** (fit a document: RAG vs stuff vs summarize, F30) and **Axis A** (compact growing history via the real memory presets, F31). The regime where compaction is *forced*; RAG wins Axis B, no universal winner on Axis A, and "keep everything" never wins on an SLM (synthesis in F32). | |
| - `data/eval/README.md` β battery schemas + glossary of terms; `review_log_v1.md` β dataset audit trail. | |
| - `evals/` β the harness code; `app/memory_presets.py`, `app/telemetry.py` β the app-side hooks. | |
| - `runs/b_report/`, `runs/c_report/` β generated tables, token curves, blinded grading workbook (Part B and the Part C screen). Re-grade / follow-up reports: `runs/kbfair_report/` (F23 KB-fair recall), `runs/d_report_profile_memory_active/` (F22 profile rerun), `runs/e2_v2_partial_report/` (v2 partial). | |
| - `runs/ds_v2_t1_report/`, `runs/ds_v2_t2_report/` β **DeepSeek-V4-Flash** v2 matrix (F25/F26: cost + contradiction + long-horizon), graded via the subagent path (`runs/ds_judge/`). Shared on HF as `eval_runs/part_f_deepseek_v2/` (the standard `mv eval_runs/part_*/* runs/` download restores them). Provider added in `app/chat_service.py` (`deepseek` / `openrouter`), pricing in `app/telemetry.py`. | |
| - `runs/deepseek_compaction_stage1_triggerfix_20260715/report/`, `runs/deepseek_structured_only_stage1_20260715/report/` β **DeepSeek long-context compaction study, stage 1** (F35/F36): the 2Γ2 {`full_history`, summarize@200k} Γ {raw, stable 10k tool-output cap} factorial, plus the follow-up **prefix-preserving compaction** arm (F37: `exp_c200_cap10k_structured`, `summarization_strategy="structured_prefix"` in `app/chat_service.py` β the summarizer call rides the provider cache instead of LangChain's cache-missing XML re-serialization; β14%/trajectory vs the XML arm). Writeups: `stage1_findings.md` / `structured_findings.md` in each report dir (incl. `audit_20260715.md`, the validity audit); runner `evals.run_compaction_experiment`; graded via the blinded subagent path into `runs/_grading/deepseek_*_20260715/`. Shared on HF as `eval_runs/deepseek_compaction_stage1/` (runs + `_grading/` judge verdicts; restore line in the download snippet above). `battery_sessions_v2.jsonl` is **deprecated for new runs** (python_colab confound β see harness corrections); the runners refuse it, use the repaired `battery_sessions_v2_1.jsonl`. | |
| - `runs/axisa_*_report/` β **SLM compaction** study (F30/F31/F32, branch `experiment/slm-compaction`): Axis B fit-a-document and Axis A compact-history on 3 Ollama SLMs at a 32k window. Shared on HF under `eval_runs/slm_compaction/` (`axis_b/<model>/`, `axis_a/axisa_<model>_<preset>/`) + batteries `eval/compaction/`; the snippet above restores Axis-A to `runs/axisa_*`, Axis-B stays under `eval_runs/slm_compaction/axis_b/` β full restore command in `evals/slm_compaction.md`. Provider `ollama` added in `app/chat_service.py`. | |