omarsol Claude Fable 5 commited on
Commit
a213ef5
Β·
1 Parent(s): 766df7d

Fix stale claims across repo docs

Browse files

Doc audit against current code found drift; align the docs with reality:

- scraping README: course setup must also add COURSE_SOURCE_KEYS (missing
it misclassifies the course as docs and fails CI); swap course workflow
steps 1/2 to match execution order; GITHUB_TOKEN is docs-workflow only;
document GEMINI_CONTEXT_MODEL and GEMINI_CONTEXT_TPM_WINDOW_SECONDS
- AGENTS.md: the two custom tools are not "always" exposed (empty source
selection removes both; evals' disable_kb keeps retrieval only); quote
the prod deploy workflow's full name (also in README.md)
- evals/graphrag.md: run_battery's no-flag default is the app default
model (now DeepSeek V4 Flash), not Gemini 3.5 Flash
- evals/compaction.md: correct the run_compaction_study default-PRESETS
note (adds summarization_only, omits context_reset + the last two)
- data/eval/README.md: compaction scripts default to data/compaction/,
not data/eval/compaction/; document the override
- evals/part_c_plan.md: drop drifted line references; correct the probe
rubric wiring note and as-built context_stats signal cells
- manual_e2e_langsmith.md: langsmith CLI is a standalone binary, not a
uv dependency; note the enabledTools discussion does not bite for the
doc's own DeepSeek payload
- url_matching_procedure.md: Chrome scripting tool is javascript_tool

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

AGENTS.md CHANGED
@@ -34,7 +34,7 @@ ChromaDB for vectors; Cohere for embeddings/rerank; chat model is provider-confi
34
 
35
  ## Architecture in brief
36
 
37
- The agent is built with `langchain.agents.create_agent()` (LangGraph), an `InMemorySaver` checkpointer keyed by `thread_id`, and middlewares assembled from the selected **memory preset** (`app/memory_presets.py`: context-editing, summarization, optional long-term student-profile memory) plus source preference. The normal frontend does not select a memory preset: the server's model-aware production policy uses `prod_v2` for DeepSeek/Gemini and the immutable legacy `prod` preset for other providers. Direct API clients and evals can override this with `memoryPreset`; the full precedence is request value β†’ `AI_TUTOR_MEMORY_PRESET` environment override β†’ model-aware production policy. `stream_chat()` is the single entry point the API calls; it yields typed `ChatEvent`s that `app/api.py` encodes into the AI SDK UI-message stream, ending each turn with a `context_stats` telemetry event (tokens incl. cache buckets, est. cost, TTFT, compaction-trigger counts β€” independent of LangSmith). It always exposes two custom tools, plus provider-native web tools when enabled:
38
 
39
  - **`retrieve_tutor_context(query)`** β€” hybrid RAG over the corpus, scoped to the user's selected sources.
40
  - **`run_kb_command(...)`** β€” read-only KB file browsing (see below).
@@ -104,7 +104,7 @@ Chat runtime: `COHERE_API_KEY` (retrieval; always required), `DEEPSEEK_API_KEY`
104
  Both HF Spaces run the same image (`Dockerfile`: FastAPI + Next.js static export, `ripgrep` installed for `run_kb_command`, port :7860), in a dev β†’ prod flow:
105
 
106
  - **Dev β€” `ai-tutor`** (private): `.github/workflows/sync-to-hf.yml` force-pushes on every push to `main` (docs/markdown-only and scraping-script-only pushes are skipped via `paths-ignore`). Verify changes here first.
107
- - **Prod β€” `ai-tutor-chatbot`** (public): `.github/workflows/deploy-prod-to-hf.yml`, **manual trigger only** (Actions tab β†’ "Deploy prod to Hugging Face" β†’ Run workflow).
108
 
109
  Both Spaces need the same runtime secrets (`COHERE_API_KEY`, model provider key, `HF_TOKEN`, optional `LANGSMITH_*`) configured in their HF settings.
110
 
 
34
 
35
  ## Architecture in brief
36
 
37
+ The agent is built with `langchain.agents.create_agent()` (LangGraph), an `InMemorySaver` checkpointer keyed by `thread_id`, and middlewares assembled from the selected **memory preset** (`app/memory_presets.py`: context-editing, summarization, optional long-term student-profile memory) plus source preference. The normal frontend does not select a memory preset: the server's model-aware production policy uses `prod_v2` for DeepSeek/Gemini and the immutable legacy `prod` preset for other providers. Direct API clients and evals can override this with `memoryPreset`; the full precedence is request value β†’ `AI_TUTOR_MEMORY_PRESET` environment override β†’ model-aware production policy. `stream_chat()` is the single entry point the API calls; it yields typed `ChatEvent`s that `app/api.py` encodes into the AI SDK UI-message stream, ending each turn with a `context_stats` telemetry event (tokens incl. cache buckets, est. cost, TTFT, compaction-trigger counts β€” independent of LangSmith). It exposes two custom tools (deselecting every source removes both for that turn, and the evals' `disable_kb` flag keeps retrieval only), plus provider-native web tools when enabled:
38
 
39
  - **`retrieve_tutor_context(query)`** β€” hybrid RAG over the corpus, scoped to the user's selected sources.
40
  - **`run_kb_command(...)`** β€” read-only KB file browsing (see below).
 
104
  Both HF Spaces run the same image (`Dockerfile`: FastAPI + Next.js static export, `ripgrep` installed for `run_kb_command`, port :7860), in a dev β†’ prod flow:
105
 
106
  - **Dev β€” `ai-tutor`** (private): `.github/workflows/sync-to-hf.yml` force-pushes on every push to `main` (docs/markdown-only and scraping-script-only pushes are skipped via `paths-ignore`). Verify changes here first.
107
+ - **Prod β€” `ai-tutor-chatbot`** (public): `.github/workflows/deploy-prod-to-hf.yml`, **manual trigger only** (Actions tab β†’ "Deploy prod to Hugging Face (ai-tutor-chatbot)" β†’ Run workflow).
108
 
109
  Both Spaces need the same runtime secrets (`COHERE_API_KEY`, model provider key, `HF_TOKEN`, optional `LANGSMITH_*`) configured in their HF settings.
110
 
README.md CHANGED
@@ -16,7 +16,7 @@ Built by [Louis-FranΓ§ois Bouchard](https://www.linkedin.com/in/whats-ai/) ([X](
16
 
17
  The live app is deployed on Hugging Face Spaces at: [AI Tutor Chatbot on Hugging Face](https://huggingface.co/spaces/towardsai-tutors/ai-tutor-chatbot) (prod).
18
 
19
- **Deployment flow:** every push to `main` (except docs/markdown-only and scraping-script-only changes) auto-deploys to the private dev Space ([ai-tutor](https://huggingface.co/spaces/towardsai-tutors/ai-tutor)) for verification; the prod Space is promoted manually via the "Deploy prod to Hugging Face" workflow in the Actions tab.
20
 
21
  ### Workshop, slides, and going deeper
22
 
 
16
 
17
  The live app is deployed on Hugging Face Spaces at: [AI Tutor Chatbot on Hugging Face](https://huggingface.co/spaces/towardsai-tutors/ai-tutor-chatbot) (prod).
18
 
19
+ **Deployment flow:** every push to `main` (except docs/markdown-only and scraping-script-only changes) auto-deploys to the private dev Space ([ai-tutor](https://huggingface.co/spaces/towardsai-tutors/ai-tutor)) for verification; the prod Space is promoted manually via the "Deploy prod to Hugging Face (ai-tutor-chatbot)" workflow in the Actions tab.
20
 
21
  ### Workshop, slides, and going deeper
22
 
data/eval/README.md CHANGED
@@ -1,6 +1,6 @@
1
  # Evaluation batteries (v1, frozen 2026-06-11)
2
 
3
- The datasets for evaluating the AI tutor (see `evals.md` at the repo root for the overall effort and results). **This README is the single reference for what every file, field, and term means.** Files added after the v1 freeze: `battery_sessions_v2.jsonl` (deprecated for new runs β€” see the v2 note in the probe-type glossary), `battery_sessions_v2_1.jsonl` (its repaired successor), and `compaction/` (the F29/F30 lesson-compaction batteries, documented in `evals/compaction.md` / `evals/slm_compaction.md`).
4
 
5
  Built from `data/academy_discussion_eval.jsonl` β€” real student posts from the academy discussion boards with real staff answers β€” plus authored content (sessions, personas).
6
 
 
1
  # Evaluation batteries (v1, frozen 2026-06-11)
2
 
3
+ The datasets for evaluating the AI tutor (see `evals.md` at the repo root for the overall effort and results). **This README is the single reference for what every file, field, and term means.** Files added after the v1 freeze: `battery_sessions_v2.jsonl` (deprecated for new runs β€” see the v2 note in the probe-type glossary), `battery_sessions_v2_1.jsonl` (its repaired successor), and `compaction/` (the F29/F30 lesson-compaction batteries, documented in `evals/compaction.md` / `evals/slm_compaction.md`; note the compaction scripts default to `data/compaction/`, not this folder β€” either regenerate there with `compaction_study build` or point `BATTERY=`/`QUESTIONS_FILE=` at these copies).
4
 
5
  Built from `data/academy_discussion_eval.jsonl` β€” real student posts from the academy discussion boards with real staff answers β€” plus authored content (sessions, personas).
6
 
data/scraping_scripts/README.md CHANGED
@@ -7,7 +7,7 @@ Make sure you have the required environment variables set:
7
  - `GEMINI_API_KEY` or `GOOGLE_API_KEY` for context generation with Gemini
8
  - `COHERE_API_KEY` for embeddings
9
  - `HF_TOKEN` for HuggingFace uploads and downloads - [access to the private HuggingFace dataset repo](https://huggingface.co/datasets/towardsai-tutors/ai-tutor-data/tree/main)
10
- - `GITHUB_TOKEN` for accessing files via the GitHub API
11
 
12
  Optional Gemini context-generation tuning:
13
 
@@ -15,6 +15,8 @@ Optional Gemini context-generation tuning:
15
  - `GEMINI_CONTEXT_TPM_SAFETY_MARGIN` - fraction of that quota to use before pausing (defaults to `0.8`)
16
  - `GEMINI_CONTEXT_CONCURRENCY` - max concurrent context requests (defaults to `50`)
17
  - `GEMINI_CONTEXT_RETRY_ATTEMPTS` - max tenacity attempts for transient Gemini API errors (defaults to `8`)
 
 
18
 
19
  ## 1. Prepare the course data
20
 
@@ -50,7 +52,11 @@ Optional Gemini context-generation tuning:
50
  8. Open `data/scraping_scripts/source_registry.py`.
51
 
52
  9. Add the new course to the `SOURCE_CONFIGS` dictionary. Sources listed in
53
- this registry are active in the knowledge base.
 
 
 
 
54
 
55
  example:
56
 
@@ -89,8 +95,8 @@ uv run -m data.scraping_scripts.add_course_workflow --courses master_ai_for_work
89
 
90
  This script will guide you through the complete process, it will:
91
 
92
- 1. Extract the markdown content from each of the lessons and create a new JSONL file for the course
93
- 2. Download the JSONL files from the other courses
94
  3. Prompt you to manually add URLs to the course content, inside the newly created JSONL file (more details in step 3 below)
95
  4. Merge the course data into the main dataset
96
  5. Rebuild the KB artifacts (raw/generated/wiki; skip with `--skip-kb`)
@@ -341,7 +347,7 @@ download small).
341
  3. By default, only new content will have context added to save time and resources. Use `--process-all-context` only if you need to regenerate context for all documents. Use `--skip-data-upload` if you don't want to upload data files to the private HuggingFace repo (they're uploaded by default).
342
 
343
  4. When adding a new course, verify that it appears in the UI:
344
- - Add the source to all four registry structures in `source_registry.py`: `SOURCE_KEY_TO_LABEL`, `SOURCE_DISPLAY_INFO` (what `/api/tools` renders), `UI_SOURCE_KEYS` (a source absent here never appears in the picker), and optionally `DEFAULT_SELECTED_SOURCE_KEYS`
345
  - Check that the new source appears in the source picker in the UI
346
  - Make sure it's properly included in the default selected sources if desired
347
  - Restart the API server to see the changes
 
7
  - `GEMINI_API_KEY` or `GOOGLE_API_KEY` for context generation with Gemini
8
  - `COHERE_API_KEY` for embeddings
9
  - `HF_TOKEN` for HuggingFace uploads and downloads - [access to the private HuggingFace dataset repo](https://huggingface.co/datasets/towardsai-tutors/ai-tutor-data/tree/main)
10
+ - `GITHUB_TOKEN` for accessing files via the GitHub API (used by the docs workflow only; the course workflow never calls GitHub)
11
 
12
  Optional Gemini context-generation tuning:
13
 
 
15
  - `GEMINI_CONTEXT_TPM_SAFETY_MARGIN` - fraction of that quota to use before pausing (defaults to `0.8`)
16
  - `GEMINI_CONTEXT_CONCURRENCY` - max concurrent context requests (defaults to `50`)
17
  - `GEMINI_CONTEXT_RETRY_ATTEMPTS` - max tenacity attempts for transient Gemini API errors (defaults to `8`)
18
+ - `GEMINI_CONTEXT_MODEL` - Gemini model used for context generation (defaults to `gemini-3.1-flash-lite`)
19
+ - `GEMINI_CONTEXT_TPM_WINDOW_SECONDS` - sliding-window length in seconds for the TPM throttle (defaults to `60`)
20
 
21
  ## 1. Prepare the course data
22
 
 
52
  8. Open `data/scraping_scripts/source_registry.py`.
53
 
54
  9. Add the new course to the `SOURCE_CONFIGS` dictionary. Sources listed in
55
+ this registry are active in the knowledge base. Also add the key to
56
+ `COURSE_SOURCE_KEYS` in the same file: a course missing from that tuple is
57
+ classified as documentation (its pages land under `raw/docs/` and
58
+ `wiki/frameworks/` with the wrong picker group) and
59
+ `test_registry_classifies_every_source` fails CI.
60
 
61
  example:
62
 
 
95
 
96
  This script will guide you through the complete process, it will:
97
 
98
+ 1. Download the JSONL files from the other courses
99
+ 2. Extract the markdown content from each of the lessons and create a new JSONL file for the course
100
  3. Prompt you to manually add URLs to the course content, inside the newly created JSONL file (more details in step 3 below)
101
  4. Merge the course data into the main dataset
102
  5. Rebuild the KB artifacts (raw/generated/wiki; skip with `--skip-kb`)
 
347
  3. By default, only new content will have context added to save time and resources. Use `--process-all-context` only if you need to regenerate context for all documents. Use `--skip-data-upload` if you don't want to upload data files to the private HuggingFace repo (they're uploaded by default).
348
 
349
  4. When adding a new course, verify that it appears in the UI:
350
+ - Add the source to the registry structures in `source_registry.py`: `COURSE_SOURCE_KEYS` (classification; a course absent here is treated as documentation and fails CI), `SOURCE_KEY_TO_LABEL`, `SOURCE_DISPLAY_INFO` (what `/api/tools` renders), `UI_SOURCE_KEYS` (a source absent here never appears in the picker), and optionally `DEFAULT_SELECTED_SOURCE_KEYS`
351
  - Check that the new source appears in the source picker in the UI
352
  - Make sure it's properly included in the default selected sources if desired
353
  - Restart the API server to see the changes
data/url_matching_procedure.md CHANGED
@@ -77,7 +77,7 @@ Array.from(document.querySelectorAll('.course-player__chapter-item__header'))
77
  .reduce((a, b) => a + +b, 0);
78
  ```
79
 
80
- Gotcha: the Chrome `javascript_exec` tool truncates any string result at
81
  about 1024 characters. Stash the full list on `window.__pairs` and read it
82
  back in small slices (`window.__pairs.slice(0, 8).join('\n')`).
83
 
@@ -218,7 +218,7 @@ open(JSONL_PATH, 'w').writelines(out)
218
 
219
  ## Operational notes that tend to bite
220
 
221
- - Chrome `javascript_exec` output is truncated at ~1024 chars.
222
  - `super.site` iframes are cross-origin β€” you can only read `iframe.src`,
223
  not their DOM.
224
  - Teachable sidebars lazy-render; always expand chapters before scraping.
 
77
  .reduce((a, b) => a + +b, 0);
78
  ```
79
 
80
+ Gotcha: the Chrome `javascript_tool` MCP tool truncates any string result at
81
  about 1024 characters. Stash the full list on `window.__pairs` and read it
82
  back in small slices (`window.__pairs.slice(0, 8).join('\n')`).
83
 
 
218
 
219
  ## Operational notes that tend to bite
220
 
221
+ - Chrome `javascript_tool` output is truncated at ~1024 chars.
222
  - `super.site` iframes are cross-origin β€” you can only read `iframe.src`,
223
  not their DOM.
224
  - Teachable sidebars lazy-render; always expand chapters before scraping.
evals/compaction.md CHANGED
@@ -126,7 +126,7 @@ Gemini numbers above); it lives in `evals/slm_compaction.md` (since merged).
126
  ```bash
127
  uv run --env-file .env -m evals.compaction_study build --questions 15
128
  PRESETS="full_history prod aggressive sliding_window prompt_compression selective_retention context_reset incontext_history_retrieval delta_summarization hierarchical_summarization" \
129
- bash evals/run_compaction_study.sh # the script's default PRESETS omits the last two
130
  uv run --env-file .env -m evals.knowledge_compaction --questions 15 \
131
  --strategies rag graphrag --out data/compaction # Family B
132
  uv run --env-file .env -m evals.compaction_study report --runs 'runs/compaction_*'
 
126
  ```bash
127
  uv run --env-file .env -m evals.compaction_study build --questions 15
128
  PRESETS="full_history prod aggressive sliding_window prompt_compression selective_retention context_reset incontext_history_retrieval delta_summarization hierarchical_summarization" \
129
+ bash evals/run_compaction_study.sh # default PRESETS differs: it adds summarization_only and omits context_reset + the last two
130
  uv run --env-file .env -m evals.knowledge_compaction --questions 15 \
131
  --strategies rag graphrag --out data/compaction # Family B
132
  uv run --env-file .env -m evals.compaction_study report --runs 'runs/compaction_*'
evals/graphrag.md CHANGED
@@ -26,7 +26,9 @@ The **only** variable is the retrieval backend behind `retrieve_tutor_context`:
26
  classical so the comparison is fair.
27
 
28
  Held constant: the **chat model is Gemini 3.5 Flash** (the agent that writes the
29
- answer, `run_battery` default), the system prompt, the battery, the token budget,
 
 
30
  and source scoping. The GraphRAG retriever is a **context provider only** -- it
31
  never runs GraphRAG's own LLM answer-synthesis, so the agent's 3.5 Flash is the
32
  sole generation model in both arms.
 
26
  classical so the comparison is fair.
27
 
28
  Held constant: the **chat model is Gemini 3.5 Flash** (the agent that writes the
29
+ answer; both arms pass `--model google-genai:gemini-3.5-flash` explicitly --
30
+ `run_battery`'s no-flag default is the app default model, since moved to
31
+ DeepSeek V4 Flash), the system prompt, the battery, the token budget,
32
  and source scoping. The GraphRAG retriever is a **context provider only** -- it
33
  never runs GraphRAG's own LLM answer-synthesis, so the agent's 3.5 Flash is the
34
  sole generation model in both arms.
evals/part_c_plan.md CHANGED
@@ -60,9 +60,9 @@ Winners (2-3) + 1-2 combinations (e.g. `profile_memory` + `clear_retrieval_kb`)
60
 
61
  How a variant is actually wired, so a fresh session can replicate it.
62
 
63
- **Middleware hook API.** Custom middlewares subclass `AgentMiddleware` and override `wrap_model_call(self, request, handler)` (+ async `awrap_model_call`): mutate the request via `request.override(messages=…, system_message=…, model_settings=…)`, then `return handler(request)`. Templates in `chat_service.py`: `SourcePreferenceMiddleware` (:803), `StudentProfileMiddleware` (:852). The built-in *state-rewriting* compaction (`SummarizationMiddleware`, `ContextEditingMiddleware`) is assembled in `build_agent_middleware` (:894) from `MemoryConfig` flags.
64
 
65
- **Telemetry-signal gotcha (now solved by the turn-signal registry).** `context_window_stats` (`app/telemetry.py`) runs over the **checkpointed** messages (called near `chat_service.py:1526`) and detects markers: `lc_source: summarization` and the `CLEARED_TOOL_OUTPUT_PLACEHOLDER`. A middleware that only reduces the *per-call* view via `wrap_model_call` leaves the checkpoint unchanged, so it would emit no signal. The general fix is in place: a per-turn signal registry in `app/telemetry.py` (`reset_turn_signals(message_id)` at turn start, `record_turn_signal(message_id, name, n)` from the middleware, `pop_turn_signals(message_id)` merged into the `context_stats` event). A new per-call-view mechanism just calls `record_turn_signal` with a name, adds that name to `COMPACTION_SIGNAL_NAMES` (app) **and** `COMPACTION_SIGNAL_KEYS` (`evals/common.py`, kept in sync by a unit test), and the gate + report pick it up. A module-level dict + lock (not a `ContextVar`) because LangChain may run sync `wrap_model_call` in a worker thread.
66
 
67
  **Worked example 1 β€” `clear_retrieval_kb` (built, verified).** One `MemoryConfig` field `clear_excludes_retrieval` (default True = prod). `build_agent_middleware` sets `ClearToolUsesEdit(exclude_tools=("retrieve_tutor_context",) if clear_excludes_retrieval else ())`. Reuses the `cleared_tool_outputs` signal β†’ zero telemetry/gate work. The F3 fix in ~5 lines. Run: `--preset clear_retrieval_kb`.
68
 
@@ -81,8 +81,8 @@ How a variant is actually wired, so a fresh session can replicate it.
81
  | Variant | Axis | Shape | Wiring | `context_stats` signal | Tests / finding |
82
  |---|---|---|---|---|---|
83
  | `sliding_window` | A | preset | keep last N messages middleware | `dropped_messages` | F9/F10 β€” recency-only memory |
84
- | `delta_summarization` | A | preset | running summary, new-only | reuse `summary_messages` + `summary_mode` | F2 cache confound |
85
- | `hierarchical_summarization` | A | preset | summarize chunks then summaries | `summary_levels` | long-session compaction |
86
  | `context_reset` | A | preset | fresh state seeded w/ summary | reuse `summary_messages` (summary-prompt variant) | F2; prefix rewrite |
87
  | `prompt_compression` | A | preset | rewrite history fewer tokens | `compressed_messages`, `chars_saved` | F2 |
88
  | `selective_retention` | A | preset | summary prompt keeps constraints/decisions | reuse `summary_messages` | quality-preserving compaction |
@@ -93,7 +93,7 @@ How a variant is actually wired, so a fresh session can replicate it.
93
  | `observation_truncation` | B | preset | head/tail tool outputs incl. KB | `truncated_tool_outputs`, `chars_saved` | F1 β€” tokens are in tool outputs |
94
  | `clear_retrieval_kb` | B | preset | extend `ClearToolUsesEdit` to retrieval + KB | `cleared_tool_outputs` (now fires) | F3 β€” clearing excluded them |
95
  | `retrieval_budget_{100k,30k,10k}` | B | knob | `DEFAULT_CONTEXT_TOKEN_BUDGET` | (token counts already captured) + `retrieval_budget` | F1 β€” direct |
96
- | `kb_off` | B | toggle | `disable_kb` flag + drop KB prompt block | `kb_enabled` | does the KB improve answers? (89%/9:1 usage) |
97
 
98
  Stretch (post-workshop, evals.md): `temporal_graph_memory` (scorecard = `fact_update` probe), `subagent_isolation` (reframed as a tool-output-token play).
99
 
@@ -157,7 +157,7 @@ Same authoring model as v1 β€” real filler + authored planted facts + a binary `
157
  1. **Freeze + data discipline.** v2 is a new file; never touch v1. Gitignored; lives only in the private `ai-tutor-data` HF dataset under `eval/`.
158
  2. **Author Tier 1 first** (moderate contradiction sessions), reusing real corpus filler; model the shape on `data/eval/sessions_generated_*.jsonl` and the schema in `data/eval/README.md`. Add Tier 2 long-horizon sessions only if we decide Q1 is worth the spend/workshop story.
159
  3. **Calibrate by tokens and observed context state, not turn count.** For Q2, confirm prod summarization/compaction fired and A is no longer in the kept-recent window before the contradiction probe; if A is still visible to prod, the probe measures nothing. For Q1, smoke-run a few lengths (for example 30/45/60 turns) and confirm `full_history` cost actually diverges before running a screen.
160
- 4. **Wire grading** β€” add `contradiction` / `longhorizon_recall` / `entity_isolation` rubric entries to `evals/judge.py` (`RUBRICS`) and the human rubric. *(Done β€” all three are in `RUBRICS`.)*
161
  5. **Test `full_history` + `prod` + `incontext` + active `profile_memory` first.** This answers the cheap questions and tells us what, if anything, to build: does `full_history` fail contradictions (Q2), does `profile_memory` recover them through the store, and does `incontext` match `full_history`'s recall at lower all-in cost once sessions are actually long (Q1)?
162
  6. **Build only what the data demands**, one at a time, single-axis vs prod (each new mechanism needs its own `context_stats` signal mirrored in `evals/common.py` or the probe gate will not see it fire):
163
  - `temporal_graph_memory` β€” **only if** `full_history` fails Q2's contradictions.
 
60
 
61
  How a variant is actually wired, so a fresh session can replicate it.
62
 
63
+ **Middleware hook API.** Custom middlewares subclass `AgentMiddleware` and override `wrap_model_call(self, request, handler)` (+ async `awrap_model_call`): mutate the request via `request.override(messages=…, system_message=…, model_settings=…)`, then `return handler(request)`. Templates in `chat_service.py`: `SourcePreferenceMiddleware`, `StudentProfileMiddleware`. The built-in *state-rewriting* compaction (`SummarizationMiddleware`, `ContextEditingMiddleware`) is assembled in `build_agent_middleware` from `MemoryConfig` flags. (Line numbers drift; grep the symbol names.)
64
 
65
+ **Telemetry-signal gotcha (now solved by the turn-signal registry).** `context_window_stats` (`app/telemetry.py`) runs over the **checkpointed** messages (called from `stream_chat` in `chat_service.py`) and detects markers: `lc_source: summarization` and the `CLEARED_TOOL_OUTPUT_PLACEHOLDER`. A middleware that only reduces the *per-call* view via `wrap_model_call` leaves the checkpoint unchanged, so it would emit no signal. The general fix is in place: a per-turn signal registry in `app/telemetry.py` (`reset_turn_signals(message_id)` at turn start, `record_turn_signal(message_id, name, n)` from the middleware, `pop_turn_signals(message_id)` merged into the `context_stats` event). A new per-call-view mechanism just calls `record_turn_signal` with a name, adds that name to `COMPACTION_SIGNAL_NAMES` (app) **and** `COMPACTION_SIGNAL_KEYS` (`evals/common.py`, kept in sync by a unit test), and the gate + report pick it up. A module-level dict + lock (not a `ContextVar`) because LangChain may run sync `wrap_model_call` in a worker thread.
66
 
67
  **Worked example 1 β€” `clear_retrieval_kb` (built, verified).** One `MemoryConfig` field `clear_excludes_retrieval` (default True = prod). `build_agent_middleware` sets `ClearToolUsesEdit(exclude_tools=("retrieve_tutor_context",) if clear_excludes_retrieval else ())`. Reuses the `cleared_tool_outputs` signal β†’ zero telemetry/gate work. The F3 fix in ~5 lines. Run: `--preset clear_retrieval_kb`.
68
 
 
81
  | Variant | Axis | Shape | Wiring | `context_stats` signal | Tests / finding |
82
  |---|---|---|---|---|---|
83
  | `sliding_window` | A | preset | keep last N messages middleware | `dropped_messages` | F9/F10 β€” recency-only memory |
84
+ | `delta_summarization` | A | preset | running summary, new-only | reuse `summary_messages` (as built; no `summary_mode` signal) | F2 cache confound |
85
+ | `hierarchical_summarization` | A | preset | summarize chunks then summaries | reuse `summary_messages` (as built; no `summary_levels` signal) | long-session compaction |
86
  | `context_reset` | A | preset | fresh state seeded w/ summary | reuse `summary_messages` (summary-prompt variant) | F2; prefix rewrite |
87
  | `prompt_compression` | A | preset | rewrite history fewer tokens | `compressed_messages`, `chars_saved` | F2 |
88
  | `selective_retention` | A | preset | summary prompt keeps constraints/decisions | reuse `summary_messages` | quality-preserving compaction |
 
93
  | `observation_truncation` | B | preset | head/tail tool outputs incl. KB | `truncated_tool_outputs`, `chars_saved` | F1 β€” tokens are in tool outputs |
94
  | `clear_retrieval_kb` | B | preset | extend `ClearToolUsesEdit` to retrieval + KB | `cleared_tool_outputs` (now fires) | F3 β€” clearing excluded them |
95
  | `retrieval_budget_{100k,30k,10k}` | B | knob | `DEFAULT_CONTEXT_TOKEN_BUDGET` | (token counts already captured) + `retrieval_budget` | F1 β€” direct |
96
+ | `kb_off` | B | toggle | `disable_kb` flag + drop KB prompt block | none as built (request flag only; no `kb_enabled` signal) | does the KB improve answers? (89%/9:1 usage) |
97
 
98
  Stretch (post-workshop, evals.md): `temporal_graph_memory` (scorecard = `fact_update` probe), `subagent_isolation` (reframed as a tool-output-token play).
99
 
 
157
  1. **Freeze + data discipline.** v2 is a new file; never touch v1. Gitignored; lives only in the private `ai-tutor-data` HF dataset under `eval/`.
158
  2. **Author Tier 1 first** (moderate contradiction sessions), reusing real corpus filler; model the shape on `data/eval/sessions_generated_*.jsonl` and the schema in `data/eval/README.md`. Add Tier 2 long-horizon sessions only if we decide Q1 is worth the spend/workshop story.
159
  3. **Calibrate by tokens and observed context state, not turn count.** For Q2, confirm prod summarization/compaction fired and A is no longer in the kept-recent window before the contradiction probe; if A is still visible to prod, the probe measures nothing. For Q1, smoke-run a few lengths (for example 30/45/60 turns) and confirm `full_history` cost actually diverges before running a screen.
160
+ 4. **Wire grading** β€” add `contradiction` / `longhorizon_recall` / `entity_isolation` grading to `evals/judge.py` and the human rubric. *(Done β€” all three route through the shared `probe` rubric via per-type instruction text, not separate `RUBRICS` keys.)*
161
  5. **Test `full_history` + `prod` + `incontext` + active `profile_memory` first.** This answers the cheap questions and tells us what, if anything, to build: does `full_history` fail contradictions (Q2), does `profile_memory` recover them through the store, and does `incontext` match `full_history`'s recall at lower all-in cost once sessions are actually long (Q1)?
162
  6. **Build only what the data demands**, one at a time, single-axis vs prod (each new mechanism needs its own `context_stats` signal mirrored in `evals/common.py` or the probe gate will not see it fire):
163
  - `temporal_graph_memory` β€” **only if** `full_history` fails Q2's contradictions.
tests/manual_e2e_langsmith.md CHANGED
@@ -20,7 +20,9 @@ Required local artifacts and environment:
20
  - `.env` has `LANGSMITH_PROJECT=ai-tutor-app`
21
  - `data/chroma-db-all_sources/` exists
22
  - `data/kb/wiki/index.md` exists
23
- - `curl`, `jq`, and the `langsmith` CLI are available through `uv run`
 
 
24
 
25
  Quick check:
26
 
@@ -103,8 +105,11 @@ JSON
103
  ```
104
 
105
  On `enabledTools`: it is an explicit allowlist of the toggle tools for the turn.
106
- `[]` (as above) disables web search and URL reading, so this run exercises only
107
- the corpus retrieval + KB shell citation paths. Two things to know:
 
 
 
108
 
109
  - **`url_context` defaults to off in the UI** (`active: False` in `_tool_catalog`,
110
  `app/api.py`), so a browser turn only sends it when the user opts in. Enabling
 
20
  - `.env` has `LANGSMITH_PROJECT=ai-tutor-app`
21
  - `data/chroma-db-all_sources/` exists
22
  - `data/kb/wiki/index.md` exists
23
+ - `curl` and `jq` are on `PATH`, and the standalone `langsmith` CLI binary is
24
+ installed (a separate Go binary, e.g. `~/.local/bin/langsmith`; `uv sync`
25
+ does not provide it β€” `uv run` merely inherits it via `PATH`)
26
 
27
  Quick check:
28
 
 
105
  ```
106
 
107
  On `enabledTools`: it is an explicit allowlist of the toggle tools for the turn.
108
+ For this payload's DeepSeek model the catalog exposes no toggle tools at all, so
109
+ `[]` and omitting the field behave identically here; the notes below bite when
110
+ the model is Gemini 3+ or Anthropic. `[]` (as above) disables web search and URL
111
+ reading, so this run exercises only the corpus retrieval + KB shell citation
112
+ paths. Two things to know:
113
 
114
  - **`url_context` defaults to off in the UI** (`active: False` in `_tool_catalog`,
115
  `app/api.py`), so a browser turn only sends it when the user opts in. Enabling