Fix stale claims across repo docs
Browse filesDoc audit against current code found drift; align the docs with reality:
- scraping README: course setup must also add COURSE_SOURCE_KEYS (missing
it misclassifies the course as docs and fails CI); swap course workflow
steps 1/2 to match execution order; GITHUB_TOKEN is docs-workflow only;
document GEMINI_CONTEXT_MODEL and GEMINI_CONTEXT_TPM_WINDOW_SECONDS
- AGENTS.md: the two custom tools are not "always" exposed (empty source
selection removes both; evals' disable_kb keeps retrieval only); quote
the prod deploy workflow's full name (also in README.md)
- evals/graphrag.md: run_battery's no-flag default is the app default
model (now DeepSeek V4 Flash), not Gemini 3.5 Flash
- evals/compaction.md: correct the run_compaction_study default-PRESETS
note (adds summarization_only, omits context_reset + the last two)
- data/eval/README.md: compaction scripts default to data/compaction/,
not data/eval/compaction/; document the override
- evals/part_c_plan.md: drop drifted line references; correct the probe
rubric wiring note and as-built context_stats signal cells
- manual_e2e_langsmith.md: langsmith CLI is a standalone binary, not a
uv dependency; note the enabledTools discussion does not bite for the
doc's own DeepSeek payload
- url_matching_procedure.md: Chrome scripting tool is javascript_tool
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- AGENTS.md +2 -2
- README.md +1 -1
- data/eval/README.md +1 -1
- data/scraping_scripts/README.md +11 -5
- data/url_matching_procedure.md +2 -2
- evals/compaction.md +1 -1
- evals/graphrag.md +3 -1
- evals/part_c_plan.md +6 -6
- tests/manual_e2e_langsmith.md +8 -3
|
@@ -34,7 +34,7 @@ ChromaDB for vectors; Cohere for embeddings/rerank; chat model is provider-confi
|
|
| 34 |
|
| 35 |
## Architecture in brief
|
| 36 |
|
| 37 |
-
The agent is built with `langchain.agents.create_agent()` (LangGraph), an `InMemorySaver` checkpointer keyed by `thread_id`, and middlewares assembled from the selected **memory preset** (`app/memory_presets.py`: context-editing, summarization, optional long-term student-profile memory) plus source preference. The normal frontend does not select a memory preset: the server's model-aware production policy uses `prod_v2` for DeepSeek/Gemini and the immutable legacy `prod` preset for other providers. Direct API clients and evals can override this with `memoryPreset`; the full precedence is request value β `AI_TUTOR_MEMORY_PRESET` environment override β model-aware production policy. `stream_chat()` is the single entry point the API calls; it yields typed `ChatEvent`s that `app/api.py` encodes into the AI SDK UI-message stream, ending each turn with a `context_stats` telemetry event (tokens incl. cache buckets, est. cost, TTFT, compaction-trigger counts β independent of LangSmith). It
|
| 38 |
|
| 39 |
- **`retrieve_tutor_context(query)`** β hybrid RAG over the corpus, scoped to the user's selected sources.
|
| 40 |
- **`run_kb_command(...)`** β read-only KB file browsing (see below).
|
|
@@ -104,7 +104,7 @@ Chat runtime: `COHERE_API_KEY` (retrieval; always required), `DEEPSEEK_API_KEY`
|
|
| 104 |
Both HF Spaces run the same image (`Dockerfile`: FastAPI + Next.js static export, `ripgrep` installed for `run_kb_command`, port :7860), in a dev β prod flow:
|
| 105 |
|
| 106 |
- **Dev β `ai-tutor`** (private): `.github/workflows/sync-to-hf.yml` force-pushes on every push to `main` (docs/markdown-only and scraping-script-only pushes are skipped via `paths-ignore`). Verify changes here first.
|
| 107 |
-
- **Prod β `ai-tutor-chatbot`** (public): `.github/workflows/deploy-prod-to-hf.yml`, **manual trigger only** (Actions tab β "Deploy prod to Hugging Face" β Run workflow).
|
| 108 |
|
| 109 |
Both Spaces need the same runtime secrets (`COHERE_API_KEY`, model provider key, `HF_TOKEN`, optional `LANGSMITH_*`) configured in their HF settings.
|
| 110 |
|
|
|
|
| 34 |
|
| 35 |
## Architecture in brief
|
| 36 |
|
| 37 |
+
The agent is built with `langchain.agents.create_agent()` (LangGraph), an `InMemorySaver` checkpointer keyed by `thread_id`, and middlewares assembled from the selected **memory preset** (`app/memory_presets.py`: context-editing, summarization, optional long-term student-profile memory) plus source preference. The normal frontend does not select a memory preset: the server's model-aware production policy uses `prod_v2` for DeepSeek/Gemini and the immutable legacy `prod` preset for other providers. Direct API clients and evals can override this with `memoryPreset`; the full precedence is request value β `AI_TUTOR_MEMORY_PRESET` environment override β model-aware production policy. `stream_chat()` is the single entry point the API calls; it yields typed `ChatEvent`s that `app/api.py` encodes into the AI SDK UI-message stream, ending each turn with a `context_stats` telemetry event (tokens incl. cache buckets, est. cost, TTFT, compaction-trigger counts β independent of LangSmith). It exposes two custom tools (deselecting every source removes both for that turn, and the evals' `disable_kb` flag keeps retrieval only), plus provider-native web tools when enabled:
|
| 38 |
|
| 39 |
- **`retrieve_tutor_context(query)`** β hybrid RAG over the corpus, scoped to the user's selected sources.
|
| 40 |
- **`run_kb_command(...)`** β read-only KB file browsing (see below).
|
|
|
|
| 104 |
Both HF Spaces run the same image (`Dockerfile`: FastAPI + Next.js static export, `ripgrep` installed for `run_kb_command`, port :7860), in a dev β prod flow:
|
| 105 |
|
| 106 |
- **Dev β `ai-tutor`** (private): `.github/workflows/sync-to-hf.yml` force-pushes on every push to `main` (docs/markdown-only and scraping-script-only pushes are skipped via `paths-ignore`). Verify changes here first.
|
| 107 |
+
- **Prod β `ai-tutor-chatbot`** (public): `.github/workflows/deploy-prod-to-hf.yml`, **manual trigger only** (Actions tab β "Deploy prod to Hugging Face (ai-tutor-chatbot)" β Run workflow).
|
| 108 |
|
| 109 |
Both Spaces need the same runtime secrets (`COHERE_API_KEY`, model provider key, `HF_TOKEN`, optional `LANGSMITH_*`) configured in their HF settings.
|
| 110 |
|
|
@@ -16,7 +16,7 @@ Built by [Louis-FranΓ§ois Bouchard](https://www.linkedin.com/in/whats-ai/) ([X](
|
|
| 16 |
|
| 17 |
The live app is deployed on Hugging Face Spaces at: [AI Tutor Chatbot on Hugging Face](https://huggingface.co/spaces/towardsai-tutors/ai-tutor-chatbot) (prod).
|
| 18 |
|
| 19 |
-
**Deployment flow:** every push to `main` (except docs/markdown-only and scraping-script-only changes) auto-deploys to the private dev Space ([ai-tutor](https://huggingface.co/spaces/towardsai-tutors/ai-tutor)) for verification; the prod Space is promoted manually via the "Deploy prod to Hugging Face" workflow in the Actions tab.
|
| 20 |
|
| 21 |
### Workshop, slides, and going deeper
|
| 22 |
|
|
|
|
| 16 |
|
| 17 |
The live app is deployed on Hugging Face Spaces at: [AI Tutor Chatbot on Hugging Face](https://huggingface.co/spaces/towardsai-tutors/ai-tutor-chatbot) (prod).
|
| 18 |
|
| 19 |
+
**Deployment flow:** every push to `main` (except docs/markdown-only and scraping-script-only changes) auto-deploys to the private dev Space ([ai-tutor](https://huggingface.co/spaces/towardsai-tutors/ai-tutor)) for verification; the prod Space is promoted manually via the "Deploy prod to Hugging Face (ai-tutor-chatbot)" workflow in the Actions tab.
|
| 20 |
|
| 21 |
### Workshop, slides, and going deeper
|
| 22 |
|
|
@@ -1,6 +1,6 @@
|
|
| 1 |
# Evaluation batteries (v1, frozen 2026-06-11)
|
| 2 |
|
| 3 |
-
The datasets for evaluating the AI tutor (see `evals.md` at the repo root for the overall effort and results). **This README is the single reference for what every file, field, and term means.** Files added after the v1 freeze: `battery_sessions_v2.jsonl` (deprecated for new runs β see the v2 note in the probe-type glossary), `battery_sessions_v2_1.jsonl` (its repaired successor), and `compaction/` (the F29/F30 lesson-compaction batteries, documented in `evals/compaction.md` / `evals/slm_compaction.md`).
|
| 4 |
|
| 5 |
Built from `data/academy_discussion_eval.jsonl` β real student posts from the academy discussion boards with real staff answers β plus authored content (sessions, personas).
|
| 6 |
|
|
|
|
| 1 |
# Evaluation batteries (v1, frozen 2026-06-11)
|
| 2 |
|
| 3 |
+
The datasets for evaluating the AI tutor (see `evals.md` at the repo root for the overall effort and results). **This README is the single reference for what every file, field, and term means.** Files added after the v1 freeze: `battery_sessions_v2.jsonl` (deprecated for new runs β see the v2 note in the probe-type glossary), `battery_sessions_v2_1.jsonl` (its repaired successor), and `compaction/` (the F29/F30 lesson-compaction batteries, documented in `evals/compaction.md` / `evals/slm_compaction.md`; note the compaction scripts default to `data/compaction/`, not this folder β either regenerate there with `compaction_study build` or point `BATTERY=`/`QUESTIONS_FILE=` at these copies).
|
| 4 |
|
| 5 |
Built from `data/academy_discussion_eval.jsonl` β real student posts from the academy discussion boards with real staff answers β plus authored content (sessions, personas).
|
| 6 |
|
|
@@ -7,7 +7,7 @@ Make sure you have the required environment variables set:
|
|
| 7 |
- `GEMINI_API_KEY` or `GOOGLE_API_KEY` for context generation with Gemini
|
| 8 |
- `COHERE_API_KEY` for embeddings
|
| 9 |
- `HF_TOKEN` for HuggingFace uploads and downloads - [access to the private HuggingFace dataset repo](https://huggingface.co/datasets/towardsai-tutors/ai-tutor-data/tree/main)
|
| 10 |
-
- `GITHUB_TOKEN` for accessing files via the GitHub API
|
| 11 |
|
| 12 |
Optional Gemini context-generation tuning:
|
| 13 |
|
|
@@ -15,6 +15,8 @@ Optional Gemini context-generation tuning:
|
|
| 15 |
- `GEMINI_CONTEXT_TPM_SAFETY_MARGIN` - fraction of that quota to use before pausing (defaults to `0.8`)
|
| 16 |
- `GEMINI_CONTEXT_CONCURRENCY` - max concurrent context requests (defaults to `50`)
|
| 17 |
- `GEMINI_CONTEXT_RETRY_ATTEMPTS` - max tenacity attempts for transient Gemini API errors (defaults to `8`)
|
|
|
|
|
|
|
| 18 |
|
| 19 |
## 1. Prepare the course data
|
| 20 |
|
|
@@ -50,7 +52,11 @@ Optional Gemini context-generation tuning:
|
|
| 50 |
8. Open `data/scraping_scripts/source_registry.py`.
|
| 51 |
|
| 52 |
9. Add the new course to the `SOURCE_CONFIGS` dictionary. Sources listed in
|
| 53 |
-
this registry are active in the knowledge base.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 54 |
|
| 55 |
example:
|
| 56 |
|
|
@@ -89,8 +95,8 @@ uv run -m data.scraping_scripts.add_course_workflow --courses master_ai_for_work
|
|
| 89 |
|
| 90 |
This script will guide you through the complete process, it will:
|
| 91 |
|
| 92 |
-
1.
|
| 93 |
-
2.
|
| 94 |
3. Prompt you to manually add URLs to the course content, inside the newly created JSONL file (more details in step 3 below)
|
| 95 |
4. Merge the course data into the main dataset
|
| 96 |
5. Rebuild the KB artifacts (raw/generated/wiki; skip with `--skip-kb`)
|
|
@@ -341,7 +347,7 @@ download small).
|
|
| 341 |
3. By default, only new content will have context added to save time and resources. Use `--process-all-context` only if you need to regenerate context for all documents. Use `--skip-data-upload` if you don't want to upload data files to the private HuggingFace repo (they're uploaded by default).
|
| 342 |
|
| 343 |
4. When adding a new course, verify that it appears in the UI:
|
| 344 |
-
- Add the source to
|
| 345 |
- Check that the new source appears in the source picker in the UI
|
| 346 |
- Make sure it's properly included in the default selected sources if desired
|
| 347 |
- Restart the API server to see the changes
|
|
|
|
| 7 |
- `GEMINI_API_KEY` or `GOOGLE_API_KEY` for context generation with Gemini
|
| 8 |
- `COHERE_API_KEY` for embeddings
|
| 9 |
- `HF_TOKEN` for HuggingFace uploads and downloads - [access to the private HuggingFace dataset repo](https://huggingface.co/datasets/towardsai-tutors/ai-tutor-data/tree/main)
|
| 10 |
+
- `GITHUB_TOKEN` for accessing files via the GitHub API (used by the docs workflow only; the course workflow never calls GitHub)
|
| 11 |
|
| 12 |
Optional Gemini context-generation tuning:
|
| 13 |
|
|
|
|
| 15 |
- `GEMINI_CONTEXT_TPM_SAFETY_MARGIN` - fraction of that quota to use before pausing (defaults to `0.8`)
|
| 16 |
- `GEMINI_CONTEXT_CONCURRENCY` - max concurrent context requests (defaults to `50`)
|
| 17 |
- `GEMINI_CONTEXT_RETRY_ATTEMPTS` - max tenacity attempts for transient Gemini API errors (defaults to `8`)
|
| 18 |
+
- `GEMINI_CONTEXT_MODEL` - Gemini model used for context generation (defaults to `gemini-3.1-flash-lite`)
|
| 19 |
+
- `GEMINI_CONTEXT_TPM_WINDOW_SECONDS` - sliding-window length in seconds for the TPM throttle (defaults to `60`)
|
| 20 |
|
| 21 |
## 1. Prepare the course data
|
| 22 |
|
|
|
|
| 52 |
8. Open `data/scraping_scripts/source_registry.py`.
|
| 53 |
|
| 54 |
9. Add the new course to the `SOURCE_CONFIGS` dictionary. Sources listed in
|
| 55 |
+
this registry are active in the knowledge base. Also add the key to
|
| 56 |
+
`COURSE_SOURCE_KEYS` in the same file: a course missing from that tuple is
|
| 57 |
+
classified as documentation (its pages land under `raw/docs/` and
|
| 58 |
+
`wiki/frameworks/` with the wrong picker group) and
|
| 59 |
+
`test_registry_classifies_every_source` fails CI.
|
| 60 |
|
| 61 |
example:
|
| 62 |
|
|
|
|
| 95 |
|
| 96 |
This script will guide you through the complete process, it will:
|
| 97 |
|
| 98 |
+
1. Download the JSONL files from the other courses
|
| 99 |
+
2. Extract the markdown content from each of the lessons and create a new JSONL file for the course
|
| 100 |
3. Prompt you to manually add URLs to the course content, inside the newly created JSONL file (more details in step 3 below)
|
| 101 |
4. Merge the course data into the main dataset
|
| 102 |
5. Rebuild the KB artifacts (raw/generated/wiki; skip with `--skip-kb`)
|
|
|
|
| 347 |
3. By default, only new content will have context added to save time and resources. Use `--process-all-context` only if you need to regenerate context for all documents. Use `--skip-data-upload` if you don't want to upload data files to the private HuggingFace repo (they're uploaded by default).
|
| 348 |
|
| 349 |
4. When adding a new course, verify that it appears in the UI:
|
| 350 |
+
- Add the source to the registry structures in `source_registry.py`: `COURSE_SOURCE_KEYS` (classification; a course absent here is treated as documentation and fails CI), `SOURCE_KEY_TO_LABEL`, `SOURCE_DISPLAY_INFO` (what `/api/tools` renders), `UI_SOURCE_KEYS` (a source absent here never appears in the picker), and optionally `DEFAULT_SELECTED_SOURCE_KEYS`
|
| 351 |
- Check that the new source appears in the source picker in the UI
|
| 352 |
- Make sure it's properly included in the default selected sources if desired
|
| 353 |
- Restart the API server to see the changes
|
|
@@ -77,7 +77,7 @@ Array.from(document.querySelectorAll('.course-player__chapter-item__header'))
|
|
| 77 |
.reduce((a, b) => a + +b, 0);
|
| 78 |
```
|
| 79 |
|
| 80 |
-
Gotcha: the Chrome `
|
| 81 |
about 1024 characters. Stash the full list on `window.__pairs` and read it
|
| 82 |
back in small slices (`window.__pairs.slice(0, 8).join('\n')`).
|
| 83 |
|
|
@@ -218,7 +218,7 @@ open(JSONL_PATH, 'w').writelines(out)
|
|
| 218 |
|
| 219 |
## Operational notes that tend to bite
|
| 220 |
|
| 221 |
-
- Chrome `
|
| 222 |
- `super.site` iframes are cross-origin β you can only read `iframe.src`,
|
| 223 |
not their DOM.
|
| 224 |
- Teachable sidebars lazy-render; always expand chapters before scraping.
|
|
|
|
| 77 |
.reduce((a, b) => a + +b, 0);
|
| 78 |
```
|
| 79 |
|
| 80 |
+
Gotcha: the Chrome `javascript_tool` MCP tool truncates any string result at
|
| 81 |
about 1024 characters. Stash the full list on `window.__pairs` and read it
|
| 82 |
back in small slices (`window.__pairs.slice(0, 8).join('\n')`).
|
| 83 |
|
|
|
|
| 218 |
|
| 219 |
## Operational notes that tend to bite
|
| 220 |
|
| 221 |
+
- Chrome `javascript_tool` output is truncated at ~1024 chars.
|
| 222 |
- `super.site` iframes are cross-origin β you can only read `iframe.src`,
|
| 223 |
not their DOM.
|
| 224 |
- Teachable sidebars lazy-render; always expand chapters before scraping.
|
|
@@ -126,7 +126,7 @@ Gemini numbers above); it lives in `evals/slm_compaction.md` (since merged).
|
|
| 126 |
```bash
|
| 127 |
uv run --env-file .env -m evals.compaction_study build --questions 15
|
| 128 |
PRESETS="full_history prod aggressive sliding_window prompt_compression selective_retention context_reset incontext_history_retrieval delta_summarization hierarchical_summarization" \
|
| 129 |
-
bash evals/run_compaction_study.sh #
|
| 130 |
uv run --env-file .env -m evals.knowledge_compaction --questions 15 \
|
| 131 |
--strategies rag graphrag --out data/compaction # Family B
|
| 132 |
uv run --env-file .env -m evals.compaction_study report --runs 'runs/compaction_*'
|
|
|
|
| 126 |
```bash
|
| 127 |
uv run --env-file .env -m evals.compaction_study build --questions 15
|
| 128 |
PRESETS="full_history prod aggressive sliding_window prompt_compression selective_retention context_reset incontext_history_retrieval delta_summarization hierarchical_summarization" \
|
| 129 |
+
bash evals/run_compaction_study.sh # default PRESETS differs: it adds summarization_only and omits context_reset + the last two
|
| 130 |
uv run --env-file .env -m evals.knowledge_compaction --questions 15 \
|
| 131 |
--strategies rag graphrag --out data/compaction # Family B
|
| 132 |
uv run --env-file .env -m evals.compaction_study report --runs 'runs/compaction_*'
|
|
@@ -26,7 +26,9 @@ The **only** variable is the retrieval backend behind `retrieve_tutor_context`:
|
|
| 26 |
classical so the comparison is fair.
|
| 27 |
|
| 28 |
Held constant: the **chat model is Gemini 3.5 Flash** (the agent that writes the
|
| 29 |
-
answer
|
|
|
|
|
|
|
| 30 |
and source scoping. The GraphRAG retriever is a **context provider only** -- it
|
| 31 |
never runs GraphRAG's own LLM answer-synthesis, so the agent's 3.5 Flash is the
|
| 32 |
sole generation model in both arms.
|
|
|
|
| 26 |
classical so the comparison is fair.
|
| 27 |
|
| 28 |
Held constant: the **chat model is Gemini 3.5 Flash** (the agent that writes the
|
| 29 |
+
answer; both arms pass `--model google-genai:gemini-3.5-flash` explicitly --
|
| 30 |
+
`run_battery`'s no-flag default is the app default model, since moved to
|
| 31 |
+
DeepSeek V4 Flash), the system prompt, the battery, the token budget,
|
| 32 |
and source scoping. The GraphRAG retriever is a **context provider only** -- it
|
| 33 |
never runs GraphRAG's own LLM answer-synthesis, so the agent's 3.5 Flash is the
|
| 34 |
sole generation model in both arms.
|
|
@@ -60,9 +60,9 @@ Winners (2-3) + 1-2 combinations (e.g. `profile_memory` + `clear_retrieval_kb`)
|
|
| 60 |
|
| 61 |
How a variant is actually wired, so a fresh session can replicate it.
|
| 62 |
|
| 63 |
-
**Middleware hook API.** Custom middlewares subclass `AgentMiddleware` and override `wrap_model_call(self, request, handler)` (+ async `awrap_model_call`): mutate the request via `request.override(messages=β¦, system_message=β¦, model_settings=β¦)`, then `return handler(request)`. Templates in `chat_service.py`: `SourcePreferenceMiddleware`
|
| 64 |
|
| 65 |
-
**Telemetry-signal gotcha (now solved by the turn-signal registry).** `context_window_stats` (`app/telemetry.py`) runs over the **checkpointed** messages (called
|
| 66 |
|
| 67 |
**Worked example 1 β `clear_retrieval_kb` (built, verified).** One `MemoryConfig` field `clear_excludes_retrieval` (default True = prod). `build_agent_middleware` sets `ClearToolUsesEdit(exclude_tools=("retrieve_tutor_context",) if clear_excludes_retrieval else ())`. Reuses the `cleared_tool_outputs` signal β zero telemetry/gate work. The F3 fix in ~5 lines. Run: `--preset clear_retrieval_kb`.
|
| 68 |
|
|
@@ -81,8 +81,8 @@ How a variant is actually wired, so a fresh session can replicate it.
|
|
| 81 |
| Variant | Axis | Shape | Wiring | `context_stats` signal | Tests / finding |
|
| 82 |
|---|---|---|---|---|---|
|
| 83 |
| `sliding_window` | A | preset | keep last N messages middleware | `dropped_messages` | F9/F10 β recency-only memory |
|
| 84 |
-
| `delta_summarization` | A | preset | running summary, new-only | reuse `summary_messages`
|
| 85 |
-
| `hierarchical_summarization` | A | preset | summarize chunks then summaries | `summary_levels` | long-session compaction |
|
| 86 |
| `context_reset` | A | preset | fresh state seeded w/ summary | reuse `summary_messages` (summary-prompt variant) | F2; prefix rewrite |
|
| 87 |
| `prompt_compression` | A | preset | rewrite history fewer tokens | `compressed_messages`, `chars_saved` | F2 |
|
| 88 |
| `selective_retention` | A | preset | summary prompt keeps constraints/decisions | reuse `summary_messages` | quality-preserving compaction |
|
|
@@ -93,7 +93,7 @@ How a variant is actually wired, so a fresh session can replicate it.
|
|
| 93 |
| `observation_truncation` | B | preset | head/tail tool outputs incl. KB | `truncated_tool_outputs`, `chars_saved` | F1 β tokens are in tool outputs |
|
| 94 |
| `clear_retrieval_kb` | B | preset | extend `ClearToolUsesEdit` to retrieval + KB | `cleared_tool_outputs` (now fires) | F3 β clearing excluded them |
|
| 95 |
| `retrieval_budget_{100k,30k,10k}` | B | knob | `DEFAULT_CONTEXT_TOKEN_BUDGET` | (token counts already captured) + `retrieval_budget` | F1 β direct |
|
| 96 |
-
| `kb_off` | B | toggle | `disable_kb` flag + drop KB prompt block | `kb_enabled` | does the KB improve answers? (89%/9:1 usage) |
|
| 97 |
|
| 98 |
Stretch (post-workshop, evals.md): `temporal_graph_memory` (scorecard = `fact_update` probe), `subagent_isolation` (reframed as a tool-output-token play).
|
| 99 |
|
|
@@ -157,7 +157,7 @@ Same authoring model as v1 β real filler + authored planted facts + a binary `
|
|
| 157 |
1. **Freeze + data discipline.** v2 is a new file; never touch v1. Gitignored; lives only in the private `ai-tutor-data` HF dataset under `eval/`.
|
| 158 |
2. **Author Tier 1 first** (moderate contradiction sessions), reusing real corpus filler; model the shape on `data/eval/sessions_generated_*.jsonl` and the schema in `data/eval/README.md`. Add Tier 2 long-horizon sessions only if we decide Q1 is worth the spend/workshop story.
|
| 159 |
3. **Calibrate by tokens and observed context state, not turn count.** For Q2, confirm prod summarization/compaction fired and A is no longer in the kept-recent window before the contradiction probe; if A is still visible to prod, the probe measures nothing. For Q1, smoke-run a few lengths (for example 30/45/60 turns) and confirm `full_history` cost actually diverges before running a screen.
|
| 160 |
-
4. **Wire grading** β add `contradiction` / `longhorizon_recall` / `entity_isolation`
|
| 161 |
5. **Test `full_history` + `prod` + `incontext` + active `profile_memory` first.** This answers the cheap questions and tells us what, if anything, to build: does `full_history` fail contradictions (Q2), does `profile_memory` recover them through the store, and does `incontext` match `full_history`'s recall at lower all-in cost once sessions are actually long (Q1)?
|
| 162 |
6. **Build only what the data demands**, one at a time, single-axis vs prod (each new mechanism needs its own `context_stats` signal mirrored in `evals/common.py` or the probe gate will not see it fire):
|
| 163 |
- `temporal_graph_memory` β **only if** `full_history` fails Q2's contradictions.
|
|
|
|
| 60 |
|
| 61 |
How a variant is actually wired, so a fresh session can replicate it.
|
| 62 |
|
| 63 |
+
**Middleware hook API.** Custom middlewares subclass `AgentMiddleware` and override `wrap_model_call(self, request, handler)` (+ async `awrap_model_call`): mutate the request via `request.override(messages=β¦, system_message=β¦, model_settings=β¦)`, then `return handler(request)`. Templates in `chat_service.py`: `SourcePreferenceMiddleware`, `StudentProfileMiddleware`. The built-in *state-rewriting* compaction (`SummarizationMiddleware`, `ContextEditingMiddleware`) is assembled in `build_agent_middleware` from `MemoryConfig` flags. (Line numbers drift; grep the symbol names.)
|
| 64 |
|
| 65 |
+
**Telemetry-signal gotcha (now solved by the turn-signal registry).** `context_window_stats` (`app/telemetry.py`) runs over the **checkpointed** messages (called from `stream_chat` in `chat_service.py`) and detects markers: `lc_source: summarization` and the `CLEARED_TOOL_OUTPUT_PLACEHOLDER`. A middleware that only reduces the *per-call* view via `wrap_model_call` leaves the checkpoint unchanged, so it would emit no signal. The general fix is in place: a per-turn signal registry in `app/telemetry.py` (`reset_turn_signals(message_id)` at turn start, `record_turn_signal(message_id, name, n)` from the middleware, `pop_turn_signals(message_id)` merged into the `context_stats` event). A new per-call-view mechanism just calls `record_turn_signal` with a name, adds that name to `COMPACTION_SIGNAL_NAMES` (app) **and** `COMPACTION_SIGNAL_KEYS` (`evals/common.py`, kept in sync by a unit test), and the gate + report pick it up. A module-level dict + lock (not a `ContextVar`) because LangChain may run sync `wrap_model_call` in a worker thread.
|
| 66 |
|
| 67 |
**Worked example 1 β `clear_retrieval_kb` (built, verified).** One `MemoryConfig` field `clear_excludes_retrieval` (default True = prod). `build_agent_middleware` sets `ClearToolUsesEdit(exclude_tools=("retrieve_tutor_context",) if clear_excludes_retrieval else ())`. Reuses the `cleared_tool_outputs` signal β zero telemetry/gate work. The F3 fix in ~5 lines. Run: `--preset clear_retrieval_kb`.
|
| 68 |
|
|
|
|
| 81 |
| Variant | Axis | Shape | Wiring | `context_stats` signal | Tests / finding |
|
| 82 |
|---|---|---|---|---|---|
|
| 83 |
| `sliding_window` | A | preset | keep last N messages middleware | `dropped_messages` | F9/F10 β recency-only memory |
|
| 84 |
+
| `delta_summarization` | A | preset | running summary, new-only | reuse `summary_messages` (as built; no `summary_mode` signal) | F2 cache confound |
|
| 85 |
+
| `hierarchical_summarization` | A | preset | summarize chunks then summaries | reuse `summary_messages` (as built; no `summary_levels` signal) | long-session compaction |
|
| 86 |
| `context_reset` | A | preset | fresh state seeded w/ summary | reuse `summary_messages` (summary-prompt variant) | F2; prefix rewrite |
|
| 87 |
| `prompt_compression` | A | preset | rewrite history fewer tokens | `compressed_messages`, `chars_saved` | F2 |
|
| 88 |
| `selective_retention` | A | preset | summary prompt keeps constraints/decisions | reuse `summary_messages` | quality-preserving compaction |
|
|
|
|
| 93 |
| `observation_truncation` | B | preset | head/tail tool outputs incl. KB | `truncated_tool_outputs`, `chars_saved` | F1 β tokens are in tool outputs |
|
| 94 |
| `clear_retrieval_kb` | B | preset | extend `ClearToolUsesEdit` to retrieval + KB | `cleared_tool_outputs` (now fires) | F3 β clearing excluded them |
|
| 95 |
| `retrieval_budget_{100k,30k,10k}` | B | knob | `DEFAULT_CONTEXT_TOKEN_BUDGET` | (token counts already captured) + `retrieval_budget` | F1 β direct |
|
| 96 |
+
| `kb_off` | B | toggle | `disable_kb` flag + drop KB prompt block | none as built (request flag only; no `kb_enabled` signal) | does the KB improve answers? (89%/9:1 usage) |
|
| 97 |
|
| 98 |
Stretch (post-workshop, evals.md): `temporal_graph_memory` (scorecard = `fact_update` probe), `subagent_isolation` (reframed as a tool-output-token play).
|
| 99 |
|
|
|
|
| 157 |
1. **Freeze + data discipline.** v2 is a new file; never touch v1. Gitignored; lives only in the private `ai-tutor-data` HF dataset under `eval/`.
|
| 158 |
2. **Author Tier 1 first** (moderate contradiction sessions), reusing real corpus filler; model the shape on `data/eval/sessions_generated_*.jsonl` and the schema in `data/eval/README.md`. Add Tier 2 long-horizon sessions only if we decide Q1 is worth the spend/workshop story.
|
| 159 |
3. **Calibrate by tokens and observed context state, not turn count.** For Q2, confirm prod summarization/compaction fired and A is no longer in the kept-recent window before the contradiction probe; if A is still visible to prod, the probe measures nothing. For Q1, smoke-run a few lengths (for example 30/45/60 turns) and confirm `full_history` cost actually diverges before running a screen.
|
| 160 |
+
4. **Wire grading** β add `contradiction` / `longhorizon_recall` / `entity_isolation` grading to `evals/judge.py` and the human rubric. *(Done β all three route through the shared `probe` rubric via per-type instruction text, not separate `RUBRICS` keys.)*
|
| 161 |
5. **Test `full_history` + `prod` + `incontext` + active `profile_memory` first.** This answers the cheap questions and tells us what, if anything, to build: does `full_history` fail contradictions (Q2), does `profile_memory` recover them through the store, and does `incontext` match `full_history`'s recall at lower all-in cost once sessions are actually long (Q1)?
|
| 162 |
6. **Build only what the data demands**, one at a time, single-axis vs prod (each new mechanism needs its own `context_stats` signal mirrored in `evals/common.py` or the probe gate will not see it fire):
|
| 163 |
- `temporal_graph_memory` β **only if** `full_history` fails Q2's contradictions.
|
|
@@ -20,7 +20,9 @@ Required local artifacts and environment:
|
|
| 20 |
- `.env` has `LANGSMITH_PROJECT=ai-tutor-app`
|
| 21 |
- `data/chroma-db-all_sources/` exists
|
| 22 |
- `data/kb/wiki/index.md` exists
|
| 23 |
-
- `curl`
|
|
|
|
|
|
|
| 24 |
|
| 25 |
Quick check:
|
| 26 |
|
|
@@ -103,8 +105,11 @@ JSON
|
|
| 103 |
```
|
| 104 |
|
| 105 |
On `enabledTools`: it is an explicit allowlist of the toggle tools for the turn.
|
| 106 |
-
|
| 107 |
-
|
|
|
|
|
|
|
|
|
|
| 108 |
|
| 109 |
- **`url_context` defaults to off in the UI** (`active: False` in `_tool_catalog`,
|
| 110 |
`app/api.py`), so a browser turn only sends it when the user opts in. Enabling
|
|
|
|
| 20 |
- `.env` has `LANGSMITH_PROJECT=ai-tutor-app`
|
| 21 |
- `data/chroma-db-all_sources/` exists
|
| 22 |
- `data/kb/wiki/index.md` exists
|
| 23 |
+
- `curl` and `jq` are on `PATH`, and the standalone `langsmith` CLI binary is
|
| 24 |
+
installed (a separate Go binary, e.g. `~/.local/bin/langsmith`; `uv sync`
|
| 25 |
+
does not provide it β `uv run` merely inherits it via `PATH`)
|
| 26 |
|
| 27 |
Quick check:
|
| 28 |
|
|
|
|
| 105 |
```
|
| 106 |
|
| 107 |
On `enabledTools`: it is an explicit allowlist of the toggle tools for the turn.
|
| 108 |
+
For this payload's DeepSeek model the catalog exposes no toggle tools at all, so
|
| 109 |
+
`[]` and omitting the field behave identically here; the notes below bite when
|
| 110 |
+
the model is Gemini 3+ or Anthropic. `[]` (as above) disables web search and URL
|
| 111 |
+
reading, so this run exercises only the corpus retrieval + KB shell citation
|
| 112 |
+
paths. Two things to know:
|
| 113 |
|
| 114 |
- **`url_context` defaults to off in the UI** (`active: False` in `_tool_catalog`,
|
| 115 |
`app/api.py`), so a browser turn only sends it when the user opts in. Enabling
|