omarsol commited on
Commit
d879dad
Β·
1 Parent(s): bd180ae

Refactor project structure and update dependencies

Browse files

- Introduced a new `config.py` file to centralize configuration and source management, replacing references to `setup.py`.
- Updated AGENTS.md to reflect changes in source configuration and clarify the single source of truth for sources.
- Removed `ipykernel` from main dependencies and added it to the development group in `pyproject.toml` and `uv.lock`.
- Adjusted import statements across various modules to utilize the new `config.py`.
- Enhanced logging practices in the application for improved clarity and reduced noise.

AGENTS.md CHANGED
@@ -23,9 +23,9 @@ ChromaDB for vectors; Cohere for embeddings/rerank; chat model is provider-confi
23
  | Hybrid retrieval | `app/chroma_rag.py` |
24
  | KB browsing sandbox + citation resolution | `app/kb_shell.py`, `app/kb_manifest.py` |
25
  | FastAPI server (`/api/chat`, `/api/tools`, `/healthz`) | `app/api.py` |
26
- | Paths, models, startup downloads | `app/setup.py` |
27
  | **Sources β€” single source of truth** | `data/scraping_scripts/source_registry.py` |
28
- | Agent tracing (LangSmith) + server logging (stdlib `logging` β†’ stdout) | `app/agent_tracing.py`, `app/setup.py` |
29
  | Data pipeline / workflows (deep guide) | `data/scraping_scripts/README.md` |
30
  | KB design + wiki maintainer workflow (deep guide) | `data/kb/MAINTAINER.md` |
31
 
@@ -59,7 +59,7 @@ Runtime guidance the agent follows is in `data/kb/AGENTS.md` (injected into the
59
 
60
  ## Sources & config
61
 
62
- `data/scraping_scripts/source_registry.py` is the **single source of truth** for sources (`SOURCE_CONFIGS`, key groupings, UI labels, defaults); `app/setup.py` re-exports them and the frontend derives the picker from it (via `/api/tools`). Docs sources ingest via the GitHub API or `llms.txt`; course sources are Notion exports. To add a source: add it to the registry (+ the relevant grouping tuples), then run the matching workflow β€” no separate UI edit needed. Models live in `setup.AVAILABLE_MODELS` (default `google-genai:gemini-3.5-flash`; also Claude Haiku 4.5; OpenAI supported in code).
63
 
64
  ## Running locally
65
 
@@ -109,4 +109,4 @@ Both Spaces need the same runtime secrets (`COHERE_API_KEY`, model provider key,
109
  - **Context generation uses Gemini**; embeddings/rerank use **Cohere**; the chat model is provider-configurable. OpenAI is required only when explicitly selected.
110
  - `data/kb/` and `data/chroma-db-all_sources/` are build artifacts β€” never commit or hand-edit; regenerate or re-download.
111
  - Two `AGENTS.md` files: this root one (repo dev guidance) vs `data/kb/AGENTS.md` (generated runtime KB rules).
112
- - Source config lives in `source_registry.py`, not `setup.py` / `process_md_files.py` (older docs were wrong).
 
23
  | Hybrid retrieval | `app/chroma_rag.py` |
24
  | KB browsing sandbox + citation resolution | `app/kb_shell.py`, `app/kb_manifest.py` |
25
  | FastAPI server (`/api/chat`, `/api/tools`, `/healthz`) | `app/api.py` |
26
+ | Paths, models, startup downloads | `app/config.py` |
27
  | **Sources β€” single source of truth** | `data/scraping_scripts/source_registry.py` |
28
+ | Agent tracing (LangSmith) + server logging (stdlib `logging` β†’ stdout) | `app/agent_tracing.py`, `app/config.py` |
29
  | Data pipeline / workflows (deep guide) | `data/scraping_scripts/README.md` |
30
  | KB design + wiki maintainer workflow (deep guide) | `data/kb/MAINTAINER.md` |
31
 
 
59
 
60
  ## Sources & config
61
 
62
+ `data/scraping_scripts/source_registry.py` is the **single source of truth** for sources (`SOURCE_CONFIGS`, key groupings, UI labels, defaults); `app/config.py` re-exports them and the frontend derives the picker from it (via `/api/tools`). Docs sources ingest via the GitHub API or `llms.txt`; course sources are Notion exports. To add a source: add it to the registry (+ the relevant grouping tuples), then run the matching workflow β€” no separate UI edit needed. Models live in `config.AVAILABLE_MODELS` (default `google-genai:gemini-3.5-flash`; also Claude Haiku 4.5; OpenAI supported in code).
63
 
64
  ## Running locally
65
 
 
109
  - **Context generation uses Gemini**; embeddings/rerank use **Cohere**; the chat model is provider-configurable. OpenAI is required only when explicitly selected.
110
  - `data/kb/` and `data/chroma-db-all_sources/` are build artifacts β€” never commit or hand-edit; regenerate or re-download.
111
  - Two `AGENTS.md` files: this root one (repo dev guidance) vs `data/kb/AGENTS.md` (generated runtime KB rules).
112
+ - Source config lives in `source_registry.py`, not `app/config.py` / `process_md_files.py` (older docs were wrong).
app/api.py CHANGED
@@ -23,7 +23,7 @@ from .chat_service import (
23
  warm_up_retriever,
24
  )
25
  from .chat_types import ChatEvent, ChatRequest, ChatTurn
26
- from .setup import (
27
  AVAILABLE_MODELS,
28
  AVAILABLE_SOURCES,
29
  AVAILABLE_SOURCES_UI,
 
23
  warm_up_retriever,
24
  )
25
  from .chat_types import ChatEvent, ChatRequest, ChatTurn
26
+ from .config import (
27
  AVAILABLE_MODELS,
28
  AVAILABLE_SOURCES,
29
  AVAILABLE_SOURCES_UI,
app/chat_service.py CHANGED
@@ -44,7 +44,12 @@ from .kb_manifest import (
44
  source_match_payload,
45
  )
46
  from .prompts import build_system_prompt
47
- from .setup import (
 
 
 
 
 
48
  BM25_INDEX_PATH,
49
  COURSE_SOURCE_KEYS,
50
  DEFAULT_SELECTED_SOURCE_KEYS,
@@ -513,28 +518,6 @@ def is_anthropic_model(model_name: str) -> bool:
513
  return provider == "anthropic"
514
 
515
 
516
- def extract_thought_summaries(content: Any) -> list[str]:
517
- if not isinstance(content, list):
518
- return []
519
-
520
- thoughts: list[str] = []
521
- for item in content:
522
- if not hasattr(item, "get"):
523
- continue
524
-
525
- item_type = item.get("type")
526
- if item_type == "thinking":
527
- thought = str(item.get("thinking", "")).strip()
528
- elif item_type == "reasoning":
529
- thought = str(item.get("reasoning", "")).strip()
530
- else:
531
- continue
532
-
533
- if thought:
534
- thoughts.append(thought)
535
- return thoughts
536
-
537
-
538
  def format_tool_args(args: Any) -> str:
539
  if isinstance(args, dict):
540
  query = str(args.get("query", "")).strip()
@@ -795,251 +778,6 @@ def agent_run_config(
795
  return config
796
 
797
 
798
- def extract_web_search_queries(response_metadata: Any) -> list[str]:
799
- """Pull the queries Gemini ran against google_search from grounding metadata."""
800
- if not isinstance(response_metadata, dict):
801
- return []
802
- grounding = response_metadata.get("grounding_metadata") or {}
803
- queries = grounding.get("web_search_queries") or []
804
- return [str(q).strip() for q in queries if isinstance(q, str) and str(q).strip()]
805
-
806
-
807
- def extract_grounding_source_matches(
808
- response_metadata: Any,
809
- matches_by_doc_id: dict[str, SourceMatch],
810
- ) -> list[SourceMatch]:
811
- """Turn Gemini grounding metadata into source matches (deduped by URI)."""
812
- if not isinstance(response_metadata, dict):
813
- return []
814
- grounding = response_metadata.get("grounding_metadata") or {}
815
- chunks = grounding.get("grounding_chunks") or []
816
- if not chunks:
817
- return []
818
-
819
- confidence_by_index: dict[int, float] = {}
820
- for support in grounding.get("grounding_supports") or []:
821
- indices = support.get("grounding_chunk_indices") or []
822
- scores = support.get("confidence_scores") or []
823
- for idx, score in zip(indices, scores):
824
- if not isinstance(idx, int):
825
- continue
826
- numeric = float(score) if isinstance(score, (int, float)) else 0.0
827
- if numeric > confidence_by_index.get(idx, 0.0):
828
- confidence_by_index[idx] = numeric
829
-
830
- updated: list[SourceMatch] = []
831
- for idx, chunk in enumerate(chunks):
832
- web = (chunk or {}).get("web") or {}
833
- uri = str(web.get("uri") or "").strip()
834
- if not uri:
835
- continue
836
- title = str(web.get("title") or uri).strip()
837
- doc_id = f"google_search::{uri}"
838
- if doc_id in matches_by_doc_id:
839
- continue
840
- score = confidence_by_index.get(idx, 1.0)
841
- source_match = SourceMatch(
842
- doc_id=doc_id,
843
- title=title,
844
- url=uri,
845
- source_key="google_search",
846
- source_label="Web",
847
- score=score,
848
- group="web",
849
- )
850
- matches_by_doc_id[doc_id] = source_match
851
- updated.append(source_match)
852
- return updated
853
-
854
-
855
- GOOGLE_SEARCH_TOOL_NAME = "google_search"
856
-
857
-
858
- class GoogleSearchActivity:
859
- """Surface Gemini's server-side google_search activity as tool events.
860
-
861
- Gemini reports search grounding via response metadata instead of tool
862
- messages, so queries and grounding results are accumulated from every
863
- metadata payload and exposed as a single synthetic tool call per turn.
864
- """
865
-
866
- def __init__(self, message_id: str, web_evidence: dict[str, SourceMatch]) -> None:
867
- self._message_id = message_id
868
- self._web_evidence = web_evidence
869
- self._call_id = ""
870
- self._queries: list[str] = []
871
- self._match_count = 0
872
-
873
- def observe(self, response_metadata: Any) -> ChatEvent | None:
874
- """Record metadata; return a tool_call_started event on first activity."""
875
- new_queries = [
876
- q
877
- for q in extract_web_search_queries(response_metadata)
878
- if q not in self._queries
879
- ]
880
- new_grounding = extract_grounding_source_matches(
881
- response_metadata,
882
- self._web_evidence,
883
- )
884
- started: ChatEvent | None = None
885
- if (new_queries or new_grounding) and not self._call_id:
886
- self._call_id = uuid4().hex
887
- joined = "; ".join(new_queries)
888
- started = ChatEvent(
889
- "tool_call_started",
890
- {
891
- "message_id": self._message_id,
892
- "call_id": self._call_id,
893
- "tool_name": GOOGLE_SEARCH_TOOL_NAME,
894
- "args": {"query": joined},
895
- "args_text": joined,
896
- },
897
- )
898
- self._queries.extend(new_queries)
899
- self._match_count += len(new_grounding)
900
- return started
901
-
902
- def completed_event(self) -> ChatEvent | None:
903
- if not self._call_id:
904
- return None
905
- joined = "; ".join(self._queries)
906
- if self._match_count == 0:
907
- output_text = "Google search ran but returned no grounding results."
908
- else:
909
- plural = "" if self._match_count == 1 else "s"
910
- output_text = (
911
- f"Google search returned {self._match_count} web result{plural}."
912
- )
913
- return ChatEvent(
914
- "tool_call_completed",
915
- {
916
- "message_id": self._message_id,
917
- "call_id": self._call_id,
918
- "tool_name": GOOGLE_SEARCH_TOOL_NAME,
919
- "args": {"query": joined},
920
- "args_text": joined,
921
- "output_text": output_text,
922
- },
923
- )
924
-
925
-
926
- ANTHROPIC_SERVER_TOOL_NAMES = frozenset({"web_search", "web_fetch"})
927
- ANTHROPIC_RESULT_BLOCK_TYPES = {
928
- "web_search_tool_result": ("web_search", "Web"),
929
- "web_fetch_tool_result": ("web_fetch", "Web page"),
930
- }
931
-
932
-
933
- def extract_anthropic_source_matches(
934
- content: Any,
935
- matches_by_doc_id: dict[str, SourceMatch],
936
- ) -> tuple[dict[str, list[SourceMatch]], dict[str, dict[str, Any]]]:
937
- """Parse Claude's server-side web tool invocations and their results.
938
-
939
- Scans ``message.content`` for three kinds of blocks emitted when Claude
940
- runs the built-in ``web_search`` / ``web_fetch`` tools:
941
-
942
- * ``tool_use`` β€” the model's call (id, name, input args)
943
- * ``web_search_tool_result`` / ``web_fetch_tool_result`` β€” the server's
944
- response, keyed by ``tool_use_id``
945
- * ``text`` blocks with ``citations`` β€” fallback for citations without a
946
- matching result block
947
-
948
- Returns ``(matches_by_tool_use_id, tool_use_index)`` where
949
- ``tool_use_index`` maps tool_use id β†’ ``{"name", "args"}`` so the caller
950
- can emit ``tool_call_started`` events with the right metadata.
951
- ``langchain-anthropic`` does not always surface server-side tool_use in
952
- ``AIMessage.tool_calls``, so we read them off the content blocks directly.
953
- """
954
- if not isinstance(content, list):
955
- return {}, {}
956
-
957
- updates: dict[str, list[SourceMatch]] = {}
958
- tool_use_index: dict[str, dict[str, Any]] = {}
959
-
960
- for block in content:
961
- if not hasattr(block, "get"):
962
- continue
963
-
964
- block_type = block.get("type")
965
-
966
- if block_type in ("server_tool_use", "tool_use"):
967
- tool_use_id = str(block.get("id") or "")
968
- tool_name = str(block.get("name") or "")
969
- if tool_use_id and tool_name in ANTHROPIC_SERVER_TOOL_NAMES:
970
- args = block.get("input") or {}
971
- if not args:
972
- partial = block.get("partial_json")
973
- if isinstance(partial, str) and partial.strip():
974
- try:
975
- parsed = json.loads(partial)
976
- except json.JSONDecodeError:
977
- parsed = None
978
- if isinstance(parsed, dict):
979
- args = parsed
980
- tool_use_index[tool_use_id] = {
981
- "id": tool_use_id,
982
- "name": tool_name,
983
- "args": args,
984
- }
985
- continue
986
-
987
- mapping = ANTHROPIC_RESULT_BLOCK_TYPES.get(block_type)
988
- if mapping:
989
- source_key, source_label = mapping
990
- tool_use_id = str(block.get("tool_use_id") or "")
991
- results = block.get("content") or []
992
- if not isinstance(results, list):
993
- continue
994
- for result in results:
995
- if not hasattr(result, "get"):
996
- continue
997
- url = str(result.get("url") or "").strip()
998
- if not url:
999
- continue
1000
- title = str(result.get("title") or url).strip()
1001
- doc_id = f"{source_key}::{url}"
1002
- if doc_id in matches_by_doc_id:
1003
- continue
1004
- source_match = SourceMatch(
1005
- doc_id=doc_id,
1006
- title=title,
1007
- url=url,
1008
- source_key=source_key,
1009
- source_label=source_label,
1010
- score=1.0,
1011
- group="web",
1012
- )
1013
- matches_by_doc_id[doc_id] = source_match
1014
- updates.setdefault(tool_use_id, []).append(source_match)
1015
- continue
1016
-
1017
- if block_type == "text":
1018
- for citation in block.get("citations") or []:
1019
- if not hasattr(citation, "get"):
1020
- continue
1021
- url = str(citation.get("url") or "").strip()
1022
- if not url:
1023
- continue
1024
- title = str(citation.get("title") or url).strip()
1025
- doc_id = f"web_search::{url}"
1026
- if doc_id in matches_by_doc_id:
1027
- continue
1028
- source_match = SourceMatch(
1029
- doc_id=doc_id,
1030
- title=title,
1031
- url=url,
1032
- source_key="web_search",
1033
- source_label="Web",
1034
- score=1.0,
1035
- group="web",
1036
- )
1037
- matches_by_doc_id[doc_id] = source_match
1038
- updates.setdefault("", []).append(source_match)
1039
-
1040
- return updates, tool_use_index
1041
-
1042
-
1043
  async def stream_chat(request: ChatRequest) -> AsyncIterator[ChatEvent]:
1044
  normalized_history = normalize_history(request.history)
1045
  retrieval_evidence: dict[str, SourceMatch] = {}
 
44
  source_match_payload,
45
  )
46
  from .prompts import build_system_prompt
47
+ from .provider_events import (
48
+ GoogleSearchActivity,
49
+ extract_anthropic_source_matches,
50
+ extract_thought_summaries,
51
+ )
52
+ from .config import (
53
  BM25_INDEX_PATH,
54
  COURSE_SOURCE_KEYS,
55
  DEFAULT_SELECTED_SOURCE_KEYS,
 
518
  return provider == "anthropic"
519
 
520
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
521
  def format_tool_args(args: Any) -> str:
522
  if isinstance(args, dict):
523
  query = str(args.get("query", "")).strip()
 
778
  return config
779
 
780
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
781
  async def stream_chat(request: ChatRequest) -> AsyncIterator[ChatEvent]:
782
  normalized_history = normalize_history(request.history)
783
  retrieval_evidence: dict[str, SourceMatch] = {}
app/{setup.py β†’ config.py} RENAMED
File without changes
app/kb_manifest.py CHANGED
@@ -8,7 +8,7 @@ from pathlib import Path
8
  from typing import Any
9
 
10
  from .chat_types import SourceMatch
11
- from .setup import COURSE_SOURCE_KEYS, KB_DIR, SOURCE_KEY_TO_LABEL
12
 
13
  KB_DOC_SCHEME_RE = re.compile(r"^kb://doc/(?P<doc_id>[^)\]\s]+)$")
14
  RAW_PATH_RE = re.compile(r"(?:data/kb/)?raw/[^\s)\]>,:]+?\.(?:mdx|md)")
 
8
  from typing import Any
9
 
10
  from .chat_types import SourceMatch
11
+ from .config import COURSE_SOURCE_KEYS, KB_DIR, SOURCE_KEY_TO_LABEL
12
 
13
  KB_DOC_SCHEME_RE = re.compile(r"^kb://doc/(?P<doc_id>[^)\]\s]+)$")
14
  RAW_PATH_RE = re.compile(r"(?:data/kb/)?raw/[^\s)\]>,:]+?\.(?:mdx|md)")
app/provider_events.py ADDED
@@ -0,0 +1,283 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Provider-specific response parsing for the chat stream.
2
+
3
+ Gemini and Anthropic surface server-side tool activity (web search, URL
4
+ fetch) and reasoning through provider-shaped response metadata and content
5
+ blocks rather than regular tool messages. This module turns those payloads
6
+ into the app's neutral `ChatEvent` / `SourceMatch` shapes so
7
+ `chat_service.stream_chat` stays provider-agnostic.
8
+ """
9
+
10
+ from __future__ import annotations
11
+
12
+ import json
13
+ from typing import Any
14
+ from uuid import uuid4
15
+
16
+ from .chat_types import ChatEvent, SourceMatch
17
+
18
+
19
+ def extract_thought_summaries(content: Any) -> list[str]:
20
+ if not isinstance(content, list):
21
+ return []
22
+
23
+ thoughts: list[str] = []
24
+ for item in content:
25
+ if not hasattr(item, "get"):
26
+ continue
27
+
28
+ item_type = item.get("type")
29
+ if item_type == "thinking":
30
+ thought = str(item.get("thinking", "")).strip()
31
+ elif item_type == "reasoning":
32
+ thought = str(item.get("reasoning", "")).strip()
33
+ else:
34
+ continue
35
+
36
+ if thought:
37
+ thoughts.append(thought)
38
+ return thoughts
39
+
40
+
41
+ def extract_web_search_queries(response_metadata: Any) -> list[str]:
42
+ """Pull the queries Gemini ran against google_search from grounding metadata."""
43
+ if not isinstance(response_metadata, dict):
44
+ return []
45
+ grounding = response_metadata.get("grounding_metadata") or {}
46
+ queries = grounding.get("web_search_queries") or []
47
+ return [str(q).strip() for q in queries if isinstance(q, str) and str(q).strip()]
48
+
49
+
50
+ def extract_grounding_source_matches(
51
+ response_metadata: Any,
52
+ matches_by_doc_id: dict[str, SourceMatch],
53
+ ) -> list[SourceMatch]:
54
+ """Turn Gemini grounding metadata into source matches (deduped by URI)."""
55
+ if not isinstance(response_metadata, dict):
56
+ return []
57
+ grounding = response_metadata.get("grounding_metadata") or {}
58
+ chunks = grounding.get("grounding_chunks") or []
59
+ if not chunks:
60
+ return []
61
+
62
+ confidence_by_index: dict[int, float] = {}
63
+ for support in grounding.get("grounding_supports") or []:
64
+ indices = support.get("grounding_chunk_indices") or []
65
+ scores = support.get("confidence_scores") or []
66
+ for idx, score in zip(indices, scores):
67
+ if not isinstance(idx, int):
68
+ continue
69
+ numeric = float(score) if isinstance(score, (int, float)) else 0.0
70
+ if numeric > confidence_by_index.get(idx, 0.0):
71
+ confidence_by_index[idx] = numeric
72
+
73
+ updated: list[SourceMatch] = []
74
+ for idx, chunk in enumerate(chunks):
75
+ web = (chunk or {}).get("web") or {}
76
+ uri = str(web.get("uri") or "").strip()
77
+ if not uri:
78
+ continue
79
+ title = str(web.get("title") or uri).strip()
80
+ doc_id = f"google_search::{uri}"
81
+ if doc_id in matches_by_doc_id:
82
+ continue
83
+ score = confidence_by_index.get(idx, 1.0)
84
+ source_match = SourceMatch(
85
+ doc_id=doc_id,
86
+ title=title,
87
+ url=uri,
88
+ source_key="google_search",
89
+ source_label="Web",
90
+ score=score,
91
+ group="web",
92
+ )
93
+ matches_by_doc_id[doc_id] = source_match
94
+ updated.append(source_match)
95
+ return updated
96
+
97
+
98
+ GOOGLE_SEARCH_TOOL_NAME = "google_search"
99
+
100
+
101
+ class GoogleSearchActivity:
102
+ """Surface Gemini's server-side google_search activity as tool events.
103
+
104
+ Gemini reports search grounding via response metadata instead of tool
105
+ messages, so queries and grounding results are accumulated from every
106
+ metadata payload and exposed as a single synthetic tool call per turn.
107
+ """
108
+
109
+ def __init__(self, message_id: str, web_evidence: dict[str, SourceMatch]) -> None:
110
+ self._message_id = message_id
111
+ self._web_evidence = web_evidence
112
+ self._call_id = ""
113
+ self._queries: list[str] = []
114
+ self._match_count = 0
115
+
116
+ def observe(self, response_metadata: Any) -> ChatEvent | None:
117
+ """Record metadata; return a tool_call_started event on first activity."""
118
+ new_queries = [
119
+ q
120
+ for q in extract_web_search_queries(response_metadata)
121
+ if q not in self._queries
122
+ ]
123
+ new_grounding = extract_grounding_source_matches(
124
+ response_metadata,
125
+ self._web_evidence,
126
+ )
127
+ started: ChatEvent | None = None
128
+ if (new_queries or new_grounding) and not self._call_id:
129
+ self._call_id = uuid4().hex
130
+ joined = "; ".join(new_queries)
131
+ started = ChatEvent(
132
+ "tool_call_started",
133
+ {
134
+ "message_id": self._message_id,
135
+ "call_id": self._call_id,
136
+ "tool_name": GOOGLE_SEARCH_TOOL_NAME,
137
+ "args": {"query": joined},
138
+ "args_text": joined,
139
+ },
140
+ )
141
+ self._queries.extend(new_queries)
142
+ self._match_count += len(new_grounding)
143
+ return started
144
+
145
+ def completed_event(self) -> ChatEvent | None:
146
+ if not self._call_id:
147
+ return None
148
+ joined = "; ".join(self._queries)
149
+ if self._match_count == 0:
150
+ output_text = "Google search ran but returned no grounding results."
151
+ else:
152
+ plural = "" if self._match_count == 1 else "s"
153
+ output_text = (
154
+ f"Google search returned {self._match_count} web result{plural}."
155
+ )
156
+ return ChatEvent(
157
+ "tool_call_completed",
158
+ {
159
+ "message_id": self._message_id,
160
+ "call_id": self._call_id,
161
+ "tool_name": GOOGLE_SEARCH_TOOL_NAME,
162
+ "args": {"query": joined},
163
+ "args_text": joined,
164
+ "output_text": output_text,
165
+ },
166
+ )
167
+
168
+
169
+ ANTHROPIC_SERVER_TOOL_NAMES = frozenset({"web_search", "web_fetch"})
170
+ ANTHROPIC_RESULT_BLOCK_TYPES = {
171
+ "web_search_tool_result": ("web_search", "Web"),
172
+ "web_fetch_tool_result": ("web_fetch", "Web page"),
173
+ }
174
+
175
+
176
+ def extract_anthropic_source_matches(
177
+ content: Any,
178
+ matches_by_doc_id: dict[str, SourceMatch],
179
+ ) -> tuple[dict[str, list[SourceMatch]], dict[str, dict[str, Any]]]:
180
+ """Parse Claude's server-side web tool invocations and their results.
181
+
182
+ Scans ``message.content`` for three kinds of blocks emitted when Claude
183
+ runs the built-in ``web_search`` / ``web_fetch`` tools:
184
+
185
+ * ``tool_use`` β€” the model's call (id, name, input args)
186
+ * ``web_search_tool_result`` / ``web_fetch_tool_result`` β€” the server's
187
+ response, keyed by ``tool_use_id``
188
+ * ``text`` blocks with ``citations`` β€” fallback for citations without a
189
+ matching result block
190
+
191
+ Returns ``(matches_by_tool_use_id, tool_use_index)`` where
192
+ ``tool_use_index`` maps tool_use id β†’ ``{"name", "args"}`` so the caller
193
+ can emit ``tool_call_started`` events with the right metadata.
194
+ ``langchain-anthropic`` does not always surface server-side tool_use in
195
+ ``AIMessage.tool_calls``, so we read them off the content blocks directly.
196
+ """
197
+ if not isinstance(content, list):
198
+ return {}, {}
199
+
200
+ updates: dict[str, list[SourceMatch]] = {}
201
+ tool_use_index: dict[str, dict[str, Any]] = {}
202
+
203
+ for block in content:
204
+ if not hasattr(block, "get"):
205
+ continue
206
+
207
+ block_type = block.get("type")
208
+
209
+ if block_type in ("server_tool_use", "tool_use"):
210
+ tool_use_id = str(block.get("id") or "")
211
+ tool_name = str(block.get("name") or "")
212
+ if tool_use_id and tool_name in ANTHROPIC_SERVER_TOOL_NAMES:
213
+ args = block.get("input") or {}
214
+ if not args:
215
+ partial = block.get("partial_json")
216
+ if isinstance(partial, str) and partial.strip():
217
+ try:
218
+ parsed = json.loads(partial)
219
+ except json.JSONDecodeError:
220
+ parsed = None
221
+ if isinstance(parsed, dict):
222
+ args = parsed
223
+ tool_use_index[tool_use_id] = {
224
+ "id": tool_use_id,
225
+ "name": tool_name,
226
+ "args": args,
227
+ }
228
+ continue
229
+
230
+ mapping = ANTHROPIC_RESULT_BLOCK_TYPES.get(block_type)
231
+ if mapping:
232
+ source_key, source_label = mapping
233
+ tool_use_id = str(block.get("tool_use_id") or "")
234
+ results = block.get("content") or []
235
+ if not isinstance(results, list):
236
+ continue
237
+ for result in results:
238
+ if not hasattr(result, "get"):
239
+ continue
240
+ url = str(result.get("url") or "").strip()
241
+ if not url:
242
+ continue
243
+ title = str(result.get("title") or url).strip()
244
+ doc_id = f"{source_key}::{url}"
245
+ if doc_id in matches_by_doc_id:
246
+ continue
247
+ source_match = SourceMatch(
248
+ doc_id=doc_id,
249
+ title=title,
250
+ url=url,
251
+ source_key=source_key,
252
+ source_label=source_label,
253
+ score=1.0,
254
+ group="web",
255
+ )
256
+ matches_by_doc_id[doc_id] = source_match
257
+ updates.setdefault(tool_use_id, []).append(source_match)
258
+ continue
259
+
260
+ if block_type == "text":
261
+ for citation in block.get("citations") or []:
262
+ if not hasattr(citation, "get"):
263
+ continue
264
+ url = str(citation.get("url") or "").strip()
265
+ if not url:
266
+ continue
267
+ title = str(citation.get("title") or url).strip()
268
+ doc_id = f"web_search::{url}"
269
+ if doc_id in matches_by_doc_id:
270
+ continue
271
+ source_match = SourceMatch(
272
+ doc_id=doc_id,
273
+ title=title,
274
+ url=url,
275
+ source_key="web_search",
276
+ source_label="Web",
277
+ score=1.0,
278
+ group="web",
279
+ )
280
+ matches_by_doc_id[doc_id] = source_match
281
+ updates.setdefault("", []).append(source_match)
282
+
283
+ return updates, tool_use_index
data/scraping_scripts/README.md CHANGED
@@ -189,7 +189,7 @@ places, both reading from the canonical template at
189
 
190
  - `data.scraping_scripts.update_kb_wiki.write_agents_md` β€” rewrites it during
191
  every `update_docs_workflow.py` run, before uploading to HuggingFace.
192
- - `app.setup.ensure_kb_agents_md` β€” rewrites it on every runtime startup,
193
  after `ensure_local_vector_db` downloads the snapshot. This catches the
194
  case where the HF snapshot has a stale AGENTS.md (e.g. uploaded before a
195
  template change landed in git) and ensures the live file always matches
@@ -207,7 +207,7 @@ place on startup either way).
207
  `data/all_sources_data.jsonl`) and `data.scraping_scripts.update_kb_wiki`.
208
  - **Uploaded** to `towardsai-tutors/ai-tutor-vector-db` by
209
  `data.scraping_scripts.upload_dbs_to_hf` (already includes `kb/**`).
210
- - **Downloaded** by `app.setup.ensure_local_vector_db` on the first
211
  chatbot start (or any start where the local KB is missing).
212
 
213
  Treat it the same way as `data/chroma-db-all_sources/`: never commit it,
 
189
 
190
  - `data.scraping_scripts.update_kb_wiki.write_agents_md` β€” rewrites it during
191
  every `update_docs_workflow.py` run, before uploading to HuggingFace.
192
+ - `app.config.ensure_kb_agents_md` β€” rewrites it on every runtime startup,
193
  after `ensure_local_vector_db` downloads the snapshot. This catches the
194
  case where the HF snapshot has a stale AGENTS.md (e.g. uploaded before a
195
  template change landed in git) and ensures the live file always matches
 
207
  `data/all_sources_data.jsonl`) and `data.scraping_scripts.update_kb_wiki`.
208
  - **Uploaded** to `towardsai-tutors/ai-tutor-vector-db` by
209
  `data.scraping_scripts.upload_dbs_to_hf` (already includes `kb/**`).
210
+ - **Downloaded** by `app.config.ensure_local_vector_db` on the first
211
  chatbot start (or any start where the local KB is missing).
212
 
213
  Treat it the same way as `data/chroma-db-all_sources/`: never commit it,
data/scraping_scripts/add_course_workflow.py CHANGED
@@ -435,7 +435,7 @@ def update_ui_files(course_name: str) -> None:
435
  return
436
 
437
  logger.info(
438
- "%s is configured in source_registry.py; no setup.py/main.py edits needed.",
439
  course_name,
440
  )
441
 
 
435
  return
436
 
437
  logger.info(
438
+ "%s is configured in source_registry.py; no app-code edits needed.",
439
  course_name,
440
  )
441
 
{app β†’ notebooks}/generate_qa_dataset.ipynb RENAMED
File without changes
pyproject.toml CHANGED
@@ -12,7 +12,6 @@ dependencies = [
12
  "google-genai",
13
  "hf-xet",
14
  "huggingface-hub",
15
- "ipykernel",
16
  "langchain",
17
  "langchain-anthropic",
18
  "langchain-google-genai",
@@ -31,6 +30,7 @@ dependencies = [
31
  [dependency-groups]
32
  dev = [
33
  "httpx",
 
34
  "pre-commit",
35
  "pytest>=9.0.3",
36
  "ruff>=0.15.10",
 
12
  "google-genai",
13
  "hf-xet",
14
  "huggingface-hub",
 
15
  "langchain",
16
  "langchain-anthropic",
17
  "langchain-google-genai",
 
30
  [dependency-groups]
31
  dev = [
32
  "httpx",
33
+ "ipykernel",
34
  "pre-commit",
35
  "pytest>=9.0.3",
36
  "ruff>=0.15.10",
uv.lock CHANGED
@@ -22,7 +22,6 @@ dependencies = [
22
  { name = "google-genai" },
23
  { name = "hf-xet" },
24
  { name = "huggingface-hub" },
25
- { name = "ipykernel" },
26
  { name = "langchain" },
27
  { name = "langchain-anthropic" },
28
  { name = "langchain-google-genai" },
@@ -41,6 +40,7 @@ dependencies = [
41
  [package.dev-dependencies]
42
  dev = [
43
  { name = "httpx" },
 
44
  { name = "pre-commit" },
45
  { name = "pytest" },
46
  { name = "ruff" },
@@ -55,7 +55,6 @@ requires-dist = [
55
  { name = "google-genai" },
56
  { name = "hf-xet" },
57
  { name = "huggingface-hub" },
58
- { name = "ipykernel" },
59
  { name = "langchain" },
60
  { name = "langchain-anthropic" },
61
  { name = "langchain-google-genai" },
@@ -74,6 +73,7 @@ requires-dist = [
74
  [package.metadata.requires-dev]
75
  dev = [
76
  { name = "httpx" },
 
77
  { name = "pre-commit" },
78
  { name = "pytest", specifier = ">=9.0.3" },
79
  { name = "ruff", specifier = ">=0.15.10" },
 
22
  { name = "google-genai" },
23
  { name = "hf-xet" },
24
  { name = "huggingface-hub" },
 
25
  { name = "langchain" },
26
  { name = "langchain-anthropic" },
27
  { name = "langchain-google-genai" },
 
40
  [package.dev-dependencies]
41
  dev = [
42
  { name = "httpx" },
43
+ { name = "ipykernel" },
44
  { name = "pre-commit" },
45
  { name = "pytest" },
46
  { name = "ruff" },
 
55
  { name = "google-genai" },
56
  { name = "hf-xet" },
57
  { name = "huggingface-hub" },
 
58
  { name = "langchain" },
59
  { name = "langchain-anthropic" },
60
  { name = "langchain-google-genai" },
 
73
  [package.metadata.requires-dev]
74
  dev = [
75
  { name = "httpx" },
76
+ { name = "ipykernel" },
77
  { name = "pre-commit" },
78
  { name = "pytest", specifier = ">=9.0.3" },
79
  { name = "ruff", specifier = ">=0.15.10" },