omarsol Claude Fable 5 commited on
Commit
f5382e2
·
1 Parent(s): 2ece3fd

Switch the default chat model to DeepSeek V4 Flash via OpenRouter

Browse files

The first-party DeepSeek account is deliberately unfunded and the team is
standardizing on OpenRouter, whose cheap endpoints also price a typical
agentic turn ~40-70% below first-party rates: output and cache-miss input
dominate our bill, and DeepSeek's ~31x first-party cache-hit discount does
not pass through OpenRouter's third-party hosts anyway (verified against
OpenRouter's endpoint listing, 2026-08-25).

- DEFAULT_MODEL_NAME / AVAILABLE_MODELS -> openrouter:deepseek/deepseek-v4-flash;
the first-party deepseek: path stays supported for experiments
- New TutorChatOpenRouter adapter preserves OpenRouter's unified streamed
`reasoning` field (mapped into the DeepSeek-style reasoning_content slot)
and the client sends the model-agnostic reasoning toggle explicitly
- "openrouter" joins PRODUCTION_LONG_CONTEXT_PROVIDERS so prod_v2 serves
the default path instead of silently downgrading to the legacy preset
- MODEL_PRICING OpenRouter row refreshed to the verified headline endpoint
rate (checked 2026-08-25); docs, runbook, and env templates updated

Verified live: reasoning streams interleaved with run_kb_command calls,
prompt caching passes through (cached_tokens populated in context_stats),
thread continuation works, and the live E2E suite passes. The Space needs
the OPENROUTER_API_KEY secret; until it is set, key substitution serves
the Gemini fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

.env.example CHANGED
@@ -1,17 +1,18 @@
1
  # To run the AI Tutor app (FastAPI backend)
2
  # COHERE_API_KEY is always required (embeddings + rerank for retrieval).
3
  COHERE_API_KEY=...
4
- # Default chat path: DeepSeek V4 Flash through the first-party DeepSeek API,
5
- # with Gemini 3.5 Flash Lite as the local application fallback if DeepSeek fails.
6
- DEEPSEEK_API_KEY=...
7
  GEMINI_API_KEY=...
8
  # Gemini also accepts GOOGLE_API_KEY instead of GEMINI_API_KEY.
9
  GOOGLE_API_KEY=...
10
 
11
- # Optional alternate providers
 
12
  OPENAI_API_KEY=...
13
  ANTHROPIC_API_KEY=...
14
- OPENROUTER_API_KEY=...
15
 
16
  # Optional: Slack-compatible incoming-webhook URL for ops alerts (provider
17
  # outages and account-state errors like an unfunded or revoked key). Alerts
 
1
  # To run the AI Tutor app (FastAPI backend)
2
  # COHERE_API_KEY is always required (embeddings + rerank for retrieval).
3
  COHERE_API_KEY=...
4
+ # Default chat path: DeepSeek V4 Flash via OpenRouter, with Gemini 3.5 Flash
5
+ # Lite as the local application fallback if OpenRouter fails.
6
+ OPENROUTER_API_KEY=...
7
  GEMINI_API_KEY=...
8
  # Gemini also accepts GOOGLE_API_KEY instead of GEMINI_API_KEY.
9
  GOOGLE_API_KEY=...
10
 
11
+ # Optional alternate providers (DEEPSEEK_API_KEY is the retired first-party
12
+ # DeepSeek path, kept for experiments)
13
  OPENAI_API_KEY=...
14
  ANTHROPIC_API_KEY=...
15
+ DEEPSEEK_API_KEY=...
16
 
17
  # Optional: Slack-compatible incoming-webhook URL for ops alerts (provider
18
  # outages and account-state errors like an unfunded or revoked key). Alerts
AGENTS.md CHANGED
@@ -6,7 +6,7 @@ This is the **canonical, tool-agnostic** instruction file for the repo. `CLAUDE.
6
 
7
  AI tutor for applied AI, LLMs, RAG, and Python. **Agentic RAG**: a LangChain/LangGraph agent grounds answers in a curated corpus of course + library docs, can browse a local file-based knowledge base, and (optionally) search the live web. One frontend: a **Next.js** UI (`frontend/`), served by a **FastAPI** backend (`app/api.py`) that streams in the Vercel AI SDK UI-message protocol. (A Gradio UI existed historically; it was removed to keep one rendering path.)
8
 
9
- ChromaDB for vectors; Cohere for embeddings/rerank; chat model is provider-configurable (DeepSeek V4 Flash through the first-party API by default, with a rescue-only in-app fallback to Gemini 3.5 Flash Lite when a Gemini key is set; Anthropic, OpenAI, and OpenRouter-compatible models are supported in code but not user-selectable). **A conversation cannot change model mid-thread** — checkpoints store provider-native message blocks that no other provider can replay; see `build_chat_model` in `app/chat_service.py`. Python ≥3.13, managed with `uv`.
10
 
11
  ## Key URLs
12
 
@@ -62,7 +62,7 @@ Runtime guidance the agent follows is in `data/kb/AGENTS.md` (injected into the
62
 
63
  ## Sources & config
64
 
65
- `data/scraping_scripts/source_registry.py` is the **single source of truth** for sources (`SOURCE_CONFIGS`, key groupings, UI labels, defaults); `app/config.py` re-exports them and the frontend derives the picker from it (via `/api/tools`). Docs sources ingest via the GitHub API or `llms.txt`; course sources are Notion exports. To add a source: add it to the registry (+ the relevant grouping tuples), then run the matching workflow — no separate UI edit needed. Models live in `config.AVAILABLE_MODELS` (currently only the default `deepseek:deepseek-v4-flash`; Anthropic Claude Haiku 4.5 and OpenAI remain supported in code but are not selectable). **`FALLBACK_MODEL_NAME` (`google-genai:gemini-3.5-flash-lite`) is deliberately NOT in `AVAILABLE_MODELS`**: it serves real traffic as the DeepSeek rescue path but is not user-selectable. Unlike the previous fallback (`gemini-2.5-flash`, a pre-Gemini-3 model that could not combine Gemini's built-in web tools with our two custom tools and so was one web-search toggle from a 400 if ever selected), 3.5 Flash Lite is a Gemini 3-series model with "tool context circulation", so keeping it fallback-only is now a product choice (one selectable model plus a rescue path), not an API constraint. Either way the fallback never receives web tools: they are bound from the requested model (DeepSeek), which has none. `build_chat_model` in `app/chat_service.py` builds exactly ONE provider's client (it accepts `deepseek:` for the first-party API, `openrouter:` for compatible experiment models, and `ollama:` for local SLM experiments, with pricing in `app/telemetry.MODEL_PRICING`); no fallback is wired at the client layer. The rescue is turn-level in `stream_chat`: an outage-class failure (connectivity/timeout/429/5xx) before the first streamed chunk retries the turn on `FALLBACK_MODEL_NAME` through the plain-text thread branch, so the pairing is free to be ANY two provider:model strings. `resolve_served_model_name` substitutes the fallback outright when the default model's provider key is missing.
66
 
67
  Memory policy is deliberately owned by `app/memory_presets.py`, not `app/config.py`. `PRODUCTION_MEMORY_PRESET` is the single primary-policy switch, `PRODUCTION_FALLBACK_MEMORY_PRESET` covers providers that cannot use it safely, and `PRODUCTION_LONG_CONTEXT_PROVIDERS` defines the model-aware allowlist. Keep historical presets immutable because saved eval runs refer to them by name.
68
 
@@ -97,7 +97,7 @@ uv run -m data.scraping_scripts.retire_source_workflow --sources KEY [--dry-run
97
 
98
  ## Environment variables
99
 
100
- Chat runtime: `COHERE_API_KEY` (retrieval; always required), `DEEPSEEK_API_KEY` for the default model plus `GEMINI_API_KEY`/`GOOGLE_API_KEY` to enable its in-app Gemini 3.5 Flash Lite fallback, and the corresponding provider key when selecting another model (`GEMINI_API_KEY`/`GOOGLE_API_KEY`, `ANTHROPIC_API_KEY`, or `OPENAI_API_KEY`; `OPENROUTER_API_KEY` is used by OpenRouter experiment models), plus `HF_TOKEN` for first-start download of the full private bundle. Without `HF_TOKEN`, cold start falls back to the public docs-only bundle (documentation sources only, course content hidden). Optional: `LANGSMITH_*` (tracing), `AI_TUTOR_API_PORT`/`HOST`/`CORS_ALLOW_ORIGINS`, `AI_TUTOR_KB_DIR`, and `NEXT_PUBLIC_AI_TUTOR_API_BASE_URL`. `AI_TUTOR_MEMORY_PRESET` is an optional deployment-wide override, not the normal default; leave it unset to retain model-aware production selection. Data workflows also need `GITHUB_TOKEN` and `GEMINI_API_KEY`/`GOOGLE_API_KEY` (context generation). See `.env.example`.
101
 
102
  ## Deployment
103
 
@@ -106,7 +106,7 @@ One HF Space runs the image (`Dockerfile`: FastAPI + Next.js static export, `rip
106
  - **Prod — `ai-tutor-chatbot`** (public): `.github/workflows/sync-to-hf.yml` force-pushes on **every push to `main`** (docs/markdown-only and scraping-script-only pushes are skipped via `paths-ignore`). Merging to `main` IS deploying to public prod: keep `main` releasable and verify before merging (`uv run pytest`, and the live E2E suite for agent/API changes).
107
  - **Manual redeploy**: `.github/workflows/deploy-prod-to-hf.yml` (Actions tab → "Redeploy prod to Hugging Face (manual)") force-pushes a chosen ref without a new commit — for secret rotations, retrying a failed sync run, or rolling back to an older branch/tag.
108
 
109
- The Space needs the runtime secrets (`COHERE_API_KEY`, model provider key, `HF_TOKEN`, optional `LANGSMITH_*`) configured in its HF settings.
110
 
111
  ## Conventions
112
 
@@ -115,7 +115,7 @@ The Space needs the runtime secrets (`COHERE_API_KEY`, model provider key, `HF_T
115
 
116
  ## Gotchas
117
 
118
- - **One thread is one provider — a conversation cannot mix models while keeping tool outputs and thought signatures.** A thread's checkpoint stores *provider-native* messages: Gemini reasoning parts carry thought signatures, Anthropic thinking blocks carry a required cryptographic `signature`, and each provider's server-side tool calls are its own block types. Replaying one provider's history to another is a hard 400, not a degraded answer — verified both ways (Gemini history → DeepSeek 400s; Gemini history → Anthropic 400s with `messages.1.content.0.thinking.signature: Field required`). So you cannot "just send the payload to another provider": a fallback, a retry-on-another-provider, or mid-conversation switching **must not** be wired at the model/client layer (a `with_fallbacks` client only works for pairings whose primary happens to write lowest-common-denominator output — an overfit trap this repo shipped once and removed). The only safe path is `sync_thread_with_history`'s branch to a fresh thread seeded with plain text, which drops the unportable state (tool outputs, summaries, thinking) on purpose — and it is the ONE mechanism for all cross-provider movement: the turn-level outage rescue in `stream_chat`, a config repoint, a model deprecation. The frontend model picker locks after the first message for this reason. The backend enforces it independently: `sync_thread_with_history` records which provider actually served each turn (`_THREAD_PROVIDERS`, taken from the served model, since rescue and key substitution mean served ≠ requested) and branches to a fresh plain-text thread when the next turn's provider differs — so a rescued turn's Gemini-owned thread returns to DeepSeek on the next healthy turn instead of stranding on the pricier provider (roughly 3x on output; see `MODEL_PRICING` in `app/telemetry.py`).
119
  - **Provider-native web tools are bound from the requested model, and not every model can mix them with our custom tools.** Gemini needs "tool context circulation" (Gemini 3+) to combine `google_search`/`url_context` with function tools, so pre-3 Gemini gets no web toggles (`supports_gemini_tool_combination`). Anthropic combines them natively, but `allowed_callers: ["direct"]` is load-bearing — without it Haiku 4.5 400s on programmatic tool calling.
120
  - **Context generation uses Gemini**; embeddings/rerank use **Cohere**; the chat model is provider-configurable. OpenAI is required only when explicitly selected.
121
  - `data/kb/` and `data/chroma-db-all_sources/` are build artifacts — never commit or hand-edit; regenerate or re-download.
 
6
 
7
  AI tutor for applied AI, LLMs, RAG, and Python. **Agentic RAG**: a LangChain/LangGraph agent grounds answers in a curated corpus of course + library docs, can browse a local file-based knowledge base, and (optionally) search the live web. One frontend: a **Next.js** UI (`frontend/`), served by a **FastAPI** backend (`app/api.py`) that streams in the Vercel AI SDK UI-message protocol. (A Gradio UI existed historically; it was removed to keep one rendering path.)
8
 
9
+ ChromaDB for vectors; Cohere for embeddings/rerank; chat model is provider-configurable (DeepSeek V4 Flash via OpenRouter by default, with a rescue-only in-app fallback to Gemini 3.5 Flash Lite when a Gemini key is set; the first-party DeepSeek API, Anthropic, and OpenAI remain supported in code but are not user-selectable). **A conversation cannot change model mid-thread** — checkpoints store provider-native message blocks that no other provider can replay; see `build_chat_model` in `app/chat_service.py`. Python ≥3.13, managed with `uv`.
10
 
11
  ## Key URLs
12
 
 
62
 
63
  ## Sources & config
64
 
65
+ `data/scraping_scripts/source_registry.py` is the **single source of truth** for sources (`SOURCE_CONFIGS`, key groupings, UI labels, defaults); `app/config.py` re-exports them and the frontend derives the picker from it (via `/api/tools`). Docs sources ingest via the GitHub API or `llms.txt`; course sources are Notion exports. To add a source: add it to the registry (+ the relevant grouping tuples), then run the matching workflow — no separate UI edit needed. Models live in `config.AVAILABLE_MODELS` (currently only the default `openrouter:deepseek/deepseek-v4-flash` — DeepSeek V4 Flash via OpenRouter, which routes across third-party hosts; the first-party `deepseek:` path, Anthropic Claude Haiku 4.5, and OpenAI remain supported in code but are not selectable. The first-party DeepSeek account is deliberately unfunded since 2026-08, when its 402s hard-failed prod and the team standardized on OpenRouter — also cheaper per agentic turn at OpenRouter's cheap endpoints, though DeepSeek's ~31x first-party cache-hit discount does not pass through; see `MODEL_PRICING`). **`FALLBACK_MODEL_NAME` (`google-genai:gemini-3.5-flash-lite`) is deliberately NOT in `AVAILABLE_MODELS`**: it serves real traffic as the default model's rescue path but is not user-selectable. Unlike the previous fallback (`gemini-2.5-flash`, a pre-Gemini-3 model that could not combine Gemini's built-in web tools with our two custom tools and so was one web-search toggle from a 400 if ever selected), 3.5 Flash Lite is a Gemini 3-series model with "tool context circulation", so keeping it fallback-only is now a product choice (one selectable model plus a rescue path), not an API constraint. Either way the fallback never receives web tools: they are bound from the requested model (DeepSeek via OpenRouter), which has none. `build_chat_model` in `app/chat_service.py` builds exactly ONE provider's client (it accepts `openrouter:` for the OpenRouter gateway — served by `TutorChatOpenRouter`, which preserves OpenRouter's unified streamed `reasoning` field and sends its model-agnostic `reasoning` toggle — plus `deepseek:` for the first-party API and `ollama:` for local SLM experiments, with pricing in `app/telemetry.MODEL_PRICING`); no fallback is wired at the client layer. The rescue is turn-level in `stream_chat`: an outage-class failure (connectivity/timeout/429/5xx) or an account-state failure (401/402/403 — an unfunded or revoked key, see `is_account_state_error`) before the first streamed chunk retries the turn on `FALLBACK_MODEL_NAME` through the plain-text thread branch, so the pairing is free to be ANY two provider:model strings. Both failure classes page ops through `app/alerts.py` (Slack-compatible webhook via `AI_TUTOR_ALERT_WEBHOOK_URL`, throttled per class), and each turn's `context_stats` reports `requested_model`/`served_model`/`rescued` so a fallback-served turn is never invisible. `resolve_served_model_name` substitutes the fallback outright when the default model's provider key is missing.
66
 
67
  Memory policy is deliberately owned by `app/memory_presets.py`, not `app/config.py`. `PRODUCTION_MEMORY_PRESET` is the single primary-policy switch, `PRODUCTION_FALLBACK_MEMORY_PRESET` covers providers that cannot use it safely, and `PRODUCTION_LONG_CONTEXT_PROVIDERS` defines the model-aware allowlist. Keep historical presets immutable because saved eval runs refer to them by name.
68
 
 
97
 
98
  ## Environment variables
99
 
100
+ Chat runtime: `COHERE_API_KEY` (retrieval; always required), `OPENROUTER_API_KEY` for the default model plus `GEMINI_API_KEY`/`GOOGLE_API_KEY` to enable its in-app Gemini 3.5 Flash Lite fallback, and the corresponding provider key when selecting another model (`GEMINI_API_KEY`/`GOOGLE_API_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, or `DEEPSEEK_API_KEY` for the retired first-party path), plus `HF_TOKEN` for first-start download of the full private bundle. Without `HF_TOKEN`, cold start falls back to the public docs-only bundle (documentation sources only, course content hidden). Optional: `AI_TUTOR_ALERT_WEBHOOK_URL` (Slack-compatible webhook for provider outage/account-state ops alerts; unset = log-only), `LANGSMITH_*` (tracing), `AI_TUTOR_API_PORT`/`HOST`/`CORS_ALLOW_ORIGINS`, `AI_TUTOR_KB_DIR`, and `NEXT_PUBLIC_AI_TUTOR_API_BASE_URL`. `AI_TUTOR_MEMORY_PRESET` is an optional deployment-wide override, not the normal default; leave it unset to retain model-aware production selection. Data workflows also need `GITHUB_TOKEN` and `GEMINI_API_KEY`/`GOOGLE_API_KEY` (context generation). See `.env.example`.
101
 
102
  ## Deployment
103
 
 
106
  - **Prod — `ai-tutor-chatbot`** (public): `.github/workflows/sync-to-hf.yml` force-pushes on **every push to `main`** (docs/markdown-only and scraping-script-only pushes are skipped via `paths-ignore`). Merging to `main` IS deploying to public prod: keep `main` releasable and verify before merging (`uv run pytest`, and the live E2E suite for agent/API changes).
107
  - **Manual redeploy**: `.github/workflows/deploy-prod-to-hf.yml` (Actions tab → "Redeploy prod to Hugging Face (manual)") force-pushes a chosen ref without a new commit — for secret rotations, retrying a failed sync run, or rolling back to an older branch/tag.
108
 
109
+ The Space needs the runtime secrets (`COHERE_API_KEY`, `OPENROUTER_API_KEY`, `GEMINI_API_KEY`, `HF_TOKEN`, optional `LANGSMITH_*` and `AI_TUTOR_ALERT_WEBHOOK_URL`) configured in its HF settings.
110
 
111
  ## Conventions
112
 
 
115
 
116
  ## Gotchas
117
 
118
+ - **One thread is one provider — a conversation cannot mix models while keeping tool outputs and thought signatures.** A thread's checkpoint stores *provider-native* messages: Gemini reasoning parts carry thought signatures, Anthropic thinking blocks carry a required cryptographic `signature`, and each provider's server-side tool calls are its own block types. Replaying one provider's history to another is a hard 400, not a degraded answer — verified both ways (Gemini history → DeepSeek 400s; Gemini history → Anthropic 400s with `messages.1.content.0.thinking.signature: Field required`). So you cannot "just send the payload to another provider": a fallback, a retry-on-another-provider, or mid-conversation switching **must not** be wired at the model/client layer (a `with_fallbacks` client only works for pairings whose primary happens to write lowest-common-denominator output — an overfit trap this repo shipped once and removed). The only safe path is `sync_thread_with_history`'s branch to a fresh thread seeded with plain text, which drops the unportable state (tool outputs, summaries, thinking) on purpose — and it is the ONE mechanism for all cross-provider movement: the turn-level outage/account-state rescue in `stream_chat`, a config repoint, a model deprecation. The frontend model picker locks after the first message for this reason. The backend enforces it independently: `sync_thread_with_history` records which provider actually served each turn (`_THREAD_PROVIDERS`, taken from the served model, since rescue and key substitution mean served ≠ requested) and branches to a fresh plain-text thread when the next turn's provider differs — so a rescued turn's Gemini-owned thread returns to the default model on the next healthy turn instead of stranding on the pricier provider (see `MODEL_PRICING` in `app/telemetry.py`).
119
  - **Provider-native web tools are bound from the requested model, and not every model can mix them with our custom tools.** Gemini needs "tool context circulation" (Gemini 3+) to combine `google_search`/`url_context` with function tools, so pre-3 Gemini gets no web toggles (`supports_gemini_tool_combination`). Anthropic combines them natively, but `allowed_callers: ["direct"]` is load-bearing — without it Haiku 4.5 400s on programmatic tool calling.
120
  - **Context generation uses Gemini**; embeddings/rerank use **Cohere**; the chat model is provider-configurable. OpenAI is required only when explicitly selected.
121
  - `data/kb/` and `data/chroma-db-all_sources/` are build artifacts — never commit or hand-edit; regenerate or re-download.
app/chat_service.py CHANGED
@@ -44,6 +44,7 @@ from .agent_tracing import langsmith_deployment_identity
44
  from .chat_types import ChatEvent, ChatRequest, ChatTurn, SourceMatch
45
  from .alerts import send_ops_alert
46
  from .deepseek_chat import TutorChatDeepSeek
 
47
  from .memory_presets import (
48
  MemoryConfig,
49
  resolve_memory_preset,
@@ -960,6 +961,12 @@ def is_deepseek_model(model_name: str) -> bool:
960
  return provider == "deepseek"
961
 
962
 
 
 
 
 
 
 
963
  def supports_gemini_tool_combination(model_name: str) -> bool:
964
  """True when a Gemini model can mix built-in tools with function tools.
965
 
@@ -1032,12 +1039,17 @@ def _build_chat_model_client(provider_model: str, include_thoughts: bool = False
1032
  # Routing inside OpenRouter is left with its own fallbacks enabled:
1033
  # pinning a single provider with allow_fallbacks=False made batches die
1034
  # on a backend's transient 429 instead of routing around it.
1035
- return ChatOpenAI(
 
 
 
 
1036
  model=actual_model,
1037
  temperature=1,
1038
  base_url="https://openrouter.ai/api/v1",
1039
  api_key=os.environ.get("OPENROUTER_API_KEY"),
1040
  stream_usage=True,
 
1041
  # Long agentic sessions fire hundreds of calls against OpenRouter's
1042
  # shared pool; bump retries (with the client's exponential backoff)
1043
  # to ride out transient upstream 429s instead of failing a whole
@@ -2876,6 +2888,9 @@ async def stream_chat(request: ChatRequest) -> AsyncIterator[ChatEvent]:
2876
  is_google_genai_model(serving_model_name)
2877
  or is_anthropic_model(serving_model_name)
2878
  or is_deepseek_model(serving_model_name)
 
 
 
2879
  )
2880
  # Gemini streams each thought summary as one complete block; Anthropic
2881
  # streams partial fragments of a single thought. The encoder uses this
 
44
  from .chat_types import ChatEvent, ChatRequest, ChatTurn, SourceMatch
45
  from .alerts import send_ops_alert
46
  from .deepseek_chat import TutorChatDeepSeek
47
+ from .openrouter_chat import TutorChatOpenRouter
48
  from .memory_presets import (
49
  MemoryConfig,
50
  resolve_memory_preset,
 
961
  return provider == "deepseek"
962
 
963
 
964
+ def is_openrouter_model(model_name: str) -> bool:
965
+ provider_model = normalize_model_name(model_name)
966
+ provider, _, _actual_model = provider_model.partition(":")
967
+ return provider == "openrouter"
968
+
969
+
970
  def supports_gemini_tool_combination(model_name: str) -> bool:
971
  """True when a Gemini model can mix built-in tools with function tools.
972
 
 
1039
  # Routing inside OpenRouter is left with its own fallbacks enabled:
1040
  # pinning a single provider with allow_fallbacks=False made batches die
1041
  # on a backend's transient 429 instead of routing around it.
1042
+ # ``reasoning`` is OpenRouter's model-agnostic thinking toggle; sending
1043
+ # it explicitly (rather than omitting it) pins the behavior across the
1044
+ # gateway's per-host defaults, and the adapter preserves the streamed
1045
+ # thought text that plain ChatOpenAI would discard.
1046
+ return TutorChatOpenRouter(
1047
  model=actual_model,
1048
  temperature=1,
1049
  base_url="https://openrouter.ai/api/v1",
1050
  api_key=os.environ.get("OPENROUTER_API_KEY"),
1051
  stream_usage=True,
1052
+ extra_body={"reasoning": {"enabled": bool(include_thoughts)}},
1053
  # Long agentic sessions fire hundreds of calls against OpenRouter's
1054
  # shared pool; bump retries (with the client's exponential backoff)
1055
  # to ride out transient upstream 429s instead of failing a whole
 
2888
  is_google_genai_model(serving_model_name)
2889
  or is_anthropic_model(serving_model_name)
2890
  or is_deepseek_model(serving_model_name)
2891
+ # OpenRouter's unified `reasoning` request field works across the
2892
+ # models it routes; hosts that don't support it ignore it.
2893
+ or is_openrouter_model(serving_model_name)
2894
  )
2895
  # Gemini streams each thought summary as one complete block; Anthropic
2896
  # streams partial fragments of a single thought. The encoder uses this
app/config.py CHANGED
@@ -70,7 +70,14 @@ KB_AGENTS_PATH = f"{KB_DIR}/AGENTS.md"
70
  # In-git template, copied into data/kb/AGENTS.md by ensure_kb_agents_md().
71
  KB_AGENTS_TEMPLATE_PATH = "data/scraping_scripts/kb_agents_template.md"
72
  DEEPSEEK_DIRECT_MODEL_NAME = "deepseek:deepseek-v4-flash"
73
- DEFAULT_MODEL_NAME = DEEPSEEK_DIRECT_MODEL_NAME
 
 
 
 
 
 
 
74
 
75
  # Two distinct roles, deliberately kept apart:
76
  #
@@ -96,7 +103,7 @@ FALLBACK_MODEL_NAME = "google-genai:gemini-3.5-flash-lite"
96
 
97
  AVAILABLE_MODELS: tuple[dict[str, str], ...] = (
98
  {
99
- "id": DEEPSEEK_DIRECT_MODEL_NAME,
100
  "label": "DeepSeek V4 Flash",
101
  },
102
  )
@@ -294,6 +301,7 @@ __all__ = [
294
  "DEFAULT_MODEL_NAME",
295
  "DEEPSEEK_DIRECT_MODEL_NAME",
296
  "FALLBACK_MODEL_NAME",
 
297
  "BM25_INDEX_PATH",
298
  "DOCUMENT_DICT_PATH",
299
  "KB_AGENTS_PATH",
 
70
  # In-git template, copied into data/kb/AGENTS.md by ensure_kb_agents_md().
71
  KB_AGENTS_TEMPLATE_PATH = "data/scraping_scripts/kb_agents_template.md"
72
  DEEPSEEK_DIRECT_MODEL_NAME = "deepseek:deepseek-v4-flash"
73
+ # DeepSeek V4 Flash via OpenRouter is the default chat path since 2026-08:
74
+ # the first-party DeepSeek account is deliberately unfunded (its 402s
75
+ # hard-failed prod for two days), and OpenRouter's cheap endpoints price a
76
+ # typical agentic turn well below first-party rates anyway. OpenRouter serves
77
+ # this model only through third-party hosts, so DeepSeek's ~31x first-party
78
+ # cache-hit discount does NOT apply there; see MODEL_PRICING in app/telemetry.
79
+ OPENROUTER_DEEPSEEK_MODEL_NAME = "openrouter:deepseek/deepseek-v4-flash"
80
+ DEFAULT_MODEL_NAME = OPENROUTER_DEEPSEEK_MODEL_NAME
81
 
82
  # Two distinct roles, deliberately kept apart:
83
  #
 
103
 
104
  AVAILABLE_MODELS: tuple[dict[str, str], ...] = (
105
  {
106
+ "id": OPENROUTER_DEEPSEEK_MODEL_NAME,
107
  "label": "DeepSeek V4 Flash",
108
  },
109
  )
 
301
  "DEFAULT_MODEL_NAME",
302
  "DEEPSEEK_DIRECT_MODEL_NAME",
303
  "FALLBACK_MODEL_NAME",
304
+ "OPENROUTER_DEEPSEEK_MODEL_NAME",
305
  "BM25_INDEX_PATH",
306
  "DOCUMENT_DICT_PATH",
307
  "KB_AGENTS_PATH",
app/memory_presets.py CHANGED
@@ -357,7 +357,13 @@ MEMORY_PRESETS: dict[str, MemoryConfig] = {
357
  # example, has a 200k input window and cannot safely wait for an 800k trigger).
358
  PRODUCTION_MEMORY_PRESET = "prod_v2"
359
  PRODUCTION_FALLBACK_MEMORY_PRESET = "prod"
360
- PRODUCTION_LONG_CONTEXT_PROVIDERS = frozenset({"deepseek", "google-genai"})
 
 
 
 
 
 
361
 
362
  # Backward-compatible import for callers that need a single default name. New
363
  # runtime code should call resolve_memory_preset(..., model_name=...) so model
 
357
  # example, has a 200k input window and cannot safely wait for an 800k trigger).
358
  PRODUCTION_MEMORY_PRESET = "prod_v2"
359
  PRODUCTION_FALLBACK_MEMORY_PRESET = "prod"
360
+ # "openrouter" is here for the default DeepSeek-via-OpenRouter chat model.
361
+ # The gate is provider-granular, so any openrouter:* model resolves to the
362
+ # long-context production preset by default; eval runs pin presets by name,
363
+ # so in practice this only decides the served chat path.
364
+ PRODUCTION_LONG_CONTEXT_PROVIDERS = frozenset(
365
+ {"deepseek", "google-genai", "openrouter"}
366
+ )
367
 
368
  # Backward-compatible import for callers that need a single default name. New
369
  # runtime code should call resolve_memory_preset(..., model_name=...) so model
app/openrouter_chat.py ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """OpenRouter chat-model compatibility helpers.
2
+
3
+ OpenRouter normalizes reasoning across the models it routes: requests take a
4
+ unified ``reasoning`` object (``{"enabled": ...}``) regardless of the
5
+ underlying model, and streamed deltas carry the thought text in a
6
+ ``reasoning`` field (verified live against ``deepseek/deepseek-v4-flash``,
7
+ 2026-08-25). The stock ``ChatOpenAI`` wrapper drops that nonstandard delta
8
+ field, so this subclass maps it into
9
+ ``additional_kwargs["reasoning_content"]`` -- the same slot the DeepSeek
10
+ adapter uses -- and ``provider_events.extract_reasoning_deltas`` then works
11
+ unchanged for either provider.
12
+ """
13
+
14
+ from __future__ import annotations
15
+
16
+ from langchain_core.messages import AIMessageChunk
17
+ from langchain_core.outputs import ChatGenerationChunk
18
+ from langchain_openai import ChatOpenAI
19
+
20
+
21
+ class TutorChatOpenRouter(ChatOpenAI):
22
+ """ChatOpenAI that preserves OpenRouter's streamed ``reasoning`` field."""
23
+
24
+ def _convert_chunk_to_generation_chunk(
25
+ self,
26
+ chunk: dict,
27
+ default_chunk_class: type,
28
+ base_generation_info: dict | None,
29
+ ) -> ChatGenerationChunk | None:
30
+ generation_chunk = super()._convert_chunk_to_generation_chunk(
31
+ chunk,
32
+ default_chunk_class,
33
+ base_generation_info,
34
+ )
35
+ if generation_chunk is None:
36
+ return None
37
+ choices = chunk.get("choices") or []
38
+ if choices and isinstance(generation_chunk.message, AIMessageChunk):
39
+ reasoning = (choices[0].get("delta") or {}).get("reasoning")
40
+ if reasoning:
41
+ generation_chunk.message.additional_kwargs["reasoning_content"] = (
42
+ reasoning
43
+ )
44
+ return generation_chunk
app/telemetry.py CHANGED
@@ -256,11 +256,18 @@ MODEL_PRICING: dict[str, ModelPricing] = {
256
  "gpt-5.6-luna": ModelPricing(
257
  input=0.20, output=1.20, cache_read=0.02, cache_write=0.25
258
  ),
259
- # DeepSeek-V4-Flash via OpenRouter (slug "deepseek/deepseek-v4-flash").
260
- # Kept for the openrouter provider path; unused now that evals run on the
261
- # first-party API. Note OpenRouter's own list price differs ($0.09/$0.18).
 
 
 
 
 
 
 
262
  "deepseek/deepseek-v4-flash": ModelPricing(
263
- input=0.14, output=0.28, cache_read=0.0028
264
  ),
265
  # DeepSeek-V4-Flash, DeepSeek FIRST-PARTY API (base https://api.deepseek.com,
266
  # model id "deepseek-v4-flash" — the id the API echoes back, so this is the
 
256
  "gpt-5.6-luna": ModelPricing(
257
  input=0.20, output=1.20, cache_read=0.02, cache_write=0.25
258
  ),
259
+ # DeepSeek-V4-Flash via OpenRouter (slug "deepseek/deepseek-v4-flash") --
260
+ # the DEFAULT chat path since 2026-08. OpenRouter routes across ~17
261
+ # third-party hosts whose prices vary widely (input $0.068-$0.44, output
262
+ # $0.168-$1.32 per 1M); default routing favors the cheap end, so this row
263
+ # stores the headline (cheapest-endpoint) rate and expensive-host turns
264
+ # are underestimated. Caching is implicit with no write surcharge; cache
265
+ # reads bill at the endpoint's own rate (~0.25x its input price), NOT
266
+ # DeepSeek first-party's ~31x discount -- DeepSeek does not serve this
267
+ # model on OpenRouter. Checked 2026-08-25 against
268
+ # openrouter.ai/api/v1/models/deepseek/deepseek-v4-flash/endpoints.
269
  "deepseek/deepseek-v4-flash": ModelPricing(
270
+ input=0.0679, output=0.168, cache_read=0.0168
271
  ),
272
  # DeepSeek-V4-Flash, DeepSeek FIRST-PARTY API (base https://api.deepseek.com,
273
  # model id "deepseek-v4-flash" — the id the API echoes back, so this is the
frontend/components/chat-shell.test.tsx CHANGED
@@ -151,7 +151,7 @@ vi.mock("@/components/chat-message", () => ({
151
 
152
  import { ChatShell } from "@/components/chat-shell";
153
 
154
- const DEFAULT_MODEL = "deepseek:deepseek-v4-flash";
155
  const CLAUDE_MODEL = "anthropic:claude-haiku-4-5";
156
 
157
  const toolsResponse = {
 
151
 
152
  import { ChatShell } from "@/components/chat-shell";
153
 
154
+ const DEFAULT_MODEL = "openrouter:deepseek/deepseek-v4-flash";
155
  const CLAUDE_MODEL = "anthropic:claude-haiku-4-5";
156
 
157
  const toolsResponse = {
tests/manual_e2e_langsmith.md CHANGED
@@ -13,7 +13,7 @@ Run commands from the repository root.
13
  Required local artifacts and environment:
14
 
15
  - `.env` contains `COHERE_API_KEY`
16
- - `.env` contains `DEEPSEEK_API_KEY`
17
  - `.env` contains `GEMINI_API_KEY` or `GOOGLE_API_KEY`
18
  - `.env` contains `LANGSMITH_API_KEY`
19
  - `.env` has `LANGSMITH_TRACING=true`
@@ -31,7 +31,7 @@ uv run dotenv -f .env run -- python - <<'PY'
31
  import os
32
  for key in [
33
  "COHERE_API_KEY",
34
- "DEEPSEEK_API_KEY",
35
  "GEMINI_API_KEY",
36
  "GOOGLE_API_KEY",
37
  "LANGSMITH_API_KEY",
@@ -97,7 +97,7 @@ cat >/tmp/ai_tutor_e2e_payload.json <<'JSON'
97
  "transformers"
98
  ],
99
  "enabledTools": [],
100
- "model": "deepseek:deepseek-v4-flash",
101
  "includeReasoning": true,
102
  "threadId": ""
103
  }
 
13
  Required local artifacts and environment:
14
 
15
  - `.env` contains `COHERE_API_KEY`
16
+ - `.env` contains `OPENROUTER_API_KEY`
17
  - `.env` contains `GEMINI_API_KEY` or `GOOGLE_API_KEY`
18
  - `.env` contains `LANGSMITH_API_KEY`
19
  - `.env` has `LANGSMITH_TRACING=true`
 
31
  import os
32
  for key in [
33
  "COHERE_API_KEY",
34
+ "OPENROUTER_API_KEY",
35
  "GEMINI_API_KEY",
36
  "GOOGLE_API_KEY",
37
  "LANGSMITH_API_KEY",
 
97
  "transformers"
98
  ],
99
  "enabledTools": [],
100
+ "model": "openrouter:deepseek/deepseek-v4-flash",
101
  "includeReasoning": true,
102
  "threadId": ""
103
  }
tests/test_api.py CHANGED
@@ -75,7 +75,7 @@ class ApiTestCase(unittest.TestCase):
75
  )
76
  self.assertEqual(transformers["label"], "Transformers Docs")
77
  self.assertEqual(transformers["shortLabel"], "Transformers")
78
- self.assertEqual(body["model"], "deepseek:deepseek-v4-flash")
79
  # DeepSeek direct is the default model, so only local KB tools
80
  # are exposed until a provider with built-in web tools is selected.
81
  tool_keys = {tool["key"] for tool in tools}
@@ -1226,7 +1226,9 @@ def live_chat_payload(
1226
  "messages": messages,
1227
  "sourceKeys": ["peft", "transformers"],
1228
  "enabledTools": enabled_tools or [],
1229
- "model": os.getenv("LIVE_API_E2E_MODEL", "deepseek:deepseek-v4-flash"),
 
 
1230
  "includeReasoning": False,
1231
  "threadId": thread_id,
1232
  }
 
75
  )
76
  self.assertEqual(transformers["label"], "Transformers Docs")
77
  self.assertEqual(transformers["shortLabel"], "Transformers")
78
+ self.assertEqual(body["model"], "openrouter:deepseek/deepseek-v4-flash")
79
  # DeepSeek direct is the default model, so only local KB tools
80
  # are exposed until a provider with built-in web tools is selected.
81
  tool_keys = {tool["key"] for tool in tools}
 
1226
  "messages": messages,
1227
  "sourceKeys": ["peft", "transformers"],
1228
  "enabledTools": enabled_tools or [],
1229
+ "model": os.getenv(
1230
+ "LIVE_API_E2E_MODEL", "openrouter:deepseek/deepseek-v4-flash"
1231
+ ),
1232
  "includeReasoning": False,
1233
  "threadId": thread_id,
1234
  }
tests/test_chat_service.py CHANGED
@@ -8,8 +8,14 @@ from unittest.mock import MagicMock, patch
8
 
9
  from langchain_core.messages import AIMessage, AIMessageChunk, HumanMessage, ToolMessage
10
 
11
- from app.config import DEEPSEEK_DIRECT_MODEL_NAME, FALLBACK_MODEL_NAME
 
 
 
 
 
12
  from app.deepseek_chat import TutorChatDeepSeek
 
13
  from app.chat_service import (
14
  THREAD_IDLE_TTL_SECONDS,
15
  _claim_kb_command_budget,
@@ -602,6 +608,50 @@ class ChatServiceTestCase(unittest.TestCase):
602
  {"thinking": {"type": "enabled"}},
603
  )
604
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
605
  def test_deepseek_thinking_tool_call_replays_reasoning_content(self) -> None:
606
  model = TutorChatDeepSeek(
607
  model="deepseek-v4-flash",
@@ -641,7 +691,7 @@ class ChatServiceTestCase(unittest.TestCase):
641
  {"GEMINI_API_KEY": "gemini-test-key"},
642
  clear=True,
643
  ):
644
- served = resolve_served_model_name(DEEPSEEK_DIRECT_MODEL_NAME)
645
  model = build_chat_model(served)
646
 
647
  self.assertEqual(served, FALLBACK_MODEL_NAME)
@@ -650,33 +700,33 @@ class ChatServiceTestCase(unittest.TestCase):
650
  def test_resolve_served_model_name_keeps_default_when_key_present(self) -> None:
651
  with patch.dict(
652
  os.environ,
653
- {"DEEPSEEK_API_KEY": "deepseek-test-key"},
654
  clear=True,
655
  ):
656
  self.assertEqual(
657
- resolve_served_model_name(DEEPSEEK_DIRECT_MODEL_NAME),
658
- DEEPSEEK_DIRECT_MODEL_NAME,
659
  )
660
 
661
  def test_rescue_model_covers_only_the_default_model(self) -> None:
662
  with patch.dict(
663
  os.environ,
664
  {
665
- "DEEPSEEK_API_KEY": "deepseek-test-key",
666
  "GEMINI_API_KEY": "gemini-test-key",
667
  },
668
  clear=True,
669
  ):
670
- self.assertEqual(
671
- rescue_model_for(DEEPSEEK_DIRECT_MODEL_NAME), FALLBACK_MODEL_NAME
672
- )
673
- # A deliberately selected non-default model must fail visibly.
674
  self.assertIsNone(rescue_model_for("anthropic:claude-haiku-4-5"))
 
675
  with patch.dict(
676
- os.environ, {"DEEPSEEK_API_KEY": "deepseek-test-key"}, clear=True
677
  ):
678
  # No fallback key, no rescue.
679
- self.assertIsNone(rescue_model_for(DEEPSEEK_DIRECT_MODEL_NAME))
680
 
681
  def test_provider_outage_classification(self) -> None:
682
  class ServerError(Exception):
@@ -1031,7 +1081,7 @@ class ChatServiceTestCase(unittest.TestCase):
1031
  request = ChatRequest(
1032
  query="ls the kb",
1033
  source_keys=("peft",),
1034
- model_name=DEEPSEEK_DIRECT_MODEL_NAME,
1035
  include_reasoning=False,
1036
  enabled_tools=(),
1037
  )
@@ -1042,7 +1092,7 @@ class ChatServiceTestCase(unittest.TestCase):
1042
  with (
1043
  patch.dict(
1044
  os.environ,
1045
- {"DEEPSEEK_API_KEY": "k1", "GEMINI_API_KEY": "k2"},
1046
  clear=False,
1047
  ),
1048
  patch("app.chat_service.build_agent", side_effect=fake_build_agent),
@@ -1051,9 +1101,7 @@ class ChatServiceTestCase(unittest.TestCase):
1051
  events = asyncio.run(collect_events())
1052
 
1053
  # The turn restarted on the fallback model and completed normally.
1054
- self.assertEqual(
1055
- built_models, [DEEPSEEK_DIRECT_MODEL_NAME, FALLBACK_MODEL_NAME]
1056
- )
1057
  types = [event.type for event in events]
1058
  self.assertIn("thread_started", types)
1059
  self.assertIn("context_stats", types)
@@ -1075,7 +1123,7 @@ class ChatServiceTestCase(unittest.TestCase):
1075
  request = ChatRequest(
1076
  query="ls the kb",
1077
  source_keys=("peft",),
1078
- model_name=DEEPSEEK_DIRECT_MODEL_NAME,
1079
  include_reasoning=False,
1080
  enabled_tools=(),
1081
  )
@@ -1086,7 +1134,7 @@ class ChatServiceTestCase(unittest.TestCase):
1086
  with (
1087
  patch.dict(
1088
  os.environ,
1089
- {"DEEPSEEK_API_KEY": "k1", "GEMINI_API_KEY": "k2"},
1090
  clear=False,
1091
  ),
1092
  patch("app.chat_service.build_agent", side_effect=fake_build_agent),
@@ -1099,14 +1147,12 @@ class ChatServiceTestCase(unittest.TestCase):
1099
  events = asyncio.run(collect_events())
1100
 
1101
  # The unfunded account rescued onto the fallback model and paged ops.
1102
- self.assertEqual(
1103
- built_models, [DEEPSEEK_DIRECT_MODEL_NAME, FALLBACK_MODEL_NAME]
1104
- )
1105
  self.assertEqual(len(alerts), 1)
1106
  self.assertEqual(alerts[0][0], "account-state")
1107
  self.assertIn("Insufficient Balance", alerts[0][1])
1108
  stats = next(event for event in events if event.type == "context_stats")
1109
- self.assertEqual(stats.data["requested_model"], DEEPSEEK_DIRECT_MODEL_NAME)
1110
  self.assertEqual(stats.data["served_model"], FALLBACK_MODEL_NAME)
1111
  self.assertTrue(stats.data["rescued"])
1112
 
@@ -1125,7 +1171,7 @@ class ChatServiceTestCase(unittest.TestCase):
1125
  request = ChatRequest(
1126
  query="ls the kb",
1127
  source_keys=("peft",),
1128
- model_name=DEEPSEEK_DIRECT_MODEL_NAME,
1129
  include_reasoning=False,
1130
  enabled_tools=(),
1131
  )
@@ -1136,7 +1182,7 @@ class ChatServiceTestCase(unittest.TestCase):
1136
  with (
1137
  patch.dict(
1138
  os.environ,
1139
- {"DEEPSEEK_API_KEY": "k1", "GEMINI_API_KEY": "k2"},
1140
  clear=False,
1141
  ),
1142
  patch("app.chat_service.build_agent", side_effect=fake_build_agent),
@@ -1146,7 +1192,7 @@ class ChatServiceTestCase(unittest.TestCase):
1146
  asyncio.run(collect_events())
1147
 
1148
  # 4xx never reaches the fallback; the primary was the only attempt.
1149
- self.assertEqual(built_models, [DEEPSEEK_DIRECT_MODEL_NAME])
1150
 
1151
  def test_stream_chat_emits_deepseek_reasoning_content(self) -> None:
1152
  agent = FakeDeepSeekReasoningAgent([])
@@ -1155,7 +1201,7 @@ class ChatServiceTestCase(unittest.TestCase):
1155
  request = ChatRequest(
1156
  query="Find the course repository",
1157
  source_keys=("peft",),
1158
- model_name=DEEPSEEK_DIRECT_MODEL_NAME,
1159
  include_reasoning=True,
1160
  enabled_tools=(),
1161
  )
@@ -1188,7 +1234,7 @@ class ChatServiceTestCase(unittest.TestCase):
1188
  request = ChatRequest(
1189
  query="Where should I store API keys?",
1190
  source_keys=("langchain",),
1191
- model_name=DEEPSEEK_DIRECT_MODEL_NAME,
1192
  include_reasoning=False,
1193
  enabled_tools=(),
1194
  )
 
8
 
9
  from langchain_core.messages import AIMessage, AIMessageChunk, HumanMessage, ToolMessage
10
 
11
+ from app.config import (
12
+ DEEPSEEK_DIRECT_MODEL_NAME,
13
+ DEFAULT_MODEL_NAME,
14
+ FALLBACK_MODEL_NAME,
15
+ OPENROUTER_DEEPSEEK_MODEL_NAME,
16
+ )
17
  from app.deepseek_chat import TutorChatDeepSeek
18
+ from app.openrouter_chat import TutorChatOpenRouter
19
  from app.chat_service import (
20
  THREAD_IDLE_TTL_SECONDS,
21
  _claim_kb_command_budget,
 
608
  {"thinking": {"type": "enabled"}},
609
  )
610
 
611
+ def test_openrouter_model_uses_adapter_with_explicit_reasoning_toggle(
612
+ self,
613
+ ) -> None:
614
+ with patch.dict(
615
+ os.environ,
616
+ {"OPENROUTER_API_KEY": "openrouter-test-key"},
617
+ clear=True,
618
+ ):
619
+ plain = build_chat_model(OPENROUTER_DEEPSEEK_MODEL_NAME)
620
+ thinking = build_chat_model(
621
+ OPENROUTER_DEEPSEEK_MODEL_NAME, include_thoughts=True
622
+ )
623
+
624
+ self.assertIsInstance(plain, TutorChatOpenRouter)
625
+ self.assertEqual(plain.model_name, "deepseek/deepseek-v4-flash")
626
+ self.assertEqual(plain.extra_body, {"reasoning": {"enabled": False}})
627
+ self.assertEqual(thinking.extra_body, {"reasoning": {"enabled": True}})
628
+
629
+ def test_openrouter_stream_preserves_reasoning_deltas(self) -> None:
630
+ model = TutorChatOpenRouter(
631
+ model="deepseek/deepseek-v4-flash",
632
+ api_key="openrouter-test-key",
633
+ base_url="https://openrouter.ai/api/v1",
634
+ )
635
+ # The delta shape OpenRouter actually streams (verified live
636
+ # 2026-08-25): thought text arrives in `reasoning`, not DeepSeek's
637
+ # `reasoning_content`.
638
+ chunk = {
639
+ "id": "gen-1",
640
+ "model": "deepseek/deepseek-v4-flash",
641
+ "choices": [
642
+ {"index": 0, "delta": {"content": "", "reasoning": "We compute "}}
643
+ ],
644
+ }
645
+
646
+ generation_chunk = model._convert_chunk_to_generation_chunk(
647
+ chunk, AIMessageChunk, None
648
+ )
649
+
650
+ self.assertEqual(
651
+ generation_chunk.message.additional_kwargs["reasoning_content"],
652
+ "We compute ",
653
+ )
654
+
655
  def test_deepseek_thinking_tool_call_replays_reasoning_content(self) -> None:
656
  model = TutorChatDeepSeek(
657
  model="deepseek-v4-flash",
 
691
  {"GEMINI_API_KEY": "gemini-test-key"},
692
  clear=True,
693
  ):
694
+ served = resolve_served_model_name(DEFAULT_MODEL_NAME)
695
  model = build_chat_model(served)
696
 
697
  self.assertEqual(served, FALLBACK_MODEL_NAME)
 
700
  def test_resolve_served_model_name_keeps_default_when_key_present(self) -> None:
701
  with patch.dict(
702
  os.environ,
703
+ {"OPENROUTER_API_KEY": "openrouter-test-key"},
704
  clear=True,
705
  ):
706
  self.assertEqual(
707
+ resolve_served_model_name(DEFAULT_MODEL_NAME),
708
+ DEFAULT_MODEL_NAME,
709
  )
710
 
711
  def test_rescue_model_covers_only_the_default_model(self) -> None:
712
  with patch.dict(
713
  os.environ,
714
  {
715
+ "OPENROUTER_API_KEY": "openrouter-test-key",
716
  "GEMINI_API_KEY": "gemini-test-key",
717
  },
718
  clear=True,
719
  ):
720
+ self.assertEqual(rescue_model_for(DEFAULT_MODEL_NAME), FALLBACK_MODEL_NAME)
721
+ # A deliberately selected non-default model must fail visibly --
722
+ # including the retired first-party DeepSeek path.
 
723
  self.assertIsNone(rescue_model_for("anthropic:claude-haiku-4-5"))
724
+ self.assertIsNone(rescue_model_for(DEEPSEEK_DIRECT_MODEL_NAME))
725
  with patch.dict(
726
+ os.environ, {"OPENROUTER_API_KEY": "openrouter-test-key"}, clear=True
727
  ):
728
  # No fallback key, no rescue.
729
+ self.assertIsNone(rescue_model_for(DEFAULT_MODEL_NAME))
730
 
731
  def test_provider_outage_classification(self) -> None:
732
  class ServerError(Exception):
 
1081
  request = ChatRequest(
1082
  query="ls the kb",
1083
  source_keys=("peft",),
1084
+ model_name=DEFAULT_MODEL_NAME,
1085
  include_reasoning=False,
1086
  enabled_tools=(),
1087
  )
 
1092
  with (
1093
  patch.dict(
1094
  os.environ,
1095
+ {"OPENROUTER_API_KEY": "k1", "GEMINI_API_KEY": "k2"},
1096
  clear=False,
1097
  ),
1098
  patch("app.chat_service.build_agent", side_effect=fake_build_agent),
 
1101
  events = asyncio.run(collect_events())
1102
 
1103
  # The turn restarted on the fallback model and completed normally.
1104
+ self.assertEqual(built_models, [DEFAULT_MODEL_NAME, FALLBACK_MODEL_NAME])
 
 
1105
  types = [event.type for event in events]
1106
  self.assertIn("thread_started", types)
1107
  self.assertIn("context_stats", types)
 
1123
  request = ChatRequest(
1124
  query="ls the kb",
1125
  source_keys=("peft",),
1126
+ model_name=DEFAULT_MODEL_NAME,
1127
  include_reasoning=False,
1128
  enabled_tools=(),
1129
  )
 
1134
  with (
1135
  patch.dict(
1136
  os.environ,
1137
+ {"OPENROUTER_API_KEY": "k1", "GEMINI_API_KEY": "k2"},
1138
  clear=False,
1139
  ),
1140
  patch("app.chat_service.build_agent", side_effect=fake_build_agent),
 
1147
  events = asyncio.run(collect_events())
1148
 
1149
  # The unfunded account rescued onto the fallback model and paged ops.
1150
+ self.assertEqual(built_models, [DEFAULT_MODEL_NAME, FALLBACK_MODEL_NAME])
 
 
1151
  self.assertEqual(len(alerts), 1)
1152
  self.assertEqual(alerts[0][0], "account-state")
1153
  self.assertIn("Insufficient Balance", alerts[0][1])
1154
  stats = next(event for event in events if event.type == "context_stats")
1155
+ self.assertEqual(stats.data["requested_model"], DEFAULT_MODEL_NAME)
1156
  self.assertEqual(stats.data["served_model"], FALLBACK_MODEL_NAME)
1157
  self.assertTrue(stats.data["rescued"])
1158
 
 
1171
  request = ChatRequest(
1172
  query="ls the kb",
1173
  source_keys=("peft",),
1174
+ model_name=DEFAULT_MODEL_NAME,
1175
  include_reasoning=False,
1176
  enabled_tools=(),
1177
  )
 
1182
  with (
1183
  patch.dict(
1184
  os.environ,
1185
+ {"OPENROUTER_API_KEY": "k1", "GEMINI_API_KEY": "k2"},
1186
  clear=False,
1187
  ),
1188
  patch("app.chat_service.build_agent", side_effect=fake_build_agent),
 
1192
  asyncio.run(collect_events())
1193
 
1194
  # 4xx never reaches the fallback; the primary was the only attempt.
1195
+ self.assertEqual(built_models, [DEFAULT_MODEL_NAME])
1196
 
1197
  def test_stream_chat_emits_deepseek_reasoning_content(self) -> None:
1198
  agent = FakeDeepSeekReasoningAgent([])
 
1201
  request = ChatRequest(
1202
  query="Find the course repository",
1203
  source_keys=("peft",),
1204
+ model_name=DEFAULT_MODEL_NAME,
1205
  include_reasoning=True,
1206
  enabled_tools=(),
1207
  )
 
1234
  request = ChatRequest(
1235
  query="Where should I store API keys?",
1236
  source_keys=("langchain",),
1237
+ model_name=DEFAULT_MODEL_NAME,
1238
  include_reasoning=False,
1239
  enabled_tools=(),
1240
  )
tests/test_config.py CHANGED
@@ -11,11 +11,14 @@ import pytest
11
  from app import config
12
 
13
 
14
- def test_default_chat_model_prefers_deepseek_direct_with_gemini_fallback() -> None:
15
- assert config.DEFAULT_MODEL_NAME == config.DEEPSEEK_DIRECT_MODEL_NAME
16
- assert config.DEFAULT_MODEL_NAME == "deepseek:deepseek-v4-flash"
17
  assert config.FALLBACK_MODEL_NAME == "google-genai:gemini-3.5-flash-lite"
18
  assert config.AVAILABLE_MODELS[0]["id"] == config.DEFAULT_MODEL_NAME
 
 
 
19
 
20
 
21
  def test_gemini_fallback_model_is_deliberately_not_selectable() -> None:
@@ -25,12 +28,12 @@ def test_gemini_fallback_model_is_deliberately_not_selectable() -> None:
25
  our custom tools) this is a product choice, one selectable model plus a
26
  rescue path, rather than the hard API constraint it was for 2.5-flash. The
27
  fallback still never receives web tools: build_agent binds them from the
28
- requested model (DeepSeek), which has none.
29
  """
30
  selectable = [model["id"] for model in config.AVAILABLE_MODELS]
31
 
32
  assert config.FALLBACK_MODEL_NAME not in selectable
33
- assert selectable == ["deepseek:deepseek-v4-flash"]
34
 
35
 
36
  def _patched_bundle(tmp_path: Path) -> ExitStack:
 
11
  from app import config
12
 
13
 
14
+ def test_default_chat_model_is_openrouter_deepseek_with_gemini_fallback() -> None:
15
+ assert config.DEFAULT_MODEL_NAME == config.OPENROUTER_DEEPSEEK_MODEL_NAME
16
+ assert config.DEFAULT_MODEL_NAME == "openrouter:deepseek/deepseek-v4-flash"
17
  assert config.FALLBACK_MODEL_NAME == "google-genai:gemini-3.5-flash-lite"
18
  assert config.AVAILABLE_MODELS[0]["id"] == config.DEFAULT_MODEL_NAME
19
+ # The retired first-party path stays addressable for experiments but is
20
+ # neither the default nor selectable.
21
+ assert config.DEEPSEEK_DIRECT_MODEL_NAME == "deepseek:deepseek-v4-flash"
22
 
23
 
24
  def test_gemini_fallback_model_is_deliberately_not_selectable() -> None:
 
28
  our custom tools) this is a product choice, one selectable model plus a
29
  rescue path, rather than the hard API constraint it was for 2.5-flash. The
30
  fallback still never receives web tools: build_agent binds them from the
31
+ requested model (DeepSeek via OpenRouter), which has none.
32
  """
33
  selectable = [model["id"] for model in config.AVAILABLE_MODELS]
34
 
35
  assert config.FALLBACK_MODEL_NAME not in selectable
36
+ assert selectable == ["openrouter:deepseek/deepseek-v4-flash"]
37
 
38
 
39
  def _patched_bundle(tmp_path: Path) -> ExitStack:
tests/test_memory_presets.py CHANGED
@@ -90,11 +90,24 @@ class MemoryPresetResolutionTests(unittest.TestCase):
90
  production_memory_preset_name("openai:gpt-5.6"),
91
  PRODUCTION_FALLBACK_MEMORY_PRESET,
92
  )
 
 
 
 
 
 
93
 
94
  def test_prod_v2_rejects_short_context_providers(self) -> None:
95
  self.assertTrue(
96
  memory_preset_supports_model("prod_v2", "deepseek:deepseek-v4-flash")
97
  )
 
 
 
 
 
 
 
98
  self.assertTrue(
99
  memory_preset_supports_model(
100
  "prod_v2", "google-genai:gemini-3.5-flash-lite"
 
90
  production_memory_preset_name("openai:gpt-5.6"),
91
  PRODUCTION_FALLBACK_MEMORY_PRESET,
92
  )
93
+ self.assertEqual(
94
+ resolve_memory_preset(
95
+ None, model_name="openrouter:deepseek/deepseek-v4-flash"
96
+ ).name,
97
+ "prod_v2",
98
+ )
99
 
100
  def test_prod_v2_rejects_short_context_providers(self) -> None:
101
  self.assertTrue(
102
  memory_preset_supports_model("prod_v2", "deepseek:deepseek-v4-flash")
103
  )
104
+ # The default chat path serves DeepSeek through OpenRouter; prod_v2
105
+ # must not silently downgrade it to the legacy preset.
106
+ self.assertTrue(
107
+ memory_preset_supports_model(
108
+ "prod_v2", "openrouter:deepseek/deepseek-v4-flash"
109
+ )
110
+ )
111
  self.assertTrue(
112
  memory_preset_supports_model(
113
  "prod_v2", "google-genai:gemini-3.5-flash-lite"