Switch the default chat model to DeepSeek V4 Flash via OpenRouter
Browse filesThe first-party DeepSeek account is deliberately unfunded and the team is
standardizing on OpenRouter, whose cheap endpoints also price a typical
agentic turn ~40-70% below first-party rates: output and cache-miss input
dominate our bill, and DeepSeek's ~31x first-party cache-hit discount does
not pass through OpenRouter's third-party hosts anyway (verified against
OpenRouter's endpoint listing, 2026-08-25).
- DEFAULT_MODEL_NAME / AVAILABLE_MODELS -> openrouter:deepseek/deepseek-v4-flash;
the first-party deepseek: path stays supported for experiments
- New TutorChatOpenRouter adapter preserves OpenRouter's unified streamed
`reasoning` field (mapped into the DeepSeek-style reasoning_content slot)
and the client sends the model-agnostic reasoning toggle explicitly
- "openrouter" joins PRODUCTION_LONG_CONTEXT_PROVIDERS so prod_v2 serves
the default path instead of silently downgrading to the legacy preset
- MODEL_PRICING OpenRouter row refreshed to the verified headline endpoint
rate (checked 2026-08-25); docs, runbook, and env templates updated
Verified live: reasoning streams interleaved with run_kb_command calls,
prompt caching passes through (cached_tokens populated in context_stats),
thread continuation works, and the live E2E suite passes. The Space needs
the OPENROUTER_API_KEY secret; until it is set, key substitution serves
the Gemini fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- .env.example +6 -5
- AGENTS.md +5 -5
- app/chat_service.py +16 -1
- app/config.py +10 -2
- app/memory_presets.py +7 -1
- app/openrouter_chat.py +44 -0
- app/telemetry.py +11 -4
- frontend/components/chat-shell.test.tsx +1 -1
- tests/manual_e2e_langsmith.md +3 -3
- tests/test_api.py +4 -2
- tests/test_chat_service.py +74 -28
- tests/test_config.py +8 -5
- tests/test_memory_presets.py +13 -0
|
@@ -1,17 +1,18 @@
|
|
| 1 |
# To run the AI Tutor app (FastAPI backend)
|
| 2 |
# COHERE_API_KEY is always required (embeddings + rerank for retrieval).
|
| 3 |
COHERE_API_KEY=...
|
| 4 |
-
# Default chat path: DeepSeek V4 Flash
|
| 5 |
-
#
|
| 6 |
-
|
| 7 |
GEMINI_API_KEY=...
|
| 8 |
# Gemini also accepts GOOGLE_API_KEY instead of GEMINI_API_KEY.
|
| 9 |
GOOGLE_API_KEY=...
|
| 10 |
|
| 11 |
-
# Optional alternate providers
|
|
|
|
| 12 |
OPENAI_API_KEY=...
|
| 13 |
ANTHROPIC_API_KEY=...
|
| 14 |
-
|
| 15 |
|
| 16 |
# Optional: Slack-compatible incoming-webhook URL for ops alerts (provider
|
| 17 |
# outages and account-state errors like an unfunded or revoked key). Alerts
|
|
|
|
| 1 |
# To run the AI Tutor app (FastAPI backend)
|
| 2 |
# COHERE_API_KEY is always required (embeddings + rerank for retrieval).
|
| 3 |
COHERE_API_KEY=...
|
| 4 |
+
# Default chat path: DeepSeek V4 Flash via OpenRouter, with Gemini 3.5 Flash
|
| 5 |
+
# Lite as the local application fallback if OpenRouter fails.
|
| 6 |
+
OPENROUTER_API_KEY=...
|
| 7 |
GEMINI_API_KEY=...
|
| 8 |
# Gemini also accepts GOOGLE_API_KEY instead of GEMINI_API_KEY.
|
| 9 |
GOOGLE_API_KEY=...
|
| 10 |
|
| 11 |
+
# Optional alternate providers (DEEPSEEK_API_KEY is the retired first-party
|
| 12 |
+
# DeepSeek path, kept for experiments)
|
| 13 |
OPENAI_API_KEY=...
|
| 14 |
ANTHROPIC_API_KEY=...
|
| 15 |
+
DEEPSEEK_API_KEY=...
|
| 16 |
|
| 17 |
# Optional: Slack-compatible incoming-webhook URL for ops alerts (provider
|
| 18 |
# outages and account-state errors like an unfunded or revoked key). Alerts
|
|
@@ -6,7 +6,7 @@ This is the **canonical, tool-agnostic** instruction file for the repo. `CLAUDE.
|
|
| 6 |
|
| 7 |
AI tutor for applied AI, LLMs, RAG, and Python. **Agentic RAG**: a LangChain/LangGraph agent grounds answers in a curated corpus of course + library docs, can browse a local file-based knowledge base, and (optionally) search the live web. One frontend: a **Next.js** UI (`frontend/`), served by a **FastAPI** backend (`app/api.py`) that streams in the Vercel AI SDK UI-message protocol. (A Gradio UI existed historically; it was removed to keep one rendering path.)
|
| 8 |
|
| 9 |
-
ChromaDB for vectors; Cohere for embeddings/rerank; chat model is provider-configurable (DeepSeek V4 Flash
|
| 10 |
|
| 11 |
## Key URLs
|
| 12 |
|
|
@@ -62,7 +62,7 @@ Runtime guidance the agent follows is in `data/kb/AGENTS.md` (injected into the
|
|
| 62 |
|
| 63 |
## Sources & config
|
| 64 |
|
| 65 |
-
`data/scraping_scripts/source_registry.py` is the **single source of truth** for sources (`SOURCE_CONFIGS`, key groupings, UI labels, defaults); `app/config.py` re-exports them and the frontend derives the picker from it (via `/api/tools`). Docs sources ingest via the GitHub API or `llms.txt`; course sources are Notion exports. To add a source: add it to the registry (+ the relevant grouping tuples), then run the matching workflow — no separate UI edit needed. Models live in `config.AVAILABLE_MODELS` (currently only the default `
|
| 66 |
|
| 67 |
Memory policy is deliberately owned by `app/memory_presets.py`, not `app/config.py`. `PRODUCTION_MEMORY_PRESET` is the single primary-policy switch, `PRODUCTION_FALLBACK_MEMORY_PRESET` covers providers that cannot use it safely, and `PRODUCTION_LONG_CONTEXT_PROVIDERS` defines the model-aware allowlist. Keep historical presets immutable because saved eval runs refer to them by name.
|
| 68 |
|
|
@@ -97,7 +97,7 @@ uv run -m data.scraping_scripts.retire_source_workflow --sources KEY [--dry-run
|
|
| 97 |
|
| 98 |
## Environment variables
|
| 99 |
|
| 100 |
-
Chat runtime: `COHERE_API_KEY` (retrieval; always required), `
|
| 101 |
|
| 102 |
## Deployment
|
| 103 |
|
|
@@ -106,7 +106,7 @@ One HF Space runs the image (`Dockerfile`: FastAPI + Next.js static export, `rip
|
|
| 106 |
- **Prod — `ai-tutor-chatbot`** (public): `.github/workflows/sync-to-hf.yml` force-pushes on **every push to `main`** (docs/markdown-only and scraping-script-only pushes are skipped via `paths-ignore`). Merging to `main` IS deploying to public prod: keep `main` releasable and verify before merging (`uv run pytest`, and the live E2E suite for agent/API changes).
|
| 107 |
- **Manual redeploy**: `.github/workflows/deploy-prod-to-hf.yml` (Actions tab → "Redeploy prod to Hugging Face (manual)") force-pushes a chosen ref without a new commit — for secret rotations, retrying a failed sync run, or rolling back to an older branch/tag.
|
| 108 |
|
| 109 |
-
The Space needs the runtime secrets (`COHERE_API_KEY`,
|
| 110 |
|
| 111 |
## Conventions
|
| 112 |
|
|
@@ -115,7 +115,7 @@ The Space needs the runtime secrets (`COHERE_API_KEY`, model provider key, `HF_T
|
|
| 115 |
|
| 116 |
## Gotchas
|
| 117 |
|
| 118 |
-
- **One thread is one provider — a conversation cannot mix models while keeping tool outputs and thought signatures.** A thread's checkpoint stores *provider-native* messages: Gemini reasoning parts carry thought signatures, Anthropic thinking blocks carry a required cryptographic `signature`, and each provider's server-side tool calls are its own block types. Replaying one provider's history to another is a hard 400, not a degraded answer — verified both ways (Gemini history → DeepSeek 400s; Gemini history → Anthropic 400s with `messages.1.content.0.thinking.signature: Field required`). So you cannot "just send the payload to another provider": a fallback, a retry-on-another-provider, or mid-conversation switching **must not** be wired at the model/client layer (a `with_fallbacks` client only works for pairings whose primary happens to write lowest-common-denominator output — an overfit trap this repo shipped once and removed). The only safe path is `sync_thread_with_history`'s branch to a fresh thread seeded with plain text, which drops the unportable state (tool outputs, summaries, thinking) on purpose — and it is the ONE mechanism for all cross-provider movement: the turn-level outage rescue in `stream_chat`, a config repoint, a model deprecation. The frontend model picker locks after the first message for this reason. The backend enforces it independently: `sync_thread_with_history` records which provider actually served each turn (`_THREAD_PROVIDERS`, taken from the served model, since rescue and key substitution mean served ≠ requested) and branches to a fresh plain-text thread when the next turn's provider differs — so a rescued turn's Gemini-owned thread returns to
|
| 119 |
- **Provider-native web tools are bound from the requested model, and not every model can mix them with our custom tools.** Gemini needs "tool context circulation" (Gemini 3+) to combine `google_search`/`url_context` with function tools, so pre-3 Gemini gets no web toggles (`supports_gemini_tool_combination`). Anthropic combines them natively, but `allowed_callers: ["direct"]` is load-bearing — without it Haiku 4.5 400s on programmatic tool calling.
|
| 120 |
- **Context generation uses Gemini**; embeddings/rerank use **Cohere**; the chat model is provider-configurable. OpenAI is required only when explicitly selected.
|
| 121 |
- `data/kb/` and `data/chroma-db-all_sources/` are build artifacts — never commit or hand-edit; regenerate or re-download.
|
|
|
|
| 6 |
|
| 7 |
AI tutor for applied AI, LLMs, RAG, and Python. **Agentic RAG**: a LangChain/LangGraph agent grounds answers in a curated corpus of course + library docs, can browse a local file-based knowledge base, and (optionally) search the live web. One frontend: a **Next.js** UI (`frontend/`), served by a **FastAPI** backend (`app/api.py`) that streams in the Vercel AI SDK UI-message protocol. (A Gradio UI existed historically; it was removed to keep one rendering path.)
|
| 8 |
|
| 9 |
+
ChromaDB for vectors; Cohere for embeddings/rerank; chat model is provider-configurable (DeepSeek V4 Flash via OpenRouter by default, with a rescue-only in-app fallback to Gemini 3.5 Flash Lite when a Gemini key is set; the first-party DeepSeek API, Anthropic, and OpenAI remain supported in code but are not user-selectable). **A conversation cannot change model mid-thread** — checkpoints store provider-native message blocks that no other provider can replay; see `build_chat_model` in `app/chat_service.py`. Python ≥3.13, managed with `uv`.
|
| 10 |
|
| 11 |
## Key URLs
|
| 12 |
|
|
|
|
| 62 |
|
| 63 |
## Sources & config
|
| 64 |
|
| 65 |
+
`data/scraping_scripts/source_registry.py` is the **single source of truth** for sources (`SOURCE_CONFIGS`, key groupings, UI labels, defaults); `app/config.py` re-exports them and the frontend derives the picker from it (via `/api/tools`). Docs sources ingest via the GitHub API or `llms.txt`; course sources are Notion exports. To add a source: add it to the registry (+ the relevant grouping tuples), then run the matching workflow — no separate UI edit needed. Models live in `config.AVAILABLE_MODELS` (currently only the default `openrouter:deepseek/deepseek-v4-flash` — DeepSeek V4 Flash via OpenRouter, which routes across third-party hosts; the first-party `deepseek:` path, Anthropic Claude Haiku 4.5, and OpenAI remain supported in code but are not selectable. The first-party DeepSeek account is deliberately unfunded since 2026-08, when its 402s hard-failed prod and the team standardized on OpenRouter — also cheaper per agentic turn at OpenRouter's cheap endpoints, though DeepSeek's ~31x first-party cache-hit discount does not pass through; see `MODEL_PRICING`). **`FALLBACK_MODEL_NAME` (`google-genai:gemini-3.5-flash-lite`) is deliberately NOT in `AVAILABLE_MODELS`**: it serves real traffic as the default model's rescue path but is not user-selectable. Unlike the previous fallback (`gemini-2.5-flash`, a pre-Gemini-3 model that could not combine Gemini's built-in web tools with our two custom tools and so was one web-search toggle from a 400 if ever selected), 3.5 Flash Lite is a Gemini 3-series model with "tool context circulation", so keeping it fallback-only is now a product choice (one selectable model plus a rescue path), not an API constraint. Either way the fallback never receives web tools: they are bound from the requested model (DeepSeek via OpenRouter), which has none. `build_chat_model` in `app/chat_service.py` builds exactly ONE provider's client (it accepts `openrouter:` for the OpenRouter gateway — served by `TutorChatOpenRouter`, which preserves OpenRouter's unified streamed `reasoning` field and sends its model-agnostic `reasoning` toggle — plus `deepseek:` for the first-party API and `ollama:` for local SLM experiments, with pricing in `app/telemetry.MODEL_PRICING`); no fallback is wired at the client layer. The rescue is turn-level in `stream_chat`: an outage-class failure (connectivity/timeout/429/5xx) or an account-state failure (401/402/403 — an unfunded or revoked key, see `is_account_state_error`) before the first streamed chunk retries the turn on `FALLBACK_MODEL_NAME` through the plain-text thread branch, so the pairing is free to be ANY two provider:model strings. Both failure classes page ops through `app/alerts.py` (Slack-compatible webhook via `AI_TUTOR_ALERT_WEBHOOK_URL`, throttled per class), and each turn's `context_stats` reports `requested_model`/`served_model`/`rescued` so a fallback-served turn is never invisible. `resolve_served_model_name` substitutes the fallback outright when the default model's provider key is missing.
|
| 66 |
|
| 67 |
Memory policy is deliberately owned by `app/memory_presets.py`, not `app/config.py`. `PRODUCTION_MEMORY_PRESET` is the single primary-policy switch, `PRODUCTION_FALLBACK_MEMORY_PRESET` covers providers that cannot use it safely, and `PRODUCTION_LONG_CONTEXT_PROVIDERS` defines the model-aware allowlist. Keep historical presets immutable because saved eval runs refer to them by name.
|
| 68 |
|
|
|
|
| 97 |
|
| 98 |
## Environment variables
|
| 99 |
|
| 100 |
+
Chat runtime: `COHERE_API_KEY` (retrieval; always required), `OPENROUTER_API_KEY` for the default model plus `GEMINI_API_KEY`/`GOOGLE_API_KEY` to enable its in-app Gemini 3.5 Flash Lite fallback, and the corresponding provider key when selecting another model (`GEMINI_API_KEY`/`GOOGLE_API_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, or `DEEPSEEK_API_KEY` for the retired first-party path), plus `HF_TOKEN` for first-start download of the full private bundle. Without `HF_TOKEN`, cold start falls back to the public docs-only bundle (documentation sources only, course content hidden). Optional: `AI_TUTOR_ALERT_WEBHOOK_URL` (Slack-compatible webhook for provider outage/account-state ops alerts; unset = log-only), `LANGSMITH_*` (tracing), `AI_TUTOR_API_PORT`/`HOST`/`CORS_ALLOW_ORIGINS`, `AI_TUTOR_KB_DIR`, and `NEXT_PUBLIC_AI_TUTOR_API_BASE_URL`. `AI_TUTOR_MEMORY_PRESET` is an optional deployment-wide override, not the normal default; leave it unset to retain model-aware production selection. Data workflows also need `GITHUB_TOKEN` and `GEMINI_API_KEY`/`GOOGLE_API_KEY` (context generation). See `.env.example`.
|
| 101 |
|
| 102 |
## Deployment
|
| 103 |
|
|
|
|
| 106 |
- **Prod — `ai-tutor-chatbot`** (public): `.github/workflows/sync-to-hf.yml` force-pushes on **every push to `main`** (docs/markdown-only and scraping-script-only pushes are skipped via `paths-ignore`). Merging to `main` IS deploying to public prod: keep `main` releasable and verify before merging (`uv run pytest`, and the live E2E suite for agent/API changes).
|
| 107 |
- **Manual redeploy**: `.github/workflows/deploy-prod-to-hf.yml` (Actions tab → "Redeploy prod to Hugging Face (manual)") force-pushes a chosen ref without a new commit — for secret rotations, retrying a failed sync run, or rolling back to an older branch/tag.
|
| 108 |
|
| 109 |
+
The Space needs the runtime secrets (`COHERE_API_KEY`, `OPENROUTER_API_KEY`, `GEMINI_API_KEY`, `HF_TOKEN`, optional `LANGSMITH_*` and `AI_TUTOR_ALERT_WEBHOOK_URL`) configured in its HF settings.
|
| 110 |
|
| 111 |
## Conventions
|
| 112 |
|
|
|
|
| 115 |
|
| 116 |
## Gotchas
|
| 117 |
|
| 118 |
+
- **One thread is one provider — a conversation cannot mix models while keeping tool outputs and thought signatures.** A thread's checkpoint stores *provider-native* messages: Gemini reasoning parts carry thought signatures, Anthropic thinking blocks carry a required cryptographic `signature`, and each provider's server-side tool calls are its own block types. Replaying one provider's history to another is a hard 400, not a degraded answer — verified both ways (Gemini history → DeepSeek 400s; Gemini history → Anthropic 400s with `messages.1.content.0.thinking.signature: Field required`). So you cannot "just send the payload to another provider": a fallback, a retry-on-another-provider, or mid-conversation switching **must not** be wired at the model/client layer (a `with_fallbacks` client only works for pairings whose primary happens to write lowest-common-denominator output — an overfit trap this repo shipped once and removed). The only safe path is `sync_thread_with_history`'s branch to a fresh thread seeded with plain text, which drops the unportable state (tool outputs, summaries, thinking) on purpose — and it is the ONE mechanism for all cross-provider movement: the turn-level outage/account-state rescue in `stream_chat`, a config repoint, a model deprecation. The frontend model picker locks after the first message for this reason. The backend enforces it independently: `sync_thread_with_history` records which provider actually served each turn (`_THREAD_PROVIDERS`, taken from the served model, since rescue and key substitution mean served ≠ requested) and branches to a fresh plain-text thread when the next turn's provider differs — so a rescued turn's Gemini-owned thread returns to the default model on the next healthy turn instead of stranding on the pricier provider (see `MODEL_PRICING` in `app/telemetry.py`).
|
| 119 |
- **Provider-native web tools are bound from the requested model, and not every model can mix them with our custom tools.** Gemini needs "tool context circulation" (Gemini 3+) to combine `google_search`/`url_context` with function tools, so pre-3 Gemini gets no web toggles (`supports_gemini_tool_combination`). Anthropic combines them natively, but `allowed_callers: ["direct"]` is load-bearing — without it Haiku 4.5 400s on programmatic tool calling.
|
| 120 |
- **Context generation uses Gemini**; embeddings/rerank use **Cohere**; the chat model is provider-configurable. OpenAI is required only when explicitly selected.
|
| 121 |
- `data/kb/` and `data/chroma-db-all_sources/` are build artifacts — never commit or hand-edit; regenerate or re-download.
|
|
@@ -44,6 +44,7 @@ from .agent_tracing import langsmith_deployment_identity
|
|
| 44 |
from .chat_types import ChatEvent, ChatRequest, ChatTurn, SourceMatch
|
| 45 |
from .alerts import send_ops_alert
|
| 46 |
from .deepseek_chat import TutorChatDeepSeek
|
|
|
|
| 47 |
from .memory_presets import (
|
| 48 |
MemoryConfig,
|
| 49 |
resolve_memory_preset,
|
|
@@ -960,6 +961,12 @@ def is_deepseek_model(model_name: str) -> bool:
|
|
| 960 |
return provider == "deepseek"
|
| 961 |
|
| 962 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 963 |
def supports_gemini_tool_combination(model_name: str) -> bool:
|
| 964 |
"""True when a Gemini model can mix built-in tools with function tools.
|
| 965 |
|
|
@@ -1032,12 +1039,17 @@ def _build_chat_model_client(provider_model: str, include_thoughts: bool = False
|
|
| 1032 |
# Routing inside OpenRouter is left with its own fallbacks enabled:
|
| 1033 |
# pinning a single provider with allow_fallbacks=False made batches die
|
| 1034 |
# on a backend's transient 429 instead of routing around it.
|
| 1035 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1036 |
model=actual_model,
|
| 1037 |
temperature=1,
|
| 1038 |
base_url="https://openrouter.ai/api/v1",
|
| 1039 |
api_key=os.environ.get("OPENROUTER_API_KEY"),
|
| 1040 |
stream_usage=True,
|
|
|
|
| 1041 |
# Long agentic sessions fire hundreds of calls against OpenRouter's
|
| 1042 |
# shared pool; bump retries (with the client's exponential backoff)
|
| 1043 |
# to ride out transient upstream 429s instead of failing a whole
|
|
@@ -2876,6 +2888,9 @@ async def stream_chat(request: ChatRequest) -> AsyncIterator[ChatEvent]:
|
|
| 2876 |
is_google_genai_model(serving_model_name)
|
| 2877 |
or is_anthropic_model(serving_model_name)
|
| 2878 |
or is_deepseek_model(serving_model_name)
|
|
|
|
|
|
|
|
|
|
| 2879 |
)
|
| 2880 |
# Gemini streams each thought summary as one complete block; Anthropic
|
| 2881 |
# streams partial fragments of a single thought. The encoder uses this
|
|
|
|
| 44 |
from .chat_types import ChatEvent, ChatRequest, ChatTurn, SourceMatch
|
| 45 |
from .alerts import send_ops_alert
|
| 46 |
from .deepseek_chat import TutorChatDeepSeek
|
| 47 |
+
from .openrouter_chat import TutorChatOpenRouter
|
| 48 |
from .memory_presets import (
|
| 49 |
MemoryConfig,
|
| 50 |
resolve_memory_preset,
|
|
|
|
| 961 |
return provider == "deepseek"
|
| 962 |
|
| 963 |
|
| 964 |
+
def is_openrouter_model(model_name: str) -> bool:
|
| 965 |
+
provider_model = normalize_model_name(model_name)
|
| 966 |
+
provider, _, _actual_model = provider_model.partition(":")
|
| 967 |
+
return provider == "openrouter"
|
| 968 |
+
|
| 969 |
+
|
| 970 |
def supports_gemini_tool_combination(model_name: str) -> bool:
|
| 971 |
"""True when a Gemini model can mix built-in tools with function tools.
|
| 972 |
|
|
|
|
| 1039 |
# Routing inside OpenRouter is left with its own fallbacks enabled:
|
| 1040 |
# pinning a single provider with allow_fallbacks=False made batches die
|
| 1041 |
# on a backend's transient 429 instead of routing around it.
|
| 1042 |
+
# ``reasoning`` is OpenRouter's model-agnostic thinking toggle; sending
|
| 1043 |
+
# it explicitly (rather than omitting it) pins the behavior across the
|
| 1044 |
+
# gateway's per-host defaults, and the adapter preserves the streamed
|
| 1045 |
+
# thought text that plain ChatOpenAI would discard.
|
| 1046 |
+
return TutorChatOpenRouter(
|
| 1047 |
model=actual_model,
|
| 1048 |
temperature=1,
|
| 1049 |
base_url="https://openrouter.ai/api/v1",
|
| 1050 |
api_key=os.environ.get("OPENROUTER_API_KEY"),
|
| 1051 |
stream_usage=True,
|
| 1052 |
+
extra_body={"reasoning": {"enabled": bool(include_thoughts)}},
|
| 1053 |
# Long agentic sessions fire hundreds of calls against OpenRouter's
|
| 1054 |
# shared pool; bump retries (with the client's exponential backoff)
|
| 1055 |
# to ride out transient upstream 429s instead of failing a whole
|
|
|
|
| 2888 |
is_google_genai_model(serving_model_name)
|
| 2889 |
or is_anthropic_model(serving_model_name)
|
| 2890 |
or is_deepseek_model(serving_model_name)
|
| 2891 |
+
# OpenRouter's unified `reasoning` request field works across the
|
| 2892 |
+
# models it routes; hosts that don't support it ignore it.
|
| 2893 |
+
or is_openrouter_model(serving_model_name)
|
| 2894 |
)
|
| 2895 |
# Gemini streams each thought summary as one complete block; Anthropic
|
| 2896 |
# streams partial fragments of a single thought. The encoder uses this
|
|
@@ -70,7 +70,14 @@ KB_AGENTS_PATH = f"{KB_DIR}/AGENTS.md"
|
|
| 70 |
# In-git template, copied into data/kb/AGENTS.md by ensure_kb_agents_md().
|
| 71 |
KB_AGENTS_TEMPLATE_PATH = "data/scraping_scripts/kb_agents_template.md"
|
| 72 |
DEEPSEEK_DIRECT_MODEL_NAME = "deepseek:deepseek-v4-flash"
|
| 73 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 74 |
|
| 75 |
# Two distinct roles, deliberately kept apart:
|
| 76 |
#
|
|
@@ -96,7 +103,7 @@ FALLBACK_MODEL_NAME = "google-genai:gemini-3.5-flash-lite"
|
|
| 96 |
|
| 97 |
AVAILABLE_MODELS: tuple[dict[str, str], ...] = (
|
| 98 |
{
|
| 99 |
-
"id":
|
| 100 |
"label": "DeepSeek V4 Flash",
|
| 101 |
},
|
| 102 |
)
|
|
@@ -294,6 +301,7 @@ __all__ = [
|
|
| 294 |
"DEFAULT_MODEL_NAME",
|
| 295 |
"DEEPSEEK_DIRECT_MODEL_NAME",
|
| 296 |
"FALLBACK_MODEL_NAME",
|
|
|
|
| 297 |
"BM25_INDEX_PATH",
|
| 298 |
"DOCUMENT_DICT_PATH",
|
| 299 |
"KB_AGENTS_PATH",
|
|
|
|
| 70 |
# In-git template, copied into data/kb/AGENTS.md by ensure_kb_agents_md().
|
| 71 |
KB_AGENTS_TEMPLATE_PATH = "data/scraping_scripts/kb_agents_template.md"
|
| 72 |
DEEPSEEK_DIRECT_MODEL_NAME = "deepseek:deepseek-v4-flash"
|
| 73 |
+
# DeepSeek V4 Flash via OpenRouter is the default chat path since 2026-08:
|
| 74 |
+
# the first-party DeepSeek account is deliberately unfunded (its 402s
|
| 75 |
+
# hard-failed prod for two days), and OpenRouter's cheap endpoints price a
|
| 76 |
+
# typical agentic turn well below first-party rates anyway. OpenRouter serves
|
| 77 |
+
# this model only through third-party hosts, so DeepSeek's ~31x first-party
|
| 78 |
+
# cache-hit discount does NOT apply there; see MODEL_PRICING in app/telemetry.
|
| 79 |
+
OPENROUTER_DEEPSEEK_MODEL_NAME = "openrouter:deepseek/deepseek-v4-flash"
|
| 80 |
+
DEFAULT_MODEL_NAME = OPENROUTER_DEEPSEEK_MODEL_NAME
|
| 81 |
|
| 82 |
# Two distinct roles, deliberately kept apart:
|
| 83 |
#
|
|
|
|
| 103 |
|
| 104 |
AVAILABLE_MODELS: tuple[dict[str, str], ...] = (
|
| 105 |
{
|
| 106 |
+
"id": OPENROUTER_DEEPSEEK_MODEL_NAME,
|
| 107 |
"label": "DeepSeek V4 Flash",
|
| 108 |
},
|
| 109 |
)
|
|
|
|
| 301 |
"DEFAULT_MODEL_NAME",
|
| 302 |
"DEEPSEEK_DIRECT_MODEL_NAME",
|
| 303 |
"FALLBACK_MODEL_NAME",
|
| 304 |
+
"OPENROUTER_DEEPSEEK_MODEL_NAME",
|
| 305 |
"BM25_INDEX_PATH",
|
| 306 |
"DOCUMENT_DICT_PATH",
|
| 307 |
"KB_AGENTS_PATH",
|
|
@@ -357,7 +357,13 @@ MEMORY_PRESETS: dict[str, MemoryConfig] = {
|
|
| 357 |
# example, has a 200k input window and cannot safely wait for an 800k trigger).
|
| 358 |
PRODUCTION_MEMORY_PRESET = "prod_v2"
|
| 359 |
PRODUCTION_FALLBACK_MEMORY_PRESET = "prod"
|
| 360 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 361 |
|
| 362 |
# Backward-compatible import for callers that need a single default name. New
|
| 363 |
# runtime code should call resolve_memory_preset(..., model_name=...) so model
|
|
|
|
| 357 |
# example, has a 200k input window and cannot safely wait for an 800k trigger).
|
| 358 |
PRODUCTION_MEMORY_PRESET = "prod_v2"
|
| 359 |
PRODUCTION_FALLBACK_MEMORY_PRESET = "prod"
|
| 360 |
+
# "openrouter" is here for the default DeepSeek-via-OpenRouter chat model.
|
| 361 |
+
# The gate is provider-granular, so any openrouter:* model resolves to the
|
| 362 |
+
# long-context production preset by default; eval runs pin presets by name,
|
| 363 |
+
# so in practice this only decides the served chat path.
|
| 364 |
+
PRODUCTION_LONG_CONTEXT_PROVIDERS = frozenset(
|
| 365 |
+
{"deepseek", "google-genai", "openrouter"}
|
| 366 |
+
)
|
| 367 |
|
| 368 |
# Backward-compatible import for callers that need a single default name. New
|
| 369 |
# runtime code should call resolve_memory_preset(..., model_name=...) so model
|
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""OpenRouter chat-model compatibility helpers.
|
| 2 |
+
|
| 3 |
+
OpenRouter normalizes reasoning across the models it routes: requests take a
|
| 4 |
+
unified ``reasoning`` object (``{"enabled": ...}``) regardless of the
|
| 5 |
+
underlying model, and streamed deltas carry the thought text in a
|
| 6 |
+
``reasoning`` field (verified live against ``deepseek/deepseek-v4-flash``,
|
| 7 |
+
2026-08-25). The stock ``ChatOpenAI`` wrapper drops that nonstandard delta
|
| 8 |
+
field, so this subclass maps it into
|
| 9 |
+
``additional_kwargs["reasoning_content"]`` -- the same slot the DeepSeek
|
| 10 |
+
adapter uses -- and ``provider_events.extract_reasoning_deltas`` then works
|
| 11 |
+
unchanged for either provider.
|
| 12 |
+
"""
|
| 13 |
+
|
| 14 |
+
from __future__ import annotations
|
| 15 |
+
|
| 16 |
+
from langchain_core.messages import AIMessageChunk
|
| 17 |
+
from langchain_core.outputs import ChatGenerationChunk
|
| 18 |
+
from langchain_openai import ChatOpenAI
|
| 19 |
+
|
| 20 |
+
|
| 21 |
+
class TutorChatOpenRouter(ChatOpenAI):
|
| 22 |
+
"""ChatOpenAI that preserves OpenRouter's streamed ``reasoning`` field."""
|
| 23 |
+
|
| 24 |
+
def _convert_chunk_to_generation_chunk(
|
| 25 |
+
self,
|
| 26 |
+
chunk: dict,
|
| 27 |
+
default_chunk_class: type,
|
| 28 |
+
base_generation_info: dict | None,
|
| 29 |
+
) -> ChatGenerationChunk | None:
|
| 30 |
+
generation_chunk = super()._convert_chunk_to_generation_chunk(
|
| 31 |
+
chunk,
|
| 32 |
+
default_chunk_class,
|
| 33 |
+
base_generation_info,
|
| 34 |
+
)
|
| 35 |
+
if generation_chunk is None:
|
| 36 |
+
return None
|
| 37 |
+
choices = chunk.get("choices") or []
|
| 38 |
+
if choices and isinstance(generation_chunk.message, AIMessageChunk):
|
| 39 |
+
reasoning = (choices[0].get("delta") or {}).get("reasoning")
|
| 40 |
+
if reasoning:
|
| 41 |
+
generation_chunk.message.additional_kwargs["reasoning_content"] = (
|
| 42 |
+
reasoning
|
| 43 |
+
)
|
| 44 |
+
return generation_chunk
|
|
@@ -256,11 +256,18 @@ MODEL_PRICING: dict[str, ModelPricing] = {
|
|
| 256 |
"gpt-5.6-luna": ModelPricing(
|
| 257 |
input=0.20, output=1.20, cache_read=0.02, cache_write=0.25
|
| 258 |
),
|
| 259 |
-
# DeepSeek-V4-Flash via OpenRouter (slug "deepseek/deepseek-v4-flash")
|
| 260 |
-
#
|
| 261 |
-
#
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 262 |
"deepseek/deepseek-v4-flash": ModelPricing(
|
| 263 |
-
input=0.
|
| 264 |
),
|
| 265 |
# DeepSeek-V4-Flash, DeepSeek FIRST-PARTY API (base https://api.deepseek.com,
|
| 266 |
# model id "deepseek-v4-flash" — the id the API echoes back, so this is the
|
|
|
|
| 256 |
"gpt-5.6-luna": ModelPricing(
|
| 257 |
input=0.20, output=1.20, cache_read=0.02, cache_write=0.25
|
| 258 |
),
|
| 259 |
+
# DeepSeek-V4-Flash via OpenRouter (slug "deepseek/deepseek-v4-flash") --
|
| 260 |
+
# the DEFAULT chat path since 2026-08. OpenRouter routes across ~17
|
| 261 |
+
# third-party hosts whose prices vary widely (input $0.068-$0.44, output
|
| 262 |
+
# $0.168-$1.32 per 1M); default routing favors the cheap end, so this row
|
| 263 |
+
# stores the headline (cheapest-endpoint) rate and expensive-host turns
|
| 264 |
+
# are underestimated. Caching is implicit with no write surcharge; cache
|
| 265 |
+
# reads bill at the endpoint's own rate (~0.25x its input price), NOT
|
| 266 |
+
# DeepSeek first-party's ~31x discount -- DeepSeek does not serve this
|
| 267 |
+
# model on OpenRouter. Checked 2026-08-25 against
|
| 268 |
+
# openrouter.ai/api/v1/models/deepseek/deepseek-v4-flash/endpoints.
|
| 269 |
"deepseek/deepseek-v4-flash": ModelPricing(
|
| 270 |
+
input=0.0679, output=0.168, cache_read=0.0168
|
| 271 |
),
|
| 272 |
# DeepSeek-V4-Flash, DeepSeek FIRST-PARTY API (base https://api.deepseek.com,
|
| 273 |
# model id "deepseek-v4-flash" — the id the API echoes back, so this is the
|
|
@@ -151,7 +151,7 @@ vi.mock("@/components/chat-message", () => ({
|
|
| 151 |
|
| 152 |
import { ChatShell } from "@/components/chat-shell";
|
| 153 |
|
| 154 |
-
const DEFAULT_MODEL = "
|
| 155 |
const CLAUDE_MODEL = "anthropic:claude-haiku-4-5";
|
| 156 |
|
| 157 |
const toolsResponse = {
|
|
|
|
| 151 |
|
| 152 |
import { ChatShell } from "@/components/chat-shell";
|
| 153 |
|
| 154 |
+
const DEFAULT_MODEL = "openrouter:deepseek/deepseek-v4-flash";
|
| 155 |
const CLAUDE_MODEL = "anthropic:claude-haiku-4-5";
|
| 156 |
|
| 157 |
const toolsResponse = {
|
|
@@ -13,7 +13,7 @@ Run commands from the repository root.
|
|
| 13 |
Required local artifacts and environment:
|
| 14 |
|
| 15 |
- `.env` contains `COHERE_API_KEY`
|
| 16 |
-
- `.env` contains `
|
| 17 |
- `.env` contains `GEMINI_API_KEY` or `GOOGLE_API_KEY`
|
| 18 |
- `.env` contains `LANGSMITH_API_KEY`
|
| 19 |
- `.env` has `LANGSMITH_TRACING=true`
|
|
@@ -31,7 +31,7 @@ uv run dotenv -f .env run -- python - <<'PY'
|
|
| 31 |
import os
|
| 32 |
for key in [
|
| 33 |
"COHERE_API_KEY",
|
| 34 |
-
"
|
| 35 |
"GEMINI_API_KEY",
|
| 36 |
"GOOGLE_API_KEY",
|
| 37 |
"LANGSMITH_API_KEY",
|
|
@@ -97,7 +97,7 @@ cat >/tmp/ai_tutor_e2e_payload.json <<'JSON'
|
|
| 97 |
"transformers"
|
| 98 |
],
|
| 99 |
"enabledTools": [],
|
| 100 |
-
"model": "
|
| 101 |
"includeReasoning": true,
|
| 102 |
"threadId": ""
|
| 103 |
}
|
|
|
|
| 13 |
Required local artifacts and environment:
|
| 14 |
|
| 15 |
- `.env` contains `COHERE_API_KEY`
|
| 16 |
+
- `.env` contains `OPENROUTER_API_KEY`
|
| 17 |
- `.env` contains `GEMINI_API_KEY` or `GOOGLE_API_KEY`
|
| 18 |
- `.env` contains `LANGSMITH_API_KEY`
|
| 19 |
- `.env` has `LANGSMITH_TRACING=true`
|
|
|
|
| 31 |
import os
|
| 32 |
for key in [
|
| 33 |
"COHERE_API_KEY",
|
| 34 |
+
"OPENROUTER_API_KEY",
|
| 35 |
"GEMINI_API_KEY",
|
| 36 |
"GOOGLE_API_KEY",
|
| 37 |
"LANGSMITH_API_KEY",
|
|
|
|
| 97 |
"transformers"
|
| 98 |
],
|
| 99 |
"enabledTools": [],
|
| 100 |
+
"model": "openrouter:deepseek/deepseek-v4-flash",
|
| 101 |
"includeReasoning": true,
|
| 102 |
"threadId": ""
|
| 103 |
}
|
|
@@ -75,7 +75,7 @@ class ApiTestCase(unittest.TestCase):
|
|
| 75 |
)
|
| 76 |
self.assertEqual(transformers["label"], "Transformers Docs")
|
| 77 |
self.assertEqual(transformers["shortLabel"], "Transformers")
|
| 78 |
-
self.assertEqual(body["model"], "
|
| 79 |
# DeepSeek direct is the default model, so only local KB tools
|
| 80 |
# are exposed until a provider with built-in web tools is selected.
|
| 81 |
tool_keys = {tool["key"] for tool in tools}
|
|
@@ -1226,7 +1226,9 @@ def live_chat_payload(
|
|
| 1226 |
"messages": messages,
|
| 1227 |
"sourceKeys": ["peft", "transformers"],
|
| 1228 |
"enabledTools": enabled_tools or [],
|
| 1229 |
-
"model": os.getenv(
|
|
|
|
|
|
|
| 1230 |
"includeReasoning": False,
|
| 1231 |
"threadId": thread_id,
|
| 1232 |
}
|
|
|
|
| 75 |
)
|
| 76 |
self.assertEqual(transformers["label"], "Transformers Docs")
|
| 77 |
self.assertEqual(transformers["shortLabel"], "Transformers")
|
| 78 |
+
self.assertEqual(body["model"], "openrouter:deepseek/deepseek-v4-flash")
|
| 79 |
# DeepSeek direct is the default model, so only local KB tools
|
| 80 |
# are exposed until a provider with built-in web tools is selected.
|
| 81 |
tool_keys = {tool["key"] for tool in tools}
|
|
|
|
| 1226 |
"messages": messages,
|
| 1227 |
"sourceKeys": ["peft", "transformers"],
|
| 1228 |
"enabledTools": enabled_tools or [],
|
| 1229 |
+
"model": os.getenv(
|
| 1230 |
+
"LIVE_API_E2E_MODEL", "openrouter:deepseek/deepseek-v4-flash"
|
| 1231 |
+
),
|
| 1232 |
"includeReasoning": False,
|
| 1233 |
"threadId": thread_id,
|
| 1234 |
}
|
|
@@ -8,8 +8,14 @@ from unittest.mock import MagicMock, patch
|
|
| 8 |
|
| 9 |
from langchain_core.messages import AIMessage, AIMessageChunk, HumanMessage, ToolMessage
|
| 10 |
|
| 11 |
-
from app.config import
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 12 |
from app.deepseek_chat import TutorChatDeepSeek
|
|
|
|
| 13 |
from app.chat_service import (
|
| 14 |
THREAD_IDLE_TTL_SECONDS,
|
| 15 |
_claim_kb_command_budget,
|
|
@@ -602,6 +608,50 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 602 |
{"thinking": {"type": "enabled"}},
|
| 603 |
)
|
| 604 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 605 |
def test_deepseek_thinking_tool_call_replays_reasoning_content(self) -> None:
|
| 606 |
model = TutorChatDeepSeek(
|
| 607 |
model="deepseek-v4-flash",
|
|
@@ -641,7 +691,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 641 |
{"GEMINI_API_KEY": "gemini-test-key"},
|
| 642 |
clear=True,
|
| 643 |
):
|
| 644 |
-
served = resolve_served_model_name(
|
| 645 |
model = build_chat_model(served)
|
| 646 |
|
| 647 |
self.assertEqual(served, FALLBACK_MODEL_NAME)
|
|
@@ -650,33 +700,33 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 650 |
def test_resolve_served_model_name_keeps_default_when_key_present(self) -> None:
|
| 651 |
with patch.dict(
|
| 652 |
os.environ,
|
| 653 |
-
{"
|
| 654 |
clear=True,
|
| 655 |
):
|
| 656 |
self.assertEqual(
|
| 657 |
-
resolve_served_model_name(
|
| 658 |
-
|
| 659 |
)
|
| 660 |
|
| 661 |
def test_rescue_model_covers_only_the_default_model(self) -> None:
|
| 662 |
with patch.dict(
|
| 663 |
os.environ,
|
| 664 |
{
|
| 665 |
-
"
|
| 666 |
"GEMINI_API_KEY": "gemini-test-key",
|
| 667 |
},
|
| 668 |
clear=True,
|
| 669 |
):
|
| 670 |
-
self.assertEqual(
|
| 671 |
-
|
| 672 |
-
|
| 673 |
-
# A deliberately selected non-default model must fail visibly.
|
| 674 |
self.assertIsNone(rescue_model_for("anthropic:claude-haiku-4-5"))
|
|
|
|
| 675 |
with patch.dict(
|
| 676 |
-
os.environ, {"
|
| 677 |
):
|
| 678 |
# No fallback key, no rescue.
|
| 679 |
-
self.assertIsNone(rescue_model_for(
|
| 680 |
|
| 681 |
def test_provider_outage_classification(self) -> None:
|
| 682 |
class ServerError(Exception):
|
|
@@ -1031,7 +1081,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1031 |
request = ChatRequest(
|
| 1032 |
query="ls the kb",
|
| 1033 |
source_keys=("peft",),
|
| 1034 |
-
model_name=
|
| 1035 |
include_reasoning=False,
|
| 1036 |
enabled_tools=(),
|
| 1037 |
)
|
|
@@ -1042,7 +1092,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1042 |
with (
|
| 1043 |
patch.dict(
|
| 1044 |
os.environ,
|
| 1045 |
-
{"
|
| 1046 |
clear=False,
|
| 1047 |
),
|
| 1048 |
patch("app.chat_service.build_agent", side_effect=fake_build_agent),
|
|
@@ -1051,9 +1101,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1051 |
events = asyncio.run(collect_events())
|
| 1052 |
|
| 1053 |
# The turn restarted on the fallback model and completed normally.
|
| 1054 |
-
self.assertEqual(
|
| 1055 |
-
built_models, [DEEPSEEK_DIRECT_MODEL_NAME, FALLBACK_MODEL_NAME]
|
| 1056 |
-
)
|
| 1057 |
types = [event.type for event in events]
|
| 1058 |
self.assertIn("thread_started", types)
|
| 1059 |
self.assertIn("context_stats", types)
|
|
@@ -1075,7 +1123,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1075 |
request = ChatRequest(
|
| 1076 |
query="ls the kb",
|
| 1077 |
source_keys=("peft",),
|
| 1078 |
-
model_name=
|
| 1079 |
include_reasoning=False,
|
| 1080 |
enabled_tools=(),
|
| 1081 |
)
|
|
@@ -1086,7 +1134,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1086 |
with (
|
| 1087 |
patch.dict(
|
| 1088 |
os.environ,
|
| 1089 |
-
{"
|
| 1090 |
clear=False,
|
| 1091 |
),
|
| 1092 |
patch("app.chat_service.build_agent", side_effect=fake_build_agent),
|
|
@@ -1099,14 +1147,12 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1099 |
events = asyncio.run(collect_events())
|
| 1100 |
|
| 1101 |
# The unfunded account rescued onto the fallback model and paged ops.
|
| 1102 |
-
self.assertEqual(
|
| 1103 |
-
built_models, [DEEPSEEK_DIRECT_MODEL_NAME, FALLBACK_MODEL_NAME]
|
| 1104 |
-
)
|
| 1105 |
self.assertEqual(len(alerts), 1)
|
| 1106 |
self.assertEqual(alerts[0][0], "account-state")
|
| 1107 |
self.assertIn("Insufficient Balance", alerts[0][1])
|
| 1108 |
stats = next(event for event in events if event.type == "context_stats")
|
| 1109 |
-
self.assertEqual(stats.data["requested_model"],
|
| 1110 |
self.assertEqual(stats.data["served_model"], FALLBACK_MODEL_NAME)
|
| 1111 |
self.assertTrue(stats.data["rescued"])
|
| 1112 |
|
|
@@ -1125,7 +1171,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1125 |
request = ChatRequest(
|
| 1126 |
query="ls the kb",
|
| 1127 |
source_keys=("peft",),
|
| 1128 |
-
model_name=
|
| 1129 |
include_reasoning=False,
|
| 1130 |
enabled_tools=(),
|
| 1131 |
)
|
|
@@ -1136,7 +1182,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1136 |
with (
|
| 1137 |
patch.dict(
|
| 1138 |
os.environ,
|
| 1139 |
-
{"
|
| 1140 |
clear=False,
|
| 1141 |
),
|
| 1142 |
patch("app.chat_service.build_agent", side_effect=fake_build_agent),
|
|
@@ -1146,7 +1192,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1146 |
asyncio.run(collect_events())
|
| 1147 |
|
| 1148 |
# 4xx never reaches the fallback; the primary was the only attempt.
|
| 1149 |
-
self.assertEqual(built_models, [
|
| 1150 |
|
| 1151 |
def test_stream_chat_emits_deepseek_reasoning_content(self) -> None:
|
| 1152 |
agent = FakeDeepSeekReasoningAgent([])
|
|
@@ -1155,7 +1201,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1155 |
request = ChatRequest(
|
| 1156 |
query="Find the course repository",
|
| 1157 |
source_keys=("peft",),
|
| 1158 |
-
model_name=
|
| 1159 |
include_reasoning=True,
|
| 1160 |
enabled_tools=(),
|
| 1161 |
)
|
|
@@ -1188,7 +1234,7 @@ class ChatServiceTestCase(unittest.TestCase):
|
|
| 1188 |
request = ChatRequest(
|
| 1189 |
query="Where should I store API keys?",
|
| 1190 |
source_keys=("langchain",),
|
| 1191 |
-
model_name=
|
| 1192 |
include_reasoning=False,
|
| 1193 |
enabled_tools=(),
|
| 1194 |
)
|
|
|
|
| 8 |
|
| 9 |
from langchain_core.messages import AIMessage, AIMessageChunk, HumanMessage, ToolMessage
|
| 10 |
|
| 11 |
+
from app.config import (
|
| 12 |
+
DEEPSEEK_DIRECT_MODEL_NAME,
|
| 13 |
+
DEFAULT_MODEL_NAME,
|
| 14 |
+
FALLBACK_MODEL_NAME,
|
| 15 |
+
OPENROUTER_DEEPSEEK_MODEL_NAME,
|
| 16 |
+
)
|
| 17 |
from app.deepseek_chat import TutorChatDeepSeek
|
| 18 |
+
from app.openrouter_chat import TutorChatOpenRouter
|
| 19 |
from app.chat_service import (
|
| 20 |
THREAD_IDLE_TTL_SECONDS,
|
| 21 |
_claim_kb_command_budget,
|
|
|
|
| 608 |
{"thinking": {"type": "enabled"}},
|
| 609 |
)
|
| 610 |
|
| 611 |
+
def test_openrouter_model_uses_adapter_with_explicit_reasoning_toggle(
|
| 612 |
+
self,
|
| 613 |
+
) -> None:
|
| 614 |
+
with patch.dict(
|
| 615 |
+
os.environ,
|
| 616 |
+
{"OPENROUTER_API_KEY": "openrouter-test-key"},
|
| 617 |
+
clear=True,
|
| 618 |
+
):
|
| 619 |
+
plain = build_chat_model(OPENROUTER_DEEPSEEK_MODEL_NAME)
|
| 620 |
+
thinking = build_chat_model(
|
| 621 |
+
OPENROUTER_DEEPSEEK_MODEL_NAME, include_thoughts=True
|
| 622 |
+
)
|
| 623 |
+
|
| 624 |
+
self.assertIsInstance(plain, TutorChatOpenRouter)
|
| 625 |
+
self.assertEqual(plain.model_name, "deepseek/deepseek-v4-flash")
|
| 626 |
+
self.assertEqual(plain.extra_body, {"reasoning": {"enabled": False}})
|
| 627 |
+
self.assertEqual(thinking.extra_body, {"reasoning": {"enabled": True}})
|
| 628 |
+
|
| 629 |
+
def test_openrouter_stream_preserves_reasoning_deltas(self) -> None:
|
| 630 |
+
model = TutorChatOpenRouter(
|
| 631 |
+
model="deepseek/deepseek-v4-flash",
|
| 632 |
+
api_key="openrouter-test-key",
|
| 633 |
+
base_url="https://openrouter.ai/api/v1",
|
| 634 |
+
)
|
| 635 |
+
# The delta shape OpenRouter actually streams (verified live
|
| 636 |
+
# 2026-08-25): thought text arrives in `reasoning`, not DeepSeek's
|
| 637 |
+
# `reasoning_content`.
|
| 638 |
+
chunk = {
|
| 639 |
+
"id": "gen-1",
|
| 640 |
+
"model": "deepseek/deepseek-v4-flash",
|
| 641 |
+
"choices": [
|
| 642 |
+
{"index": 0, "delta": {"content": "", "reasoning": "We compute "}}
|
| 643 |
+
],
|
| 644 |
+
}
|
| 645 |
+
|
| 646 |
+
generation_chunk = model._convert_chunk_to_generation_chunk(
|
| 647 |
+
chunk, AIMessageChunk, None
|
| 648 |
+
)
|
| 649 |
+
|
| 650 |
+
self.assertEqual(
|
| 651 |
+
generation_chunk.message.additional_kwargs["reasoning_content"],
|
| 652 |
+
"We compute ",
|
| 653 |
+
)
|
| 654 |
+
|
| 655 |
def test_deepseek_thinking_tool_call_replays_reasoning_content(self) -> None:
|
| 656 |
model = TutorChatDeepSeek(
|
| 657 |
model="deepseek-v4-flash",
|
|
|
|
| 691 |
{"GEMINI_API_KEY": "gemini-test-key"},
|
| 692 |
clear=True,
|
| 693 |
):
|
| 694 |
+
served = resolve_served_model_name(DEFAULT_MODEL_NAME)
|
| 695 |
model = build_chat_model(served)
|
| 696 |
|
| 697 |
self.assertEqual(served, FALLBACK_MODEL_NAME)
|
|
|
|
| 700 |
def test_resolve_served_model_name_keeps_default_when_key_present(self) -> None:
|
| 701 |
with patch.dict(
|
| 702 |
os.environ,
|
| 703 |
+
{"OPENROUTER_API_KEY": "openrouter-test-key"},
|
| 704 |
clear=True,
|
| 705 |
):
|
| 706 |
self.assertEqual(
|
| 707 |
+
resolve_served_model_name(DEFAULT_MODEL_NAME),
|
| 708 |
+
DEFAULT_MODEL_NAME,
|
| 709 |
)
|
| 710 |
|
| 711 |
def test_rescue_model_covers_only_the_default_model(self) -> None:
|
| 712 |
with patch.dict(
|
| 713 |
os.environ,
|
| 714 |
{
|
| 715 |
+
"OPENROUTER_API_KEY": "openrouter-test-key",
|
| 716 |
"GEMINI_API_KEY": "gemini-test-key",
|
| 717 |
},
|
| 718 |
clear=True,
|
| 719 |
):
|
| 720 |
+
self.assertEqual(rescue_model_for(DEFAULT_MODEL_NAME), FALLBACK_MODEL_NAME)
|
| 721 |
+
# A deliberately selected non-default model must fail visibly --
|
| 722 |
+
# including the retired first-party DeepSeek path.
|
|
|
|
| 723 |
self.assertIsNone(rescue_model_for("anthropic:claude-haiku-4-5"))
|
| 724 |
+
self.assertIsNone(rescue_model_for(DEEPSEEK_DIRECT_MODEL_NAME))
|
| 725 |
with patch.dict(
|
| 726 |
+
os.environ, {"OPENROUTER_API_KEY": "openrouter-test-key"}, clear=True
|
| 727 |
):
|
| 728 |
# No fallback key, no rescue.
|
| 729 |
+
self.assertIsNone(rescue_model_for(DEFAULT_MODEL_NAME))
|
| 730 |
|
| 731 |
def test_provider_outage_classification(self) -> None:
|
| 732 |
class ServerError(Exception):
|
|
|
|
| 1081 |
request = ChatRequest(
|
| 1082 |
query="ls the kb",
|
| 1083 |
source_keys=("peft",),
|
| 1084 |
+
model_name=DEFAULT_MODEL_NAME,
|
| 1085 |
include_reasoning=False,
|
| 1086 |
enabled_tools=(),
|
| 1087 |
)
|
|
|
|
| 1092 |
with (
|
| 1093 |
patch.dict(
|
| 1094 |
os.environ,
|
| 1095 |
+
{"OPENROUTER_API_KEY": "k1", "GEMINI_API_KEY": "k2"},
|
| 1096 |
clear=False,
|
| 1097 |
),
|
| 1098 |
patch("app.chat_service.build_agent", side_effect=fake_build_agent),
|
|
|
|
| 1101 |
events = asyncio.run(collect_events())
|
| 1102 |
|
| 1103 |
# The turn restarted on the fallback model and completed normally.
|
| 1104 |
+
self.assertEqual(built_models, [DEFAULT_MODEL_NAME, FALLBACK_MODEL_NAME])
|
|
|
|
|
|
|
| 1105 |
types = [event.type for event in events]
|
| 1106 |
self.assertIn("thread_started", types)
|
| 1107 |
self.assertIn("context_stats", types)
|
|
|
|
| 1123 |
request = ChatRequest(
|
| 1124 |
query="ls the kb",
|
| 1125 |
source_keys=("peft",),
|
| 1126 |
+
model_name=DEFAULT_MODEL_NAME,
|
| 1127 |
include_reasoning=False,
|
| 1128 |
enabled_tools=(),
|
| 1129 |
)
|
|
|
|
| 1134 |
with (
|
| 1135 |
patch.dict(
|
| 1136 |
os.environ,
|
| 1137 |
+
{"OPENROUTER_API_KEY": "k1", "GEMINI_API_KEY": "k2"},
|
| 1138 |
clear=False,
|
| 1139 |
),
|
| 1140 |
patch("app.chat_service.build_agent", side_effect=fake_build_agent),
|
|
|
|
| 1147 |
events = asyncio.run(collect_events())
|
| 1148 |
|
| 1149 |
# The unfunded account rescued onto the fallback model and paged ops.
|
| 1150 |
+
self.assertEqual(built_models, [DEFAULT_MODEL_NAME, FALLBACK_MODEL_NAME])
|
|
|
|
|
|
|
| 1151 |
self.assertEqual(len(alerts), 1)
|
| 1152 |
self.assertEqual(alerts[0][0], "account-state")
|
| 1153 |
self.assertIn("Insufficient Balance", alerts[0][1])
|
| 1154 |
stats = next(event for event in events if event.type == "context_stats")
|
| 1155 |
+
self.assertEqual(stats.data["requested_model"], DEFAULT_MODEL_NAME)
|
| 1156 |
self.assertEqual(stats.data["served_model"], FALLBACK_MODEL_NAME)
|
| 1157 |
self.assertTrue(stats.data["rescued"])
|
| 1158 |
|
|
|
|
| 1171 |
request = ChatRequest(
|
| 1172 |
query="ls the kb",
|
| 1173 |
source_keys=("peft",),
|
| 1174 |
+
model_name=DEFAULT_MODEL_NAME,
|
| 1175 |
include_reasoning=False,
|
| 1176 |
enabled_tools=(),
|
| 1177 |
)
|
|
|
|
| 1182 |
with (
|
| 1183 |
patch.dict(
|
| 1184 |
os.environ,
|
| 1185 |
+
{"OPENROUTER_API_KEY": "k1", "GEMINI_API_KEY": "k2"},
|
| 1186 |
clear=False,
|
| 1187 |
),
|
| 1188 |
patch("app.chat_service.build_agent", side_effect=fake_build_agent),
|
|
|
|
| 1192 |
asyncio.run(collect_events())
|
| 1193 |
|
| 1194 |
# 4xx never reaches the fallback; the primary was the only attempt.
|
| 1195 |
+
self.assertEqual(built_models, [DEFAULT_MODEL_NAME])
|
| 1196 |
|
| 1197 |
def test_stream_chat_emits_deepseek_reasoning_content(self) -> None:
|
| 1198 |
agent = FakeDeepSeekReasoningAgent([])
|
|
|
|
| 1201 |
request = ChatRequest(
|
| 1202 |
query="Find the course repository",
|
| 1203 |
source_keys=("peft",),
|
| 1204 |
+
model_name=DEFAULT_MODEL_NAME,
|
| 1205 |
include_reasoning=True,
|
| 1206 |
enabled_tools=(),
|
| 1207 |
)
|
|
|
|
| 1234 |
request = ChatRequest(
|
| 1235 |
query="Where should I store API keys?",
|
| 1236 |
source_keys=("langchain",),
|
| 1237 |
+
model_name=DEFAULT_MODEL_NAME,
|
| 1238 |
include_reasoning=False,
|
| 1239 |
enabled_tools=(),
|
| 1240 |
)
|
|
@@ -11,11 +11,14 @@ import pytest
|
|
| 11 |
from app import config
|
| 12 |
|
| 13 |
|
| 14 |
-
def
|
| 15 |
-
assert config.DEFAULT_MODEL_NAME == config.
|
| 16 |
-
assert config.DEFAULT_MODEL_NAME == "
|
| 17 |
assert config.FALLBACK_MODEL_NAME == "google-genai:gemini-3.5-flash-lite"
|
| 18 |
assert config.AVAILABLE_MODELS[0]["id"] == config.DEFAULT_MODEL_NAME
|
|
|
|
|
|
|
|
|
|
| 19 |
|
| 20 |
|
| 21 |
def test_gemini_fallback_model_is_deliberately_not_selectable() -> None:
|
|
@@ -25,12 +28,12 @@ def test_gemini_fallback_model_is_deliberately_not_selectable() -> None:
|
|
| 25 |
our custom tools) this is a product choice, one selectable model plus a
|
| 26 |
rescue path, rather than the hard API constraint it was for 2.5-flash. The
|
| 27 |
fallback still never receives web tools: build_agent binds them from the
|
| 28 |
-
requested model (DeepSeek), which has none.
|
| 29 |
"""
|
| 30 |
selectable = [model["id"] for model in config.AVAILABLE_MODELS]
|
| 31 |
|
| 32 |
assert config.FALLBACK_MODEL_NAME not in selectable
|
| 33 |
-
assert selectable == ["
|
| 34 |
|
| 35 |
|
| 36 |
def _patched_bundle(tmp_path: Path) -> ExitStack:
|
|
|
|
| 11 |
from app import config
|
| 12 |
|
| 13 |
|
| 14 |
+
def test_default_chat_model_is_openrouter_deepseek_with_gemini_fallback() -> None:
|
| 15 |
+
assert config.DEFAULT_MODEL_NAME == config.OPENROUTER_DEEPSEEK_MODEL_NAME
|
| 16 |
+
assert config.DEFAULT_MODEL_NAME == "openrouter:deepseek/deepseek-v4-flash"
|
| 17 |
assert config.FALLBACK_MODEL_NAME == "google-genai:gemini-3.5-flash-lite"
|
| 18 |
assert config.AVAILABLE_MODELS[0]["id"] == config.DEFAULT_MODEL_NAME
|
| 19 |
+
# The retired first-party path stays addressable for experiments but is
|
| 20 |
+
# neither the default nor selectable.
|
| 21 |
+
assert config.DEEPSEEK_DIRECT_MODEL_NAME == "deepseek:deepseek-v4-flash"
|
| 22 |
|
| 23 |
|
| 24 |
def test_gemini_fallback_model_is_deliberately_not_selectable() -> None:
|
|
|
|
| 28 |
our custom tools) this is a product choice, one selectable model plus a
|
| 29 |
rescue path, rather than the hard API constraint it was for 2.5-flash. The
|
| 30 |
fallback still never receives web tools: build_agent binds them from the
|
| 31 |
+
requested model (DeepSeek via OpenRouter), which has none.
|
| 32 |
"""
|
| 33 |
selectable = [model["id"] for model in config.AVAILABLE_MODELS]
|
| 34 |
|
| 35 |
assert config.FALLBACK_MODEL_NAME not in selectable
|
| 36 |
+
assert selectable == ["openrouter:deepseek/deepseek-v4-flash"]
|
| 37 |
|
| 38 |
|
| 39 |
def _patched_bundle(tmp_path: Path) -> ExitStack:
|
|
@@ -90,11 +90,24 @@ class MemoryPresetResolutionTests(unittest.TestCase):
|
|
| 90 |
production_memory_preset_name("openai:gpt-5.6"),
|
| 91 |
PRODUCTION_FALLBACK_MEMORY_PRESET,
|
| 92 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 93 |
|
| 94 |
def test_prod_v2_rejects_short_context_providers(self) -> None:
|
| 95 |
self.assertTrue(
|
| 96 |
memory_preset_supports_model("prod_v2", "deepseek:deepseek-v4-flash")
|
| 97 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 98 |
self.assertTrue(
|
| 99 |
memory_preset_supports_model(
|
| 100 |
"prod_v2", "google-genai:gemini-3.5-flash-lite"
|
|
|
|
| 90 |
production_memory_preset_name("openai:gpt-5.6"),
|
| 91 |
PRODUCTION_FALLBACK_MEMORY_PRESET,
|
| 92 |
)
|
| 93 |
+
self.assertEqual(
|
| 94 |
+
resolve_memory_preset(
|
| 95 |
+
None, model_name="openrouter:deepseek/deepseek-v4-flash"
|
| 96 |
+
).name,
|
| 97 |
+
"prod_v2",
|
| 98 |
+
)
|
| 99 |
|
| 100 |
def test_prod_v2_rejects_short_context_providers(self) -> None:
|
| 101 |
self.assertTrue(
|
| 102 |
memory_preset_supports_model("prod_v2", "deepseek:deepseek-v4-flash")
|
| 103 |
)
|
| 104 |
+
# The default chat path serves DeepSeek through OpenRouter; prod_v2
|
| 105 |
+
# must not silently downgrade it to the legacy preset.
|
| 106 |
+
self.assertTrue(
|
| 107 |
+
memory_preset_supports_model(
|
| 108 |
+
"prod_v2", "openrouter:deepseek/deepseek-v4-flash"
|
| 109 |
+
)
|
| 110 |
+
)
|
| 111 |
self.assertTrue(
|
| 112 |
memory_preset_supports_model(
|
| 113 |
"prod_v2", "google-genai:gemini-3.5-flash-lite"
|