omarsol Claude Fable 5 commited on
Commit
6cd03bd
·
1 Parent(s): 117a59c

Remove Claude Haiku 4.5 from the selectable model list

Browse files

DeepSeek V4 Flash is now the only entry in AVAILABLE_MODELS; the picker
and /api/chat validation both derive from it, so Haiku is gone from the
UI and 422s ("Unknown model") for direct API callers. Anthropic support
stays in code (build_chat_model, tool catalog, pricing, preset rules),
so re-adding it later is a one-line revert. Tests that needed a
selectable model with web toggles re-admit Haiku via a patched
AVAILABLE_MODELS.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Files changed (4) hide show
  1. AGENTS.md +2 -2
  2. app/config.py +0 -1
  3. tests/test_api.py +32 -12
  4. tests/test_config.py +1 -4
AGENTS.md CHANGED
@@ -6,7 +6,7 @@ This is the **canonical, tool-agnostic** instruction file for the repo. `CLAUDE.
6
 
7
  AI tutor for applied AI, LLMs, RAG, and Python. **Agentic RAG**: a LangChain/LangGraph agent grounds answers in a curated corpus of course + library docs, can browse a local file-based knowledge base, and (optionally) search the live web. One frontend: a **Next.js** UI (`frontend/`), served by a **FastAPI** backend (`app/api.py`) that streams in the Vercel AI SDK UI-message protocol. (A Gradio UI existed historically; it was removed to keep one rendering path.)
8
 
9
- ChromaDB for vectors; Cohere for embeddings/rerank; chat model is provider-configurable (DeepSeek V4 Flash through the first-party API by default, with a rescue-only in-app fallback to Gemini 2.5 Flash when a Gemini key is set; Claude Haiku 4.5 is also selectable; OpenAI and OpenRouter-compatible models are supported in code). **A conversation cannot change model mid-thread** — checkpoints store provider-native message blocks that no other provider can replay; see `build_chat_model` in `app/chat_service.py`. Python ≥3.13, managed with `uv`.
10
 
11
  ## Key URLs
12
 
@@ -62,7 +62,7 @@ Runtime guidance the agent follows is in `data/kb/AGENTS.md` (injected into the
62
 
63
  ## Sources & config
64
 
65
- `data/scraping_scripts/source_registry.py` is the **single source of truth** for sources (`SOURCE_CONFIGS`, key groupings, UI labels, defaults); `app/config.py` re-exports them and the frontend derives the picker from it (via `/api/tools`). Docs sources ingest via the GitHub API or `llms.txt`; course sources are Notion exports. To add a source: add it to the registry (+ the relevant grouping tuples), then run the matching workflow — no separate UI edit needed. Models live in `config.AVAILABLE_MODELS` (default `deepseek:deepseek-v4-flash`; Claude Haiku 4.5 is also selectable; OpenAI is supported in code). **`GEMINI_FALLBACK_MODEL_NAME` (`google-genai:gemini-2.5-flash`) is deliberately NOT in `AVAILABLE_MODELS`**: it serves real traffic as the DeepSeek rescue path but must never be user-selectable, because pre-Gemini-3 models cannot combine Gemini's built-in web tools with our two custom tools (that needs Gemini 3+ "tool context circulation"), so selecting it would be one web-search toggle away from a 400. The fallback is safe only because it never receives web tools. `build_chat_model` in `app/chat_service.py` accepts `deepseek:` for the first-party API, `openrouter:` for compatible experiment models, and `ollama:` for local SLM experiments, with pricing in `app/telemetry.MODEL_PRICING`; for the default DeepSeek model it also wires an in-app fallback to `google-genai:gemini-2.5-flash` whenever a Gemini key is configured (and substitutes it outright if `DEEPSEEK_API_KEY` is missing).
66
 
67
  Memory policy is deliberately owned by `app/memory_presets.py`, not `app/config.py`. `PRODUCTION_MEMORY_PRESET` is the single primary-policy switch, `PRODUCTION_FALLBACK_MEMORY_PRESET` covers providers that cannot use it safely, and `PRODUCTION_LONG_CONTEXT_PROVIDERS` defines the model-aware allowlist. Keep historical presets immutable because saved eval runs refer to them by name.
68
 
 
6
 
7
  AI tutor for applied AI, LLMs, RAG, and Python. **Agentic RAG**: a LangChain/LangGraph agent grounds answers in a curated corpus of course + library docs, can browse a local file-based knowledge base, and (optionally) search the live web. One frontend: a **Next.js** UI (`frontend/`), served by a **FastAPI** backend (`app/api.py`) that streams in the Vercel AI SDK UI-message protocol. (A Gradio UI existed historically; it was removed to keep one rendering path.)
8
 
9
+ ChromaDB for vectors; Cohere for embeddings/rerank; chat model is provider-configurable (DeepSeek V4 Flash through the first-party API by default, with a rescue-only in-app fallback to Gemini 2.5 Flash when a Gemini key is set; Anthropic, OpenAI, and OpenRouter-compatible models are supported in code but not user-selectable). **A conversation cannot change model mid-thread** — checkpoints store provider-native message blocks that no other provider can replay; see `build_chat_model` in `app/chat_service.py`. Python ≥3.13, managed with `uv`.
10
 
11
  ## Key URLs
12
 
 
62
 
63
  ## Sources & config
64
 
65
+ `data/scraping_scripts/source_registry.py` is the **single source of truth** for sources (`SOURCE_CONFIGS`, key groupings, UI labels, defaults); `app/config.py` re-exports them and the frontend derives the picker from it (via `/api/tools`). Docs sources ingest via the GitHub API or `llms.txt`; course sources are Notion exports. To add a source: add it to the registry (+ the relevant grouping tuples), then run the matching workflow — no separate UI edit needed. Models live in `config.AVAILABLE_MODELS` (currently only the default `deepseek:deepseek-v4-flash`; Anthropic Claude Haiku 4.5 and OpenAI remain supported in code but are not selectable). **`GEMINI_FALLBACK_MODEL_NAME` (`google-genai:gemini-2.5-flash`) is deliberately NOT in `AVAILABLE_MODELS`**: it serves real traffic as the DeepSeek rescue path but must never be user-selectable, because pre-Gemini-3 models cannot combine Gemini's built-in web tools with our two custom tools (that needs Gemini 3+ "tool context circulation"), so selecting it would be one web-search toggle away from a 400. The fallback is safe only because it never receives web tools. `build_chat_model` in `app/chat_service.py` accepts `deepseek:` for the first-party API, `openrouter:` for compatible experiment models, and `ollama:` for local SLM experiments, with pricing in `app/telemetry.MODEL_PRICING`; for the default DeepSeek model it also wires an in-app fallback to `google-genai:gemini-2.5-flash` whenever a Gemini key is configured (and substitutes it outright if `DEEPSEEK_API_KEY` is missing).
66
 
67
  Memory policy is deliberately owned by `app/memory_presets.py`, not `app/config.py`. `PRODUCTION_MEMORY_PRESET` is the single primary-policy switch, `PRODUCTION_FALLBACK_MEMORY_PRESET` covers providers that cannot use it safely, and `PRODUCTION_LONG_CONTEXT_PROVIDERS` defines the model-aware allowlist. Keep historical presets immutable because saved eval runs refer to them by name.
68
 
app/config.py CHANGED
@@ -99,7 +99,6 @@ AVAILABLE_MODELS: tuple[dict[str, str], ...] = (
99
  "id": DEEPSEEK_DIRECT_MODEL_NAME,
100
  "label": "DeepSeek V4 Flash",
101
  },
102
- {"id": "anthropic:claude-haiku-4-5", "label": "Claude Haiku 4.5"},
103
  )
104
 
105
 
 
99
  "id": DEEPSEEK_DIRECT_MODEL_NAME,
100
  "label": "DeepSeek V4 Flash",
101
  },
 
102
  )
103
 
104
 
tests/test_api.py CHANGED
@@ -208,16 +208,22 @@ class ApiTestCase(unittest.TestCase):
208
  ],
209
  "sourceKeys": ["langchain", "transformers"],
210
  "enabledTools": ["web_search", "not_a_real_tool"],
211
- # A selectable model that offers web_search; gemini-2.5-flash is
212
- # fallback-only and now 422s here (see test_config.py).
 
 
213
  "model": "anthropic:claude-haiku-4-5",
214
  "threadId": "thread_0",
215
  }
 
 
 
216
 
217
- with patch("app.api.stream_chat", fake_stream_chat):
218
- with TestClient(app) as client:
219
- with client.stream("POST", "/api/chat", json=payload) as response:
220
- body = "".join(response.iter_text())
 
221
 
222
  self.assertEqual(response.status_code, 200)
223
  self.assertEqual(
@@ -1100,14 +1106,28 @@ class ApiTestCase(unittest.TestCase):
1100
  )
1101
  self.assertEqual(raised.exception.status_code, 422)
1102
 
1103
- with self.assertRaises(HTTPException) as incompatible:
 
 
1104
  build_chat_request(
1105
- ApiChatRequest(
1106
- query="What is RAG?",
1107
- model="anthropic:claude-haiku-4-5",
1108
- memoryPreset="prod_v2",
1109
- )
1110
  )
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1111
  self.assertEqual(incompatible.exception.status_code, 422)
1112
  self.assertIn("does not support provider", incompatible.exception.detail)
1113
 
 
208
  ],
209
  "sourceKeys": ["langchain", "transformers"],
210
  "enabledTools": ["web_search", "not_a_real_tool"],
211
+ # The toggle passthrough below needs a model that offers web_search,
212
+ # and no selectable model does today (DeepSeek has no web toggles;
213
+ # gemini-2.5-flash is fallback-only). Re-admit Claude Haiku for this
214
+ # request only.
215
  "model": "anthropic:claude-haiku-4-5",
216
  "threadId": "thread_0",
217
  }
218
+ selectable_with_web_search = (
219
+ {"id": "anthropic:claude-haiku-4-5", "label": "Claude Haiku 4.5"},
220
+ )
221
 
222
+ with patch("app.api.AVAILABLE_MODELS", selectable_with_web_search):
223
+ with patch("app.api.stream_chat", fake_stream_chat):
224
+ with TestClient(app) as client:
225
+ with client.stream("POST", "/api/chat", json=payload) as response:
226
+ body = "".join(response.iter_text())
227
 
228
  self.assertEqual(response.status_code, 200)
229
  self.assertEqual(
 
1106
  )
1107
  self.assertEqual(raised.exception.status_code, 422)
1108
 
1109
+ # Claude Haiku was removed from AVAILABLE_MODELS, so the model check
1110
+ # rejects it before the preset check gets a say.
1111
+ with self.assertRaises(HTTPException) as unselectable:
1112
  build_chat_request(
1113
+ ApiChatRequest(query="What is RAG?", model="anthropic:claude-haiku-4-5")
 
 
 
 
1114
  )
1115
+ self.assertEqual(unselectable.exception.status_code, 422)
1116
+ self.assertEqual(unselectable.exception.detail, "Unknown model")
1117
+
1118
+ # The preset-provider incompatibility 422 needs a selectable model
1119
+ # outside the long-context allowlist; none exists today, so re-admit
1120
+ # Haiku for this request only.
1121
+ haiku = ({"id": "anthropic:claude-haiku-4-5", "label": "Claude Haiku 4.5"},)
1122
+ with patch("app.api.AVAILABLE_MODELS", haiku):
1123
+ with self.assertRaises(HTTPException) as incompatible:
1124
+ build_chat_request(
1125
+ ApiChatRequest(
1126
+ query="What is RAG?",
1127
+ model="anthropic:claude-haiku-4-5",
1128
+ memoryPreset="prod_v2",
1129
+ )
1130
+ )
1131
  self.assertEqual(incompatible.exception.status_code, 422)
1132
  self.assertIn("does not support provider", incompatible.exception.detail)
1133
 
tests/test_config.py CHANGED
@@ -30,10 +30,7 @@ def test_gemini_fallback_model_is_deliberately_not_selectable() -> None:
30
  selectable = [model["id"] for model in config.AVAILABLE_MODELS]
31
 
32
  assert config.GEMINI_FALLBACK_MODEL_NAME not in selectable
33
- assert selectable == [
34
- "deepseek:deepseek-v4-flash",
35
- "anthropic:claude-haiku-4-5",
36
- ]
37
 
38
 
39
  def _patched_bundle(tmp_path: Path) -> ExitStack:
 
30
  selectable = [model["id"] for model in config.AVAILABLE_MODELS]
31
 
32
  assert config.GEMINI_FALLBACK_MODEL_NAME not in selectable
33
+ assert selectable == ["deepseek:deepseek-v4-flash"]
 
 
 
34
 
35
 
36
  def _patched_bundle(tmp_path: Path) -> ExitStack: