Henry Ndubuaku commited on
Commit
6df988a
·
verified ·
1 Parent(s): 3504e19

Model card: add handoff benchmarks and routing-quality (AUROC) results

Browse files
Files changed (1) hide show
  1. README.md +48 -2
README.md CHANGED
@@ -15,8 +15,7 @@ A small, on-device model is fast and private, but sometimes wrong. At Cactus we
15
  post-train models to *know when they are wrong*: we ship probes inside the
16
  checkpoint that score every answer with a **confidence** between 0 and 1,
17
  returned as structured data (never parsed out of the answer text). Answer
18
- on-device when confidence is high; re-route to a bigger model when it's low —
19
- `0.85` is a good threshold:
20
 
21
  ```python
22
  if confidence < 0.85:
@@ -30,6 +29,26 @@ checkpoint (same keys); the repo adds eleven probe tensors, a remote-code model
30
  class and a `custom_generate` recipe. **Stock engine commands work unchanged** —
31
  you only add `--trust-remote-code` / `trust_remote_code=True`.
32
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
33
  ## Quickstart
34
 
35
  ```python
@@ -135,6 +154,33 @@ sequences = model.generate(**inputs, max_new_tokens=512, emit_trailer=False)
135
  | `model*.safetensors` | base weights (identical keys) + `handoff_probe.*` tensors |
136
  | `gemma_4_e2b_it_hybrid.py` | single-file `mlx-lm` model, wired via config.json's `model_file` |
137
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
138
  ## All formats
139
 
140
  All Cactus Hybrid builds live in the
 
15
  post-train models to *know when they are wrong*: we ship probes inside the
16
  checkpoint that score every answer with a **confidence** between 0 and 1,
17
  returned as structured data (never parsed out of the answer text). Answer
18
+ on-device when confidence is high; re-route to a bigger model when it's low:
 
19
 
20
  ```python
21
  if confidence < 0.85:
 
29
  class and a `custom_generate` recipe. **Stock engine commands work unchanged** —
30
  you only add `--trust-remote-code` / `trust_remote_code=True`.
31
 
32
+ ## Benchmarks
33
+
34
+ Gemma 4 E2B Hybrid, the smallest Gemma model, matches Gemini 3.1 Flash-Lite on
35
+ most benchmarks by routing only 15–35% of queries to Flash-Lite and running the
36
+ rest itself:
37
+
38
+ | Benchmark | Handoff to match Flash-Lite (FP16) | At 4-bit | At 3-bit |
39
+ |---|---|---|---|
40
+ | ChartQA | 15–20% | 25–30% | 40–50% |
41
+ | MMBench | 30–35% | 40–45% | 50–55% |
42
+ | LibriSpeech | 25–30% | 35–40% | 55–65% |
43
+ | GigaSpeech | 30–35% | 40–45% | 50–55% |
44
+ | MMAU | 30–35% | 35–40% | 50–55% |
45
+ | MMLU-Pro | 45–55% | ~90% | n/a |
46
+
47
+ Quantisation quality is measured on
48
+ [Cactus Quants](https://github.com/cactus-compute/cactus/blob/main/docs/cactus_quants.md),
49
+ which performs well at uniform quantization; developers are encouraged to
50
+ benchmark Unsloth, GGUF, and MLX quantization independently.
51
+
52
  ## Quickstart
53
 
54
  ```python
 
154
  | `model*.safetensors` | base weights (identical keys) + `handoff_probe.*` tensors |
155
  | `gemma_4_e2b_it_hybrid.py` | single-file `mlx-lm` model, wired via config.json's `model_file` |
156
 
157
+ ## Routing quality (AUROC)
158
+
159
+ AUROC measures how well the probe separates wrong answers from right ones
160
+ (higher = better, 0.5 is random, 1.0 is perfect):
161
+
162
+ | Hold-out | Modality | Cactus Hybrid | Token Entropy |
163
+ |---|---|---|---|
164
+ | MMLU | text MCQ | **0.770** | 0.697 |
165
+ | MMLU-Pro | text MCQ | **0.771** | 0.692 |
166
+ | ARC-Easy | text MCQ | **0.888** | 0.655 |
167
+ | ARC-Challenge | text MCQ | **0.834** | 0.646 |
168
+ | GSM8K (3-shot) | text gen | **0.782** | 0.731 |
169
+ | MMBench-EN-Dev | vision MCQ | **0.840** | 0.435 |
170
+ | ChartQA | vision QA | **0.779** | 0.615 |
171
+ | DocVQA | vision QA | **0.781** | 0.512 |
172
+ | MMAU | audio MCQ | **0.789** | 0.517 |
173
+ | GigaSpeech | audio | **0.876** | 0.343 |
174
+ | Earnings-22 | audio | **0.839** | 0.323 |
175
+ | LibriSpeech | audio | **0.822** | 0.427 |
176
+ | **Mean** | | **0.814** | **0.549** |
177
+
178
+ The strongest result: the probe was trained on **zero audio data**, yet achieves
179
+ 0.79–0.88 AUROC on four audio benchmarks (two transcription, one audio MCQ, one
180
+ out-of-domain transcription). This rules out surface-level explanations: the
181
+ probe is reading a modality-independent correctness signal from the hidden
182
+ state, not memorizing patterns from training data.
183
+
184
  ## All formats
185
 
186
  All Cactus Hybrid builds live in the