System One for Gemma 4 12B: answer now, reason first or ask back (work in progress)

A router in front of an unmodified Gemma 4 12B that decides, per request, whether to answer now in one line, reason first with a worked reply, or ask back one short question when the request does not settle the answer. In the spirit of Laya and Jev, but in free text and on top of an unchanged Gemma.

Everything runs with Elixir/Nx from olafura/gemma-4-mic-transcribe (mix gemma.system_one), measured on an AMD Strix Halo through EXLA/ROCm and on Nvidia L4 and A100 GPUs through EXLA/CUDA. This is a WIP snapshot, not a release. The base weights are not in this repo and are never modified.

The repo holds two things:

path what size
ask-probe/ the router's ask-back probe (recommended) 15 kB
parameters.safetensors, manifest.etf the earlier round-3 expert, added inside the last three decoder layers 283 MB

Quick start

From a checkout of the repo with the packed 12B prefix/tail artifacts built (see its README, "Extracting a decoder block"), on a machine with the ROCm or CUDA EXLA client:

export XLA_FLAGS='--xla_gpu_autotune_level=0 --xla_gpu_enable_command_buffer= --xla_gpu_enable_triton_gemm=false'
hf download olafura/gemma4-12b-system-one --include 'ask-probe/*' --local-dir artifacts/system-one

# a file of requests: System One items or {"id", "prompt", "answer": "number"}, one per line
mix gemma.system_one route --input requests.jsonl --output routed.jsonl

# or a server that streams each reply as it is written
mix gemma.system_one serve --port 7860
curl -N -d '{"prompt": "What is 17 * 23?", "answer": "number"}' localhost:7860/route

A System One item is a state, a question and options:

{"id": "voi01", "state": {"nlu_log": "text=\"send it to the printer upstairs\" | printers=[HP-2200@first-floor, Brother-L@second-floor] | job=\"contract.pdf\""},
 "question": "Pick the printer.", "options": {"hp_2200": "The HP laser printer", "brother_l": "The Brother laser printer"}}

The question can be spoken instead: a row with "audio": "q.wav" (16 kHz mono, relative to the input file) keeps the state and options as text and puts the WAV in the audio slot after them, through Gemma's own audio encoder. --audio-seconds (default 8) is the audio bucket the WAV is cut and padded to.

How the router decides

  1. The request is rendered in a direct form that asks for one line, Answer: <x>, and the prefix (layers 0–44) is run over it.
  2. Ask back when the ask-back probe scores above 0.8 (next section).
  3. Otherwise the direct answer is generated, and its confidence is the lowest probability among the answer tokens. Answer now at 0.9 or above.
  4. Otherwise reason: the request again, ending in "End your reply with a line of the form 'Answer: '", with room for up to 768 tokens.

How ask back works

Before Gemma answers, a small probe reads Gemma's internal state for the request, and if it scores above 0.8, Gemma writes one clarifying question instead of an answer.

  1. Read. The prefix (layers 0–44) runs over the direct prompt. The output at the last prompt token is a vector of 3840 numbers: what Gemma has made of the request just before it would start answering.
  2. Score. The probe is a logistic regression on that vector: score = sigmoid(vector · weight + bias), between 0 and 1. It reads Gemma's state, not the text, so it costs one dot product on top of the prefix pass.
  3. Decide. Above 0.8 the router asks back; no answer is generated.
  4. Ask. The same request is rendered again, ending in "The information above does not settle this. Reply with only the one short question you would ask to settle it, and nothing else." Gemma writes the question (up to 64 tokens), in the request's language.

Nothing in Gemma is changed: the probe only reads the residual stream, and the question is Gemma's own ordinary generation.

Where the probe comes from. It was trained on 3,689 prompts in the direct form, 1,494 of which should be asked back:

source rows
twin pairs: a decidable item and its twin with one detail removed 2,389
spoken twins, the question as a WAV 600
replay prompts, ordinary requests to base Gemma 400
GSM8K and ARC questions, which should never ask 300

At 0.8, 0.9% of the training prompts that should not ask score above it, and 50% of the ones that should. The threshold trades missed asks for needless ones: a needless question costs the user a turn, a missed one lets Gemma guess.

Where it fails.

  1. It misses unclear requests in domains it has not seen. Those go on to answer or reason. Retraining on your own domains is the fix.
  2. It keys on the wording it was trained on (see Accuracy).
  3. Scores near 0.8 can flip between GPUs or builds, since int4 kernels round differently: one request scored 0.78 before a kernel change and 0.81 after it.
  4. With --expert, requests just under 0.8 also get a second opinion (below).

Examples

Real rows from route on held-out requests (never trained on). Each output row repeats the input's fields and adds the route, the reply and what the decision was based on.

A twin pair: answer now, and ask back. The two requests differ in one detail: whether the printers are on different floors.

{"id": "ho-voi01-d", "state": {"nlu_log": "text=\"send it to the printer upstairs\" | printers=[HP-2200@first-floor, Brother-L@second-floor] | job=\"contract.pdf\""},
 "question": "Pick the printer.", "options": {"hp_2200": "The HP laser printer", "brother_l": "The Brother laser printer"}}
{"id": "ho-voi01-u", "state": {"nlu_log": "text=\"send it to the printer upstairs\" | printers=[HP-2200@second-floor, Brother-L@second-floor] | job=\"contract.pdf\""},
 "question": "Pick the printer.", "options": {"hp_2200": "The HP laser printer", "brother_l": "The Brother laser printer"}}
{"id": "ho-voi01-d", "route": "answer", "reply": "Answer: brother_l", "ask_score": 0.21, "confidence": 0.9997, "asked_by": null}
{"id": "ho-voi01-u", "route": "ask", "reply": "Which printer should I use, the Brother or the HP?", "ask_score": 0.999, "confidence": null, "asked_by": "probe"}

The first scores 0.21 on the probe, so the direct answer is generated; its least likely token has probability 0.9997, above 0.9, so it is kept. The second scores 0.999, so no answer is generated at all: Gemma is asked for the question that would settle it.

Reason when the direct answer is unsure. A plain question works too, with "answer": "number" asking for a number instead of an option:

{"id": "g2-011", "prompt": "A bag of buttons had 21 buttons in it. Seven buttons had two holes and the rest had four holes. How many holes were in all the buttons in the bag?", "answer": "number"}
{"id": "g2-011", "route": "reason", "ask_score": 0.00009, "direct_reply": "Answer: 63", "confidence": 0.68,
 "reply": "To find the total number of holes in all the buttons, we need to calculate ...\n\n4.  **Calculate the total number of holes:**\n    Add the holes from both types of buttons together:\n    $14 \\text{ holes} + 56 \\text{ holes} = 70 \\text{ holes}$.\n\nAnswer: 70"}

The direct answer was wrong (63) and unsure (0.68), so the request was run again with room to work; the worked reply gets 70. An easy one stays on the fast path: "Tracy used a piece of wire 4 feet long … cut into pieces 6 inches long. How many pieces?" goes to answer with Answer: 8 at 0.9999.

Spoken, in Spanish. The state and options are text, the question is a WAV (the spoken form of "¿A qué lista se añade la leche?"):

{"id": "ho-voi09-u", "state": {"orden": "añade leche a la lista", "listas": {"compra_semanal": {"activa": true}, "fiesta": {"activa": true}}},
 "options": {"compra_semanal": "La lista de la compra semanal", "fiesta": "La lista para la fiesta"}, "audio": "ho-voi09.wav"}
{"id": "ho-voi09-u", "route": "ask", "reply": "¿A qué lista quieres añadir la leche?", "ask_score": 0.97, "ms": 2850}

A miss. The probe's errors look like this: both interviewers are free on Tuesday morning, but the candidate's note ("mornings, some exceptions") leaves it open. The probe scores 0.67, under 0.8, and Gemma answers confidently:

{"id": "ho-sch16-u", "state": {"interview": {"stage": "final", "panel": ["ana", "rob"]}, "availability": {"ana": "Tue am", "rob": "Tue am"}, "candidate": {"note": "mornings, some exceptions"}},
 "question": "Is the Tuesday morning slot used?", "options": {"tuesday_am": "Hold it Tuesday morning", "find_new_slot": "Look for another slot"}}
{"id": "ho-sch16-u", "route": "answer", "reply": "Answer: tuesday_am", "ask_score": 0.67, "confidence": 0.988}

The output fields:

field meaning
route answer, ask or reason
reply what to show the user: the one-line answer, the follow-up question, or the worked reply ending in Answer:
ask_score the probe's score; above the threshold (0.8) means ask back
asked_by probe, expert (with --expert) or null
direct_reply, confidence the direct answer and its lowest token probability; null when the request was asked back
ms, probe_ms, answer_ms, followup_ms wall time in total and per stage
tokens, prompt_tokens tokens generated for reply, and in the prompt

Latency and streaming

The same router, packed int4 weights, on Hugging Face Jobs (35 written and 35 spoken requests, median per route, first request excluded as it compiles). The whole request, from arrival to the last token:

hardware answer now ask back reason probe pass decode step
AMD Strix Halo (gfx1151, ROCm) 2.5 s 3.0 s 23 s 0.84 s 97 ms
Nvidia L4, 24 GB 0.9 s 1.3 s 16 s 0.23 s 73 ms
Nvidia A100, 80 GB 0.42 s 0.64 s 7.0 s 0.10 s 30 ms

The router's own cost is the probe pass column, one extra run of layers 0–44 over the prompt. The rest is Gemma generating the reply (the one-line answer, the follow-up question, or the direct try followed by the worked reply).

Routes on the L4 match the Strix Halo on 34–35 of 35 requests (int4 kernels round differently per GPU). On a 24 GB card, run with GEMMA_Q4_CUDA_PREFILL=packed (the default prefill needs a 14 GB temporary) and --bf16-embedding; both are in hf-space/system-one/job.sh, which is how these were measured.

A reasoned reply takes 7–23 s to finish, but it does not have to finish before it is shown. With streaming (serve, or route --stop-early), the reply is sent piece by piece as Gemma writes it, so the wait that matters is the time to its first text:

hardware answer now ask back reason reason, whole reply
AMD Strix Halo (gfx1151, ROCm) 2.6 s 2.0 s 3.6 s 24 s
Nvidia L4, 24 GB 0.9 s 0.5 s 1.1 s 16 s
Nvidia A100, 80 GB 0.5 s 0.3 s 0.6 s 7.1 s

These are medians over the 35 written requests; the 35 spoken ones are within 0.1 s on the L4 and the A100. A reason starts sooner because the direct try stops at the first answer token whose probability is under 0.9, which already decides the route. An answer is not streamed: it is only known to be an answer once all of it is there, and it is one short line. Routes and final replies are the same as without streaming.

serve puts the router behind HTTP:

mix gemma.system_one serve --port 7860

curl -N -d '{"prompt": "What is 17 * 23?", "answer": "number"}' localhost:7860/route

The response is one JSON event per line: route once the route is chosen (with ask_score and confidence), text with each new piece of the reply, and done with the finished row as route writes it. A spoken request sends its WAV as base64 in audio_wav. GET / is a page to try it from a browser. Requests run one at a time on the GPU. Before it listens, the server compiles every graph a request can need (about 5 minutes on the Strix Halo), so no request waits on a compile.

Accuracy

The probe only works on the wording it was trained on (3,689 prompts in the router's direct form: twins, replay prompts, GSM8K, ARC and 600 spoken twins). On System-One-template prompts its AUROC drops from 0.83 to 0.64. The tables below were measured with the first, written-only probe (3,089 prompts); on written requests the current one asks the same 67/100 unclear twins with 8 needless asks instead of 6.

on 100 + 100 held-out twins decidable right needless asks unclear asked s/q
direct answer only 87 0 0 1.5
ask option in the prompt 67 9 54 1.5
router, cutoff 0.9 85 8 64 4.5
router, cutoff 0.99 88 8 64 7.9

On 200 unseen GSM8K + ARC-Challenge questions the probe never asks (highest score 0.29). The confidence cutoff sends 70 to reasoning and gets 189/200 right at 12.1 s per question, against 139/200 answering directly and 29.9 s per question reasoning on everything. The follow-up questions come back in the request's language ("Qual dos dois Joões você quer que eu ligue?"). The probe costs one extra prefix pass, 0.88 s per request.

Compared with the expert, the router asks needlessly a third as often (8% of decidable twins) and cannot change an ordinary reply, but guesses on more unclear requests (36 of 100). Its limit is the probe: a nonlinear probe (MLP, RBF SVM) does no better, and neither do more training domains at a fixed number of rows. Pooling over the prompt and reading an earlier layer (the output of layer 31 instead of 44) do not move it either: every variant lands at held-out AUROC 0.85–0.86, and recall gains of 5–7 points have 95% bootstrap intervals that include zero. The unclear requests it misses, mostly in unseen domains, are not linearly separated in the prompt's hidden states. Asking Gemma itself whether the request is settled is worse (AUROC 0.78): it says No on half the clear requests and on half of plain GSM8K and ARC questions.

Spoken requests

A request can be spoken: a row with "audio": "q.wav" (16 kHz mono, relative to the input file) keeps the state and options as text and puts the WAV in the audio slot after them, through Gemma's own audio encoder. Measured on the 200 held-out twins, with each question spoken in its language by edge-tts voices (1.9–3.7 s):

200 held-out twins AUROC needless asks unclear asked
written, written-only probe 0.885 6 67
spoken, written-only probe 0.854 3 48
spoken, current probe (with 600 spoken training twins) 0.883 4 60

Spoken requests route as fast as written ones: answer now has a median of 2.5 s, ask back 2.9 s and reason 23 s. The follow-up question is written from the audio ("¿A qué lista quieres añadir la leche?"). The written-only probe scores speech lower, so it asks less. Adding spoken twins to its training recovers 12 of those asks (95% CI +6 to +19) without more needless ones. The gain is in English; the 17 non-English unclear requests move from 7 to 8. The table used a 4 s audio bucket (--audio-seconds 4); the default 8 s bucket gives the same routes but a slower probe pass (1.3 s instead of 0.84 s).

The router and the expert together

The two fail differently. The probe rarely asks needlessly but misses unclear requests in unseen domains, while the expert asks more often either way. With --expert, a System One item that the probe is unsure about (a score above 0.4 and at most 0.8) is put to the expert first. If the expert replies with a question, that question is the ask-back. Otherwise the item is routed as before, on the pipeline without the expert, so the expert never changes an answer, only whether one is given.

mix gemma.system_one route --expert artifacts/system-one/round3 --input requests.jsonl --output routed.jsonl
on 100 + 100 held-out twins (served) needless asks unclear asked unseen domains (of 52) seen domains (of 48)
probe alone 6 67 29 38
expert alone, floor 0.8 32 79
probe + expert in 0.4–0.8 14 83 41 42

The expert runs on 43 of the 200 twins, with a median time of 1.8 s. It never runs on the 200 GSM8K and ARC questions, whose routes are unchanged. Most of the extra needless asks are for a number that is already in the state. This is a trade-off, not a win: use it when a guess costs more than an extra question.

Retraining the probe on your own requests

Since the probe's weak spot is domains it has not seen, the practical fix is training rows from the domains it will serve. scripts/system_one/retrain_ask_probe.sh OUT mine.jsonl [-- eval.jsonl]:

  1. Takes rows with an id, an item or a bare prompt, and a boolean decidable. Twin pairs are best.
  2. Renders them as route serves them and checks that rendering against the router.
  3. Caches only the new rows and exports OUT/ask-probe.
  4. Scores that probe against the shipped one on a 400-row held-out set.

Serve the result with route --ask-probe OUT/ask-probe.

The round-3 expert

The first approach, before the router: a routed "quick decision" expert added beside the FFN of the last three decoder layers. Given a state and a question it answers in one line, and when the state does not settle the answer it asks one follow-up question instead of guessing. It is 70.79M new parameters after three training rounds.

In layers 45, 46 and 47 of the 48-layer decoder,

ffn_out = base_ffn(x) + g(x) * expert(x)

where expert is a GELU FFN 3840 → 2048 → 3840 and g is a per-token linear sigmoid router on the same input. A gate under the serving floor is clamped to exactly 0, so with the router shut the output is bit-identical to base Gemma. The router is trained as the decision itself (round 3, --gate-mode classifier): it opens only when something is missing from the state, base Gemma answers everything else.

file contents
parameters.safetensors 15 tensors, 70.79M parameters, 283 MB (expert gate/intermediate/output kernels and router kernel/bias per layer)
manifest.etf layers [45, 46, 47], expert_size 2048, gate_floor 0.5, hidden_size 3840, parameter paths, training meta (Erlang external term format)

The manifest's gate_floor is the training floor; serve at 0.8 (see the results below).

Running the expert

With the same setup as the quick start:

hf download olafura/gemma4-12b-system-one --local-dir artifacts/system-one/round3

# items.jsonl: System One items as above, or a bare {"prompt"}; spoken items work the same way
mix gemma.system_one generate --input items.jsonl --output replies.jsonl \
  --expert artifacts/system-one/round3 --gate-floor 0.8 \
  --system-message 'Answer in one short line. Name exactly one of the options you are given. If the state you are given does not determine the answer, do not guess: ask one short question for the missing detail instead.'

Spoken questions, first check (12 items)

Six held-out twin pairs (data/system-one/spoken12/ in the repo), each question spoken with one of two edge-tts voices (1.5–3 s), 4 s audio bucket, the same system message as above.

condition decidable correct underspecified asked
base Gemma (router closed) 6/6 5/6
expert, floor 0.8 5/6 6/6

The same profile as with written questions: the expert stops the one commit base Gemma makes on an underspecified item, and asks once where it should not. Its follow-ups get more specific ("Which printer upstairs?", "Which of the two came first?"). The expert was never trained on audio; it acts on the residual stream after the audio has been read, which is why it carries over. The same audio in a 6 s bucket decodes token-identical, so the padding is masked correctly. Twelve English TTS items show that the path works, not how good it is.

Each output row carries the reply, the token ids, and gate_last, the router's gate at the decision position (a prefill probe; it cannot see gates opened later in the reply). --gate-floor 1.0 forces every router closed and gives base Gemma back; mix gemma.system_one regress --expert … --gate-floor 0.8 checks token identity on ordinary prompts.

Results (round 3, 600 held-out items, judged by a Laya cascade)

dec_acc is accuracy on decidable items, commit_u the share of underspecified items the model answered anyway (the failure this exists to fix), needless the share of decidable items it asked about instead of answering.

condition headline dec_acc commit_u needless
base Gemma + system message −0.106 0.853 0.637 0.02
expert, floor 0.5 +0.434 0.540 0.170 0.41
expert, floor 0.8 (recommended) +0.386 0.663 0.247 0.26

Paired bootstrap against base + system message: +0.49 [0.39, 0.60] at floor 0.8. The non-English items (120 of the 600: German, Spanish, French, Japanese, Icelandic, Polish, Portuguese, Dutch, Korean) improve in every language with more than eight items; Korean and Portuguese have too few to read. On a 400-prompt regression set of ordinary requests the floor-0.8 model is token-identical to base Gemma on 97–100% of rows (390 of 400); the rest reword mid-reply, none turns into a question.

Known limit: the router is linear. Its gate at the decision position has median 0.59 on decidable and 0.79 on underspecified unseen items, and it reacts to surface cues of ambiguity, so no floor gives both high decidable accuracy and low commit_u. The next step is a router with a hidden layer, not more data. Full write-up, all three rounds and the scorecard design: docs/system-one-expert-plan.md.

Training data

4,944 rows: twin pairs (a decidable item and its blurred twin with one detail removed) across scheduling, configuration, support, finance, clinic, education, logistics and more, 82% English and the rest across eleven other languages, plus 1,000 replay rows of base Gemma's own replies to ordinary prompts as closed-gate targets. Held-out: 600 items, never trained on, in near and far splits. All of it is in data/system-one/ of the repo.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for olafura/gemma4-12b-system-one