Instructions to use olafura/gemma4-12b-system-one with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use olafura/gemma4-12b-system-one with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
System One for Gemma 4 12B: answer now, reason first or ask back (work in progress)
A router in front of an unmodified Gemma 4 12B that decides, per request, whether to answer now in one line, reason first with a worked reply, or ask back one short question when the request does not settle the answer. In the spirit of Laya and Jev, but in free text and on top of an unchanged Gemma.
Everything runs with Elixir/Nx from
olafura/gemma-4-mic-transcribe
(mix gemma.system_one), measured on an AMD Strix Halo through EXLA/ROCm
and on Nvidia L4 and A100 GPUs through EXLA/CUDA. This is a WIP snapshot,
not a release. The base weights are not in this repo and are never
modified.
The repo holds two things:
| path | what | size |
|---|---|---|
ask-probe/ |
the router's ask-back probe (recommended) | 15 kB |
parameters.safetensors, manifest.etf |
the earlier round-3 expert, added inside the last three decoder layers | 283 MB |
Quick start
From a checkout of the repo with the packed 12B prefix/tail artifacts built (see its README, "Extracting a decoder block"), on a machine with the ROCm or CUDA EXLA client:
export XLA_FLAGS='--xla_gpu_autotune_level=0 --xla_gpu_enable_command_buffer= --xla_gpu_enable_triton_gemm=false'
hf download olafura/gemma4-12b-system-one --include 'ask-probe/*' --local-dir artifacts/system-one
# a file of requests: System One items or {"id", "prompt", "answer": "number"}, one per line
mix gemma.system_one route --input requests.jsonl --output routed.jsonl
# or a server that streams each reply as it is written
mix gemma.system_one serve --port 7860
curl -N -d '{"prompt": "What is 17 * 23?", "answer": "number"}' localhost:7860/route
A System One item is a state, a question and options:
{"id": "voi01", "state": {"nlu_log": "text=\"send it to the printer upstairs\" | printers=[HP-2200@first-floor, Brother-L@second-floor] | job=\"contract.pdf\""},
"question": "Pick the printer.", "options": {"hp_2200": "The HP laser printer", "brother_l": "The Brother laser printer"}}
The question can be spoken instead: a row with "audio": "q.wav" (16 kHz
mono, relative to the input file) keeps the state and options as text and
puts the WAV in the audio slot after them, through Gemma's own audio
encoder. --audio-seconds (default 8) is the audio bucket the WAV is cut
and padded to.
How the router decides
- The request is rendered in a direct form that asks for one line,
Answer: <x>, and the prefix (layers 0–44) is run over it. - Ask back when the ask-back probe scores above 0.8 (next section).
- Otherwise the direct answer is generated, and its confidence is the lowest probability among the answer tokens. Answer now at 0.9 or above.
- Otherwise reason: the request again, ending in "End your reply with a line of the form 'Answer: '", with room for up to 768 tokens.
How ask back works
Before Gemma answers, a small probe reads Gemma's internal state for the request, and if it scores above 0.8, Gemma writes one clarifying question instead of an answer.
- Read. The prefix (layers 0–44) runs over the direct prompt. The output at the last prompt token is a vector of 3840 numbers: what Gemma has made of the request just before it would start answering.
- Score. The probe is a logistic regression on that vector:
score = sigmoid(vector · weight + bias), between 0 and 1. It reads Gemma's state, not the text, so it costs one dot product on top of the prefix pass. - Decide. Above 0.8 the router asks back; no answer is generated.
- Ask. The same request is rendered again, ending in "The information above does not settle this. Reply with only the one short question you would ask to settle it, and nothing else." Gemma writes the question (up to 64 tokens), in the request's language.
Nothing in Gemma is changed: the probe only reads the residual stream, and the question is Gemma's own ordinary generation.
Where the probe comes from. It was trained on 3,689 prompts in the direct form, 1,494 of which should be asked back:
| source | rows |
|---|---|
| twin pairs: a decidable item and its twin with one detail removed | 2,389 |
| spoken twins, the question as a WAV | 600 |
| replay prompts, ordinary requests to base Gemma | 400 |
| GSM8K and ARC questions, which should never ask | 300 |
At 0.8, 0.9% of the training prompts that should not ask score above it, and 50% of the ones that should. The threshold trades missed asks for needless ones: a needless question costs the user a turn, a missed one lets Gemma guess.
Where it fails.
- It misses unclear requests in domains it has not seen. Those go on to answer or reason. Retraining on your own domains is the fix.
- It keys on the wording it was trained on (see Accuracy).
- Scores near 0.8 can flip between GPUs or builds, since int4 kernels round differently: one request scored 0.78 before a kernel change and 0.81 after it.
- With
--expert, requests just under 0.8 also get a second opinion (below).
Examples
Real rows from route on held-out requests (never trained on).
Each output row repeats the input's fields and adds the
route, the reply and what the decision was based on.
A twin pair: answer now, and ask back. The two requests differ in one detail: whether the printers are on different floors.
{"id": "ho-voi01-d", "state": {"nlu_log": "text=\"send it to the printer upstairs\" | printers=[HP-2200@first-floor, Brother-L@second-floor] | job=\"contract.pdf\""},
"question": "Pick the printer.", "options": {"hp_2200": "The HP laser printer", "brother_l": "The Brother laser printer"}}
{"id": "ho-voi01-u", "state": {"nlu_log": "text=\"send it to the printer upstairs\" | printers=[HP-2200@second-floor, Brother-L@second-floor] | job=\"contract.pdf\""},
"question": "Pick the printer.", "options": {"hp_2200": "The HP laser printer", "brother_l": "The Brother laser printer"}}
{"id": "ho-voi01-d", "route": "answer", "reply": "Answer: brother_l", "ask_score": 0.21, "confidence": 0.9997, "asked_by": null}
{"id": "ho-voi01-u", "route": "ask", "reply": "Which printer should I use, the Brother or the HP?", "ask_score": 0.999, "confidence": null, "asked_by": "probe"}
The first scores 0.21 on the probe, so the direct answer is generated; its least likely token has probability 0.9997, above 0.9, so it is kept. The second scores 0.999, so no answer is generated at all: Gemma is asked for the question that would settle it.
Reason when the direct answer is unsure. A plain question works too,
with "answer": "number" asking for a number instead of an option:
{"id": "g2-011", "prompt": "A bag of buttons had 21 buttons in it. Seven buttons had two holes and the rest had four holes. How many holes were in all the buttons in the bag?", "answer": "number"}
{"id": "g2-011", "route": "reason", "ask_score": 0.00009, "direct_reply": "Answer: 63", "confidence": 0.68,
"reply": "To find the total number of holes in all the buttons, we need to calculate ...\n\n4. **Calculate the total number of holes:**\n Add the holes from both types of buttons together:\n $14 \\text{ holes} + 56 \\text{ holes} = 70 \\text{ holes}$.\n\nAnswer: 70"}
The direct answer was wrong (63) and unsure (0.68), so the request was run
again with room to work; the worked reply gets 70. An easy one stays on the
fast path: "Tracy used a piece of wire 4 feet long … cut into pieces 6
inches long. How many pieces?" goes to answer with Answer: 8 at 0.9999.
Spoken, in Spanish. The state and options are text, the question is a WAV (the spoken form of "¿A qué lista se añade la leche?"):
{"id": "ho-voi09-u", "state": {"orden": "añade leche a la lista", "listas": {"compra_semanal": {"activa": true}, "fiesta": {"activa": true}}},
"options": {"compra_semanal": "La lista de la compra semanal", "fiesta": "La lista para la fiesta"}, "audio": "ho-voi09.wav"}
{"id": "ho-voi09-u", "route": "ask", "reply": "¿A qué lista quieres añadir la leche?", "ask_score": 0.97, "ms": 2850}
A miss. The probe's errors look like this: both interviewers are free on Tuesday morning, but the candidate's note ("mornings, some exceptions") leaves it open. The probe scores 0.67, under 0.8, and Gemma answers confidently:
{"id": "ho-sch16-u", "state": {"interview": {"stage": "final", "panel": ["ana", "rob"]}, "availability": {"ana": "Tue am", "rob": "Tue am"}, "candidate": {"note": "mornings, some exceptions"}},
"question": "Is the Tuesday morning slot used?", "options": {"tuesday_am": "Hold it Tuesday morning", "find_new_slot": "Look for another slot"}}
{"id": "ho-sch16-u", "route": "answer", "reply": "Answer: tuesday_am", "ask_score": 0.67, "confidence": 0.988}
The output fields:
| field | meaning |
|---|---|
route |
answer, ask or reason |
reply |
what to show the user: the one-line answer, the follow-up question, or the worked reply ending in Answer: |
ask_score |
the probe's score; above the threshold (0.8) means ask back |
asked_by |
probe, expert (with --expert) or null |
direct_reply, confidence |
the direct answer and its lowest token probability; null when the request was asked back |
ms, probe_ms, answer_ms, followup_ms |
wall time in total and per stage |
tokens, prompt_tokens |
tokens generated for reply, and in the prompt |
Latency and streaming
The same router, packed int4 weights, on Hugging Face Jobs (35 written and 35 spoken requests, median per route, first request excluded as it compiles). The whole request, from arrival to the last token:
| hardware | answer now | ask back | reason | probe pass | decode step |
|---|---|---|---|---|---|
| AMD Strix Halo (gfx1151, ROCm) | 2.5 s | 3.0 s | 23 s | 0.84 s | 97 ms |
| Nvidia L4, 24 GB | 0.9 s | 1.3 s | 16 s | 0.23 s | 73 ms |
| Nvidia A100, 80 GB | 0.42 s | 0.64 s | 7.0 s | 0.10 s | 30 ms |
The router's own cost is the probe pass column, one extra run of layers 0–44 over the prompt. The rest is Gemma generating the reply (the one-line answer, the follow-up question, or the direct try followed by the worked reply).
Routes on the L4 match the Strix Halo on 34–35 of 35 requests (int4
kernels round differently per GPU). On a 24 GB card, run with
GEMMA_Q4_CUDA_PREFILL=packed (the default prefill needs a 14 GB
temporary) and --bf16-embedding; both are in
hf-space/system-one/job.sh,
which is how these were measured.
A reasoned reply takes 7–23 s to finish, but it does not have to finish
before it is shown. With streaming (serve, or route --stop-early),
the reply is sent piece by piece as Gemma writes it, so the wait that
matters is the time to its first text:
| hardware | answer now | ask back | reason | reason, whole reply |
|---|---|---|---|---|
| AMD Strix Halo (gfx1151, ROCm) | 2.6 s | 2.0 s | 3.6 s | 24 s |
| Nvidia L4, 24 GB | 0.9 s | 0.5 s | 1.1 s | 16 s |
| Nvidia A100, 80 GB | 0.5 s | 0.3 s | 0.6 s | 7.1 s |
These are medians over the 35 written requests; the 35 spoken ones are within 0.1 s on the L4 and the A100. A reason starts sooner because the direct try stops at the first answer token whose probability is under 0.9, which already decides the route. An answer is not streamed: it is only known to be an answer once all of it is there, and it is one short line. Routes and final replies are the same as without streaming.
serve puts the router behind HTTP:
mix gemma.system_one serve --port 7860
curl -N -d '{"prompt": "What is 17 * 23?", "answer": "number"}' localhost:7860/route
The response is one JSON event per line: route once the route is chosen
(with ask_score and confidence), text with each new piece of the
reply, and done with the finished row as route writes it. A spoken
request sends its WAV as base64 in audio_wav. GET / is a page to try
it from a browser. Requests run one at a time on the GPU. Before it
listens, the server compiles every graph a request can need (about 5
minutes on the Strix Halo), so no request waits on a compile.
Accuracy
The probe only works on the wording it was trained on (3,689 prompts in the router's direct form: twins, replay prompts, GSM8K, ARC and 600 spoken twins). On System-One-template prompts its AUROC drops from 0.83 to 0.64. The tables below were measured with the first, written-only probe (3,089 prompts); on written requests the current one asks the same 67/100 unclear twins with 8 needless asks instead of 6.
| on 100 + 100 held-out twins | decidable right | needless asks | unclear asked | s/q |
|---|---|---|---|---|
| direct answer only | 87 | 0 | 0 | 1.5 |
ask option in the prompt |
67 | 9 | 54 | 1.5 |
| router, cutoff 0.9 | 85 | 8 | 64 | 4.5 |
| router, cutoff 0.99 | 88 | 8 | 64 | 7.9 |
On 200 unseen GSM8K + ARC-Challenge questions the probe never asks (highest score 0.29). The confidence cutoff sends 70 to reasoning and gets 189/200 right at 12.1 s per question, against 139/200 answering directly and 29.9 s per question reasoning on everything. The follow-up questions come back in the request's language ("Qual dos dois Joões você quer que eu ligue?"). The probe costs one extra prefix pass, 0.88 s per request.
Compared with the expert, the router asks needlessly a third as often (8% of decidable twins) and cannot change an ordinary reply, but guesses on more unclear requests (36 of 100). Its limit is the probe: a nonlinear probe (MLP, RBF SVM) does no better, and neither do more training domains at a fixed number of rows. Pooling over the prompt and reading an earlier layer (the output of layer 31 instead of 44) do not move it either: every variant lands at held-out AUROC 0.85–0.86, and recall gains of 5–7 points have 95% bootstrap intervals that include zero. The unclear requests it misses, mostly in unseen domains, are not linearly separated in the prompt's hidden states. Asking Gemma itself whether the request is settled is worse (AUROC 0.78): it says No on half the clear requests and on half of plain GSM8K and ARC questions.
Spoken requests
A request can be spoken: a row with "audio": "q.wav" (16 kHz mono,
relative to the input file) keeps the state and options as text and puts
the WAV in the audio slot after them, through Gemma's own audio encoder.
Measured on the 200 held-out twins, with each question spoken in its
language by edge-tts voices (1.9–3.7 s):
| 200 held-out twins | AUROC | needless asks | unclear asked |
|---|---|---|---|
| written, written-only probe | 0.885 | 6 | 67 |
| spoken, written-only probe | 0.854 | 3 | 48 |
| spoken, current probe (with 600 spoken training twins) | 0.883 | 4 | 60 |
Spoken requests route as fast as written ones: answer now has a median of 2.5 s,
ask back 2.9 s and reason 23 s. The follow-up question is written from
the audio ("¿A qué lista quieres añadir la leche?"). The written-only
probe scores speech lower, so it asks less. Adding spoken twins to its
training recovers 12 of those asks (95% CI +6 to +19) without more
needless ones. The gain is in English; the 17 non-English unclear
requests move from 7 to 8. The table used a 4 s audio bucket (--audio-seconds 4); the
default 8 s bucket gives the same routes but a slower probe pass (1.3 s
instead of 0.84 s).
The router and the expert together
The two fail differently. The probe rarely asks needlessly but misses
unclear requests in unseen domains, while the expert asks more often
either way. With --expert, a System One item that the probe is unsure
about (a score above 0.4 and at most 0.8) is put to the expert first. If
the expert replies with a question, that question is the ask-back.
Otherwise the item is routed as before, on the pipeline without the
expert, so the expert never changes an answer, only whether one is
given.
mix gemma.system_one route --expert artifacts/system-one/round3 --input requests.jsonl --output routed.jsonl
| on 100 + 100 held-out twins (served) | needless asks | unclear asked | unseen domains (of 52) | seen domains (of 48) |
|---|---|---|---|---|
| probe alone | 6 | 67 | 29 | 38 |
| expert alone, floor 0.8 | 32 | 79 | ||
| probe + expert in 0.4–0.8 | 14 | 83 | 41 | 42 |
The expert runs on 43 of the 200 twins, with a median time of 1.8 s. It never runs on the 200 GSM8K and ARC questions, whose routes are unchanged. Most of the extra needless asks are for a number that is already in the state. This is a trade-off, not a win: use it when a guess costs more than an extra question.
Retraining the probe on your own requests
Since the probe's weak spot is domains it has not seen, the practical fix
is training rows from the domains it will serve.
scripts/system_one/retrain_ask_probe.sh OUT mine.jsonl [-- eval.jsonl]:
- Takes rows with an
id, an item or a bareprompt, and a booleandecidable. Twin pairs are best. - Renders them as
routeserves them and checks that rendering against the router. - Caches only the new rows and exports
OUT/ask-probe. - Scores that probe against the shipped one on a 400-row held-out set.
Serve the result with route --ask-probe OUT/ask-probe.
The round-3 expert
The first approach, before the router: a routed "quick decision" expert added beside the FFN of the last three decoder layers. Given a state and a question it answers in one line, and when the state does not settle the answer it asks one follow-up question instead of guessing. It is 70.79M new parameters after three training rounds.
In layers 45, 46 and 47 of the 48-layer decoder,
ffn_out = base_ffn(x) + g(x) * expert(x)
where expert is a GELU FFN 3840 → 2048 → 3840 and g is a per-token
linear sigmoid router on the same input. A gate under the serving floor is
clamped to exactly 0, so with the router shut the output is bit-identical
to base Gemma. The router is trained as the decision itself (round 3,
--gate-mode classifier): it opens only when something is missing from
the state, base Gemma answers everything else.
| file | contents |
|---|---|
parameters.safetensors |
15 tensors, 70.79M parameters, 283 MB (expert gate/intermediate/output kernels and router kernel/bias per layer) |
manifest.etf |
layers [45, 46, 47], expert_size 2048, gate_floor 0.5, hidden_size 3840, parameter paths, training meta (Erlang external term format) |
The manifest's gate_floor is the training floor; serve at 0.8 (see
the results below).
Running the expert
With the same setup as the quick start:
hf download olafura/gemma4-12b-system-one --local-dir artifacts/system-one/round3
# items.jsonl: System One items as above, or a bare {"prompt"}; spoken items work the same way
mix gemma.system_one generate --input items.jsonl --output replies.jsonl \
--expert artifacts/system-one/round3 --gate-floor 0.8 \
--system-message 'Answer in one short line. Name exactly one of the options you are given. If the state you are given does not determine the answer, do not guess: ask one short question for the missing detail instead.'
Spoken questions, first check (12 items)
Six held-out twin pairs (data/system-one/spoken12/ in the repo), each
question spoken with one of two edge-tts voices (1.5–3 s), 4 s audio
bucket, the same system message as above.
| condition | decidable correct | underspecified asked |
|---|---|---|
| base Gemma (router closed) | 6/6 | 5/6 |
| expert, floor 0.8 | 5/6 | 6/6 |
The same profile as with written questions: the expert stops the one commit base Gemma makes on an underspecified item, and asks once where it should not. Its follow-ups get more specific ("Which printer upstairs?", "Which of the two came first?"). The expert was never trained on audio; it acts on the residual stream after the audio has been read, which is why it carries over. The same audio in a 6 s bucket decodes token-identical, so the padding is masked correctly. Twelve English TTS items show that the path works, not how good it is.
Each output row carries the reply, the token ids, and gate_last, the
router's gate at the decision position (a prefill probe; it cannot see
gates opened later in the reply). --gate-floor 1.0 forces every router
closed and gives base Gemma back; mix gemma.system_one regress --expert … --gate-floor 0.8 checks token identity on ordinary prompts.
Results (round 3, 600 held-out items, judged by a Laya cascade)
dec_acc is accuracy on decidable items, commit_u the share of
underspecified items the model answered anyway (the failure this exists
to fix), needless the share of decidable items it asked about instead
of answering.
| condition | headline | dec_acc | commit_u | needless |
|---|---|---|---|---|
| base Gemma + system message | −0.106 | 0.853 | 0.637 | 0.02 |
| expert, floor 0.5 | +0.434 | 0.540 | 0.170 | 0.41 |
| expert, floor 0.8 (recommended) | +0.386 | 0.663 | 0.247 | 0.26 |
Paired bootstrap against base + system message: +0.49 [0.39, 0.60] at floor 0.8. The non-English items (120 of the 600: German, Spanish, French, Japanese, Icelandic, Polish, Portuguese, Dutch, Korean) improve in every language with more than eight items; Korean and Portuguese have too few to read. On a 400-prompt regression set of ordinary requests the floor-0.8 model is token-identical to base Gemma on 97–100% of rows (390 of 400); the rest reword mid-reply, none turns into a question.
Known limit: the router is linear. Its gate at the decision position has
median 0.59 on decidable and 0.79 on underspecified unseen items, and it
reacts to surface cues of ambiguity, so no floor gives both high decidable
accuracy and low commit_u. The next step is a router with a hidden layer,
not more data. Full write-up, all three rounds and the scorecard design:
docs/system-one-expert-plan.md.
Training data
4,944 rows: twin pairs (a decidable item and its blurred twin with one
detail removed) across scheduling, configuration, support, finance,
clinic, education, logistics and more, 82% English and the rest across
eleven other languages, plus 1,000
replay rows of base Gemma's own replies to ordinary prompts as closed-gate
targets. Held-out: 600 items, never trained on, in near and far splits.
All of it is in data/system-one/ of the repo.
- Downloads last month
- -
Model tree for olafura/gemma4-12b-system-one
Base model
google/gemma-4-12B