Instructions to use EryriLabs/DiffusionGemma-26B-A4B-Hev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use EryriLabs/DiffusionGemma-26B-A4B-Hev with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("EryriLabs/DiffusionGemma-26B-A4B-Hev", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Hev: DiffusionGemma with a Hydra decision head
By Dwain Barnes. First published 2026-09-24.
What it does
Give Hev a lump of text (a support ticket, an agent log, a security alert, whatever your program is looking at) and a list of typed questions, and it gives you back probabilities rather than prose. "Which team should handle this?" comes back as billing 0.88, technical 0.12, sales 0.00. "Is this urgent?" comes back as 0.67. That is the job TypeSafe's Jev does; Hev does it with open weights you can run yourself.
There are two parts:
- DiffusionGemma, Google's open text-diffusion model, frozen. I never fine-tuned it. One pass of its encoder reads your text and one pass of its bidirectional denoiser reads a small JSON answer canvas.
- A Hydra head I trained on top. Hydra is a bidirectional state-space model, the two-way cousin of Mamba. It reads every token of your text and every answer slot together, in linear time, and turns that into calibrated probabilities. It is an ensemble of 3 heads (seeds 42, 43, 44) whose probabilities are averaged, 44,476,866 parameters in total, and it trains in minutes.
What is new here
As far as I can tell (checked 2026-09-24), nobody had done any of these three things before:
- put Hydra on top of a diffusion language model's representations;
- used a state-space model anywhere inside a System-One or Jev-style decision model. Every open clone I found is an autoregressive Qwen or Gemma with LoRA, or a BERT-style encoder with a linear head;
- published an open-weights specialist that leads the specialist rows of the typed-decisions leaderboard (fitted on the train split, as they are). How it stands next to the zero-shot generalists, Jev included, is spelled out under Results, with the dataset's own warning that the two kinds of number are not comparable.
The idea, the name and the direction are mine. If you build on it, please cite this repo; the commit history is the timestamp.
Results
Scored on the test split of LocalLLaMA/typed-decisions (400 cases, 2,000 decisions) in the specialist setting (the head is fitted on the train split), using the metrics defined on that dataset's card (see the evaluation notes under Limits). The "zero-shot canvas read" row is DiffusionGemma with no head, reading its own option probabilities off the canvas, which is what the vLLM structured-diffusion mode does.
| Model | n | acc | soft_acc | macro_f1 | kl | tv | brier | nll | ece | score_mae | within_1 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DiffusionGemma zero-shot canvas read | 2000 | 0.708 | 0.573 | 0.551 | 1.333 | 0.362 | 0.304 | 2.100 | 0.213 | 0.411 | 0.934 |
| Hev | 2000 | 0.731 | 0.584 | 0.626 | 0.185 | 0.213 | 0.094 | 0.952 | 0.162 | 0.303 | 0.976 |
Hev is a specialist: its head was fitted on the train split of the same four workflows it is scored on, and it cannot answer a question schema it was not fitted for. On the dataset's leaderboard (read 2026-09-25) the comparable rows are the specialists fitted per workflow: ModernBERT-base at accuracy 0.646 and Brier 0.119, and MiniLM-L6 at 0.587 and 0.143. Hev is 8.5 points above the best of them on accuracy and its Brier is 0.025 lower. The generalist rows, scored zero-shot on workflows they never saw, are TypeSafe Jev 1.13.0 at 0.727 and 0.148 and meraGPT Decider 1 at 0.768 and 0.052. Hev's accuracy is 0.4 points above Jev's (less than one standard error, 1.0 points, on 2,000 decisions) and its Brier is 0.054 lower, but the dataset card is explicit that a specialist number next to a generalist one is not a ranking, and Decider 1 leads both on accuracy and on Brier. Hev's head lifts the frozen backbone's zero-shot accuracy from 0.708 to 0.731 and reaches Brier 0.094.
How to run it
You need a GPU with 80 GB of memory for the bf16 backbone. The head itself is tiny and runs anywhere; the backbone is the heavy part. Loading with load_in_4bit=True is untested: DiffusionGemma's fused expert parameters are not nn.Linear, so bitsandbytes leaves most weights in bf16 and the memory saving is small.
Install:
pip install "transformers>=5.17" accelerate datasets safetensors
from transformers import AutoModel
model = AutoModel.from_pretrained("EryriLabs/DiffusionGemma-26B-A4B-Hev", trust_remote_code=True)
answer = model.decide(
state="Help! My payouts have been failing for 3 days.",
questions={
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payments, invoicing, refunds",
"technical": "Bugs, outages, integrations",
"sales": "Pricing, upgrades, new accounts"}},
"urgent": {"type": "noul", "instructions": "This needs attention today."},
"severity": {"type": "score", "instructions": "How bad is the customer impact?",
"criteria": ["none", "minor", "major", "critical"]},
},
)
print(answer["answers"])
The request and response shapes match Jev's /v1/systemone API, so existing client code needs no changes. To serve that endpoint locally, pip install fastapi uvicorn and run:
import uvicorn
from transformers import AutoModel
from transformers.dynamic_module_utils import get_class_from_dynamic_module
model = AutoModel.from_pretrained("EryriLabs/DiffusionGemma-26B-A4B-Hev", trust_remote_code=True)
build_app = get_class_from_dynamic_module("client.build_app", "EryriLabs/DiffusionGemma-26B-A4B-Hev")
uvicorn.run(build_app(model), host="127.0.0.1", port=8080)
The .py files in this repo are a flattened copy of the hev package and use relative imports, so they cannot be run as scripts.
How it was trained
- Features: one forward pass per example through the frozen backbone, canvas length 64, slot seed mode
pad, hidden states from the final encoder and decoder norms, plus the unit-RMS outputs of layers 14 and 22 of the encoder for the prompt rows and of the decoder for the answer rows, so each row carries 8448 features. - Head: linear projection 8448 to 512, 3 Hydra blocks, pointer readout (each answer slot scored against each option's description span), plus a learned residual on the backbone's own option log-probabilities, plus fitted temperatures per question type for each member (seed 42: [1.339, 1.776, 1.816]; seed 43: [1.401, 1.899, 1.832]; seed 44: [1.39, 1.616, 1.888]).
- Loss: KL to the gold distribution plus 0.1 times Brier. AdamW, cosine schedule, early stopping on validation log loss.
- Data: typed-decisions train split plus samples from 19 jev-bench sources. Dataset licences are as declared by each source on the Hub.
- Hydra is implemented in plain PyTorch in
hydra_torch.py. Check against the official CUDA kernels: skipped (session 4 reuses session 2's kernel check).
Limits
- Accuracy on typed-decisions measures agreement with a synthetic teacher, not ground truth. The teacher's own self-agreement is 0.735.
- Version one supports 2 to 64 options per choice question, 2 to 10 score levels, and prompts up to 1,024 tokens. English only. Text only.
- Request size: every question takes a few canvas tokens (its id plus the JSON around the answer slot), so the 64-token answer canvas holds about 8 questions with short ids. A longer request fails with a clear error (HTTP 422 from the server).
- Evaluation notes:
ecein the table is top-label expected calibration error with 15 equal-width bins, and is not comparable with the ECE column on the dataset card.acccounts a hit when the top option matches the card's discretelabelfor each typed-decisions question, so ties in the gold distribution are broken the way the dataset breaks them;soft_accis the gold probability of the top option. - No GGUF yet. llama.cpp's DiffusionGemma support is still a draft branch and cannot expose the hidden states the head needs, and there is no Hydra kernel in ggml. If that changes I will port it. Until then this ships as safetensors.
- On fastino/fast-decisions, 17 domains the head never saw, it scores 0.575 exact match on the 2,600 single-label tasks against 0.611 for the backbone's own zero-shot read and 0.670 for GLiNER2.5-Decide run the same way, so as a generalist the head is a step back from the backbone; it is a specialist for the schema it was trained on.
- The backbone is Google's, unchanged, Apache-2.0. Everything else here is Apache-2.0 too.
Citation
@misc{barnes2026hev,
title = {Hev: A Hydra decision head on frozen DiffusionGemma for calibrated System-One decisions},
author = {Barnes, Dwain},
year = {2026},
url = {https://huggingface.co/EryriLabs/DiffusionGemma-26B-A4B-Hev}
}
Thanks to Google DeepMind for DiffusionGemma, to Hwang, Lahoti, Dao and Gu for Hydra, to Matt Mastracci for showing that a diffusion model's canvas is a decision engine, and to the LocalLLaMA community for the typed-decisions benchmark.
- Downloads last month
- 12
Model tree for EryriLabs/DiffusionGemma-26B-A4B-Hev
Base model
google/diffusiongemma-26B-A4B-it