Hev: DiffusionGemma with a Hydra decision head

By Dwain Barnes. First published 2026-09-24.

What it does

Give Hev a lump of text (a support ticket, an agent log, a security alert, whatever your program is looking at) and a list of typed questions, and it gives you back probabilities rather than prose. "Which team should handle this?" comes back as billing 0.88, technical 0.12, sales 0.00. "Is this urgent?" comes back as 0.67. That is the job TypeSafe's Jev does; Hev does it with open weights you can run yourself.

There are two parts:

  1. DiffusionGemma, Google's open text-diffusion model, frozen. I never fine-tuned it. One pass of its encoder reads your text and one pass of its bidirectional denoiser reads a small JSON answer canvas.
  2. A Hydra head I trained on top. Hydra is a bidirectional state-space model, the two-way cousin of Mamba. It reads every token of your text and every answer slot together, in linear time, and turns that into calibrated probabilities. It is an ensemble of 3 heads (seeds 42, 43, 44) whose probabilities are averaged, 44,476,866 parameters in total, and it trains in minutes.

What is new here

As far as I can tell (checked 2026-09-24), nobody had done any of these three things before:

  • put Hydra on top of a diffusion language model's representations;
  • used a state-space model anywhere inside a System-One or Jev-style decision model. Every open clone I found is an autoregressive Qwen or Gemma with LoRA, or a BERT-style encoder with a linear head;
  • published an open-weights specialist that leads the specialist rows of the typed-decisions leaderboard (fitted on the train split, as they are). How it stands next to the zero-shot generalists, Jev included, is spelled out under Results, with the dataset's own warning that the two kinds of number are not comparable.

The idea, the name and the direction are mine. If you build on it, please cite this repo; the commit history is the timestamp.

Results

Scored on the test split of LocalLLaMA/typed-decisions (400 cases, 2,000 decisions) in the specialist setting (the head is fitted on the train split), using the metrics defined on that dataset's card (see the evaluation notes under Limits). The "zero-shot canvas read" row is DiffusionGemma with no head, reading its own option probabilities off the canvas, which is what the vLLM structured-diffusion mode does.

Model n acc soft_acc macro_f1 kl tv brier nll ece score_mae within_1
DiffusionGemma zero-shot canvas read 2000 0.708 0.573 0.551 1.333 0.362 0.304 2.100 0.213 0.411 0.934
Hev 2000 0.731 0.584 0.626 0.185 0.213 0.094 0.952 0.162 0.303 0.976

Hev is a specialist: its head was fitted on the train split of the same four workflows it is scored on, and it cannot answer a question schema it was not fitted for. On the dataset's leaderboard (read 2026-09-25) the comparable rows are the specialists fitted per workflow: ModernBERT-base at accuracy 0.646 and Brier 0.119, and MiniLM-L6 at 0.587 and 0.143. Hev is 8.5 points above the best of them on accuracy and its Brier is 0.025 lower. The generalist rows, scored zero-shot on workflows they never saw, are TypeSafe Jev 1.13.0 at 0.727 and 0.148 and meraGPT Decider 1 at 0.768 and 0.052. Hev's accuracy is 0.4 points above Jev's (less than one standard error, 1.0 points, on 2,000 decisions) and its Brier is 0.054 lower, but the dataset card is explicit that a specialist number next to a generalist one is not a ranking, and Decider 1 leads both on accuracy and on Brier. Hev's head lifts the frozen backbone's zero-shot accuracy from 0.708 to 0.731 and reaches Brier 0.094.

How to run it

You need a GPU with 80 GB of memory for the bf16 backbone. The head itself is tiny and runs anywhere; the backbone is the heavy part. Loading with load_in_4bit=True is untested: DiffusionGemma's fused expert parameters are not nn.Linear, so bitsandbytes leaves most weights in bf16 and the memory saving is small.

Install:

pip install "transformers>=5.17" accelerate datasets safetensors
from transformers import AutoModel

model = AutoModel.from_pretrained("EryriLabs/DiffusionGemma-26B-A4B-Hev", trust_remote_code=True)
answer = model.decide(
    state="Help! My payouts have been failing for 3 days.",
    questions={
        "department": {"type": "choice", "instructions": "Which team should handle this?",
                        "criteria": {"billing": "Payments, invoicing, refunds",
                                     "technical": "Bugs, outages, integrations",
                                     "sales": "Pricing, upgrades, new accounts"}},
        "urgent": {"type": "noul", "instructions": "This needs attention today."},
        "severity": {"type": "score", "instructions": "How bad is the customer impact?",
                      "criteria": ["none", "minor", "major", "critical"]},
    },
)
print(answer["answers"])

The request and response shapes match Jev's /v1/systemone API, so existing client code needs no changes. To serve that endpoint locally, pip install fastapi uvicorn and run:

import uvicorn
from transformers import AutoModel
from transformers.dynamic_module_utils import get_class_from_dynamic_module

model = AutoModel.from_pretrained("EryriLabs/DiffusionGemma-26B-A4B-Hev", trust_remote_code=True)
build_app = get_class_from_dynamic_module("client.build_app", "EryriLabs/DiffusionGemma-26B-A4B-Hev")
uvicorn.run(build_app(model), host="127.0.0.1", port=8080)

The .py files in this repo are a flattened copy of the hev package and use relative imports, so they cannot be run as scripts.

How it was trained

  • Features: one forward pass per example through the frozen backbone, canvas length 64, slot seed mode pad, hidden states from the final encoder and decoder norms, plus the unit-RMS outputs of layers 14 and 22 of the encoder for the prompt rows and of the decoder for the answer rows, so each row carries 8448 features.
  • Head: linear projection 8448 to 512, 3 Hydra blocks, pointer readout (each answer slot scored against each option's description span), plus a learned residual on the backbone's own option log-probabilities, plus fitted temperatures per question type for each member (seed 42: [1.339, 1.776, 1.816]; seed 43: [1.401, 1.899, 1.832]; seed 44: [1.39, 1.616, 1.888]).
  • Loss: KL to the gold distribution plus 0.1 times Brier. AdamW, cosine schedule, early stopping on validation log loss.
  • Data: typed-decisions train split plus samples from 19 jev-bench sources. Dataset licences are as declared by each source on the Hub.
  • Hydra is implemented in plain PyTorch in hydra_torch.py. Check against the official CUDA kernels: skipped (session 4 reuses session 2's kernel check).

Limits

  • Accuracy on typed-decisions measures agreement with a synthetic teacher, not ground truth. The teacher's own self-agreement is 0.735.
  • Version one supports 2 to 64 options per choice question, 2 to 10 score levels, and prompts up to 1,024 tokens. English only. Text only.
  • Request size: every question takes a few canvas tokens (its id plus the JSON around the answer slot), so the 64-token answer canvas holds about 8 questions with short ids. A longer request fails with a clear error (HTTP 422 from the server).
  • Evaluation notes: ece in the table is top-label expected calibration error with 15 equal-width bins, and is not comparable with the ECE column on the dataset card. acc counts a hit when the top option matches the card's discrete label for each typed-decisions question, so ties in the gold distribution are broken the way the dataset breaks them; soft_acc is the gold probability of the top option.
  • No GGUF yet. llama.cpp's DiffusionGemma support is still a draft branch and cannot expose the hidden states the head needs, and there is no Hydra kernel in ggml. If that changes I will port it. Until then this ships as safetensors.
  • On fastino/fast-decisions, 17 domains the head never saw, it scores 0.575 exact match on the 2,600 single-label tasks against 0.611 for the backbone's own zero-shot read and 0.670 for GLiNER2.5-Decide run the same way, so as a generalist the head is a step back from the backbone; it is a specialist for the schema it was trained on.
  • The backbone is Google's, unchanged, Apache-2.0. Everything else here is Apache-2.0 too.

Citation

@misc{barnes2026hev,
  title  = {Hev: A Hydra decision head on frozen DiffusionGemma for calibrated System-One decisions},
  author = {Barnes, Dwain},
  year   = {2026},
  url    = {https://huggingface.co/EryriLabs/DiffusionGemma-26B-A4B-Hev}
}

Thanks to Google DeepMind for DiffusionGemma, to Hwang, Lahoti, Dao and Gu for Hydra, to Matt Mastracci for showing that a diffusion model's canvas is a decision engine, and to the LocalLLaMA community for the typed-decisions benchmark.

Downloads last month
12
Safetensors
Model size
26B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EryriLabs/DiffusionGemma-26B-A4B-Hev

Finetuned
(29)
this model

Datasets used to train EryriLabs/DiffusionGemma-26B-A4B-Hev