Multiple Choice
PEFT
Safetensors
English
lora
decision-model
calibration
gemma4
How to use from the
Use from the
PEFT library
from peft import PeftModel
from transformers import AutoModelForCausalLM

base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-31B-it")
model = PeftModel.from_pretrained(base_model, "moganai/mogan-decision-31B-it")

Mogan-Decision

Blog Dataset Decision Index Buy Me a Coffee

mogan-decision-31B-it

mogan-decision-31B-it is a decision model. You give it a state (text or JSON), a question and a set of options, and it returns a probability for every option in a single forward pass, with no generated reasoning. It supports two question types:

  • choice: pick one of 2–255 options,
  • noul: the probability that a statement is true.

It is a LoRA (r = 128) on google/gemma-4-31B-it @ 842da37, trained with soft labels on 473k decision questions. At inference the adapter is merged into the bf16 weights.

Results: Decision Index

Scores are from our full run of the Decision Index suite (151,476 requests, all answered, none unsupported), scored with the official kit. They are self-reported until the board maintainers reproduce them.

Edition Decision Index Raw index
0.2.1 (kit 87d4650) 65.42 73.29
0.2 (kit 19ad28e) 60.17 69.93

Area skill on 0.2.1, with the board's reference system Jev 1.13 (0.2.1 index 57.91):

Area mogan-decision-31B-it Jev 1.13
Knowledge & Reasoning 51.2 51.4
Language Understanding 72.0 62.0
Retrieval & Classification 69.6 55.4
Tools & Automation 79.7 75.1
Arts & Human Taste 50.7 37.7

Results files: moganai/mogan-decision-31B-it-decision-index-results. Median in-process latency is 107 ms per request (one process per GH200, one request at a time). A single-GPU rerun of 3,000 random suite requests with the published adapter and code reproduces the submitted choices on 100% of 6,447 questions.

This model is not zero-shot on many of these benchmarks. It was trained on the public training splits of several benchmark families; the full list is in Training data. The table shows our 0.2.1 skill per benchmark next to the board's reference system, Jev 1.13 (0.2.1 index 57.91). βœ“ marks families with training data (a train split, or generated items of the same task type).

Benchmark mogan-decision-31B-it Jev 1.13 Training data
ACOS 45.7 27.3 βœ“
Amazon ESCI 50.5 43.8 βœ“
ANLI 69.4 62.2 βœ“
API-Bank 84.4 88.0
BANKING77 91.9 79.5 βœ“
BBH 84.3 89.7 βœ“
BFCL 95.8 94.3
BPoMP 87.0 81.8
BRIGHT 40.1 40.6
cfcolor 29.0 28.8
ChessBench 23.6 9.8 βœ“
CLadder 75.2 45.3 βœ“
CLINC150+OOS 95.8 89.2 βœ“
ContractNLI 75.5 59.1 βœ“
CRUXEval 59.9 57.1 βœ“
FinEntity 87.9 80.8
ForecastBench 23.8 30.6
GPQA-Diamond 32.6 71.4
GSM8K 97.1 75.6 βœ“
Habermas 18.5 21.5
HellaSwag 95.9 92.7 βœ“
HLE 0.0 4.7
Home appliances 60.2 52.3
HoVer 78.1 45.7 βœ“
Humicroedit 39.1 23.7 βœ“
iSarcasmEval 58.7 36.3 βœ“
MMLU-Pro 66.7 80.5
MuSR 45.9 46.1
New Yorker 84.9 62.6 βœ“
NLI4CT 53.6 69.0 βœ“
PhishNChips 57.4 25.1
POP909 77.9 15.9 βœ“
RAGTruth 67.9 51.3 βœ“
SATA-Bench 31.0 25.4
ToolRet 63.0 59.9
VAST 70.8 46.9 βœ“
When2Call 90.8 74.6 βœ“
WinoGrande 87.7 83.9 βœ“

Skills are on a 0–100 scale: 0 is chance and 100 is perfect.

We are ahead of Jev on 28 of the 38 index benchmarks and behind on 10. The largest gaps are on reasoning-heavy knowledge benchmarks: GPQA-Diamond βˆ’38.8, NLI4CT βˆ’15.3 and MMLU-Pro βˆ’13.8.

Usage

code/ holds the inference code: di_gemma_engine.py and the prompt protocol gemma_protokol.py. The protocol renders each question as one chat with thinking off, then reads the softmax over the option-label logits at the first generated position, divided by the temperature in temperatures.json. The engine class builds on the Decision Index kit, so install the kit too.

pip install "git+https://github.com/apolinario/decision-index@87d4650" \
    "torch>=2.8" "transformers>=5.17" "peft>=0.21" huggingface_hub
hf download moganai/mogan-decision-31B-it --include "code/*" --local-dir mogan
export PYTHONPATH=$PWD/mogan/code:$PYTHONPATH
from di_gemma_engine import GemmaEngine

engine = GemmaEngine("google/gemma-4-31B-it", adapter="moganai/mogan-decision-31B-it")
response, _ = engine(
    {"ticket": "I was charged twice for my subscription this month."},
    {"route": {"type": "choice", "instructions": "Which team should handle this ticket?",
               "criteria": {"billing": "Billing and payments", "tech": "Technical support",
                            "sales": "Sales"}}},
)
print(response["answers"]["route"])   # {"type": "choice", "choice": "billing", "probabilities": {...}}

Reproducing the Decision Index run

This needs a local copy of the suite. Build it with the kit's suite rebuild and suite import commands (see the kit's README; the sources' terms apply), then point --suite-dir at it:

python -m decision_index pipeline --edition 0.2.1 --suite-dir <your suite-0.2 dir> \
    --engine di_gemma_engine:GemmaEngine \
    --option model=google/gemma-4-31B-it --option model_revision=842da3794eaa0b77d5f08bae87a17459d91ff475 \
    --option adapter=moganai/mogan-decision-31B-it --out runs/mogan-decision-31B-it

The model needs one GPU with about 70 GB free (bf16). A prompt longer than 32,768 tokens raises Unsupported; nothing is truncated.

Training

  • Objective: soft cross-entropy between the model's softmax over the option-label logits (first generated position, same protocol as inference) and a target distribution. Option order is shuffled per example.
  • Hyperparameters:
    • LoRA r = 128, Ξ± = 128, dropout 0.05, on all attention and MLP projections of the language model.
    • LR 1e-4 with 50 warm-up steps, then cosine to 1e-5.
    • Global batch 64, one epoch (7,392 steps), up to 8,192 prompt tokens.
  • Hardware: 32 GH200 GPUs (8 nodes), about 5 hours.
  • Temperature: after training, log T = a + b Β· ln(n_options) was fitted per question type on a held-out 2% split of the training mix (8,022 questions), with no benchmark test data. The fitted T is about 0.9–0.96, so the model is close to calibrated without it.

Training data

473,141 questions in six parts:

Part Questions What
Benchmark-family train splits 223,117 Public training data of benchmark families, in the same question format
Teacher soft labels 123,094 moganai/mogan-decision-distill: Qwen3.8-Flash-Next answer distributions on MMLU auxiliary_train, SciQ, MedMCQA, OpenBookQA, SuperGPQA and GSM8K train
Own decision data 58,664 Public classification/NLI sets reformatted (MNLI, SNLI, BoolQ, AG News, IMDB, SGD, MS MARCO, TweetEval, …), LocalLLaMA/typed-decisions workflows, and scenarios written by Qwen3.6-35B-A3B from those workflows
Code reasoning 38,313 Python functions written by DeepSeek-V4-Flash; the answer options come from executing them
Base-model replay 24,998 The base model's own answer distributions on held-out questions, to limit forgetting
Tool-call decisions 4,955 NVIDIA When2Call train split; distractor options written by DeepSeek-V4-Flash

Benchmark families and the splits used. In every case the evaluation suite uses a different split:

Family Split used for training Split the suite uses
ANLI, VAST, ContractNLI, NLI4CT-2024, ACOS train + dev test
Humicroedit (SemEval-2020 Task 7, subtask 2) train + dev test
CLINC150+OOS train + val (+ oos train/val) test
New Yorker caption contest (matching) train + validation test
BANKING77, iSarcasmEval (En), Amazon ESCI, RAGTruth train test
HellaSwag, WinoGrande train validation
HoVer train dev
ARC-Easy, ARC-Challenge train test
ChessBench (searchless_chess behavioral cloning) train test
MMLU auxiliary_train test
GSM8K train test
When2Call train test
SGD train (inside our own decision data) test
POP909 (no official split) all 115 songs outside the suite (909 βˆ’ 792, minus songs 518 and 620, which the builder excludes for everyone) 792 other songs
BBH, CLadder (no train split) rule-generated items of the same task types test
CRUXEval LLM-written functions (no CRUXEval items) test

Not used: cfcolor, Home appliances, BRIGHT, ToolRet, BFCL, API-Bank, RouterBench, MuSR, SATA-Bench, GPQA, HLE, MMLU-Pro, FinEntity, Habermas, ForecastBench, BPoMP, PhishNChips and SimpleBench. For BFCL and API-Bank, we did train on tool-selection items from NVIDIA Nemotron tool-calling train data. Some of them were phrased the way BFCL asks its question and used the base model's own answers as targets; others used a "which API" format.

Decontamination

  • Method: every training item was checked against the frozen Decision Index 0.2 suite, including its excluded rows. The suite row content is sha256 b2b56d6f… (selected) and 7429f3c9… (added).
    • An item was dropped for a normalized exact match of any state or instruction text of at least 12 characters, or for any shared 13-gram.
    • ChessBench positions were matched on FEN.
  • Removed: 21,645 training items. We reclassified the 20,810 removed items of the main parts against the whole suite request:
    • 24 were identical requests,
    • 1,250 shared the full context with a suite request but asked a different question (mostly SuperGPQA questions that also appear in MMLU-Pro),
    • 19,536 shared only a text span. These spans are shared premises (ANLI), shared source documents (RAGTruth, HoVer), shared tool descriptions (ToolRet), and GSM8K training problems that also appear in RouterBench prompts.
  • Suite runs: this checkpoint was evaluated on the full suite once, plus one attempt that stopped after ~2,500 requests on an out-of-memory error. No setting was changed after either run.

Limitations

  • The model is specialized to single-step decisions with explicit options. It is not a chat model, and thinking is off.
  • Its strength on the Decision Index comes mostly from the benchmark families it was trained on; treat those scores as in-distribution. It is weakest on hard reasoning (GPQA-Diamond, MMLU-Pro, HLE), where answering in one forward pass without reasoning is a limit.
  • It is English-centric.

License

The adapter is released under CC BY-NC 4.0, because part of the training data is non-commercial:

  • ANLI, SciQ and ContractNLI (NC-SA),
  • MS MARCO,
  • RACE (inside MMLU auxiliary_train).

The base model, Gemma-4-31B-it, is Apache 2.0. The teacher labels come from Qwen3.8-Flash-Next (Qwen Community License 1.0).

Downloads last month
33
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for moganai/mogan-decision-31B-it

Adapter
(320)
this model

Datasets used to train moganai/mogan-decision-31B-it