Instructions to use moganai/mogan-decision-31B-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use moganai/mogan-decision-31B-it with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-31B-it") model = PeftModel.from_pretrained(base_model, "moganai/mogan-decision-31B-it") - Notebooks
- Google Colab
- Kaggle
mogan-decision-31B-it
mogan-decision-31B-it is a decision model. You give it a state (text or JSON), a question and a set of options, and it returns a probability for every option in a single forward pass, with no generated reasoning. It supports two question types:
choice: pick one of 2β255 options,noul: the probability that a statement is true.
It is a LoRA (r = 128) on google/gemma-4-31B-it @ 842da37, trained with soft labels on 473k decision questions. At inference the adapter is merged into the bf16 weights.
Results: Decision Index
Scores are from our full run of the Decision Index suite (151,476 requests, all answered, none unsupported), scored with the official kit. They are self-reported until the board maintainers reproduce them.
| Edition | Decision Index | Raw index |
|---|---|---|
0.2.1 (kit 87d4650) |
65.42 | 73.29 |
0.2 (kit 19ad28e) |
60.17 | 69.93 |
Area skill on 0.2.1, with the board's reference system Jev 1.13 (0.2.1 index 57.91):
| Area | mogan-decision-31B-it | Jev 1.13 |
|---|---|---|
| Knowledge & Reasoning | 51.2 | 51.4 |
| Language Understanding | 72.0 | 62.0 |
| Retrieval & Classification | 69.6 | 55.4 |
| Tools & Automation | 79.7 | 75.1 |
| Arts & Human Taste | 50.7 | 37.7 |
Results files: moganai/mogan-decision-31B-it-decision-index-results. Median in-process latency is 107 ms per request (one process per GH200, one request at a time). A single-GPU rerun of 3,000 random suite requests with the published adapter and code reproduces the submitted choices on 100% of 6,447 questions.
This model is not zero-shot on many of these benchmarks. It was trained on the public training splits of several benchmark families; the full list is in Training data. The table shows our 0.2.1 skill per benchmark next to the board's reference system, Jev 1.13 (0.2.1 index 57.91). β marks families with training data (a train split, or generated items of the same task type).
| Benchmark | mogan-decision-31B-it | Jev 1.13 | Training data |
|---|---|---|---|
| ACOS | 45.7 | 27.3 | β |
| Amazon ESCI | 50.5 | 43.8 | β |
| ANLI | 69.4 | 62.2 | β |
| API-Bank | 84.4 | 88.0 | |
| BANKING77 | 91.9 | 79.5 | β |
| BBH | 84.3 | 89.7 | β |
| BFCL | 95.8 | 94.3 | |
| BPoMP | 87.0 | 81.8 | |
| BRIGHT | 40.1 | 40.6 | |
| cfcolor | 29.0 | 28.8 | |
| ChessBench | 23.6 | 9.8 | β |
| CLadder | 75.2 | 45.3 | β |
| CLINC150+OOS | 95.8 | 89.2 | β |
| ContractNLI | 75.5 | 59.1 | β |
| CRUXEval | 59.9 | 57.1 | β |
| FinEntity | 87.9 | 80.8 | |
| ForecastBench | 23.8 | 30.6 | |
| GPQA-Diamond | 32.6 | 71.4 | |
| GSM8K | 97.1 | 75.6 | β |
| Habermas | 18.5 | 21.5 | |
| HellaSwag | 95.9 | 92.7 | β |
| HLE | 0.0 | 4.7 | |
| Home appliances | 60.2 | 52.3 | |
| HoVer | 78.1 | 45.7 | β |
| Humicroedit | 39.1 | 23.7 | β |
| iSarcasmEval | 58.7 | 36.3 | β |
| MMLU-Pro | 66.7 | 80.5 | |
| MuSR | 45.9 | 46.1 | |
| New Yorker | 84.9 | 62.6 | β |
| NLI4CT | 53.6 | 69.0 | β |
| PhishNChips | 57.4 | 25.1 | |
| POP909 | 77.9 | 15.9 | β |
| RAGTruth | 67.9 | 51.3 | β |
| SATA-Bench | 31.0 | 25.4 | |
| ToolRet | 63.0 | 59.9 | |
| VAST | 70.8 | 46.9 | β |
| When2Call | 90.8 | 74.6 | β |
| WinoGrande | 87.7 | 83.9 | β |
Skills are on a 0β100 scale: 0 is chance and 100 is perfect.
We are ahead of Jev on 28 of the 38 index benchmarks and behind on 10. The largest gaps are on reasoning-heavy knowledge benchmarks: GPQA-Diamond β38.8, NLI4CT β15.3 and MMLU-Pro β13.8.
Usage
code/ holds the inference code: di_gemma_engine.py and the prompt protocol gemma_protokol.py. The
protocol renders each question as one chat with thinking off, then reads the softmax over the option-label
logits at the first generated position, divided by the temperature in temperatures.json. The engine class
builds on the Decision Index kit, so install the kit too.
pip install "git+https://github.com/apolinario/decision-index@87d4650" \
"torch>=2.8" "transformers>=5.17" "peft>=0.21" huggingface_hub
hf download moganai/mogan-decision-31B-it --include "code/*" --local-dir mogan
export PYTHONPATH=$PWD/mogan/code:$PYTHONPATH
from di_gemma_engine import GemmaEngine
engine = GemmaEngine("google/gemma-4-31B-it", adapter="moganai/mogan-decision-31B-it")
response, _ = engine(
{"ticket": "I was charged twice for my subscription this month."},
{"route": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Billing and payments", "tech": "Technical support",
"sales": "Sales"}}},
)
print(response["answers"]["route"]) # {"type": "choice", "choice": "billing", "probabilities": {...}}
Reproducing the Decision Index run
This needs a local copy of the suite. Build it with the kit's suite rebuild and suite import commands (see
the kit's README; the sources' terms apply), then point --suite-dir at it:
python -m decision_index pipeline --edition 0.2.1 --suite-dir <your suite-0.2 dir> \
--engine di_gemma_engine:GemmaEngine \
--option model=google/gemma-4-31B-it --option model_revision=842da3794eaa0b77d5f08bae87a17459d91ff475 \
--option adapter=moganai/mogan-decision-31B-it --out runs/mogan-decision-31B-it
The model needs one GPU with about 70 GB free (bf16). A prompt longer than 32,768 tokens raises
Unsupported; nothing is truncated.
Training
- Objective: soft cross-entropy between the model's softmax over the option-label logits (first generated position, same protocol as inference) and a target distribution. Option order is shuffled per example.
- Hyperparameters:
- LoRA r = 128, Ξ± = 128, dropout 0.05, on all attention and MLP projections of the language model.
- LR 1e-4 with 50 warm-up steps, then cosine to 1e-5.
- Global batch 64, one epoch (7,392 steps), up to 8,192 prompt tokens.
- Hardware: 32 GH200 GPUs (8 nodes), about 5 hours.
- Temperature: after training, log T = a + b Β· ln(n_options) was fitted per question type on a held-out 2% split of the training mix (8,022 questions), with no benchmark test data. The fitted T is about 0.9β0.96, so the model is close to calibrated without it.
Training data
473,141 questions in six parts:
| Part | Questions | What |
|---|---|---|
| Benchmark-family train splits | 223,117 | Public training data of benchmark families, in the same question format |
| Teacher soft labels | 123,094 | moganai/mogan-decision-distill: Qwen3.8-Flash-Next answer distributions on MMLU auxiliary_train, SciQ, MedMCQA, OpenBookQA, SuperGPQA and GSM8K train |
| Own decision data | 58,664 | Public classification/NLI sets reformatted (MNLI, SNLI, BoolQ, AG News, IMDB, SGD, MS MARCO, TweetEval, β¦), LocalLLaMA/typed-decisions workflows, and scenarios written by Qwen3.6-35B-A3B from those workflows |
| Code reasoning | 38,313 | Python functions written by DeepSeek-V4-Flash; the answer options come from executing them |
| Base-model replay | 24,998 | The base model's own answer distributions on held-out questions, to limit forgetting |
| Tool-call decisions | 4,955 | NVIDIA When2Call train split; distractor options written by DeepSeek-V4-Flash |
Benchmark families and the splits used. In every case the evaluation suite uses a different split:
| Family | Split used for training | Split the suite uses |
|---|---|---|
| ANLI, VAST, ContractNLI, NLI4CT-2024, ACOS | train + dev | test |
| Humicroedit (SemEval-2020 Task 7, subtask 2) | train + dev | test |
| CLINC150+OOS | train + val (+ oos train/val) | test |
| New Yorker caption contest (matching) | train + validation | test |
| BANKING77, iSarcasmEval (En), Amazon ESCI, RAGTruth | train | test |
| HellaSwag, WinoGrande | train | validation |
| HoVer | train | dev |
| ARC-Easy, ARC-Challenge | train | test |
| ChessBench (searchless_chess behavioral cloning) | train | test |
| MMLU | auxiliary_train | test |
| GSM8K | train | test |
| When2Call | train | test |
| SGD | train (inside our own decision data) | test |
| POP909 (no official split) | all 115 songs outside the suite (909 β 792, minus songs 518 and 620, which the builder excludes for everyone) | 792 other songs |
| BBH, CLadder (no train split) | rule-generated items of the same task types | test |
| CRUXEval | LLM-written functions (no CRUXEval items) | test |
Not used: cfcolor, Home appliances, BRIGHT, ToolRet, BFCL, API-Bank, RouterBench, MuSR, SATA-Bench, GPQA, HLE, MMLU-Pro, FinEntity, Habermas, ForecastBench, BPoMP, PhishNChips and SimpleBench. For BFCL and API-Bank, we did train on tool-selection items from NVIDIA Nemotron tool-calling train data. Some of them were phrased the way BFCL asks its question and used the base model's own answers as targets; others used a "which API" format.
Decontamination
- Method: every training item was checked against the frozen Decision Index 0.2 suite, including its
excluded rows. The suite row content is sha256
b2b56d6fβ¦(selected) and7429f3c9β¦(added).- An item was dropped for a normalized exact match of any state or instruction text of at least 12 characters, or for any shared 13-gram.
- ChessBench positions were matched on FEN.
- Removed: 21,645 training items. We reclassified the 20,810 removed items of the main parts against the
whole suite request:
- 24 were identical requests,
- 1,250 shared the full context with a suite request but asked a different question (mostly SuperGPQA questions that also appear in MMLU-Pro),
- 19,536 shared only a text span. These spans are shared premises (ANLI), shared source documents (RAGTruth, HoVer), shared tool descriptions (ToolRet), and GSM8K training problems that also appear in RouterBench prompts.
- Suite runs: this checkpoint was evaluated on the full suite once, plus one attempt that stopped after ~2,500 requests on an out-of-memory error. No setting was changed after either run.
Limitations
- The model is specialized to single-step decisions with explicit options. It is not a chat model, and thinking is off.
- Its strength on the Decision Index comes mostly from the benchmark families it was trained on; treat those scores as in-distribution. It is weakest on hard reasoning (GPQA-Diamond, MMLU-Pro, HLE), where answering in one forward pass without reasoning is a limit.
- It is English-centric.
License
The adapter is released under CC BY-NC 4.0, because part of the training data is non-commercial:
- ANLI, SciQ and ContractNLI (NC-SA),
- MS MARCO,
- RACE (inside MMLU auxiliary_train).
The base model, Gemma-4-31B-it, is Apache 2.0. The teacher labels come from Qwen3.8-Flash-Next (Qwen Community License 1.0).
- Downloads last month
- 33
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-31B-it") model = PeftModel.from_pretrained(base_model, "moganai/mogan-decision-31B-it")