Configuration Parsing Warning:In config.json: "num_experts" must be a number

TermGrade: gemma-4-31B-it band-central, merged

Part of TermGrade: graded environments and trajectories for terminal agents. Read the blog post.

46.1 on Terminal-Bench 2.1 · +3.1 over base · no LoRA flags at serving time

google/gemma-4-31B-it with the band-central adapter already merged at bf16.

Adapter, methodology and caveats: ai-and/termgrade-gemma4-31b-lora.


Quick start

vllm serve ai-and/termgrade-gemma4-31b-merged \
  --max-model-len 262144 \
  --reasoning-parser gemma4 --tool-call-parser gemma4 --enable-auto-tool-choice \
  --tensor-parallel-size 8 --trust-remote-code

Two settings that matter. Serve at 262144 context: at 128K, long agentic rollouts get cut off and fail to parse, which costs roughly 10 points. And use the native gemma4 parser; a qwen3_coder-style text parser costs roughly 20 points here.


Results

46.1 against a 43.0 base, so +3.1, the mean of seven evaluation runs across these weights and the adapter. These weights are that adapter merged into the base, so we pool the evaluations. Across five independent training runs of the same recipe the mean is 45.1 (+2.1), SE 0.6, with all five above the base model. That is the method's result; +3.1 is this checkpoint's.

The Artificial Analysis methodology we follow, the training recipe and the limitations are on the adapter card.

Citation

If you use TermGrade, please cite:

@misc{calik2026termgrade,
  title        = {{TermGrade}: 1k Graded Terminal Environments, 36k Trajectories, and the {RL} Run They Trained},
  author       = {Calik, Yagiz and Wu, Jianbo and Hara, Shimpei},
  year         = {2026},
  month        = oct,
  howpublished = {\url{https://www.aiand.com/newsroom/termgrade}},
  note         = {ai\& Research blog post}
}
Downloads last month
18
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ai-and/termgrade-gemma4-31b-merged

Finetuned
(293)
this model

Dataset used to train ai-and/termgrade-gemma4-31b-merged

Collection including ai-and/termgrade-gemma4-31b-merged