Instructions to use ai-and/termgrade-gemma4-31b-merged with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ai-and/termgrade-gemma4-31b-merged with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ai-and/termgrade-gemma4-31b-merged") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ai-and/termgrade-gemma4-31b-merged") model = AutoModelForCausalLM.from_pretrained("ai-and/termgrade-gemma4-31b-merged", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ai-and/termgrade-gemma4-31b-merged with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ai-and/termgrade-gemma4-31b-merged" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-and/termgrade-gemma4-31b-merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ai-and/termgrade-gemma4-31b-merged
- SGLang
How to use ai-and/termgrade-gemma4-31b-merged with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ai-and/termgrade-gemma4-31b-merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-and/termgrade-gemma4-31b-merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ai-and/termgrade-gemma4-31b-merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-and/termgrade-gemma4-31b-merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ai-and/termgrade-gemma4-31b-merged with Docker Model Runner:
docker model run hf.co/ai-and/termgrade-gemma4-31b-merged
Configuration Parsing Warning:In config.json: "num_experts" must be a number
TermGrade: gemma-4-31B-it band-central, merged
Part of TermGrade: graded environments and trajectories for terminal agents. Read the blog post.
46.1 on Terminal-Bench 2.1 · +3.1 over base · no LoRA flags at serving time
google/gemma-4-31B-it with the band-central adapter already merged at bf16.
Adapter, methodology and caveats: ai-and/termgrade-gemma4-31b-lora.
Quick start
vllm serve ai-and/termgrade-gemma4-31b-merged \
--max-model-len 262144 \
--reasoning-parser gemma4 --tool-call-parser gemma4 --enable-auto-tool-choice \
--tensor-parallel-size 8 --trust-remote-code
Two settings that matter. Serve at 262144 context: at 128K, long agentic rollouts get cut off and fail to parse, which costs roughly 10 points. And use the native
gemma4parser; aqwen3_coder-style text parser costs roughly 20 points here.
Results
46.1 against a 43.0 base, so +3.1, the mean of seven evaluation runs across these weights and the adapter. These weights are that adapter merged into the base, so we pool the evaluations. Across five independent training runs of the same recipe the mean is 45.1 (+2.1), SE 0.6, with all five above the base model. That is the method's result; +3.1 is this checkpoint's.
The Artificial Analysis methodology we follow, the training recipe and the limitations are on the adapter card.
Citation
If you use TermGrade, please cite:
@misc{calik2026termgrade,
title = {{TermGrade}: 1k Graded Terminal Environments, 36k Trajectories, and the {RL} Run They Trained},
author = {Calik, Yagiz and Wu, Jianbo and Hara, Shimpei},
year = {2026},
month = oct,
howpublished = {\url{https://www.aiand.com/newsroom/termgrade}},
note = {ai\& Research blog post}
}
- Downloads last month
- 18