Instructions to use GotoAI-Inc/gemma-4-31B-it-W4A16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GotoAI-Inc/gemma-4-31B-it-W4A16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="GotoAI-Inc/gemma-4-31B-it-W4A16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("GotoAI-Inc/gemma-4-31B-it-W4A16") model = AutoModelForMultimodalLM.from_pretrained("GotoAI-Inc/gemma-4-31B-it-W4A16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GotoAI-Inc/gemma-4-31B-it-W4A16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GotoAI-Inc/gemma-4-31B-it-W4A16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GotoAI-Inc/gemma-4-31B-it-W4A16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/GotoAI-Inc/gemma-4-31B-it-W4A16
- SGLang
How to use GotoAI-Inc/gemma-4-31B-it-W4A16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GotoAI-Inc/gemma-4-31B-it-W4A16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GotoAI-Inc/gemma-4-31B-it-W4A16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GotoAI-Inc/gemma-4-31B-it-W4A16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GotoAI-Inc/gemma-4-31B-it-W4A16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use GotoAI-Inc/gemma-4-31B-it-W4A16 with Docker Model Runner:
docker model run hf.co/GotoAI-Inc/gemma-4-31B-it-W4A16
gemma-4-31B-it-W4A16
Int4 weight-only quantization of google/gemma-4-31B-it, in compressed-tensors format for vLLM. 19.07 GB, down from 62.55 GB.
Unofficial and unaffiliated with Google. All model capabilities, evaluations and limitations belong to the original model card — see the base model for those.
If you want the best 4-bit Gemma 4 31B, use Google's own gemma-4-31B-it-qat-w4a16-ct instead. It is quantization-aware trained, this one is post-training round-to-nearest. The reason to reach for this build is size: 19.07 GB against 23.27 GB, which is the difference between fitting a 24 GB card at short context and not.
What was changed
Weights were quantized from bfloat16 to int4, group size 128, symmetric, weight-only
(activations stay 16-bit) using llmcompressor.model_free_ptq. No calibration data was
used and the model was never loaded — the quantizer operates directly on the safetensors.
Architecture, tokenizer, chat template and processor config are the vendor's, unmodified.
410 Linear modules were converted, covering 79.2% of the checkpoint's bytes:
| component | precision | size |
|---|---|---|
| language-model linears (60 layers) | int4 g128 | 15.10 GB (79.2%) |
embed_tokens (tied to the output head) |
bfloat16 | 2.82 GB (14.8%) |
vision tower + embed_vision projection |
bfloat16 | 1.15 GB (6.0%) |
| norms, layer scalars | bfloat16 | 0.003 GB |
| total | 19.07 GB |
Left at bfloat16:
vision_tower,embed_vision— the tower'sintermediate_sizeis 4304, which is not divisible by 64, so int4 Marlin-style kernels cannot serve it. vLLM's Gemma 4 loader has an explicit guard for exactly this case and builds the towers unquantized; a checkpoint carrying quantized vision weights is asking for a version-dependent load failure.embed_tokens— precision-sensitive, and it is the output head here (tie_word_embeddings: true, and nolm_headtensor exists in the checkpoint).
Unlike the 12B, this model has no audio tower (audio_config: null, zero audio
tensors — Gemma 4 ships audio only on E2B, E4B and 12B), so the profile's re:.*audio.*
pattern matches nothing here. It also uses a conventional dedicated vision encoder rather
than the 12B's encoder-free "Unified" design, which is why the tower shows up as 1.15 GB of
separate weights.
Against Google's QAT build
| this repo | google/…-qat-w4a16-ct | |
|---|---|---|
| total size | 19.07 GB | 23.27 GB |
| method | data-free RTN (PTQ) | quantization-aware training |
| group size | 128 | 32 |
| quantized modules | 410 (79.2% of bytes) | 410 (70.8% of bytes) |
| int4 payload | 14.64 GB | 14.64 GB |
| scales | 0.46 GB | 1.83 GB |
| output head | tied to embed_tokens |
separate lm_head (2.82 GB, bf16) |
Both builds quantize the same 410 modules to an identical 14.64 GB of packed int4. The
entire 4.20 GB difference is the other two rows: group-32 scales cost 1.37 GB more than
group-128, and Google's build materializes an untied lm_head (2.82 GB) even though its
config still says tie_word_embeddings: true.
Usage
vllm serve GotoAI-Inc/gemma-4-31B-it-W4A16 \
--max-model-len 65536 \
--enable-auto-tool-choice --tool-call-parser gemma4 \
--reasoning-parser gemma4
Do not pass --quantization; compressed-tensors is detected from config.json. The int4
W4A16 scheme uses Marlin kernels and runs on compute capability 7.5 and above.
Version note. Gemma 4 needs transformers >= 5.10.1, but vLLM ≤ 0.27.1 cannot serve
Gemma 4 with transformers >= 5.15: at 5.15 the config turns global_head_dim into a
per-layer head_dim override, and older vLLM reads head_dim globally, raising
AmbiguousGlobalPerLayerAttributeError at engine-config time. Use either
transformers < 5.15 with vLLM 0.25.1–0.27, or vLLM >= 0.28, which handles both
layouts. The architecture itself (Gemma4ForConditionalGeneration) and both gemma4
parsers are present from 0.25.1 on.
Fitting the card
Gemma 4 interleaves 50 sliding-attention layers (window 1024, 16 KV heads, head_dim 256)
with 10 global layers (every 6th, 4 KV heads, global_head_dim 512). The sliding layers'
cache is bounded by the window at roughly 0.8 GB per sequence regardless of context
length; only the global layers grow, at about 80 KB/token:
| context | KV cache | + weights |
|---|---|---|
| 32k | ~3.4 GB | ~22.5 GB |
| 128k | ~11.3 GB | ~30.4 GB |
| 256k (max) | ~21.8 GB | ~40.9 GB |
So a 48 GB card runs this comfortably at 128k. A 24 GB card is marginal even at 32k once
activations and CUDA graphs are counted — add --language-model-only, which frees the
1.15 GB tower (vLLM skips tower weights entirely when every modality limit is zero) plus
the multimodal profiling headroom, and keep the context modest. This is arithmetic from
config.json, not a measured deployment.
Note that the global layers use unified keys and values (attention_k_eq_v: true, and the
checkpoint has no v_proj on those layers). That saves weight bytes, but vLLM loads the K
weights into both the K and V slots, so the cache still holds both copies — the table above
already assumes that.
Thinking
The chat template defaults enable_thinking to false, so this model does not think
unless asked. Both knobs are template variables passed through chat_template_kwargs:
{"chat_template_kwargs": {"enable_thinking": true}} // injects <|think|> into the system turn
{"chat_template_kwargs": {"preserve_thinking": true}} // keep thinking on tool-call turns
With thinking on, reasoning is emitted as <|channel>thought … <channel|>, which
--reasoning-parser gemma4 splits into reasoning_content. Per the base model card,
thinking from earlier turns should not be replayed into history — except on tool-call
turns, which is exactly what preserve_thinking keeps.
Sampling, per the base model card: temperature=1.0, top_p=0.95, top_k=64 for all use
cases. Place image content before the text in a prompt. The visual token budget is
configurable (70/140/280/560/1120, default 280) — lower it for video and captioning,
raise it for OCR and document parsing.
Reproducing this checkpoint
Built with llm-quantizer:
./llmq.py run --profile gemma-4-31b-it
which re-shards the source — it ships as 2 shards, the larger 49.78 GB, which no consumer GPU can hold — into 17 pieces of ~4 GB, then:
# llmcompressor==0.13.1a20260814, compressed-tensors==0.18.1a20260818,
# transformers==5.15.1, torch==2.13.0
from llmcompressor import model_free_ptq
model_free_ptq(
model_stub="gemma-4-31B-it-resharded",
save_directory="gemma-4-31B-it-W4A16",
scheme="W4A16",
ignore=["re:.*vision.*", "re:.*audio.*", "lm_head", "re:.*embed_tokens.*"],
device="cuda:0",
)
A job holds one shard at a time, so after re-sharding the build peaks at a few GB of VRAM rather than the model size — no large GPU required.
Evaluation
No benchmarks have been run. Data-free round-to-nearest quantization degrades quality more than the vendor's QAT build; how much, for your task, is unmeasured here. Treat published Gemma 4 benchmark numbers as describing the bf16 model, not this one — and note that a directly comparable QAT checkpoint exists, so if quality matters more than the 4.2 GB, use Google's.
License
Apache 2.0, inherited from the base model — see LICENSE and Google's
Gemma 4 license page. The base
repository ships no LICENSE file, so the Apache-2.0 text is included here for
redistribution. "Gemma" is Google's mark; this repository is not endorsed by or
affiliated with Google.
- Downloads last month
- 14