henrybravo/gemma-4-31b-8bit

This repository provides an MLX quantized conversion of google/gemma-4-31b with an added chat template for tool calling support, prepared for local inference on Apple Silicon.

This variant is based on the mlx-community/gemma-4-31b-8bit conversion (using mlx-vlm 0.4.3), with the addition of a chat template that was missing from the original conversion.

Published by henrybravo. All model capabilities, limitations, and licensing inherit from the original Google Gemma release.

What this variant adds

The mlx-community conversion shipped without a chat template — neither tokenizer_config.json nor any separate template file contained one. Google's original google/gemma-4-31b also has no chat template in its published configs.

This variant adds a Jinja2 chat template that:

  • Uses Gemma 4's native special tokens (<|turn> / <turn|> for turn boundaries)
  • Supports tool/function calling via <|tool_call> / <tool_call|> and <|tool_response> / <tool_response|> tokens
  • Handles system messages, multi-turn conversations, and multimodal content
  • Is injected into both tokenizer_config.json and provided as a standalone chat_template.jinja

Model summary

  • Base model: google/gemma-4-31b
  • Architecture: Gemma4ForConditionalGeneration (gemma4)
  • Quantization: 8-bit (from mlx-community/gemma-4-31b-8bit)
  • Weight format: safetensors (7 shards)
  • Intended runtime: mlx-vlm / MLX ecosystem on macOS

Install

pip install -U mlx-vlm>=0.4.3

For OpenAI-compatible local serving and routing on top of MLX models, see mlx-router.

Quick usage (mlx-vlm)

You can run the examples below directly with mlx_vlm.generate, or serve the model through mlx-router.

Text prompt

python -m mlx_vlm.generate \
  --model henrybravo/gemma-4-31b-8bit \
  --max-tokens 100 \
  --temperature 0.0 \
  --prompt "Hello, what model are you?"

Vision prompt

python -m mlx_vlm.generate \
  --model henrybravo/gemma-4-31b-8bit \
  --max-tokens 200 \
  --temperature 0.0 \
  --prompt "Describe this image in detail." \
  --image https://upload.wikimedia.org/wikipedia/commons/thumb/a/a7/Camponotus_flavomarginatus_ant.jpg/320px-Camponotus_flavomarginatus_ant.jpg

Notes

  • This is a VLM model — load with mlx-vlm, not mlx-lm (which does not support the gemma4 architecture yet)
  • The chat template was authored based on Gemma 4's native special tokens found in tokenizer_config.json and the Gemma 3 template structure
  • has_tool_calling is not set on the tokenizer; tool call parsing should be handled by the serving layer (e.g. mlx-router parses <|tool_call>...<tool_call|> blocks)

Upstream references

Downloads last month
9
Safetensors
Model size
31B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support