Instructions to use henrybravo/gemma-4-31b-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use henrybravo/gemma-4-31b-8bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("henrybravo/gemma-4-31b-8bit") config = load_config("henrybravo/gemma-4-31b-8bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use henrybravo/gemma-4-31b-8bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "henrybravo/gemma-4-31b-8bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "henrybravo/gemma-4-31b-8bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use henrybravo/gemma-4-31b-8bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "henrybravo/gemma-4-31b-8bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default henrybravo/gemma-4-31b-8bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use henrybravo/gemma-4-31b-8bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "henrybravo/gemma-4-31b-8bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "henrybravo/gemma-4-31b-8bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
henrybravo/gemma-4-31b-8bit
This repository provides an MLX quantized conversion of google/gemma-4-31b with an added chat template for tool calling support, prepared for local inference on Apple Silicon.
This variant is based on the mlx-community/gemma-4-31b-8bit conversion (using mlx-vlm 0.4.3), with the addition of a chat template that was missing from the original conversion.
Published by henrybravo. All model capabilities, limitations, and licensing inherit from the original Google Gemma release.
What this variant adds
The mlx-community conversion shipped without a chat template — neither tokenizer_config.json nor any separate template file contained one. Google's original google/gemma-4-31b also has no chat template in its published configs.
This variant adds a Jinja2 chat template that:
- Uses Gemma 4's native special tokens (
<|turn>/<turn|>for turn boundaries) - Supports tool/function calling via
<|tool_call>/<tool_call|>and<|tool_response>/<tool_response|>tokens - Handles system messages, multi-turn conversations, and multimodal content
- Is injected into both
tokenizer_config.jsonand provided as a standalonechat_template.jinja
Model summary
- Base model:
google/gemma-4-31b - Architecture:
Gemma4ForConditionalGeneration(gemma4) - Quantization: 8-bit (from
mlx-community/gemma-4-31b-8bit) - Weight format:
safetensors(7 shards) - Intended runtime:
mlx-vlm/ MLX ecosystem on macOS
Install
pip install -U mlx-vlm>=0.4.3
For OpenAI-compatible local serving and routing on top of MLX models, see mlx-router.
Quick usage (mlx-vlm)
You can run the examples below directly with mlx_vlm.generate, or serve the model through mlx-router.
Text prompt
python -m mlx_vlm.generate \
--model henrybravo/gemma-4-31b-8bit \
--max-tokens 100 \
--temperature 0.0 \
--prompt "Hello, what model are you?"
Vision prompt
python -m mlx_vlm.generate \
--model henrybravo/gemma-4-31b-8bit \
--max-tokens 200 \
--temperature 0.0 \
--prompt "Describe this image in detail." \
--image https://upload.wikimedia.org/wikipedia/commons/thumb/a/a7/Camponotus_flavomarginatus_ant.jpg/320px-Camponotus_flavomarginatus_ant.jpg
Notes
- This is a VLM model — load with
mlx-vlm, notmlx-lm(which does not support thegemma4architecture yet) - The chat template was authored based on Gemma 4's native special tokens found in
tokenizer_config.jsonand the Gemma 3 template structure has_tool_callingis not set on the tokenizer; tool call parsing should be handled by the serving layer (e.g.mlx-routerparses<|tool_call>...<tool_call|>blocks)
Upstream references
- Base model: https://huggingface.co/google/gemma-4-31b
- MLX conversion source: https://huggingface.co/mlx-community/gemma-4-31b-8bit
- Gemma license: https://ai.google.dev/gemma/terms
- Downloads last month
- 9
8-bit