Instructions to use llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic") model = AutoModelForMultimodalLM.from_pretrained("llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic
- SGLang
How to use llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic with Docker Model Runner:
docker model run hf.co/llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic
Model works with Ollama (MLX, safetensors) on Apple Silicon
✅ Model works with Ollama (MLX, safetensors) on Apple Silicon
Sharing that this model runs successfully in Ollama directly from safetensors — not via GGUF conversion, but through the experimental MLX import path Ollama added for Apple Silicon.
Hardware
MacBook Pro M3 Max, 128GB unified memory
macOS
Ollama 0.32.0
How I downloaded it
The unquantized version weighs ~40GB (this is the qat-q4_0-unquantized variant — i.e., the full-precision model left over after quantization-aware training, without the final quantization step applied).
hf download llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic \
--local-dir /Volumes/Lexar4TB/original
(hf is the current CLI from huggingface_hub, replacing the deprecated huggingface-cli)
How I imported it into Ollama
cd /Volumes/Lexar4TB/original
echo "FROM ." > Modelfile
ollama create --experimental gemma4-31b-unquant -f Modelfile
The --experimental flag is required — without it, the safetensors import won't run; the MLX path in Ollama is still marked experimental.
Running it
ollama run gemma4-31b-unquant
The model loaded and responds at 100% GPU (confirmed via ollama ps), with no errors during import or launch.
First impressions
Subjectively feels faster than the GGUF version of the same model — but this isn't a proper benchmark yet, just a first impression. I plan to run a real comparison (tokens/sec on identical prompts) and will update this post with numbers.
Important note for anyone trying to reproduce this
If the model responds very slowly after import and ollama ps shows mixed CPU/GPU usage — check the context window. Ollama sometimes defaults to the model's maximum context length (which can be very large for Gemma 4), and the KV cache at that context size may not fully fit in GPU memory for a model this size, pushing some layers to CPU and drastically slowing everything down. Explicitly limiting the context at launch fixes this.