Text Generation
Transformers
Safetensors
gemma4_unified
image-text-to-text
gemma4
coding
code
reasoning
thinking
conversational
Instructions to use yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1") model = AutoModelForMultimodalLM.from_pretrained("yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1
- SGLang
How to use yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1 with Docker Model Runner:
docker model run hf.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1
File size: 6,911 Bytes
237e5d8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 | ---
license: apache-2.0
base_model: google/gemma-4-12B-it
library_name: transformers
pipeline_tag: text-generation
tags: [gemma4, coding, code, reasoning, thinking, safetensors, transformers]
---
# π» Gemma4-12B-Coder β **safetensors master (full precision)** β¨
### Composer 2.5 Γ Fable 5 Β· v1 / code edition
> **This is the full-precision `safetensors` master** for my Gemma 4 12B coding fine-tune β the same model many of
> you have been running as GGUF, now in its original weights. π§ π» A focused fine-tune of Gemma 4 12B on
> **verifiable Python coding** data: it reasons in the open (edge cases, complexity, approach) and then writes a
> clean, runnable solution.
---
## π― What this repo is for
This repo holds the **un-quantized master weights** (`model.safetensors`, bf16). Use it to:
- π§ **Roll your own quants** β make custom GGUF / **MLX** / AWQ / GPTQ builds from full precision.
- π§ͺ **Fine-tune further** β it's a clean base for your own LoRA / continued training.
- π€ **Run it in `transformers`** (needs a recent build with `gemma4_unified` support).
> π **Just want to run it?** You don't need this repo β grab a ready-made quant from the
> **[GGUF repo β](https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF)** (runs in ~4.5 GB of
> VRAM / unified memory in LM Studio, Ollama, llama.cpp, Janβ¦). This master is for *builders*. π
---
## π Announcements
**π v2 is almost here!** Initial training of **v2 is done** and it's in **benchmarking + final QA**. So many of you
flagged the **agentic** behavior β so this round I **significantly grew the dataset (especially agentic data)**.
**v2 is focused on agentic + coding.** Targeting a release **this Friday or Saturday (US Pacific).** π
**π£ Context length is 256K.** This master ships with the corrected `max_position_embeddings = 262144` (256K) β the
well-known upstream Gemma 4 metadata bug (`config.json` once said `131072`) is **already fixed here**, so anything you
quantize/convert from these weights inherits the full 256K. π Thanks to the community member who spotted it!
---
## π€ Run it in transformers
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
repo = "yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map="auto")
msgs = [{"role": "user", "content": "Write a Python function to check if a string is a valid IPv4 address."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=1024)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
```
> π§ **Thinking mode:** it thinks in Gemma's native thought channel before answering (keep `enable_thinking=true`,
> the default chat template handles it). Recommended sampling: `temp 1.0, top_p 0.95, top_k 64`; for coding you can
> also go greedy (`temp 0`) for more deterministic solutions. Needs a **recent `transformers`** that knows the
> `gemma4_unified` architecture.
---
## π¦ Ready-made GGUF quants
All from the **[GGUF repo](https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF)**:
| Quant | Size | Vibe |
|------|------|------|
| π’ [**Q2_K**](https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF/blob/main/gemma4-coding-Q2_K.gguf) | **4.5 GB** | tiniest β runs almost anywhere |
| π‘ [**Q3_K_M**](https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF/blob/main/gemma4-coding-Q3_K_M.gguf) | **5.7 GB** | great for 8 GB VRAM |
| π΅ [**Q4_K_M**](https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF/blob/main/gemma4-coding-Q4_K_M.gguf) | **6.87 GB** | the sweet spot π (recommended) |
| π£ [**Q6_K**](https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF/blob/main/gemma4-coding-Q6_K.gguf) | **9.11 GB** | near-lossless |
| βͺ [**Q8_0**](https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF/blob/main/gemma4-coding-Q8_0.gguf) | **11.8 GB** | basically full quality |
> β οΈ GGUF needs a **recent llama.cpp** β this is the `gemma4_unified` architecture, older builds won't load it.
---
## β‘ Optional: free speed with MTP (lossless)
There's a tiny **Gemma 4 MTP draft model** in my main reasoning repo β
**[`MTP/` folder](https://huggingface.co/yuxinlu1/gemma-4-12B-it-Claude-4.6-4.8-Opus-GGUF/tree/main/MTP)**. It's the
**stock Gemma 4 drafter**, so it pairs with **any** Gemma 4 12B quant β including these coder quants β for
**lossless speculative decoding** (byte-for-byte identical output, just faster). Because it's trained on base Gemma 4,
the hit-rate on this fine-tune is a bit lower than on vanilla Gemma 4, but it's free and has no downside. Add three
flags (`--model-draft`, `--spec-type draft-mtp`, `--n-gpu-layers-draft`); see the
[main repo](https://huggingface.co/yuxinlu1/gemma-4-12B-it-Claude-4.6-4.8-Opus-GGUF) for the full command. ποΈ
---
## π Training data (the interesting part π³)
A **distillation** of two complementary chain-of-thought sources over verifiable Python coding tasks
(algorithmic / function-level problems with deterministic tests):
- **π₯ Main β Composer 2.5 *real* CoT.** Genuine model-authored reasoning traces; each solution was **run against the
task's tests and only passing ones were kept**. The reasoning you learn from leads to code that *actually works*.
- **π₯ Aux β Fable 5 redo.** The problems where Composer 2.5 got it **wrong**, handed to Fable 5 to *re-derive* a fresh,
self-consistent CoT and a correct solution β again **gated on passing the tests**. Recovers the hard cases the main
teacher missed. These are synthetic (rationalized) CoT and are tagged separately.
Real CoT for solid coverage + synthetic "second-attempt" CoT to patch the failures β **all verified by execution**
before training. β
---
## β οΈ Good to know
- **Reduced refusals:** task-focused training with no safety hedging, so it refuses less than the base model. It is
**not** safety-aligned β add your own guardrails for production. Use responsibly. π
- Specialized for **Python / algorithmic** coding; general-knowledge facts/numbers should still be double-checked.
- English-centric.
---
## π Base & License
- **License: Apache 2.0.** Gemma 4 is released by Google under
**[Apache 2.0](https://ai.google.dev/gemma/apache_2)** (unlike the older Gemma 1/2/3 terms), so this fine-tune is
**Apache 2.0** too β free to use, modify, and redistribute. π
- **Base model:** [`google/gemma-4-12B-it`](https://huggingface.co/google/gemma-4-12B-it).
- Personal/hobby project β shared as-is, no warranty. Have fun, and happy hacking! πΎβ¨
|