Instructions to use nghodki/Gemma4-Instinct-12B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use nghodki/Gemma4-Instinct-12B with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-12B-it") model = PeftModel.from_pretrained(base_model, "nghodki/Gemma4-Instinct-12B") - Notebooks
- Google Colab
- Kaggle
Gemma4-Instinct-12B
LoRA adapter for google/gemma-4-12B-it — a fast, structured decision model that makes typed judgments in a single forward pass without generating any tokens.
Instinct reads a policy context + evidence + schema, then scores answer codes via logit extraction. No chain-of-thought, no token generation — just one prefill pass and a probability distribution over the allowed choices.
How It Works
Input: Policy + Evidence + Schema {approve, deny, escalate}
|
Model: Single forward pass (prefill only)
|
Output: {approve: 0.94, deny: 0.04, escalate: 0.02} <- ~135ms on H100
No tokens generated. The model reads its own logits at the answer position and returns calibrated probabilities.
Results
| Metric | Score |
|---|---|
| Eval accuracy | 93.2% |
| Base accuracy (no training) | 77.5% |
| Improvement | +15.7% |
| Bespoke-Nimble-9B (reference) | 90.1% |
Evaluated on 324 held-out contrastive decision examples across 10 domains (commerce, education, travel, software, public services, etc.) and 3 task types (choice, boolean, ordered rubric).
Latency
| Condition | Latency |
|---|---|
| H100 (warm, single example) | ~135 ms |
| H100 (cold / first inference) | ~570 ms |
Task Types
- Choice — pick one from N options (e.g., triage routing, defect severity)
- Boolean (noul) — true/false with calibrated P(true)
- Score — ordered rubric levels with expected-score computation
Training Config
- Method: LoRA (rank=32, alpha=32)
- Target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
- Learning rate: 5e-5 (cosine schedule, 20% warmup)
- Epochs: 3
- Batch size: 2 (gradient accumulation 8, effective batch 16)
- Training data: 2,676 contrastive decision pairs
- Hardware: Single H100 80GB, ~55 minutes
- Trainable params: 131M / 12B total (1.08%)
Key Technical Detail
Gemma 4's chat template includes a thinking channel (<|channel>thought<channel|>) that must be stripped for logit-based scoring to work. Without this fix, training degrades the model. The thinking channel tokens occupy the position where the answer code logits should be read.
# Required fix when building prompts for Gemma 4
template = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
template = template.replace("<|channel>thought\n<channel|>", "")
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
base = AutoModelForCausalLM.from_pretrained("google/gemma-4-12B-it", torch_dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, "nghodki/Gemma4-Instinct-12B")
License
Apache-2.0 (same as base model)
- Downloads last month
- 27