File size: 4,751 Bytes
e36a69c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
57ae30c
 
 
e36a69c
 
 
 
 
 
 
57ae30c
 
 
 
 
 
 
e36a69c
 
57ae30c
 
 
 
 
e36a69c
 
 
 
57ae30c
 
 
 
 
 
 
12d7cbb
 
57ae30c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e36a69c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
---
pipeline_tag: image-text-to-text
library_name: mlx-vlm
license: other
base_model: kai-os/Grug-12B
base_model_relation: quantized
tags:
- mlx
- mlx-vlm
- gemma4_unified
- gemma-4
- vision-language
- image-text-to-text
- quantized
- 4-bit
- 6-bit
- 8-bit
- apple-silicon
datasets:
- hotdogs/uka-glm-5.2
- Scale-or-Reason/general-reasoning-ift-pairs
- samcheng0/lumia-reasoning-sft-v1
- HSH-Intelligence/verified-math-reasoning-3k
- kd13/CodeDebug-Instruct-v2-Reasoning
- Madarabr/cortex-adaptive-thinking
- CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
---

# Grug-12B VLM MLX

Apple Silicon MLX VLM quantizations of
[`kai-os/Grug-12B`](https://huggingface.co/kai-os/Grug-12B), packaged as a
single Hugging Face repo with one folder per quantization level.

`Grug-12B` is a compact-reasoning fine-tune of
[`google/gemma-4-12B-it`](https://huggingface.co/google/gemma-4-12B-it). The
source model was released as merged Transformers/safetensors weights after
QLoRA training. This repo only provides MLX quantized derivatives for Apple
Silicon inference and keeps the original vision-language model structure.

## Highlights

- Vision-language support is preserved through the Gemma 4 unified VLM config.
- Three MLX affine quantizations are available in one repo: 8-bit, 6-bit, and 4-bit.
- Benchmarked with oMLX on the MLX LM engine; screenshots are included below.
- The original BF16 Transformers weights remain in the source repo.

## Available variants

| Variant | Folder | Quantization | Size | Best fit |
| --- | --- | --- | ---: | --- |
| MLX 8-bit | `mlx-8bit/` | affine, group size 64 | 12 GB | Highest-quality local MLX run. |
| MLX 6-bit | `mlx-6bit/` | affine, group size 64 | 9.1 GB | Balanced quality, memory, and speed. |
| MLX 4-bit | `mlx-4bit/` | affine, group size 64 | 6.3 GB | Smallest footprint and best peak memory. |

These are not GGUF files and are not llama.cpp quants. They are MLX safetensors
folders intended for `mlx-vlm`.

## Benchmarks

Benchmarks were run with [oMLX](https://github.com/jundot/omlx), using the
`Force mlx-lm` engine. Each run used prompt prefill sizes of 1024, 4096, and
8192 tokens with 128 generated tokens. Values below are copied from the
captured benchmark output.

Hardware: Apple Mac Studio with M4 Max and 64 GB unified memory.

| Variant | pp1024 tg TPS | pp4096 tg TPS | pp8192 tg TPS | pp8192 E2E | Peak mem |
| --- | ---: | ---: | ---: | ---: | ---: |
| `mlx-8bit` | 30.3 tok/s | 20.4 tok/s | 31.6 tok/s | 20.189 s | 13.80 GB |
| `mlx-6bit` | 38.9 tok/s | 38.7 tok/s | 37.8 tok/s | 19.795 s | 11.03 GB |
| `mlx-4bit` | 21.7 tok/s | 15.7 tok/s | 50.9 tok/s | 18.540 s | 8.26 GB |

Continuous batching at `pp1024 / tg128`:

| Variant | Batch 1 tg TPS | Batch 2 tg TPS | Batch 2 speedup |
| --- | ---: | ---: | ---: |
| `mlx-8bit` | 30.3 tok/s | 34.2 tok/s | 1.13x |
| `mlx-6bit` | 38.9 tok/s | 40.5 tok/s | 1.04x |
| `mlx-4bit` | 21.7 tok/s | 56.1 tok/s | 2.59x |

<details>
<summary>Benchmark screenshots</summary>

### MLX 8-bit

![oMLX benchmark for Grug-12B VLM 8-bit](assets/benchmarks/omlx-8bit.png)

### MLX 6-bit

![oMLX benchmark for Grug-12B VLM 6-bit](assets/benchmarks/omlx-6bit.png)

### MLX 4-bit

![oMLX benchmark for Grug-12B VLM 4-bit](assets/benchmarks/omlx-4bit.png)

</details>

## Usage

Download only the variant you want:

```python
from pathlib import Path
from huggingface_hub import snapshot_download

repo_id = "chanderbalaji/Grug-12B-VLM-MLX"
variant = "mlx-4bit"

snapshot = snapshot_download(
    repo_id,
    allow_patterns=[f"{variant}/*"],
)
model_path = Path(snapshot) / variant
print(model_path)
```

Run with `mlx-vlm`:

```bash
python -m mlx_vlm.generate \
  --model /path/to/downloaded/snapshot/mlx-4bit \
  --prompt "Describe this image." \
  --image /path/to/image.jpg \
  --max-tokens 256
```

For text-only prompts, omit the `--image` argument.

## Provenance and attribution

- Source model: [`kai-os/Grug-12B`](https://huggingface.co/kai-os/Grug-12B)
- Base model: [`google/gemma-4-12B-it`](https://huggingface.co/google/gemma-4-12B-it)
- Relationship: MLX quantized derivatives of the source model
- Source revision used locally: `ad3feab42542e3361dcaf0ebe795d55009765918`
- Conversion target: Gemma 4 unified VLM with `vision_config` preserved

The source model card describes the original training recipe, datasets, local
evaluation, limitations, and acknowledgements. Please refer to that card for
the full model provenance and license context.

## Limitations

Quantization can change output quality, numerical behavior, and edge-case
performance. These files are intended for local MLX inference on Apple Silicon.
Use the source model repo for the original BF16 Transformers weights.