ashxhart commited on
Commit
8f50ae2
·
verified ·
1 Parent(s): a13613b

Add Mac memory chooser and reproducible demo prompt

Browse files
Files changed (1) hide show
  1. README.md +71 -0
README.md CHANGED
@@ -18,22 +18,27 @@ tags:
18
  - 4-bit
19
  ---
20
 
 
21
  <p align="center">
22
  <img src="https://huggingface.co/front/assets/huggingface_logo-noborder.svg" width="92" alt="Hugging Face logo">
23
  </p>
24
 
 
25
  <p align="center">
26
  <img src="https://img.shields.io/badge/Meta-Muse_Glimmer-0467DF?style=for-the-badge&logo=meta&logoColor=white" alt="Meta Muse Glimmer">
27
  <img src="https://img.shields.io/badge/Apple_Silicon-MLX-000000?style=for-the-badge&logo=apple&logoColor=white" alt="Apple silicon MLX">
28
  <img src="https://img.shields.io/badge/Vontra-MLX_VLM-6E56CF?style=for-the-badge&logo=huggingface&logoColor=white" alt="Vontra MLX VLM">
29
  </p>
30
 
 
31
  <h1 align="center">Muse Glimmer 30B — MLX 4-bit</h1>
32
 
 
33
  <p align="center">
34
  A native Apple-silicon conversion of <a href="https://huggingface.co/meta-models/Muse-Glimmer-30B">meta-models/Muse-Glimmer-30B</a>, converted to uniform affine 4-bit MLX weights for text-and-image inference with MLX-VLM.
35
  </p>
36
 
 
37
  <p align="center">
38
  <a href="https://huggingface.co/meta-models/Muse-Glimmer-30B">Original model</a> ·
39
  <a href="https://github.com/Blaizzy/mlx-vlm">MLX-VLM</a> ·
@@ -41,10 +46,13 @@ tags:
41
  <a href="https://huggingface.co/Vontra">More Vontra conversions</a>
42
  </p>
43
 
 
44
  ## About this conversion
45
 
 
46
  This repository contains a **uniform MLX 4-bit conversion** of Muse Glimmer 30B, a dense agentic language model with a dedicated perception encoder. It preserves the upstream tokenizer, ATEM chat template, image processor configuration, generation configuration, licence, and usage policy.
47
 
 
48
  | Item | Value |
49
  | --- | --- |
50
  | Base model | [`meta-models/Muse-Glimmer-30B`](https://huggingface.co/meta-models/Muse-Glimmer-30B) |
@@ -58,15 +66,20 @@ This repository contains a **uniform MLX 4-bit conversion** of Muse Glimmer 30B,
58
  | Maximum visual tokens | 4,096 per image |
59
  | Architecture | `muse_glimmer` |
60
 
 
61
  Modules that MLX-VLM does not classify as quantizable remain at their source-compatible precision, so the effective whole-checkpoint bits-per-weight is higher than four. The `quantization` metadata records the actual eligible-module recipe.
62
 
 
63
  > [!IMPORTANT]
64
  > Muse Glimmer support was validated against the official MLX-VLM source at commit [`5262cb6`](https://github.com/Blaizzy/mlx-vlm/commit/5262cb6a27c797c5c3daf64b17d757924c5c474f), reporting package version 0.6.12. An older MLX-VLM or oMLX bundle may report `Model type muse_glimmer not supported`; update to a build containing the upstream Muse implementation before loading this checkpoint.
65
 
 
66
  ## Apple-silicon performance
67
 
 
68
  This checkpoint was load-tested, text-generation tested, and benchmarked on:
69
 
 
70
  | Hardware | Configuration |
71
  | --- | --- |
72
  | Host | Mac Studio |
@@ -75,8 +88,10 @@ This checkpoint was load-tested, text-generation tested, and benchmarked on:
75
  | Unified memory | 256 GB |
76
  | Runtime | MLX-VLM 0.6.12 source revision `5262cb6` with the oMLX MLX runtime |
77
 
 
78
  A warmed local text-only test produced:
79
 
 
80
  | Measurement | Result |
81
  | --- | ---: |
82
  | Decode (median) | **37.70 tokens/s** |
@@ -86,29 +101,38 @@ A warmed local text-only test produced:
86
  | Warm-up | 256 generated tokens |
87
  | Prompt | 83 tokens after chat templating |
88
 
 
89
  The decode figure is the median of three greedy 256-token runs after a full 256-token Metal-kernel warm-up. It is a practical local reference, not a controlled cross-platform benchmark. Prompt length, images, context growth, sampler settings, memory pressure, thermal state, and runtime versions can materially change performance. Image prefill is not included in this decode benchmark.
90
 
 
91
  ## Runtime setup
92
 
 
93
  Until Muse Glimmer support reaches the MLX-VLM build supplied by your application, install the exact official source revision used for validation:
94
 
 
95
  ```bash
96
  python -m pip install \
97
  "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm.git@5262cb6a27c797c5c3daf64b17d757924c5c474f"
98
  ```
99
 
 
100
  Download the checkpoint if a local copy is preferred:
101
 
 
102
  ```bash
103
  hf download Vontra/Muse-Glimmer-30B-MLX-4bit \
104
  --local-dir ~/.omlx/models/Vontra/Muse-Glimmer-30B-MLX-4bit
105
  ```
106
 
 
107
  ## Text generation
108
 
 
109
  ```python
110
  from mlx_vlm import generate, load
111
 
 
112
  model, processor = load("Vontra/Muse-Glimmer-30B-MLX-4bit")
113
  tokenizer = processor.tokenizer if hasattr(processor, "tokenizer") else processor
114
  messages = [
@@ -130,15 +154,20 @@ result = generate(
130
  print(result.text)
131
  ```
132
 
 
133
  Muse Glimmer supports `low`, `medium`, `high`, and `xhigh` reasoning strengths. Higher settings can spend more tokens reasoning before returning the final answer.
134
 
 
135
  ## Image and text generation
136
 
 
137
  The multimodal path was smoke-tested locally with a real image. Include an image content item so the ATEM template emits the required patch token:
138
 
 
139
  ```python
140
  from mlx_vlm import generate, load
141
 
 
142
  model, processor = load("Vontra/Muse-Glimmer-30B-MLX-4bit")
143
  tokenizer = processor.tokenizer if hasattr(processor, "tokenizer") else processor
144
  messages = [
@@ -166,10 +195,13 @@ result = generate(
166
  print(result.text)
167
  ```
168
 
 
169
  ## Architecture
170
 
 
171
  Muse Glimmer combines a dense causal transformer with a dedicated perception encoder for interleaved text and image input.
172
 
 
173
  | Architecture detail | Upstream value |
174
  | --- | ---: |
175
  | Total parameters | ~29.6B |
@@ -184,10 +216,13 @@ Muse Glimmer combines a dense causal transformer with a dedicated perception enc
184
  | Perception encoder | ~1.8B-parameter ViT-G/14, 50 layers |
185
  | Context length | 131,072 tokens |
186
 
 
187
  The model supports agentic task completion, tool use through the upstream ATEM protocol, controllable reasoning effort, multilingual input, failure recovery, and multimodal understanding. See the [original model card](https://huggingface.co/meta-models/Muse-Glimmer-30B) for upstream benchmarks, training details, intended uses, limitations, and safety guidance.
188
 
 
189
  ## Conversion and validation notes
190
 
 
191
  - Source weights: upstream BF16 checkpoint at revision `a4e59da52a7bc87ae7251dd5545c0dd437c44b68`.
192
  - Quantization mode: affine, 4-bit, group size 64, without mixed-precision overrides.
193
  - The upstream tokenizer, ATEM chat template, processor configuration, generation configuration, licence, and usage policy are included.
@@ -196,10 +231,46 @@ The model supports agentic task completion, tool use through the upstream ATEM p
196
  - Quantization can reduce output quality relative to BF16. Use a higher-precision variant when quality matters more than memory use.
197
  - This release does not include or claim support for the upstream speculative drafter.
198
 
 
199
  This is a community conversion, not an official Meta release. Validate quality, safety, and numerical behaviour on representative workloads before production use.
200
 
 
201
  ## Licence, usage policy, and attribution
202
 
 
203
  The upstream model is released under the **Apache License 2.0**. The upstream `LICENSE` and `USAGE_POLICY.md` files are included in this repository; use is subject to both the licence and the upstream usage policy.
204
 
 
205
  All model design, training, benchmark, and upstream documentation credit belongs to Meta and the original contributors. The MLX conversion, Apple-silicon validation, compatibility work, and model card are provided by [Vontra](https://huggingface.co/Vontra).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
18
  - 4-bit
19
  ---
20
 
21
+
22
  <p align="center">
23
  <img src="https://huggingface.co/front/assets/huggingface_logo-noborder.svg" width="92" alt="Hugging Face logo">
24
  </p>
25
 
26
+
27
  <p align="center">
28
  <img src="https://img.shields.io/badge/Meta-Muse_Glimmer-0467DF?style=for-the-badge&logo=meta&logoColor=white" alt="Meta Muse Glimmer">
29
  <img src="https://img.shields.io/badge/Apple_Silicon-MLX-000000?style=for-the-badge&logo=apple&logoColor=white" alt="Apple silicon MLX">
30
  <img src="https://img.shields.io/badge/Vontra-MLX_VLM-6E56CF?style=for-the-badge&logo=huggingface&logoColor=white" alt="Vontra MLX VLM">
31
  </p>
32
 
33
+
34
  <h1 align="center">Muse Glimmer 30B — MLX 4-bit</h1>
35
 
36
+
37
  <p align="center">
38
  A native Apple-silicon conversion of <a href="https://huggingface.co/meta-models/Muse-Glimmer-30B">meta-models/Muse-Glimmer-30B</a>, converted to uniform affine 4-bit MLX weights for text-and-image inference with MLX-VLM.
39
  </p>
40
 
41
+
42
  <p align="center">
43
  <a href="https://huggingface.co/meta-models/Muse-Glimmer-30B">Original model</a> ·
44
  <a href="https://github.com/Blaizzy/mlx-vlm">MLX-VLM</a> ·
 
46
  <a href="https://huggingface.co/Vontra">More Vontra conversions</a>
47
  </p>
48
 
49
+
50
  ## About this conversion
51
 
52
+
53
  This repository contains a **uniform MLX 4-bit conversion** of Muse Glimmer 30B, a dense agentic language model with a dedicated perception encoder. It preserves the upstream tokenizer, ATEM chat template, image processor configuration, generation configuration, licence, and usage policy.
54
 
55
+
56
  | Item | Value |
57
  | --- | --- |
58
  | Base model | [`meta-models/Muse-Glimmer-30B`](https://huggingface.co/meta-models/Muse-Glimmer-30B) |
 
66
  | Maximum visual tokens | 4,096 per image |
67
  | Architecture | `muse_glimmer` |
68
 
69
+
70
  Modules that MLX-VLM does not classify as quantizable remain at their source-compatible precision, so the effective whole-checkpoint bits-per-weight is higher than four. The `quantization` metadata records the actual eligible-module recipe.
71
 
72
+
73
  > [!IMPORTANT]
74
  > Muse Glimmer support was validated against the official MLX-VLM source at commit [`5262cb6`](https://github.com/Blaizzy/mlx-vlm/commit/5262cb6a27c797c5c3daf64b17d757924c5c474f), reporting package version 0.6.12. An older MLX-VLM or oMLX bundle may report `Model type muse_glimmer not supported`; update to a build containing the upstream Muse implementation before loading this checkpoint.
75
 
76
+
77
  ## Apple-silicon performance
78
 
79
+
80
  This checkpoint was load-tested, text-generation tested, and benchmarked on:
81
 
82
+
83
  | Hardware | Configuration |
84
  | --- | --- |
85
  | Host | Mac Studio |
 
88
  | Unified memory | 256 GB |
89
  | Runtime | MLX-VLM 0.6.12 source revision `5262cb6` with the oMLX MLX runtime |
90
 
91
+
92
  A warmed local text-only test produced:
93
 
94
+
95
  | Measurement | Result |
96
  | --- | ---: |
97
  | Decode (median) | **37.70 tokens/s** |
 
101
  | Warm-up | 256 generated tokens |
102
  | Prompt | 83 tokens after chat templating |
103
 
104
+
105
  The decode figure is the median of three greedy 256-token runs after a full 256-token Metal-kernel warm-up. It is a practical local reference, not a controlled cross-platform benchmark. Prompt length, images, context growth, sampler settings, memory pressure, thermal state, and runtime versions can materially change performance. Image prefill is not included in this decode benchmark.
106
 
107
+
108
  ## Runtime setup
109
 
110
+
111
  Until Muse Glimmer support reaches the MLX-VLM build supplied by your application, install the exact official source revision used for validation:
112
 
113
+
114
  ```bash
115
  python -m pip install \
116
  "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm.git@5262cb6a27c797c5c3daf64b17d757924c5c474f"
117
  ```
118
 
119
+
120
  Download the checkpoint if a local copy is preferred:
121
 
122
+
123
  ```bash
124
  hf download Vontra/Muse-Glimmer-30B-MLX-4bit \
125
  --local-dir ~/.omlx/models/Vontra/Muse-Glimmer-30B-MLX-4bit
126
  ```
127
 
128
+
129
  ## Text generation
130
 
131
+
132
  ```python
133
  from mlx_vlm import generate, load
134
 
135
+
136
  model, processor = load("Vontra/Muse-Glimmer-30B-MLX-4bit")
137
  tokenizer = processor.tokenizer if hasattr(processor, "tokenizer") else processor
138
  messages = [
 
154
  print(result.text)
155
  ```
156
 
157
+
158
  Muse Glimmer supports `low`, `medium`, `high`, and `xhigh` reasoning strengths. Higher settings can spend more tokens reasoning before returning the final answer.
159
 
160
+
161
  ## Image and text generation
162
 
163
+
164
  The multimodal path was smoke-tested locally with a real image. Include an image content item so the ATEM template emits the required patch token:
165
 
166
+
167
  ```python
168
  from mlx_vlm import generate, load
169
 
170
+
171
  model, processor = load("Vontra/Muse-Glimmer-30B-MLX-4bit")
172
  tokenizer = processor.tokenizer if hasattr(processor, "tokenizer") else processor
173
  messages = [
 
195
  print(result.text)
196
  ```
197
 
198
+
199
  ## Architecture
200
 
201
+
202
  Muse Glimmer combines a dense causal transformer with a dedicated perception encoder for interleaved text and image input.
203
 
204
+
205
  | Architecture detail | Upstream value |
206
  | --- | ---: |
207
  | Total parameters | ~29.6B |
 
216
  | Perception encoder | ~1.8B-parameter ViT-G/14, 50 layers |
217
  | Context length | 131,072 tokens |
218
 
219
+
220
  The model supports agentic task completion, tool use through the upstream ATEM protocol, controllable reasoning effort, multilingual input, failure recovery, and multimodal understanding. See the [original model card](https://huggingface.co/meta-models/Muse-Glimmer-30B) for upstream benchmarks, training details, intended uses, limitations, and safety guidance.
221
 
222
+
223
  ## Conversion and validation notes
224
 
225
+
226
  - Source weights: upstream BF16 checkpoint at revision `a4e59da52a7bc87ae7251dd5545c0dd437c44b68`.
227
  - Quantization mode: affine, 4-bit, group size 64, without mixed-precision overrides.
228
  - The upstream tokenizer, ATEM chat template, processor configuration, generation configuration, licence, and usage policy are included.
 
231
  - Quantization can reduce output quality relative to BF16. Use a higher-precision variant when quality matters more than memory use.
232
  - This release does not include or claim support for the upstream speculative drafter.
233
 
234
+
235
  This is a community conversion, not an official Meta release. Validate quality, safety, and numerical behaviour on representative workloads before production use.
236
 
237
+
238
  ## Licence, usage policy, and attribution
239
 
240
+
241
  The upstream model is released under the **Apache License 2.0**. The upstream `LICENSE` and `USAGE_POLICY.md` files are included in this repository; use is subject to both the licence and the upstream usage policy.
242
 
243
+
244
  All model design, training, benchmark, and upstream documentation credit belongs to Meta and the original contributors. The MLX conversion, Apple-silicon validation, compatibility work, and model card are provided by [Vontra](https://huggingface.co/Vontra).
245
+
246
+
247
+
248
+ <!-- vontra-chooser-start -->
249
+ ## Choose for your Mac
250
+
251
+ [64GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-64gb-macs-6a9fefda17932216ec9ab457) · [128GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-128gb-macs-6a9ff0abd31bc9abbe7922d7) · [256GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-256gb-macs-6a9ff0ef9fed7c5bdca15e9b)
252
+
253
+ Published peak memory: **21.33 GB**; estimated starting tier: **64GB**, leaving about **42 GB** nominal headroom. The collections use published M3 Studio peaks with at least 25% nominal headroom; fit on other Macs is an estimate, and full context is not guaranteed. Start with short context and one request.
254
+
255
+ ### Runtime and evidence
256
+
257
+ The exact tested oMLX application version is not recorded here; a library version is not an app version. The original performance tables retain their benchmark conditions and speed figures; this documentation update adds no new test results.
258
+
259
+ ### Quick start and demo prompt
260
+
261
+ ```bash
262
+ hf download Vontra/Muse-Glimmer-30B-MLX-4bit --local-dir ./models/Muse-Glimmer-30B-MLX-4bit
263
+ ```
264
+
265
+ Add the downloaded folder to oMLX model directories, refresh the list, and follow this card's architecture and MTP compatibility requirements before loading.
266
+
267
+ Try this in a new chat with a 128-token output limit:
268
+
269
+ ```text
270
+ Explain why the sky looks blue in three short sentences.
271
+ ```
272
+
273
+ This is a demo prompt to try, not a recorded successful run; a captured demonstration for this documentation update is not yet available.
274
+
275
+ [Follow Vontra for new Apple Silicon releases and fixes.](https://huggingface.co/Vontra)
276
+ <!-- vontra-chooser-end -->