AutomatosX commited on
Commit
886785f
·
verified ·
1 Parent(s): d172f31

Publish AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP

Browse files
README.md CHANGED
@@ -24,13 +24,9 @@ tags:
24
  An **AXQuant (AXQ)** mixed-precision MLX checkpoint for Apple Silicon, converted directly from
25
  the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved at BF16 in the checkpoint (or a bound sidecar when present).
26
 
27
- > **Checkpoint Tier 1 certified** on `df-macstudio-m2` (2026-08-15) at Hub commit
28
- > `594de6507dc9` — measured size against a matched uniform baseline, quality retention,
29
- > and conversion integrity. Current `main` preserves that revision's exact Safetensors
30
- > payloads while allowing metadata-only compatibility fixes. Tier 1 is a checkpoint claim,
31
- > **not** a speed claim: MTP
32
- > acceleration is **not certified**; no MTP speedup claim for this checkpoint.
33
- > See the [checkpoint Tier 1 certificate](https://github.com/defai-digital/axquant/blob/main/docs/certifications/qwen38-27b-axq-mxfp4-mtp-tier1.md) for the bound evidence and thresholds.
34
 
35
 
36
  ## Model details
@@ -42,17 +38,17 @@ the BF16 source model. The language path is quantized while the multi-token-pred
42
  | Product family | `qwen3.8` |
43
  | Source architecture | `Qwen3_5ForConditionalGeneration` (dense); text path optimized |
44
  | Main-model parameters | 27.36B logical parameters |
45
- | Quantizer | AXQuant `1.8.1` |
46
  | Hub budget class | `MXFP4` |
47
  | AXQuant base precision class | `6bit` |
48
  | Planned storage-adjusted BPW | 5.6720 |
49
  | Measured main-model BPW | 4.8441 |
50
  | Measured total BPW, including MTP | **5.0147** |
51
  | Safetensors weight size | 17.41 GB |
52
- | Approximate complete download | 17.44 GB |
53
  | Configured maximum context | 262,144 tokens; practical limits depend on unified memory |
54
  | Primary MLX runtime | MLX-LM |
55
- | AX Engine native execution | Native manifest included; execution still requires a runtime check |
56
  | MTP present | `True` |
57
  | Vision present | `True` |
58
  | Audio present | `False` |
@@ -85,7 +81,7 @@ python -m pip install -U huggingface_hub
85
  hf download AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP --local-dir ./AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP
86
  ```
87
 
88
- Allow at least 17.44 GB of free disk space. Pin the resulting Hub commit in reproducible
89
  deployments rather than relying indefinitely on `main`.
90
 
91
  ## Run with MLX-LM
@@ -102,29 +98,25 @@ mlx_lm.generate \
102
  MLX-LM compatibility covers standard **text/backbone inference**. It may ignore AXQuant runtime
103
  metadata and optional sidecars (`vision.safetensors`, `mtp.safetensors`); this command therefore
104
  does not establish MTP acceleration or vision-language quality. The artifact records MLX
105
- `0.32.0` and MLX-LM `0.31.3` from conversion.
106
 
107
- ## Serve with AX Engine and MTP
108
 
109
- After installing AX Engine, download the complete repository (see
110
- [AXQuant](https://github.com/defai-digital/axquant) for conversion, certificates, and
111
- model-card tooling) and serve the local directory:
112
-
113
- ```bash
114
- ax-engine serve ./AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP --port 31418
115
- ```
116
-
117
- AX Engine is the authority for the AXQ runtime contract and native MTP sidecar.
118
- This development package does not claim runtime speedups until identical-checkpoint benchmarks are
119
- published. The artifact records AX Engine version `not recorded`. Native
120
- `model-manifest.json` status: included as `model-manifest.json`.
121
 
122
  ## Use the packaged Qwen MTP head with oMLX or MTPLX
123
 
124
- Download the complete repository to a writable local directory. In oMLX `0.6.3rc2` or newer, add
125
- that directory, open **Model Settings**, choose **Import MTP side-car**, and then enable
126
- **Lightning MTP**. The import changes only the local copy so the sidecar tensors become visible
127
- through the checkpoint index.
 
 
 
128
 
129
  MTPLX can consume the packaged sidecar directly:
130
 
@@ -137,7 +129,8 @@ mtplx quickstart \
137
  ```
138
 
139
  `mtplx_runtime.json` declares the canonical `qwen3-next-mtp` execution contract. This enables
140
- strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.
 
141
 
142
  ## Quantization layout
143
 
@@ -147,7 +140,7 @@ strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed
147
  | `8bit` | 2.54B | 9.15% |
148
  | `bf16` | 888.07M | 3.20% |
149
 
150
- - Quantization methods: `affine, bf16`.
151
  - Group sizes used by quantized assignments: `32, 64`.
152
  - MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
153
  - Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
@@ -165,23 +158,14 @@ establish MTP acceleration or vision-language quality.
165
  | Planning evidence | `architecture_prior` |
166
  | Calibration | none; the allocation is based on architecture priors |
167
  | Quantizer execution | 498/498 recorded module conversions succeeded; 0 fallbacks |
168
- | AX Engine native manifest | included as `model-manifest.json` |
169
  | Quality versus BF16 or uniform baselines | Not published; no quality-retention claim |
170
- | MTP acceptance and speed | **not certified** on `df-macstudio-m2` / AX Engine 6.16.1 ([Tier 2 record](https://github.com/defai-digital/axquant/blob/main/docs/certifications/qwen38-27b-axq-mxfp4-mtp-tier2.md)); greedy exactness failed and measured speedups were 0.0. Product default remains direct fallback. |
171
  | AX Engine kernel evidence | `unmeasured` |
172
- | Vision-language quality | Present, not certified; text Tier 1 does not imply VLM quality |
173
- | Speech-recognition quality | Not applicable (audio disabled for this pack) |
174
  | Long-context quality | 262,144-token capacity is config metadata, not a validated claim |
175
- | Release certification | **Checkpoint Tier 1 certified** on `df-macstudio-m2` (2026-08-15), Hub commit `594de6507dc9`; the formal AXQuant M0-M8 release campaign is a separate process and is not implied |
176
-
177
- ## Modalities (capability-gated)
178
-
179
- Text checkpoint Tier 1 does **not** imply vision or audio quality. `Vision present=true` on a pack is not a quality pass.
180
-
181
- | Modality | Claim | Supported | Reason |
182
- | --- | --- | --- | --- |
183
- | Vision | `present-not-certified` | `true` | vision sidecar present; mlx-vlm smoke not a quality pass (prefixes=['model.visual']) |
184
- | Audio | `not-applicable` | `false` | audio not supported on this pack |
185
 
186
  ## Intended use and limitations
187
 
@@ -190,10 +174,10 @@ Text checkpoint Tier 1 does **not** imply vision or audio quality. `Vision prese
190
  KV-cache policy, runtime buffers, and other processes using unified memory.
191
  - Architecture-prior allocation is not measured sensitivity. It must not be presented as measured
192
  model quality.
193
- - MTP requires a sidecar-aware runtime. oMLX/MTPLX discovery compatibility does not establish
194
- exactness or speed certification for those runtimes.
195
  - Vision weights are preserved at BF16, but this release does not claim validated VLM quality.
196
  - The configured context window can require substantially more memory as the KV cache grows.
 
197
 
198
  - Upstream capabilities, limitations, biases, and responsible-use guidance still apply.
199
 
@@ -207,7 +191,6 @@ Text checkpoint Tier 1 does **not** imply vision or audio quality. `Vision prese
207
  - [`axquant_runtime.json`](axquant_runtime.json): declared AX Engine and MLX compatibility metadata; runtime checks remain separate evidence.
208
  - [`axquant_mtp_sidecar_manifest.json`](axquant_mtp_sidecar_manifest.json): MTP tensor provenance.
209
  - [`axquant_vision_sidecar_manifest.json`](axquant_vision_sidecar_manifest.json): protected vision tensor provenance.
210
- - [`model-manifest.json`](model-manifest.json): AX Engine native tensor manifest.
211
 
212
  All published provenance uses repository-relative paths. Local source paths are stripped before
213
  publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ
 
24
  An **AXQuant (AXQ)** mixed-precision MLX checkpoint for Apple Silicon, converted directly from
25
  the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved at BF16 in the checkpoint (or a bound sidecar when present).
26
 
27
+ > **Development evidence — not a certified AXQuant release.** This package has conversion and
28
+ > artifact-integrity records, but it does not publish measured quality, long-context, kernel-speed,
29
+ > or MTP-speed evidence. Do not interpret the AXQ product label as a benchmark claim.
 
 
 
 
30
 
31
 
32
  ## Model details
 
38
  | Product family | `qwen3.8` |
39
  | Source architecture | `Qwen3_5ForConditionalGeneration` (dense); text path optimized |
40
  | Main-model parameters | 27.36B logical parameters |
41
+ | Quantizer | AXQuant `1.9.0` |
42
  | Hub budget class | `MXFP4` |
43
  | AXQuant base precision class | `6bit` |
44
  | Planned storage-adjusted BPW | 5.6720 |
45
  | Measured main-model BPW | 4.8441 |
46
  | Measured total BPW, including MTP | **5.0147** |
47
  | Safetensors weight size | 17.41 GB |
48
+ | Approximate complete download | 17.45 GB |
49
  | Configured maximum context | 262,144 tokens; practical limits depend on unified memory |
50
  | Primary MLX runtime | MLX-LM |
51
+ | AX Engine native execution | Not established; no validated native manifest is included |
52
  | MTP present | `True` |
53
  | Vision present | `True` |
54
  | Audio present | `False` |
 
81
  hf download AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP --local-dir ./AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP
82
  ```
83
 
84
+ Allow at least 17.45 GB of free disk space. Pin the resulting Hub commit in reproducible
85
  deployments rather than relying indefinitely on `main`.
86
 
87
  ## Run with MLX-LM
 
98
  MLX-LM compatibility covers standard **text/backbone inference**. It may ignore AXQuant runtime
99
  metadata and optional sidecars (`vision.safetensors`, `mtp.safetensors`); this command therefore
100
  does not establish MTP acceleration or vision-language quality. The artifact records MLX
101
+ `0.32.1` and MLX-LM `0.31.3` from conversion.
102
 
103
+ ## AX Engine status
104
 
105
+ This package does **not** include a validated native `model-manifest.json`, so AX Engine execution
106
+ is not established by this release. The AX Engine fields in `axquant_runtime.json` describe the
107
+ intended compatibility contract, not observed runtime evidence. Use the architecture-specific MLX
108
+ runtime path above. The artifact records AX Engine version
109
+ `7.5.7`, but version discovery alone is not a runtime check.
 
 
 
 
 
 
 
110
 
111
  ## Use the packaged Qwen MTP head with oMLX or MTPLX
112
 
113
+ This repository keeps the Qwen MTP head in `mtp.safetensors`; stock MLX-LM does not load that
114
+ sidecar by itself. Download the complete repository to a writable local directory before using a
115
+ sidecar-aware runtime.
116
+
117
+ In oMLX `0.6.3rc2` or newer, add the local directory, open **Model Settings**, choose
118
+ **Import MTP side-car**, and then enable **Lightning MTP**. The import changes only the local copy
119
+ so the MTP tensors become visible through the checkpoint index.
120
 
121
  MTPLX can consume the packaged sidecar directly:
122
 
 
129
  ```
130
 
131
  `mtplx_runtime.json` declares the canonical `qwen3-next-mtp` execution contract. This enables
132
+ strict runtime discovery; it does not extend AXQuant quality, exactness, or speed certification to
133
+ oMLX or MTPLX.
134
 
135
  ## Quantization layout
136
 
 
140
  | `8bit` | 2.54B | 9.15% |
141
  | `bf16` | 888.07M | 3.20% |
142
 
143
+ - Quantization methods: `affine, bf16, mxfp4`.
144
  - Group sizes used by quantized assignments: `32, 64`.
145
  - MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
146
  - Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
 
158
  | Planning evidence | `architecture_prior` |
159
  | Calibration | none; the allocation is based on architecture priors |
160
  | Quantizer execution | 498/498 recorded module conversions succeeded; 0 fallbacks |
161
+ | AX Engine native manifest | not included |
162
  | Quality versus BF16 or uniform baselines | Not published; no quality-retention claim |
163
+ | MTP acceptance and speed | not measured; no MTP speedup claim |
164
  | AX Engine kernel evidence | `unmeasured` |
165
+ | Vision-language quality | Not evaluated or claimed; vision tensors are preserved at BF16 |
166
+ | Speech-recognition quality | Not applicable |
167
  | Long-context quality | 262,144-token capacity is config metadata, not a validated claim |
168
+ | Release certification | **Not certified**; formal AXQuant M0-M8 gates are not closed |
 
 
 
 
 
 
 
 
 
169
 
170
  ## Intended use and limitations
171
 
 
174
  KV-cache policy, runtime buffers, and other processes using unified memory.
175
  - Architecture-prior allocation is not measured sensitivity. It must not be presented as measured
176
  model quality.
177
+ - MTP requires a sidecar-aware runtime. oMLX/MTPLX discovery compatibility does not establish exactness or speed certification for those runtimes.
 
178
  - Vision weights are preserved at BF16, but this release does not claim validated VLM quality.
179
  - The configured context window can require substantially more memory as the KV cache grows.
180
+ - AX Engine execution is not established because this package has no validated native manifest.
181
 
182
  - Upstream capabilities, limitations, biases, and responsible-use guidance still apply.
183
 
 
191
  - [`axquant_runtime.json`](axquant_runtime.json): declared AX Engine and MLX compatibility metadata; runtime checks remain separate evidence.
192
  - [`axquant_mtp_sidecar_manifest.json`](axquant_mtp_sidecar_manifest.json): MTP tensor provenance.
193
  - [`axquant_vision_sidecar_manifest.json`](axquant_vision_sidecar_manifest.json): protected vision tensor provenance.
 
194
 
195
  All published provenance uses repository-relative paths. Local source paths are stripped before
196
  publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ
axquant_manifest.json CHANGED
@@ -1,37 +1,47 @@
1
  {
2
- "axquant_version": "1.8.1",
3
  "calibration": null,
4
- "created_at": "2026-08-15T16:16:21.884670Z",
5
  "effective_bpw": 5.6719817910099914,
6
  "files": [
7
  {
8
  "path": "README.md",
9
- "sha256": "1989a02ef81ae864b2efc2cf92536efdeb54f98bc5fc32217b21d1c05a41a1f0",
10
- "size_bytes": 10490
11
  },
12
  {
13
  "path": "axquant_mtp_sidecar_manifest.json",
14
- "sha256": "11cc9ea99289a7021240024358df53674102c828a8b930930af0a83074e93b87",
15
  "size_bytes": 880
16
  },
 
 
 
 
 
17
  {
18
  "path": "axquant_plan.json",
19
- "sha256": "fce5a93299e4ccc15b7d5c4d861572f3ded153b2d76d8da2e8da5919b7435b68",
20
- "size_bytes": 1091360
21
  },
22
  {
23
  "path": "axquant_quantizer_execution.json",
24
- "sha256": "c389d287323a820bbbd56cc49d29e0add0011db9c5cc75bc012aa5a43c6eb39b",
25
- "size_bytes": 130100
26
  },
27
  {
28
  "path": "axquant_runtime.json",
29
- "sha256": "35b5323bb62130be059d863dd684c17fab25157657e4788ba0a6a103192390c8",
30
- "size_bytes": 1676
 
 
 
 
 
31
  },
32
  {
33
  "path": "axquant_vision_sidecar_manifest.json",
34
- "sha256": "aa54f3fc80257ea5433e7e161e6a66d4cb25335f4fb2a00f69c2b013e2b0ef28",
35
  "size_bytes": 887
36
  },
37
  {
@@ -49,6 +59,11 @@
49
  "sha256": "e70c136c1b78ddc1fb0905bac8e733a4dc448d4f852a5dd75143fffc70be550e",
50
  "size_bytes": 202
51
  },
 
 
 
 
 
52
  {
53
  "path": "model-00001-of-00003.safetensors",
54
  "sha256": "cf4901a75d3c819b69df7002d8134d81982c94d20ba600522add8a2df858033b",
@@ -64,11 +79,6 @@
64
  "sha256": "5cd45370f41b0d6bf5b8f10def7b0be1cbaf76cdc392320c5778726763a639f5",
65
  "size_bytes": 4927126499
66
  },
67
- {
68
- "path": "model-manifest.json",
69
- "sha256": "524163237ca2b879bff41f782a5eee6a1266b4f80f5d2c99e07117b7b899127a",
70
- "size_bytes": 433842
71
- },
72
  {
73
  "path": "model.safetensors.index.json",
74
  "sha256": "e20dbf6b96058c7b6ccfc1d79745c3f5f4cf8ea4d287ed408da95bdaa4833911",
@@ -84,6 +94,11 @@
84
  "sha256": "840c099003ecd5e25bcf8cbad13c83ff2335bfb9aebf5d743bfd82b4c0c058ae",
85
  "size_bytes": 919
86
  },
 
 
 
 
 
87
  {
88
  "path": "tokenizer.json",
89
  "sha256": "06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523",
@@ -98,6 +113,11 @@
98
  "path": "vision.safetensors",
99
  "sha256": "d0d927c489250588557d5761dcd47ca267ef64c95f7e7d58dd6a0cccf54ac68f",
100
  "size_bytes": 921497320
 
 
 
 
 
101
  }
102
  ],
103
  "format": "mlx",
@@ -128,7 +148,7 @@
128
  },
129
  "mtp_present": true,
130
  "mtp_weight_file_size_bytes": 849400520,
131
- "plan_sha256": "cbfbf9fbc718c89dfbd5241f97cd784bdb62bd8e570c137a0c3fc2c764d8bcb1",
132
  "profile": "agent-coding",
133
  "protected_weight_file_size_bytes": 921497320,
134
  "quantizer": "axquant",
@@ -155,9 +175,10 @@
155
  "support_level": "standard-inference"
156
  }
157
  ],
158
- "created_at": "2026-08-15T16:16:13.776072Z",
159
  "kv_cache": null,
160
  "memory_policy": {
 
161
  "kv_cache_precision": "runtime-default",
162
  "mtp_buffers": "preallocate-when-enabled",
163
  "prefix_cache": "runtime-managed",
@@ -192,12 +213,12 @@
192
  },
193
  "schema_version": "axquant.artifact.v2",
194
  "software_versions": {
195
- "ax_engine": null,
196
- "axquant": "1.8.1",
197
- "mlx": "0.32.0",
198
  "mlx_lm": "0.31.3",
199
  "pydantic": "2.13.4",
200
- "python": "3.13.14",
201
  "safetensors": "0.8.0"
202
  },
203
  "source_model": {
 
1
  {
2
+ "axquant_version": "1.9.0",
3
  "calibration": null,
4
+ "created_at": "2026-10-04T19:41:26.446575Z",
5
  "effective_bpw": 5.6719817910099914,
6
  "files": [
7
  {
8
  "path": "README.md",
9
+ "sha256": "a1c50ddf3102bc0496925805324d98952b15c85b8c88b87664e528b7578a4ea1",
10
+ "size_bytes": 9264
11
  },
12
  {
13
  "path": "axquant_mtp_sidecar_manifest.json",
14
+ "sha256": "65b5b43ba1951fbe2be29db9113643ed04a968aef4f36f93b3b3f7819e400504",
15
  "size_bytes": 880
16
  },
17
+ {
18
+ "path": "axquant_omlx_compat.json",
19
+ "sha256": "211a3b00bde87ce684d2817c6250738488fa308a9af21e199f9e00b5360d70a5",
20
+ "size_bytes": 1034
21
+ },
22
  {
23
  "path": "axquant_plan.json",
24
+ "sha256": "04cefba998b61c0938e571b4dc143354060aca9196245acced6cd212b60ba9d0",
25
+ "size_bytes": 1090819
26
  },
27
  {
28
  "path": "axquant_quantizer_execution.json",
29
+ "sha256": "f65d9c8333b146d1dee416450d959e9f403a0ca48262be7917b39030ac18a44f",
30
+ "size_bytes": 129604
31
  },
32
  {
33
  "path": "axquant_runtime.json",
34
+ "sha256": "f0cd02cce04d192eced3dbd34169edb41b41b1b757fd2d7f6539a0d647a5d1aa",
35
+ "size_bytes": 1704
36
+ },
37
+ {
38
+ "path": "axquant_source_binding.json",
39
+ "sha256": "b25d5ef759649db215772b8535b5870c1130fe3cd39f619f548415d2337ac76c",
40
+ "size_bytes": 2293
41
  },
42
  {
43
  "path": "axquant_vision_sidecar_manifest.json",
44
+ "sha256": "634c27a2e02b48dd6e3634c5972f17a78111e475c7e5a4a8a356a44f1b19092d",
45
  "size_bytes": 887
46
  },
47
  {
 
59
  "sha256": "e70c136c1b78ddc1fb0905bac8e733a4dc448d4f852a5dd75143fffc70be550e",
60
  "size_bytes": 202
61
  },
62
+ {
63
+ "path": "merges.txt",
64
+ "sha256": "a9d356d7bdf1ef4949e3e748e95b8e10ad9d4e2e838eddc38a0a7b6b94d1db8d",
65
+ "size_bytes": 3353259
66
+ },
67
  {
68
  "path": "model-00001-of-00003.safetensors",
69
  "sha256": "cf4901a75d3c819b69df7002d8134d81982c94d20ba600522add8a2df858033b",
 
79
  "sha256": "5cd45370f41b0d6bf5b8f10def7b0be1cbaf76cdc392320c5778726763a639f5",
80
  "size_bytes": 4927126499
81
  },
 
 
 
 
 
82
  {
83
  "path": "model.safetensors.index.json",
84
  "sha256": "e20dbf6b96058c7b6ccfc1d79745c3f5f4cf8ea4d287ed408da95bdaa4833911",
 
94
  "sha256": "840c099003ecd5e25bcf8cbad13c83ff2335bfb9aebf5d743bfd82b4c0c058ae",
95
  "size_bytes": 919
96
  },
97
+ {
98
+ "path": "preprocessor_config.json",
99
+ "sha256": "27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516",
100
+ "size_bytes": 390
101
+ },
102
  {
103
  "path": "tokenizer.json",
104
  "sha256": "06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523",
 
113
  "path": "vision.safetensors",
114
  "sha256": "d0d927c489250588557d5761dcd47ca267ef64c95f7e7d58dd6a0cccf54ac68f",
115
  "size_bytes": 921497320
116
+ },
117
+ {
118
+ "path": "vocab.json",
119
+ "sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003",
120
+ "size_bytes": 6722759
121
  }
122
  ],
123
  "format": "mlx",
 
148
  },
149
  "mtp_present": true,
150
  "mtp_weight_file_size_bytes": 849400520,
151
+ "plan_sha256": "c80f0982facd6110ddd9355e025599081b1cfca641eda241fc09a5e1157f9b81",
152
  "profile": "agent-coding",
153
  "protected_weight_file_size_bytes": 921497320,
154
  "quantizer": "axquant",
 
175
  "support_level": "standard-inference"
176
  }
177
  ],
178
+ "created_at": "2026-10-04T19:41:17.898482Z",
179
  "kv_cache": null,
180
  "memory_policy": {
181
+ "expert_stream": "off",
182
  "kv_cache_precision": "runtime-default",
183
  "mtp_buffers": "preallocate-when-enabled",
184
  "prefix_cache": "runtime-managed",
 
213
  },
214
  "schema_version": "axquant.artifact.v2",
215
  "software_versions": {
216
+ "ax_engine": "7.5.7",
217
+ "axquant": "1.9.0",
218
+ "mlx": "0.32.1",
219
  "mlx_lm": "0.31.3",
220
  "pydantic": "2.13.4",
221
+ "python": "3.12.13",
222
  "safetensors": "0.8.0"
223
  },
224
  "source_model": {
axquant_mtp_sidecar_manifest.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "created_at": "2026-08-15T16:16:13.564577Z",
3
  "dtypes": [
4
  "BF16"
5
  ],
 
1
  {
2
+ "created_at": "2026-10-04T19:41:17.848062Z",
3
  "dtypes": [
4
  "BF16"
5
  ],
axquant_plan.json CHANGED
The diff for this file is too large to render. See raw diff
 
axquant_quantizer_execution.json CHANGED
The diff for this file is too large to render. See raw diff
 
axquant_runtime.json CHANGED
@@ -21,9 +21,10 @@
21
  "support_level": "standard-inference"
22
  }
23
  ],
24
- "created_at": "2026-08-15T16:16:13.776072Z",
25
  "kv_cache": null,
26
  "memory_policy": {
 
27
  "kv_cache_precision": "runtime-default",
28
  "mtp_buffers": "preallocate-when-enabled",
29
  "prefix_cache": "runtime-managed",
 
21
  "support_level": "standard-inference"
22
  }
23
  ],
24
+ "created_at": "2026-10-04T19:41:17.898482Z",
25
  "kv_cache": null,
26
  "memory_policy": {
27
+ "expert_stream": "off",
28
  "kv_cache_precision": "runtime-default",
29
  "mtp_buffers": "preallocate-when-enabled",
30
  "prefix_cache": "runtime-managed",
axquant_source_binding.json ADDED
@@ -0,0 +1,88 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "config_sha256": "191e0af232104ed8b65258cf3fb2b842e288008baca7633c11b82a1ac7203aab",
3
+ "created_at": "2026-10-04T19:25:40.842603Z",
4
+ "index_sha256": "77042094076611b69791a610065f28b7013b8c621795fa86ddccc8bac7d1b9df",
5
+ "members": [
6
+ {
7
+ "path": "model-00001-of-00018.safetensors",
8
+ "size_bytes": 3966730552
9
+ },
10
+ {
11
+ "path": "model-00002-of-00018.safetensors",
12
+ "size_bytes": 3043080328
13
+ },
14
+ {
15
+ "path": "model-00003-of-00018.safetensors",
16
+ "size_bytes": 2542796952
17
+ },
18
+ {
19
+ "path": "model-00004-of-00018.safetensors",
20
+ "size_bytes": 3988973152
21
+ },
22
+ {
23
+ "path": "model-00005-of-00018.safetensors",
24
+ "size_bytes": 2099339864
25
+ },
26
+ {
27
+ "path": "model-00006-of-00018.safetensors",
28
+ "size_bytes": 3979553696
29
+ },
30
+ {
31
+ "path": "model-00007-of-00018.safetensors",
32
+ "size_bytes": 2108759344
33
+ },
34
+ {
35
+ "path": "model-00008-of-00018.safetensors",
36
+ "size_bytes": 3979553696
37
+ },
38
+ {
39
+ "path": "model-00009-of-00018.safetensors",
40
+ "size_bytes": 2108759344
41
+ },
42
+ {
43
+ "path": "model-00010-of-00018.safetensors",
44
+ "size_bytes": 3979553696
45
+ },
46
+ {
47
+ "path": "model-00011-of-00018.safetensors",
48
+ "size_bytes": 2108759344
49
+ },
50
+ {
51
+ "path": "model-00012-of-00018.safetensors",
52
+ "size_bytes": 3979553696
53
+ },
54
+ {
55
+ "path": "model-00013-of-00018.safetensors",
56
+ "size_bytes": 2108759344
57
+ },
58
+ {
59
+ "path": "model-00014-of-00018.safetensors",
60
+ "size_bytes": 3979553696
61
+ },
62
+ {
63
+ "path": "model-00015-of-00018.safetensors",
64
+ "size_bytes": 2108759344
65
+ },
66
+ {
67
+ "path": "model-00016-of-00018.safetensors",
68
+ "size_bytes": 3979564040
69
+ },
70
+ {
71
+ "path": "model-00017-of-00018.safetensors",
72
+ "size_bytes": 2108759344
73
+ },
74
+ {
75
+ "path": "model-00018-of-00018.safetensors",
76
+ "size_bytes": 3392197344
77
+ }
78
+ ],
79
+ "plan_sha256": "c80f0982facd6110ddd9355e025599081b1cfca641eda241fc09a5e1157f9b81",
80
+ "schema_version": "axquant.source-plan-binding.v1",
81
+ "source_model": {
82
+ "architecture": "Qwen3_5ForConditionalGeneration",
83
+ "format": "mlx",
84
+ "local_path": null,
85
+ "model_id": "Qwen/Qwen3.8-27B",
86
+ "revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0"
87
+ }
88
+ }
axquant_vision_sidecar_manifest.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "created_at": "2026-08-15T16:16:11.098120Z",
3
  "dtypes": [
4
  "BF16"
5
  ],
 
1
  {
2
+ "created_at": "2026-10-04T19:41:15.156713Z",
3
  "dtypes": [
4
  "BF16"
5
  ],
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
preprocessor_config.json ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "size": {
3
+ "longest_edge": 16777216,
4
+ "shortest_edge": 65536
5
+ },
6
+ "patch_size": 16,
7
+ "temporal_patch_size": 2,
8
+ "merge_size": 2,
9
+ "image_mean": [
10
+ 0.5,
11
+ 0.5,
12
+ 0.5
13
+ ],
14
+ "image_std": [
15
+ 0.5,
16
+ 0.5,
17
+ 0.5
18
+ ],
19
+ "processor_class": "Qwen3VLProcessor",
20
+ "image_processor_type": "Qwen2VLImageProcessorFast"
21
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff