Text Generation
MLX
Safetensors
qwen3_5
apple-silicon
quantized
mixed-precision
axquant
axq
development
qwen3.8
MXFP4
mtp
vision
conversational
4-bit precision
Instructions to use AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Publish AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP
Browse files- README.md +31 -48
- axquant_manifest.json +44 -23
- axquant_mtp_sidecar_manifest.json +1 -1
- axquant_plan.json +0 -0
- axquant_quantizer_execution.json +0 -0
- axquant_runtime.json +2 -1
- axquant_source_binding.json +88 -0
- axquant_vision_sidecar_manifest.json +1 -1
- merges.txt +0 -0
- preprocessor_config.json +21 -0
- vocab.json +0 -0
README.md
CHANGED
|
@@ -24,13 +24,9 @@ tags:
|
|
| 24 |
An **AXQuant (AXQ)** mixed-precision MLX checkpoint for Apple Silicon, converted directly from
|
| 25 |
the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved at BF16 in the checkpoint (or a bound sidecar when present).
|
| 26 |
|
| 27 |
-
> **
|
| 28 |
-
>
|
| 29 |
-
>
|
| 30 |
-
> payloads while allowing metadata-only compatibility fixes. Tier 1 is a checkpoint claim,
|
| 31 |
-
> **not** a speed claim: MTP
|
| 32 |
-
> acceleration is **not certified**; no MTP speedup claim for this checkpoint.
|
| 33 |
-
> See the [checkpoint Tier 1 certificate](https://github.com/defai-digital/axquant/blob/main/docs/certifications/qwen38-27b-axq-mxfp4-mtp-tier1.md) for the bound evidence and thresholds.
|
| 34 |
|
| 35 |
|
| 36 |
## Model details
|
|
@@ -42,17 +38,17 @@ the BF16 source model. The language path is quantized while the multi-token-pred
|
|
| 42 |
| Product family | `qwen3.8` |
|
| 43 |
| Source architecture | `Qwen3_5ForConditionalGeneration` (dense); text path optimized |
|
| 44 |
| Main-model parameters | 27.36B logical parameters |
|
| 45 |
-
| Quantizer | AXQuant `1.
|
| 46 |
| Hub budget class | `MXFP4` |
|
| 47 |
| AXQuant base precision class | `6bit` |
|
| 48 |
| Planned storage-adjusted BPW | 5.6720 |
|
| 49 |
| Measured main-model BPW | 4.8441 |
|
| 50 |
| Measured total BPW, including MTP | **5.0147** |
|
| 51 |
| Safetensors weight size | 17.41 GB |
|
| 52 |
-
| Approximate complete download | 17.
|
| 53 |
| Configured maximum context | 262,144 tokens; practical limits depend on unified memory |
|
| 54 |
| Primary MLX runtime | MLX-LM |
|
| 55 |
-
| AX Engine native execution |
|
| 56 |
| MTP present | `True` |
|
| 57 |
| Vision present | `True` |
|
| 58 |
| Audio present | `False` |
|
|
@@ -85,7 +81,7 @@ python -m pip install -U huggingface_hub
|
|
| 85 |
hf download AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP --local-dir ./AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP
|
| 86 |
```
|
| 87 |
|
| 88 |
-
Allow at least 17.
|
| 89 |
deployments rather than relying indefinitely on `main`.
|
| 90 |
|
| 91 |
## Run with MLX-LM
|
|
@@ -102,29 +98,25 @@ mlx_lm.generate \
|
|
| 102 |
MLX-LM compatibility covers standard **text/backbone inference**. It may ignore AXQuant runtime
|
| 103 |
metadata and optional sidecars (`vision.safetensors`, `mtp.safetensors`); this command therefore
|
| 104 |
does not establish MTP acceleration or vision-language quality. The artifact records MLX
|
| 105 |
-
`0.32.
|
| 106 |
|
| 107 |
-
##
|
| 108 |
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
``
|
| 114 |
-
ax-engine serve ./AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP --port 31418
|
| 115 |
-
```
|
| 116 |
-
|
| 117 |
-
AX Engine is the authority for the AXQ runtime contract and native MTP sidecar.
|
| 118 |
-
This development package does not claim runtime speedups until identical-checkpoint benchmarks are
|
| 119 |
-
published. The artifact records AX Engine version `not recorded`. Native
|
| 120 |
-
`model-manifest.json` status: included as `model-manifest.json`.
|
| 121 |
|
| 122 |
## Use the packaged Qwen MTP head with oMLX or MTPLX
|
| 123 |
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
|
|
|
|
|
|
|
|
|
| 128 |
|
| 129 |
MTPLX can consume the packaged sidecar directly:
|
| 130 |
|
|
@@ -137,7 +129,8 @@ mtplx quickstart \
|
|
| 137 |
```
|
| 138 |
|
| 139 |
`mtplx_runtime.json` declares the canonical `qwen3-next-mtp` execution contract. This enables
|
| 140 |
-
strict runtime discovery; it does not
|
|
|
|
| 141 |
|
| 142 |
## Quantization layout
|
| 143 |
|
|
@@ -147,7 +140,7 @@ strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed
|
|
| 147 |
| `8bit` | 2.54B | 9.15% |
|
| 148 |
| `bf16` | 888.07M | 3.20% |
|
| 149 |
|
| 150 |
-
- Quantization methods: `affine, bf16`.
|
| 151 |
- Group sizes used by quantized assignments: `32, 64`.
|
| 152 |
- MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
|
| 153 |
- Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
|
|
@@ -165,23 +158,14 @@ establish MTP acceleration or vision-language quality.
|
|
| 165 |
| Planning evidence | `architecture_prior` |
|
| 166 |
| Calibration | none; the allocation is based on architecture priors |
|
| 167 |
| Quantizer execution | 498/498 recorded module conversions succeeded; 0 fallbacks |
|
| 168 |
-
| AX Engine native manifest | included
|
| 169 |
| Quality versus BF16 or uniform baselines | Not published; no quality-retention claim |
|
| 170 |
-
| MTP acceptance and speed |
|
| 171 |
| AX Engine kernel evidence | `unmeasured` |
|
| 172 |
-
| Vision-language quality |
|
| 173 |
-
| Speech-recognition quality | Not applicable
|
| 174 |
| Long-context quality | 262,144-token capacity is config metadata, not a validated claim |
|
| 175 |
-
| Release certification | **
|
| 176 |
-
|
| 177 |
-
## Modalities (capability-gated)
|
| 178 |
-
|
| 179 |
-
Text checkpoint Tier 1 does **not** imply vision or audio quality. `Vision present=true` on a pack is not a quality pass.
|
| 180 |
-
|
| 181 |
-
| Modality | Claim | Supported | Reason |
|
| 182 |
-
| --- | --- | --- | --- |
|
| 183 |
-
| Vision | `present-not-certified` | `true` | vision sidecar present; mlx-vlm smoke not a quality pass (prefixes=['model.visual']) |
|
| 184 |
-
| Audio | `not-applicable` | `false` | audio not supported on this pack |
|
| 185 |
|
| 186 |
## Intended use and limitations
|
| 187 |
|
|
@@ -190,10 +174,10 @@ Text checkpoint Tier 1 does **not** imply vision or audio quality. `Vision prese
|
|
| 190 |
KV-cache policy, runtime buffers, and other processes using unified memory.
|
| 191 |
- Architecture-prior allocation is not measured sensitivity. It must not be presented as measured
|
| 192 |
model quality.
|
| 193 |
-
- MTP requires a sidecar-aware runtime. oMLX/MTPLX discovery compatibility does not establish
|
| 194 |
-
exactness or speed certification for those runtimes.
|
| 195 |
- Vision weights are preserved at BF16, but this release does not claim validated VLM quality.
|
| 196 |
- The configured context window can require substantially more memory as the KV cache grows.
|
|
|
|
| 197 |
|
| 198 |
- Upstream capabilities, limitations, biases, and responsible-use guidance still apply.
|
| 199 |
|
|
@@ -207,7 +191,6 @@ Text checkpoint Tier 1 does **not** imply vision or audio quality. `Vision prese
|
|
| 207 |
- [`axquant_runtime.json`](axquant_runtime.json): declared AX Engine and MLX compatibility metadata; runtime checks remain separate evidence.
|
| 208 |
- [`axquant_mtp_sidecar_manifest.json`](axquant_mtp_sidecar_manifest.json): MTP tensor provenance.
|
| 209 |
- [`axquant_vision_sidecar_manifest.json`](axquant_vision_sidecar_manifest.json): protected vision tensor provenance.
|
| 210 |
-
- [`model-manifest.json`](model-manifest.json): AX Engine native tensor manifest.
|
| 211 |
|
| 212 |
All published provenance uses repository-relative paths. Local source paths are stripped before
|
| 213 |
publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ
|
|
|
|
| 24 |
An **AXQuant (AXQ)** mixed-precision MLX checkpoint for Apple Silicon, converted directly from
|
| 25 |
the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved at BF16 in the checkpoint (or a bound sidecar when present).
|
| 26 |
|
| 27 |
+
> **Development evidence — not a certified AXQuant release.** This package has conversion and
|
| 28 |
+
> artifact-integrity records, but it does not publish measured quality, long-context, kernel-speed,
|
| 29 |
+
> or MTP-speed evidence. Do not interpret the AXQ product label as a benchmark claim.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 30 |
|
| 31 |
|
| 32 |
## Model details
|
|
|
|
| 38 |
| Product family | `qwen3.8` |
|
| 39 |
| Source architecture | `Qwen3_5ForConditionalGeneration` (dense); text path optimized |
|
| 40 |
| Main-model parameters | 27.36B logical parameters |
|
| 41 |
+
| Quantizer | AXQuant `1.9.0` |
|
| 42 |
| Hub budget class | `MXFP4` |
|
| 43 |
| AXQuant base precision class | `6bit` |
|
| 44 |
| Planned storage-adjusted BPW | 5.6720 |
|
| 45 |
| Measured main-model BPW | 4.8441 |
|
| 46 |
| Measured total BPW, including MTP | **5.0147** |
|
| 47 |
| Safetensors weight size | 17.41 GB |
|
| 48 |
+
| Approximate complete download | 17.45 GB |
|
| 49 |
| Configured maximum context | 262,144 tokens; practical limits depend on unified memory |
|
| 50 |
| Primary MLX runtime | MLX-LM |
|
| 51 |
+
| AX Engine native execution | Not established; no validated native manifest is included |
|
| 52 |
| MTP present | `True` |
|
| 53 |
| Vision present | `True` |
|
| 54 |
| Audio present | `False` |
|
|
|
|
| 81 |
hf download AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP --local-dir ./AX-Qwen3.8-27B-MLX-AXQ-MXFP4-MTP
|
| 82 |
```
|
| 83 |
|
| 84 |
+
Allow at least 17.45 GB of free disk space. Pin the resulting Hub commit in reproducible
|
| 85 |
deployments rather than relying indefinitely on `main`.
|
| 86 |
|
| 87 |
## Run with MLX-LM
|
|
|
|
| 98 |
MLX-LM compatibility covers standard **text/backbone inference**. It may ignore AXQuant runtime
|
| 99 |
metadata and optional sidecars (`vision.safetensors`, `mtp.safetensors`); this command therefore
|
| 100 |
does not establish MTP acceleration or vision-language quality. The artifact records MLX
|
| 101 |
+
`0.32.1` and MLX-LM `0.31.3` from conversion.
|
| 102 |
|
| 103 |
+
## AX Engine status
|
| 104 |
|
| 105 |
+
This package does **not** include a validated native `model-manifest.json`, so AX Engine execution
|
| 106 |
+
is not established by this release. The AX Engine fields in `axquant_runtime.json` describe the
|
| 107 |
+
intended compatibility contract, not observed runtime evidence. Use the architecture-specific MLX
|
| 108 |
+
runtime path above. The artifact records AX Engine version
|
| 109 |
+
`7.5.7`, but version discovery alone is not a runtime check.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 110 |
|
| 111 |
## Use the packaged Qwen MTP head with oMLX or MTPLX
|
| 112 |
|
| 113 |
+
This repository keeps the Qwen MTP head in `mtp.safetensors`; stock MLX-LM does not load that
|
| 114 |
+
sidecar by itself. Download the complete repository to a writable local directory before using a
|
| 115 |
+
sidecar-aware runtime.
|
| 116 |
+
|
| 117 |
+
In oMLX `0.6.3rc2` or newer, add the local directory, open **Model Settings**, choose
|
| 118 |
+
**Import MTP side-car**, and then enable **Lightning MTP**. The import changes only the local copy
|
| 119 |
+
so the MTP tensors become visible through the checkpoint index.
|
| 120 |
|
| 121 |
MTPLX can consume the packaged sidecar directly:
|
| 122 |
|
|
|
|
| 129 |
```
|
| 130 |
|
| 131 |
`mtplx_runtime.json` declares the canonical `qwen3-next-mtp` execution contract. This enables
|
| 132 |
+
strict runtime discovery; it does not extend AXQuant quality, exactness, or speed certification to
|
| 133 |
+
oMLX or MTPLX.
|
| 134 |
|
| 135 |
## Quantization layout
|
| 136 |
|
|
|
|
| 140 |
| `8bit` | 2.54B | 9.15% |
|
| 141 |
| `bf16` | 888.07M | 3.20% |
|
| 142 |
|
| 143 |
+
- Quantization methods: `affine, bf16, mxfp4`.
|
| 144 |
- Group sizes used by quantized assignments: `32, 64`.
|
| 145 |
- MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
|
| 146 |
- Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
|
|
|
|
| 158 |
| Planning evidence | `architecture_prior` |
|
| 159 |
| Calibration | none; the allocation is based on architecture priors |
|
| 160 |
| Quantizer execution | 498/498 recorded module conversions succeeded; 0 fallbacks |
|
| 161 |
+
| AX Engine native manifest | not included |
|
| 162 |
| Quality versus BF16 or uniform baselines | Not published; no quality-retention claim |
|
| 163 |
+
| MTP acceptance and speed | not measured; no MTP speedup claim |
|
| 164 |
| AX Engine kernel evidence | `unmeasured` |
|
| 165 |
+
| Vision-language quality | Not evaluated or claimed; vision tensors are preserved at BF16 |
|
| 166 |
+
| Speech-recognition quality | Not applicable |
|
| 167 |
| Long-context quality | 262,144-token capacity is config metadata, not a validated claim |
|
| 168 |
+
| Release certification | **Not certified**; formal AXQuant M0-M8 gates are not closed |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 169 |
|
| 170 |
## Intended use and limitations
|
| 171 |
|
|
|
|
| 174 |
KV-cache policy, runtime buffers, and other processes using unified memory.
|
| 175 |
- Architecture-prior allocation is not measured sensitivity. It must not be presented as measured
|
| 176 |
model quality.
|
| 177 |
+
- MTP requires a sidecar-aware runtime. oMLX/MTPLX discovery compatibility does not establish exactness or speed certification for those runtimes.
|
|
|
|
| 178 |
- Vision weights are preserved at BF16, but this release does not claim validated VLM quality.
|
| 179 |
- The configured context window can require substantially more memory as the KV cache grows.
|
| 180 |
+
- AX Engine execution is not established because this package has no validated native manifest.
|
| 181 |
|
| 182 |
- Upstream capabilities, limitations, biases, and responsible-use guidance still apply.
|
| 183 |
|
|
|
|
| 191 |
- [`axquant_runtime.json`](axquant_runtime.json): declared AX Engine and MLX compatibility metadata; runtime checks remain separate evidence.
|
| 192 |
- [`axquant_mtp_sidecar_manifest.json`](axquant_mtp_sidecar_manifest.json): MTP tensor provenance.
|
| 193 |
- [`axquant_vision_sidecar_manifest.json`](axquant_vision_sidecar_manifest.json): protected vision tensor provenance.
|
|
|
|
| 194 |
|
| 195 |
All published provenance uses repository-relative paths. Local source paths are stripped before
|
| 196 |
publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ
|
axquant_manifest.json
CHANGED
|
@@ -1,37 +1,47 @@
|
|
| 1 |
{
|
| 2 |
-
"axquant_version": "1.
|
| 3 |
"calibration": null,
|
| 4 |
-
"created_at": "2026-
|
| 5 |
"effective_bpw": 5.6719817910099914,
|
| 6 |
"files": [
|
| 7 |
{
|
| 8 |
"path": "README.md",
|
| 9 |
-
"sha256": "
|
| 10 |
-
"size_bytes":
|
| 11 |
},
|
| 12 |
{
|
| 13 |
"path": "axquant_mtp_sidecar_manifest.json",
|
| 14 |
-
"sha256": "
|
| 15 |
"size_bytes": 880
|
| 16 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
{
|
| 18 |
"path": "axquant_plan.json",
|
| 19 |
-
"sha256": "
|
| 20 |
-
"size_bytes":
|
| 21 |
},
|
| 22 |
{
|
| 23 |
"path": "axquant_quantizer_execution.json",
|
| 24 |
-
"sha256": "
|
| 25 |
-
"size_bytes":
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"path": "axquant_runtime.json",
|
| 29 |
-
"sha256": "
|
| 30 |
-
"size_bytes":
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 31 |
},
|
| 32 |
{
|
| 33 |
"path": "axquant_vision_sidecar_manifest.json",
|
| 34 |
-
"sha256": "
|
| 35 |
"size_bytes": 887
|
| 36 |
},
|
| 37 |
{
|
|
@@ -49,6 +59,11 @@
|
|
| 49 |
"sha256": "e70c136c1b78ddc1fb0905bac8e733a4dc448d4f852a5dd75143fffc70be550e",
|
| 50 |
"size_bytes": 202
|
| 51 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
{
|
| 53 |
"path": "model-00001-of-00003.safetensors",
|
| 54 |
"sha256": "cf4901a75d3c819b69df7002d8134d81982c94d20ba600522add8a2df858033b",
|
|
@@ -64,11 +79,6 @@
|
|
| 64 |
"sha256": "5cd45370f41b0d6bf5b8f10def7b0be1cbaf76cdc392320c5778726763a639f5",
|
| 65 |
"size_bytes": 4927126499
|
| 66 |
},
|
| 67 |
-
{
|
| 68 |
-
"path": "model-manifest.json",
|
| 69 |
-
"sha256": "524163237ca2b879bff41f782a5eee6a1266b4f80f5d2c99e07117b7b899127a",
|
| 70 |
-
"size_bytes": 433842
|
| 71 |
-
},
|
| 72 |
{
|
| 73 |
"path": "model.safetensors.index.json",
|
| 74 |
"sha256": "e20dbf6b96058c7b6ccfc1d79745c3f5f4cf8ea4d287ed408da95bdaa4833911",
|
|
@@ -84,6 +94,11 @@
|
|
| 84 |
"sha256": "840c099003ecd5e25bcf8cbad13c83ff2335bfb9aebf5d743bfd82b4c0c058ae",
|
| 85 |
"size_bytes": 919
|
| 86 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 87 |
{
|
| 88 |
"path": "tokenizer.json",
|
| 89 |
"sha256": "06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523",
|
|
@@ -98,6 +113,11 @@
|
|
| 98 |
"path": "vision.safetensors",
|
| 99 |
"sha256": "d0d927c489250588557d5761dcd47ca267ef64c95f7e7d58dd6a0cccf54ac68f",
|
| 100 |
"size_bytes": 921497320
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 101 |
}
|
| 102 |
],
|
| 103 |
"format": "mlx",
|
|
@@ -128,7 +148,7 @@
|
|
| 128 |
},
|
| 129 |
"mtp_present": true,
|
| 130 |
"mtp_weight_file_size_bytes": 849400520,
|
| 131 |
-
"plan_sha256": "
|
| 132 |
"profile": "agent-coding",
|
| 133 |
"protected_weight_file_size_bytes": 921497320,
|
| 134 |
"quantizer": "axquant",
|
|
@@ -155,9 +175,10 @@
|
|
| 155 |
"support_level": "standard-inference"
|
| 156 |
}
|
| 157 |
],
|
| 158 |
-
"created_at": "2026-
|
| 159 |
"kv_cache": null,
|
| 160 |
"memory_policy": {
|
|
|
|
| 161 |
"kv_cache_precision": "runtime-default",
|
| 162 |
"mtp_buffers": "preallocate-when-enabled",
|
| 163 |
"prefix_cache": "runtime-managed",
|
|
@@ -192,12 +213,12 @@
|
|
| 192 |
},
|
| 193 |
"schema_version": "axquant.artifact.v2",
|
| 194 |
"software_versions": {
|
| 195 |
-
"ax_engine":
|
| 196 |
-
"axquant": "1.
|
| 197 |
-
"mlx": "0.32.
|
| 198 |
"mlx_lm": "0.31.3",
|
| 199 |
"pydantic": "2.13.4",
|
| 200 |
-
"python": "3.
|
| 201 |
"safetensors": "0.8.0"
|
| 202 |
},
|
| 203 |
"source_model": {
|
|
|
|
| 1 |
{
|
| 2 |
+
"axquant_version": "1.9.0",
|
| 3 |
"calibration": null,
|
| 4 |
+
"created_at": "2026-10-04T19:41:26.446575Z",
|
| 5 |
"effective_bpw": 5.6719817910099914,
|
| 6 |
"files": [
|
| 7 |
{
|
| 8 |
"path": "README.md",
|
| 9 |
+
"sha256": "a1c50ddf3102bc0496925805324d98952b15c85b8c88b87664e528b7578a4ea1",
|
| 10 |
+
"size_bytes": 9264
|
| 11 |
},
|
| 12 |
{
|
| 13 |
"path": "axquant_mtp_sidecar_manifest.json",
|
| 14 |
+
"sha256": "65b5b43ba1951fbe2be29db9113643ed04a968aef4f36f93b3b3f7819e400504",
|
| 15 |
"size_bytes": 880
|
| 16 |
},
|
| 17 |
+
{
|
| 18 |
+
"path": "axquant_omlx_compat.json",
|
| 19 |
+
"sha256": "211a3b00bde87ce684d2817c6250738488fa308a9af21e199f9e00b5360d70a5",
|
| 20 |
+
"size_bytes": 1034
|
| 21 |
+
},
|
| 22 |
{
|
| 23 |
"path": "axquant_plan.json",
|
| 24 |
+
"sha256": "04cefba998b61c0938e571b4dc143354060aca9196245acced6cd212b60ba9d0",
|
| 25 |
+
"size_bytes": 1090819
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"path": "axquant_quantizer_execution.json",
|
| 29 |
+
"sha256": "f65d9c8333b146d1dee416450d959e9f403a0ca48262be7917b39030ac18a44f",
|
| 30 |
+
"size_bytes": 129604
|
| 31 |
},
|
| 32 |
{
|
| 33 |
"path": "axquant_runtime.json",
|
| 34 |
+
"sha256": "f0cd02cce04d192eced3dbd34169edb41b41b1b757fd2d7f6539a0d647a5d1aa",
|
| 35 |
+
"size_bytes": 1704
|
| 36 |
+
},
|
| 37 |
+
{
|
| 38 |
+
"path": "axquant_source_binding.json",
|
| 39 |
+
"sha256": "b25d5ef759649db215772b8535b5870c1130fe3cd39f619f548415d2337ac76c",
|
| 40 |
+
"size_bytes": 2293
|
| 41 |
},
|
| 42 |
{
|
| 43 |
"path": "axquant_vision_sidecar_manifest.json",
|
| 44 |
+
"sha256": "634c27a2e02b48dd6e3634c5972f17a78111e475c7e5a4a8a356a44f1b19092d",
|
| 45 |
"size_bytes": 887
|
| 46 |
},
|
| 47 |
{
|
|
|
|
| 59 |
"sha256": "e70c136c1b78ddc1fb0905bac8e733a4dc448d4f852a5dd75143fffc70be550e",
|
| 60 |
"size_bytes": 202
|
| 61 |
},
|
| 62 |
+
{
|
| 63 |
+
"path": "merges.txt",
|
| 64 |
+
"sha256": "a9d356d7bdf1ef4949e3e748e95b8e10ad9d4e2e838eddc38a0a7b6b94d1db8d",
|
| 65 |
+
"size_bytes": 3353259
|
| 66 |
+
},
|
| 67 |
{
|
| 68 |
"path": "model-00001-of-00003.safetensors",
|
| 69 |
"sha256": "cf4901a75d3c819b69df7002d8134d81982c94d20ba600522add8a2df858033b",
|
|
|
|
| 79 |
"sha256": "5cd45370f41b0d6bf5b8f10def7b0be1cbaf76cdc392320c5778726763a639f5",
|
| 80 |
"size_bytes": 4927126499
|
| 81 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
{
|
| 83 |
"path": "model.safetensors.index.json",
|
| 84 |
"sha256": "e20dbf6b96058c7b6ccfc1d79745c3f5f4cf8ea4d287ed408da95bdaa4833911",
|
|
|
|
| 94 |
"sha256": "840c099003ecd5e25bcf8cbad13c83ff2335bfb9aebf5d743bfd82b4c0c058ae",
|
| 95 |
"size_bytes": 919
|
| 96 |
},
|
| 97 |
+
{
|
| 98 |
+
"path": "preprocessor_config.json",
|
| 99 |
+
"sha256": "27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516",
|
| 100 |
+
"size_bytes": 390
|
| 101 |
+
},
|
| 102 |
{
|
| 103 |
"path": "tokenizer.json",
|
| 104 |
"sha256": "06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523",
|
|
|
|
| 113 |
"path": "vision.safetensors",
|
| 114 |
"sha256": "d0d927c489250588557d5761dcd47ca267ef64c95f7e7d58dd6a0cccf54ac68f",
|
| 115 |
"size_bytes": 921497320
|
| 116 |
+
},
|
| 117 |
+
{
|
| 118 |
+
"path": "vocab.json",
|
| 119 |
+
"sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003",
|
| 120 |
+
"size_bytes": 6722759
|
| 121 |
}
|
| 122 |
],
|
| 123 |
"format": "mlx",
|
|
|
|
| 148 |
},
|
| 149 |
"mtp_present": true,
|
| 150 |
"mtp_weight_file_size_bytes": 849400520,
|
| 151 |
+
"plan_sha256": "c80f0982facd6110ddd9355e025599081b1cfca641eda241fc09a5e1157f9b81",
|
| 152 |
"profile": "agent-coding",
|
| 153 |
"protected_weight_file_size_bytes": 921497320,
|
| 154 |
"quantizer": "axquant",
|
|
|
|
| 175 |
"support_level": "standard-inference"
|
| 176 |
}
|
| 177 |
],
|
| 178 |
+
"created_at": "2026-10-04T19:41:17.898482Z",
|
| 179 |
"kv_cache": null,
|
| 180 |
"memory_policy": {
|
| 181 |
+
"expert_stream": "off",
|
| 182 |
"kv_cache_precision": "runtime-default",
|
| 183 |
"mtp_buffers": "preallocate-when-enabled",
|
| 184 |
"prefix_cache": "runtime-managed",
|
|
|
|
| 213 |
},
|
| 214 |
"schema_version": "axquant.artifact.v2",
|
| 215 |
"software_versions": {
|
| 216 |
+
"ax_engine": "7.5.7",
|
| 217 |
+
"axquant": "1.9.0",
|
| 218 |
+
"mlx": "0.32.1",
|
| 219 |
"mlx_lm": "0.31.3",
|
| 220 |
"pydantic": "2.13.4",
|
| 221 |
+
"python": "3.12.13",
|
| 222 |
"safetensors": "0.8.0"
|
| 223 |
},
|
| 224 |
"source_model": {
|
axquant_mtp_sidecar_manifest.json
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
{
|
| 2 |
-
"created_at": "2026-
|
| 3 |
"dtypes": [
|
| 4 |
"BF16"
|
| 5 |
],
|
|
|
|
| 1 |
{
|
| 2 |
+
"created_at": "2026-10-04T19:41:17.848062Z",
|
| 3 |
"dtypes": [
|
| 4 |
"BF16"
|
| 5 |
],
|
axquant_plan.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
axquant_quantizer_execution.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
axquant_runtime.json
CHANGED
|
@@ -21,9 +21,10 @@
|
|
| 21 |
"support_level": "standard-inference"
|
| 22 |
}
|
| 23 |
],
|
| 24 |
-
"created_at": "2026-
|
| 25 |
"kv_cache": null,
|
| 26 |
"memory_policy": {
|
|
|
|
| 27 |
"kv_cache_precision": "runtime-default",
|
| 28 |
"mtp_buffers": "preallocate-when-enabled",
|
| 29 |
"prefix_cache": "runtime-managed",
|
|
|
|
| 21 |
"support_level": "standard-inference"
|
| 22 |
}
|
| 23 |
],
|
| 24 |
+
"created_at": "2026-10-04T19:41:17.898482Z",
|
| 25 |
"kv_cache": null,
|
| 26 |
"memory_policy": {
|
| 27 |
+
"expert_stream": "off",
|
| 28 |
"kv_cache_precision": "runtime-default",
|
| 29 |
"mtp_buffers": "preallocate-when-enabled",
|
| 30 |
"prefix_cache": "runtime-managed",
|
axquant_source_binding.json
ADDED
|
@@ -0,0 +1,88 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"config_sha256": "191e0af232104ed8b65258cf3fb2b842e288008baca7633c11b82a1ac7203aab",
|
| 3 |
+
"created_at": "2026-10-04T19:25:40.842603Z",
|
| 4 |
+
"index_sha256": "77042094076611b69791a610065f28b7013b8c621795fa86ddccc8bac7d1b9df",
|
| 5 |
+
"members": [
|
| 6 |
+
{
|
| 7 |
+
"path": "model-00001-of-00018.safetensors",
|
| 8 |
+
"size_bytes": 3966730552
|
| 9 |
+
},
|
| 10 |
+
{
|
| 11 |
+
"path": "model-00002-of-00018.safetensors",
|
| 12 |
+
"size_bytes": 3043080328
|
| 13 |
+
},
|
| 14 |
+
{
|
| 15 |
+
"path": "model-00003-of-00018.safetensors",
|
| 16 |
+
"size_bytes": 2542796952
|
| 17 |
+
},
|
| 18 |
+
{
|
| 19 |
+
"path": "model-00004-of-00018.safetensors",
|
| 20 |
+
"size_bytes": 3988973152
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"path": "model-00005-of-00018.safetensors",
|
| 24 |
+
"size_bytes": 2099339864
|
| 25 |
+
},
|
| 26 |
+
{
|
| 27 |
+
"path": "model-00006-of-00018.safetensors",
|
| 28 |
+
"size_bytes": 3979553696
|
| 29 |
+
},
|
| 30 |
+
{
|
| 31 |
+
"path": "model-00007-of-00018.safetensors",
|
| 32 |
+
"size_bytes": 2108759344
|
| 33 |
+
},
|
| 34 |
+
{
|
| 35 |
+
"path": "model-00008-of-00018.safetensors",
|
| 36 |
+
"size_bytes": 3979553696
|
| 37 |
+
},
|
| 38 |
+
{
|
| 39 |
+
"path": "model-00009-of-00018.safetensors",
|
| 40 |
+
"size_bytes": 2108759344
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"path": "model-00010-of-00018.safetensors",
|
| 44 |
+
"size_bytes": 3979553696
|
| 45 |
+
},
|
| 46 |
+
{
|
| 47 |
+
"path": "model-00011-of-00018.safetensors",
|
| 48 |
+
"size_bytes": 2108759344
|
| 49 |
+
},
|
| 50 |
+
{
|
| 51 |
+
"path": "model-00012-of-00018.safetensors",
|
| 52 |
+
"size_bytes": 3979553696
|
| 53 |
+
},
|
| 54 |
+
{
|
| 55 |
+
"path": "model-00013-of-00018.safetensors",
|
| 56 |
+
"size_bytes": 2108759344
|
| 57 |
+
},
|
| 58 |
+
{
|
| 59 |
+
"path": "model-00014-of-00018.safetensors",
|
| 60 |
+
"size_bytes": 3979553696
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"path": "model-00015-of-00018.safetensors",
|
| 64 |
+
"size_bytes": 2108759344
|
| 65 |
+
},
|
| 66 |
+
{
|
| 67 |
+
"path": "model-00016-of-00018.safetensors",
|
| 68 |
+
"size_bytes": 3979564040
|
| 69 |
+
},
|
| 70 |
+
{
|
| 71 |
+
"path": "model-00017-of-00018.safetensors",
|
| 72 |
+
"size_bytes": 2108759344
|
| 73 |
+
},
|
| 74 |
+
{
|
| 75 |
+
"path": "model-00018-of-00018.safetensors",
|
| 76 |
+
"size_bytes": 3392197344
|
| 77 |
+
}
|
| 78 |
+
],
|
| 79 |
+
"plan_sha256": "c80f0982facd6110ddd9355e025599081b1cfca641eda241fc09a5e1157f9b81",
|
| 80 |
+
"schema_version": "axquant.source-plan-binding.v1",
|
| 81 |
+
"source_model": {
|
| 82 |
+
"architecture": "Qwen3_5ForConditionalGeneration",
|
| 83 |
+
"format": "mlx",
|
| 84 |
+
"local_path": null,
|
| 85 |
+
"model_id": "Qwen/Qwen3.8-27B",
|
| 86 |
+
"revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0"
|
| 87 |
+
}
|
| 88 |
+
}
|
axquant_vision_sidecar_manifest.json
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
{
|
| 2 |
-
"created_at": "2026-
|
| 3 |
"dtypes": [
|
| 4 |
"BF16"
|
| 5 |
],
|
|
|
|
| 1 |
{
|
| 2 |
+
"created_at": "2026-10-04T19:41:15.156713Z",
|
| 3 |
"dtypes": [
|
| 4 |
"BF16"
|
| 5 |
],
|
merges.txt
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
preprocessor_config.json
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"size": {
|
| 3 |
+
"longest_edge": 16777216,
|
| 4 |
+
"shortest_edge": 65536
|
| 5 |
+
},
|
| 6 |
+
"patch_size": 16,
|
| 7 |
+
"temporal_patch_size": 2,
|
| 8 |
+
"merge_size": 2,
|
| 9 |
+
"image_mean": [
|
| 10 |
+
0.5,
|
| 11 |
+
0.5,
|
| 12 |
+
0.5
|
| 13 |
+
],
|
| 14 |
+
"image_std": [
|
| 15 |
+
0.5,
|
| 16 |
+
0.5,
|
| 17 |
+
0.5
|
| 18 |
+
],
|
| 19 |
+
"processor_class": "Qwen3VLProcessor",
|
| 20 |
+
"image_processor_type": "Qwen2VLImageProcessorFast"
|
| 21 |
+
}
|
vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|