Instructions to use AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Record architecture-specific runtime format audit and n-gram scope
Browse files- README.md +11 -0
- axquant_manifest.json +8 -3
- runtime_audit.json +63 -0
README.md
CHANGED
|
@@ -12,6 +12,17 @@ tags:
|
|
| 12 |
- mtp
|
| 13 |
---
|
| 14 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
# AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP
|
| 16 |
|
| 17 |
This repository contains an AXQuant development conversion of the pinned NVIDIA BF16 source. The main checkpoint uses standard MLX-LM config, tokenizer, index, and safetensors files. MXFP quantization is applied to eligible backbone weights; protected tensors keep their declared precision.
|
|
|
|
| 12 |
- mtp
|
| 13 |
---
|
| 14 |
|
| 15 |
+
## Runtime format audit (2026-10-06)
|
| 16 |
+
|
| 17 |
+
No quantization-container correction was needed.
|
| 18 |
+
This family has no n-gram tensors; no n-gram file or declaration was added.
|
| 19 |
+
This remote format audit does not grant an oMLX/MTPLX runtime profile or a successful load/generation claim.
|
| 20 |
+
|
| 21 |
+
See [runtime_audit.json](runtime_audit.json) for pinned config/index/header
|
| 22 |
+
bindings, architecture, physical-format findings, and applied corrections.
|
| 23 |
+
This is development evidence; no quality, MTP exactness, speed, or certification
|
| 24 |
+
claim is added. Historical evidence stays bound to its original revision.
|
| 25 |
+
|
| 26 |
# AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP
|
| 27 |
|
| 28 |
This repository contains an AXQuant development conversion of the pinned NVIDIA BF16 source. The main checkpoint uses standard MLX-LM config, tokenizer, index, and safetensors files. MXFP quantization is applied to eligible backbone weights; protected tensors keep their declared precision.
|
axquant_manifest.json
CHANGED
|
@@ -1,13 +1,13 @@
|
|
| 1 |
{
|
| 2 |
"axquant_version": "1.9.0",
|
| 3 |
"calibration": null,
|
| 4 |
-
"created_at": "2026-10-
|
| 5 |
"effective_bpw": 5.609812258365519,
|
| 6 |
"files": [
|
| 7 |
{
|
| 8 |
"path": "README.md",
|
| 9 |
-
"sha256": "
|
| 10 |
-
"size_bytes":
|
| 11 |
},
|
| 12 |
{
|
| 13 |
"path": "ax_expert_stream.json",
|
|
@@ -89,6 +89,11 @@
|
|
| 89 |
"sha256": "40dd606b285acd5045bd4246a916463eeb1d5f4cc7a4ebe051555985ec436e72",
|
| 90 |
"size_bytes": 2670685376
|
| 91 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
{
|
| 93 |
"path": "special_tokens_map.json",
|
| 94 |
"sha256": "e9435fefd6d838fd9fcbbc44b97a8e3ff322be7f6dfb7e4fd2468586574bb52b",
|
|
|
|
| 1 |
{
|
| 2 |
"axquant_version": "1.9.0",
|
| 3 |
"calibration": null,
|
| 4 |
+
"created_at": "2026-10-06T21:01:05.546176+00:00",
|
| 5 |
"effective_bpw": 5.609812258365519,
|
| 6 |
"files": [
|
| 7 |
{
|
| 8 |
"path": "README.md",
|
| 9 |
+
"sha256": "cf665799010644c67d9072ad1b0dbca7825fe230a32540a06979e9590b8c3938",
|
| 10 |
+
"size_bytes": 2492
|
| 11 |
},
|
| 12 |
{
|
| 13 |
"path": "ax_expert_stream.json",
|
|
|
|
| 89 |
"sha256": "40dd606b285acd5045bd4246a916463eeb1d5f4cc7a4ebe051555985ec436e72",
|
| 90 |
"size_bytes": 2670685376
|
| 91 |
},
|
| 92 |
+
{
|
| 93 |
+
"path": "runtime_audit.json",
|
| 94 |
+
"sha256": "dcde3c91fa4c1fa8bd4a3b27e3bae0986a345eff8b4d26f9004c0f7c091f9f22",
|
| 95 |
+
"size_bytes": 2893
|
| 96 |
+
},
|
| 97 |
{
|
| 98 |
"path": "special_tokens_map.json",
|
| 99 |
"sha256": "e9435fefd6d838fd9fcbbc44b97a8e3ff322be7f6dfb7e4fd2468586574bb52b",
|
runtime_audit.json
ADDED
|
@@ -0,0 +1,63 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"applied_config_corrections": [],
|
| 3 |
+
"current_config_sha256": "6a6d3246c4db834f09eae87deff809e157ca26617a23bded100de42a0ff5a1fa",
|
| 4 |
+
"date": "2026-10-06",
|
| 5 |
+
"header_sha256": {
|
| 6 |
+
"model-00001-of-00004.safetensors": "ac490936b49fb441500f2f91d68dfe3c9d098559dad6303d024ba4e76c2c2b76",
|
| 7 |
+
"model-00002-of-00004.safetensors": "b313f2840234aff995f100eef741083b5ae1af9d253661b76b8aabc4e42ca3b2",
|
| 8 |
+
"model-00003-of-00004.safetensors": "eee80c2f961f201a8536e7a680944e9faaeb05d32c398f8778df6fd3e6f51650",
|
| 9 |
+
"model-00004-of-00004.safetensors": "6028d48a118e3dd7bee2750d68937ea805510b73983af01ae61322e0c94f6b69",
|
| 10 |
+
"mtp.safetensors": "dfc474d9015b22f6524ff0cb2a88fbad72dde8493cc169e4e3267ef444a78696"
|
| 11 |
+
},
|
| 12 |
+
"input_bindings": {
|
| 13 |
+
"ax_nemotron_mtp_manifest.json": "9338ce8221bd67be7172acf465172fdf0fd023722bd0034e579e9e97dd359af5",
|
| 14 |
+
"axquant_manifest.json": "84a4060ece33dd6677e600ce32e942908c8ff98f1063303a433f8aa99e21a06d",
|
| 15 |
+
"axquant_mtp_sidecar_manifest.json": "70868bd7b9910518b65bb4de2cd9d2119cb54db0d3b4db4c44a21b1b8233bf08",
|
| 16 |
+
"config.json": "6a6d3246c4db834f09eae87deff809e157ca26617a23bded100de42a0ff5a1fa",
|
| 17 |
+
"model.safetensors.index.json": "3e068938b04a3279ca3e6e1197882741cd5f67d5f7bce238a099f9cbb4da429e"
|
| 18 |
+
},
|
| 19 |
+
"issues": [],
|
| 20 |
+
"model_type": "nemotron_h",
|
| 21 |
+
"mtp_files": [
|
| 22 |
+
"mtp.safetensors"
|
| 23 |
+
],
|
| 24 |
+
"ngram_action": "none; do not invent n-gram data",
|
| 25 |
+
"ngram_files": [],
|
| 26 |
+
"ngram_quantization": [],
|
| 27 |
+
"ngram_table_metadata": null,
|
| 28 |
+
"ngram_tensor_count": 0,
|
| 29 |
+
"quality_certified": false,
|
| 30 |
+
"quantization": {
|
| 31 |
+
"quantization": {
|
| 32 |
+
"container_mode": "affine",
|
| 33 |
+
"per_module_modes": {
|
| 34 |
+
"affine": 1,
|
| 35 |
+
"mxfp4": 162
|
| 36 |
+
},
|
| 37 |
+
"physical_recipe_issues": []
|
| 38 |
+
},
|
| 39 |
+
"quantization_config": {
|
| 40 |
+
"container_mode": "affine",
|
| 41 |
+
"per_module_modes": {
|
| 42 |
+
"affine": 1,
|
| 43 |
+
"mxfp4": 162
|
| 44 |
+
},
|
| 45 |
+
"physical_recipe_issues": []
|
| 46 |
+
}
|
| 47 |
+
},
|
| 48 |
+
"removed_source_evidence": [],
|
| 49 |
+
"repo_id": "AutomatosX/AX-Nemotron-3.5-Lightning-30B-A3B-MLX-AXQ-MXFP4-MTP",
|
| 50 |
+
"runtime_arch_id": null,
|
| 51 |
+
"runtime_verified": false,
|
| 52 |
+
"schema_version": "axquant.hub-runtime-audit.v1",
|
| 53 |
+
"scope": "Pinned remote config/index/Safetensors header audit; no runtime load or generation claim.",
|
| 54 |
+
"source_revision": "5ab1716c9885e9bce30c063d159f5da4e0bb8c6f",
|
| 55 |
+
"status": "no-audited-peer-export-profile",
|
| 56 |
+
"unchanged_weight_sha256": {
|
| 57 |
+
"model-00001-of-00004.safetensors": "a8b6891cdb0d7a97316fe841c48fc0b926588a555ecd5b08ed471e7468d58f51",
|
| 58 |
+
"model-00002-of-00004.safetensors": "033922fe5c61923583d6cd8e72028b5d5fb1c5b0e788556356492581c7b7d70f",
|
| 59 |
+
"model-00003-of-00004.safetensors": "35999dfcfdc58c5084e50df422b1b5ad6fb5a2f80c643f159c61fd4cd5ab75d8",
|
| 60 |
+
"model-00004-of-00004.safetensors": "3c28f34687c69b0011ac52de0c557540d967d8df37442ebddc69dc71027c5c03",
|
| 61 |
+
"mtp.safetensors": "40dd606b285acd5045bd4246a916463eeb1d5f4cc7a4ebe051555985ec436e72"
|
| 62 |
+
}
|
| 63 |
+
}
|