--- library_name: mlx license: other license_name: openmdw-1.1 license_link: https://openmdw.ai/license/1-1/ base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 base_model_relation: quantized pipeline_tag: text-generation language: - en - es - fr - de - it - ja tags: - mlx - mlx-lm - omlx - nvidia - nemotron-3.5 - mixture-of-experts - mamba - quantized - apple-silicon - 4-bit ---

NVIDIA Nemotron Apple silicon MLX Vontra oMLX

NVIDIA Nemotron 3.5 Lightning 30B-A3B — MLX 4-bit

A native Apple-silicon conversion of nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, quantized with stock 4-bit affine weights and packaged for MLX-LM and oMLX.

Original model · NVIDIA Nemotron · MLX-LM · OpenMDW 1.1 license

## About this conversion This repository contains a stock 4-bit affine MLX conversion of NVIDIA Nemotron 3.5 Lightning. The source is a 30B-total / 3B-active hybrid mixture-of-experts model that interleaves Mamba-2, sparse MoE, and attention layers. The upstream tokenizer, chat template, and generation configuration are preserved. | Item | Value | | --- | --- | | Base model | [`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16`](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16) | | Format | MLX safetensors | | Quantization | 4-bit affine, group size 64 | | Conversion stack | `mlx-lm 0.31.3`, `mlx 0.32.0` | | Weight shards | 4 | | Weight size | 17.78 GB (16.55 GiB) | | Maximum configured context | 262,144 tokens | | Architecture | `nemotron_h` — Mamba-2 + sparse MoE + attention | > [!NOTE] MLX-LM reported an effective precision of 4.503 bits per weight. ## Apple-silicon performance This checkpoint was load-tested and generation-tested on the following machine: | Hardware | Configuration | | --- | --- | | Host | Mac Studio | | Chip | Apple M3 Ultra | | CPU | 32 cores (24 performance + 8 efficiency) | | Unified memory | 256 GB | | Runtime | MLX-LM 0.31.3 / MLX 0.32.0 | A warmed local test produced: | Measurement | Result | | --- | ---: | | Decode (median) | **168.41 tokens/s** | | Reported peak memory | **17.95 GB** | | Timed runs | 3 × 256 generated tokens | | Warm-up | 32 generated tokens | | Prompt | 36 tokens after chat templating | The decode figure is the median of three greedy 256-token runs after a 32-token Metal-kernel warm-up. It is a practical local reference, not a controlled cross-platform benchmark. Prompt length, context growth, sampler settings, memory pressure, thermal state, and MLX/oMLX versions can materially change performance. ## Quick start with MLX-LM Install recent MLX-LM and Hugging Face tooling: ```bash python -m pip install -U mlx-lm huggingface_hub ``` Run directly from the Hub: ```bash mlx_lm.generate \ --model Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit \ --prompt "Explain why hybrid Mamba and MoE architectures are efficient." \ --max-tokens 512 \ --temp 1.0 \ --top-p 0.95 ``` Reasoning mode is enabled by the upstream chat template by default. To disable it: ```bash mlx_lm.generate \ --model Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit \ --chat-template-config '{"enable_thinking": false}' \ --prompt "Write a short hello-world program in Swift." \ --max-tokens 256 ``` Python usage: ```python from mlx_lm import load, generate model, tokenizer = load("Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit") messages = [ {"role": "user", "content": "Explain sparse mixture-of-experts routing."} ] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=False, enable_thinking=False, ) print(generate(model, tokenizer, prompt=prompt, max_tokens=512)) ``` To download the repository first: ```bash hf download Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit \ --local-dir ~/.omlx/models/Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit ``` ## Using it with oMLX 1. Place the downloaded model at `~/.omlx/models/Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit`. 2. Refresh the oMLX model registry. 3. Load `NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit` and use the normal chat UI or OpenAI-compatible endpoint. Example request: ```bash curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OMLX_API_KEY" \ -d '{ "model": "NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit", "messages": [{"role": "user", "content": "Say hello from Nemotron on MLX."}], "temperature": 1.0, "top_p": 0.95, "max_tokens": 128 }' ``` For long prompts, begin with a conservative context limit and increase it while watching memory pressure. The configured 256K context is a model capability, not a guarantee that every host can prefill that context within its available unified memory. ## Architecture Nemotron 3.5 Lightning is a hybrid sparse model designed for efficient agentic and reasoning workloads. | Architecture detail | Upstream value | | --- | ---: | | Total / active parameters | 30B / 3B | | Layers | 52 | | Routed / shared experts | 128 / 1 | | Active routed experts | 6 | | Attention heads / KV heads | 32 / 2 | | Hidden size | 2,688 | | Expert intermediate size | 1,856 | | Vocabulary size | 131,072 | | Configured context | 262,144 tokens | The upstream release is intended for coding, tool use, reasoning, research, and customization. For NVIDIA's evaluations, deployment guidance, intended use, limitations, safety information, and full architecture discussion, see the [original model card](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16). ## Conversion and validation notes - Source weights: NVIDIA's BF16 checkpoint. - Quantization group size: 64. - Quantization mode: affine. - The upstream `chat_template.jinja` is preserved. - All 729 converted tensors and every indexed shard were checked locally. - The model was loaded and exercised through end-to-end generation on Apple silicon. - Quantization can reduce output quality relative to BF16; use a higher-precision variant when quality matters more than memory use. This is a community conversion, not an official NVIDIA release. Validate quality and numerical behavior on your own representative workload before production use. ## License and attribution The upstream model is released under the **OpenMDW License Agreement, version 1.1**. A copy is included in this repository; review it before use or redistribution. All model design, training, benchmark, and upstream documentation credit belongs to NVIDIA and the original contributors. The MLX conversion, Apple-silicon validation, and packaging are provided by [Vontra](https://huggingface.co/Vontra). ## Choose for your Mac [64GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-64gb-macs-6a9fefda17932216ec9ab457) · [128GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-128gb-macs-6a9ff0abd31bc9abbe7922d7) · [256GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-256gb-macs-6a9ff0ef9fed7c5bdca15e9b) Published peak memory: **17.95 GB**; estimated starting tier: **64GB**, leaving about **46 GB** nominal headroom. The collections use published M3 Studio peaks with at least 25% nominal headroom; fit on other Macs is an estimate, and full context is not guaranteed. Start with short context and one request. ### Runtime and evidence The exact tested oMLX application version is not recorded here; a library version is not an app version. The original performance tables retain their benchmark conditions and speed figures; this documentation update adds no new test results. ### Quick start and demo prompt ```bash hf download Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit --local-dir ./models/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit ``` Add the downloaded folder to oMLX model directories, refresh the list, and follow this card's architecture and MTP compatibility requirements before loading. Try this in a new chat with a 128-token output limit: ```text Explain why the sky looks blue in three short sentences. ``` This is a demo prompt to try, not a recorded successful run; a captured demonstration for this documentation update is not yet available. [Follow Vontra for new Apple Silicon releases and fixes.](https://huggingface.co/Vontra)