Instructions to use mlx-community/Qwen3.6-35B-A3B-DFlash2-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/Qwen3.6-35B-A3B-DFlash2-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("mlx-community/Qwen3.6-35B-A3B-DFlash2-4bit") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use mlx-community/Qwen3.6-35B-A3B-DFlash2-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "mlx-community/Qwen3.6-35B-A3B-DFlash2-4bit" --prompt "Once upon a time"
- Atomic Chat
|
Download README.md from mlx-community/Qwen3.6-35B-A3B-DFlash2-4bit: direct link, hf CLI and curl.
- Browser
- Download file 1.33 kB
-
https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-DFlash2-4bit/resolve/main/README.md
- Command line
-
hf download hf://mlx-community/Qwen3.6-35B-A3B-DFlash2-4bit/README.md
-
curl -L -o README.md https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-DFlash2-4bit/resolve/main/README.md
1.33 kB
| license: apache-2.0 | |
| library_name: mlx | |
| pipeline_tag: text-generation | |
| base_model: | |
| - incoai/Qwen3.6-35B-A3B-Splash | |
| tags: | |
| - mlx | |
| - dflash2 | |
| - speculative-decoding | |
| - qwen3.6 | |
| - 4-bit | |
| # Qwen3.6-35B-A3B-DFlash2-4bit | |
| The DFlash 2 draft model for Qwen3.6-35B-A3B, as a standard safetensors checkpoint with MLX affine | |
| 4-bit weights (group size 64). | |
| ## Source and changes | |
| - Source: the `draft/` files of [`incoai/Qwen3.6-35B-A3B-Splash`](https://huggingface.co/incoai/Qwen3.6-35B-A3B-Splash), | |
| revision `0f4714b2db37b5f3c42a10de07281e74f88e4adc` (Apache-2.0). That package names its draft source as | |
| `incoai/Qwen3.6-35B-A3B-DFlash2`, revision `8e713508f0bb02f03b5cb5cabbc8d9604f924be2`. | |
| - The draft model was made by Inco AI. This repository is not published or endorsed by Inco AI. | |
| - Changes: the Splash runtime storage (tiled q4 integers with bf16 scales and biases) was converted to row-major | |
| MLX affine storage in one `model.safetensors`, and a `config.json` was written. The integers, scales and biases | |
| are moved bit for bit; no tensor was requantized (208 tensors). The original unquantized weights are not | |
| recovered. Mask token id and RoPE parameters come from the target's configuration. | |
| `source.json` lists the SHA-256 of every source file. | |
| ## License | |
| Apache-2.0, as the source package. See `LICENSE`. | |