How to use from the
Use from the
MLX library
# Make sure mlx-lm is installed
# pip install --upgrade mlx-lm

# Generate text with mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("SirSahOl/Qwen3-0.6B-chat-mlx-4bit")

prompt = "Write a story about Einstein"
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True
)

text = generate(model, tokenizer, prompt=prompt, verbose=True)

Qwen3-0.6B-mlx-4bit

4-bit MLX conversion of Qwen/Qwen3-0.6B for Apple Silicon.

Converted by: SirSahOl Source model: Qwen/Qwen3-0.6B Framework: MLX by Apple Quantization: 4-bit Format: safetensors License: apache-2.0


Quick Start

Installation

pip install mlx-lm

CLI Usage

# Chat interactively
mlx_lm.chat --model SirSahOl/Qwen3-0.6B-chat-mlx-4bit

# Generate text
mlx_lm.generate --model SirSahOl/Qwen3-0.6B-chat-mlx-4bit --prompt "Your prompt here"

Python Usage

from mlx_lm import load, generate

model, tokenizer = load("SirSahOl/Qwen3-0.6B-chat-mlx-4bit")
response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256)
print(response)

Performance Benchmarks

| Metric | 4-bit | 8-bit | 16-bit | |--------|--------|--------|--------|| Tokens/sec | 114.11 | 70.97 | 42.97 | | TTFT | 8.77 ms | 14.09 ms | 23.28 ms | | Peak Memory | 449.8 MB | 393.5 MB | 171.6 MB |

Benchmarked on Apple M1 with 8GB unified memory. Average over 5 runs with 256 max tokens.


Who Should Use This?

Your Hardware Recommended Quantization
M1/M2 (8GB) 4-bit — Best balance of quality and memory usage
M1/M2 Pro/Max (16-32GB) 8-bit — Higher quality with reasonable memory
M2/M3/M4 Ultra (64GB+) 16-bit — Full precision, no quality loss

General guidance:

  • Use 4-bit if you want to run this model alongside other applications
  • Use 8-bit if you have the memory and want better quality
  • Use 16-bit for research, evaluation, or if memory isn't a concern

Other Quantization Variants


Conversion Details

Property Value
Source Model Qwen/Qwen3-0.6B
Quantization 4-bit
mlx-lm Version 0.31.3
Conversion Time 4.66s
Output Size 330.9 MB
Date 2026-09-10T17:45:38.042701+00:00

Reproduction

To reproduce this conversion:

pip install mlx-lm==0.31.3
python3 -m mlx_lm.convert --hf-path Qwen/Qwen3-0.6B --mlx-path output/Qwen3-0.6B-mlx-4bit -q --q-bits 4

Limitations & Known Issues

  • Performance may degrade with very long contexts (>8K tokens) at lower quantization levels.
  • This is a weight-only conversion; the model architecture and behavior are inherited from the source model.
  • Quantization introduces a small quality loss compared to the original model. Lower bit counts = more loss.
  • This model requires Apple Silicon (M1 or later) to run with MLX.

License

This model conversion inherits the license of the source model: apache-2.0.

See the original model card for full license details.


Changelog

Version Date Changes
v1.0 2026-09-10 Initial conversion

Converted with MLX Foundry — a professional pipeline for converting models to Apple MLX format.

Downloads last month
41
Safetensors
Model size
0.6B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SirSahOl/Qwen3-0.6B-chat-mlx-4bit

Finetuned
Qwen/Qwen3-0.6B
Quantized
(436)
this model

Collection including SirSahOl/Qwen3-0.6B-chat-mlx-4bit