Instructions to use aufklarer/Qwen3.5-0.8B-Chat-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use aufklarer/Qwen3.5-0.8B-Chat-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("aufklarer/Qwen3.5-0.8B-Chat-MLX") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use aufklarer/Qwen3.5-0.8B-Chat-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "aufklarer/Qwen3.5-0.8B-Chat-MLX" --prompt "Once upon a time"
- Atomic Chat
Clarify the runtime bit-width requirement and what is dropped from the export
Browse files
README.md
CHANGED
|
@@ -78,8 +78,9 @@ quantized; they stay float. Quantized tensors are stored as MLX's
|
|
| 78 |
| `int8/tokenizer_config.json` | 16.7 kB | Chat template and special tokens |
|
| 79 |
|
| 80 |
Each variant is self-contained: one directory is everything needed to run it.
|
| 81 |
-
The multi-token-prediction draft head (`mtp.*`)
|
| 82 |
-
included — no runtime here reads them
|
|
|
|
| 83 |
|
| 84 |
## Measured quality
|
| 85 |
|
|
@@ -140,10 +141,11 @@ let response = try model.generate(
|
|
| 140 |
```
|
| 141 |
|
| 142 |
> **INT5 and INT8 need a runtime that reads the bit width from `config.json`.**
|
| 143 |
-
> Older speech-swift versions build every `QuantizedLinear`
|
| 144 |
-
>
|
| 145 |
-
>
|
| 146 |
-
>
|
|
|
|
| 147 |
|
| 148 |
### Python
|
| 149 |
|
|
|
|
| 78 |
| `int8/tokenizer_config.json` | 16.7 kB | Chat template and special tokens |
|
| 79 |
|
| 80 |
Each variant is self-contained: one directory is everything needed to run it.
|
| 81 |
+
The vision tower and the multi-token-prediction draft head (`mtp.*`) are not
|
| 82 |
+
included — no runtime here reads them. Dropping the draft head alone took 14 MB
|
| 83 |
+
off the INT4 download compared with the previous revision.
|
| 84 |
|
| 85 |
## Measured quality
|
| 86 |
|
|
|
|
| 141 |
```
|
| 142 |
|
| 143 |
> **INT5 and INT8 need a runtime that reads the bit width from `config.json`.**
|
| 144 |
+
> Older speech-swift versions build every `QuantizedLinear` and
|
| 145 |
+
> `PreQuantizedEmbedding` with `bits = 4` hardcoded, which fits the INT4 file
|
| 146 |
+
> only; they will not read an INT5 or INT8 file correctly. Use a speech-swift
|
| 147 |
+
> version that takes `quantization_bits` and `quantization_group_size` from
|
| 148 |
+
> `config.json`, or stay on INT4.
|
| 149 |
|
| 150 |
### Python
|
| 151 |
|