N8Programs's picture
Add MLX-VLM runtime, configuration, and model card
c179c12 verified
|
Raw
History Blame
9.95 kB
---
library_name: mlx
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/995ad96eacd98c81ed38be0c5b274b04031597b0/LICENSE
base_model:
- nvidia/Qwen3.6-35B-A3B-NVFP4
- N8Programs/Qwen3.6-35B-A3B-AntiLoop
pipeline_tag: image-text-to-text
tags:
- mlx
- mlx-vlm
- image-text-to-text
- qwen3.6
- moe
- modelopt
- quantized
- nvfp4
- fp4
- fp8
- lora
- merged
- antidoom
---
# Qwen3.6-35B-A3B-AntiLoop-NVFP4 for MLX-VLM
![AntiLoop looping rate versus GPQA capability](assets/looping_vs_gpqa.png)
![LoopHard judged-loop rates across four inference settings](assets/loophard_four_settings.png)
This is the Apple-silicon MLX-VLM conversion of
[`N8Programs/Qwen3.6-35B-A3B-AntiLoop-NVFP4`](https://huggingface.co/N8Programs/Qwen3.6-35B-A3B-AntiLoop-NVFP4).
It supports both text and vision inputs, including Qwen thinking controls.
The conversion does not re-quantize the mixed NVIDIA ModelOpt language
checkpoint. Its FP8 and NVFP4 payloads and scales are re-expressed for MLX's
native MXFP8/NVFP4 kernels, while the source tensor-level scales are applied by
the included runtime. Activations remain in the model dtype. The vision tower
is the original BF16 data, byte-for-byte rather than quantized.
- 1,808 tensors across 42 safetensors shards
- 130 scaled MXFP8 dense modules
- 121 scaled NVFP4 dense modules
- 120 scaled NVFP4 expert projections
- 333 BF16 vision tensors (893,142,496 tensor-data bytes)
The checkpoint was converted and smoke-tested with `mlx==0.31.2`,
`mlx-lm==0.31.3`, and `mlx-vlm==0.6.4`.
## Use with MLX-VLM
Install MLX-VLM and download the repository:
```bash
pip install -U "mlx-vlm==0.6.4"
hf download mlx-community/Qwen3.6-35B-A3B-AntiLoop-NVFP4 \
--local-dir Qwen3.6-35B-A3B-AntiLoop-NVFP4
cd Qwen3.6-35B-A3B-AntiLoop-NVFP4
```
Image + text generation:
```bash
python run_mlx_vlm.py \
--model . \
--trust-remote-code \
--image /path/to/image.png \
--prompt "Describe this image." \
--max-tokens 256
```
Text-only generation with thinking enabled:
```bash
python run_mlx_vlm.py \
--model . \
--trust-remote-code \
--enable-thinking \
--prompt "Solve: 27 * 43" \
--max-tokens 512
```
For programmatic loading from the downloaded repository:
```python
from mlx_vlm_model_file_loader import load
model, processor = load(".")
```
MLX-VLM 0.6.4 does not yet natively consult a model-local
`vlm_model_file`. `run_mlx_vlm.py` installs that single lookup inside the
current process without modifying the installed package. The local runtime is
executed only when `--trust-remote-code` (or `trust_remote_code=True`) is
explicitly enabled. Review the included Python files before trusting them.
An HTTP server can be started similarly:
```bash
python run_mlx_vlm_server.py \
--model . \
--trust-remote-code \
--enable-thinking
```
## Original model card
This is a mixed-precision NVIDIA ModelOpt deployment checkpoint for
[`Qwen3.6-35B-A3B-AntiLoop`](https://huggingface.co/N8Programs/Qwen3.6-35B-A3B-AntiLoop),
a narrow fine-tune intended to recover from pathological self-verification and
enumeration loops while preserving ordinary long-form reasoning.
No PEFT adapter is required at inference time. The MLX conversion retains the
upstream multimodal architecture, tokenizer, chat template, and 262,144-token
native context configuration. It does not include the source checkpoint's MTP
draft weights.
## Training data
The exact 178 masked supervised targets used for the final AntiLoop training
round are published in the
[`Qwen3.6-35B-A3B-AntiLoop-SFT` dataset](https://huggingface.co/datasets/N8Programs/Qwen3.6-35B-A3B-AntiLoop-SFT).
The dataset preserves each `loss_start_char` boundary so the pathological loop
prefix remains conditioning context rather than a supervised target. It
intentionally excludes the separately generated KL-regularization anchors.
## Training procedure
The AntiLoop adapter was trained on the 178 masked supervised targets using a
standard supervised fine-tuning procedure, but regularized via KL-loss on separately generated non-loop anchors from the base model on everyday prompts.
## Benchmark results
### LoopHard
**LoopHard** is our held-out set of 285 enumeration prompts designed to elicit
futile recall, recounting, and self-verification loops. The primary metric is
**judged loops**: whether the model's reasoning trace remains stuck in a futile
cycle when generation ends.
| Model | Judged loops | Loop rate |
|---|---:|---:|
| NVIDIA NVFP4 | 72 / 285 | 25.26% |
| **AntiLoop NVFP4** | **10 / 285** | **3.51%** |
| NVIDIA NVFP4 + `presence_penalty=1.5` | 30 / 285 | 10.53% |
| **AntiLoop NVFP4 + `presence_penalty=1.5`** | **1 / 285** | **0.35%** |
The matched `presence_penalty=1.5` comparison converted all 30 control loops to
clean completions while introducing one different loop. Exact two-sided
McNemar `p = 2.98e-8`.
Generation used thinking mode, `temperature=0.7`, `top_p=0.95`, `top_k=20`, a
6,144-token completion limit, and concurrency 24. The two
`presence_penalty=1.5` arms used the exact original and AntiLoop NVFP4
checkpoints with the same vLLM build and serving configuration: TP1, FP8 KV
cache, FlashInfer attention, Marlin NVFP4 MoE, and MTP speculative decoding
with three draft tokens.
LoopHard is judged by GLM-5.2 using a convergence-aware rubric: systematic
reasoning and verification that reaches a conclusion are not loops, and a trace
that notices its own circling and exits is classified as recovered. The
calibration set contained 42 manually labeled traces. Across three judge runs,
accuracy was 88.1%, 92.9%, and 95.2%; all three runs identified all 17 labeled
loops, with 2–5 false positives among the 25 non-loop traces.
The 285 prompts, metadata, and GLM-5.2 evaluation code are published in the
[LoopHard dataset](https://huggingface.co/datasets/N8Programs/LoopHard) on
Hugging Face.
### Capability preservation
The capability checks below compare the same round-2 AntiLoop adapter against
its FP8 reference model under a matched runtime-LoRA setup. These runs used the
default presence penalty and should not be interpreted as evaluations of the
exact mixed-precision artifact at `presence_penalty=1.5`.
#### GPQA Diamond
| Model | Accuracy |
|---|---:|
| Qwen3.6-35B-A3B official model card | 86.0% |
| FP8 reference, our matched harness | 167 / 198 (84.34%) |
| **AntiLoop FP8, our matched harness** | **166 / 198 (83.84%)** |
The official-model-card number is included for context and was not produced by
our harness.
Our GPQA run used thinking mode, paired per-question seeds,
`temperature=0.7`, `top_p=0.95`, `top_k=20`, MTP3 speculative decoding, a
65,536-token reasoning budget, and 4,096 tokens of answer headroom. The difference was not significant.
Source for the published 86.0% result:
[`Qwen/Qwen3.6-35B-A3B` model card](https://huggingface.co/Qwen/Qwen3.6-35B-A3B).
#### GSM8K
| Model | Accuracy |
|---|---:|
| FP8 reference | 1273 / 1319 (96.51%) |
| **AntiLoop FP8** | **1270 / 1319 (96.29%)** |
The GSM8K run used the exact 1,319-example `openai/gsm8k` `main` test split,
thinking mode, paired seeds, `temperature=0.7`, `top_p=0.95`, `top_k=20`, MTP3,
an 8,192-token reasoning budget and 1,024 tokens of answer headroom. The difference was not significant.
Taken together, the matched GPQA and GSM8K results show no material or
statistically detectable capability loss at these sample sizes. They do not
establish equivalence across other tasks, modalities, or sampling settings.
## Original NVIDIA ModelOpt usage
Use a recent vLLM build with ModelOpt mixed-precision support:
```bash
vllm serve N8Programs/Qwen3.6-35B-A3B-AntiLoop-NVFP4 \
--quantization modelopt \
--trust-remote-code \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3
```
MTP speculative decoding can be enabled on a compatible build with:
```bash
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
```
For the measured LoopHard setting, send the following sampling parameters:
```json
{
"temperature": 0.7,
"top_p": 0.95,
"top_k": 20,
"presence_penalty": 1.5
}
```
The capability-preservation results above used the default presence penalty;
`presence_penalty=1.5` has not yet been evaluated on GPQA or GSM8K.
Follow the
[`nvidia/Qwen3.6-35B-A3B-NVFP4` model card](https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4)
for deployment requirements and the
[`Qwen/Qwen3.6-35B-A3B` model card](https://huggingface.co/Qwen/Qwen3.6-35B-A3B)
for chat templating, thinking controls, multimodal inputs, and base-model
limitations.
## Limitations
- This is a narrow behavioral fine-tune, not a general alignment or safety model.
- The MLX conversion does not include an MTP speculative drafter.
- LoopHard is a task-specific, judge-based benchmark; its loop rate should not
be interpreted as a general safety, truthfulness, or factuality score.
- The GPQA and GSM8K checks used runtime LoRA on an FP8 base, not this exact
mixed-precision artifact.
- Capability preservation has not been tested at `presence_penalty=1.5`.
- Fixed-scale FP8 re-quantization approximates the exact BF16 LoRA merge; small
adapter updates can round away or clip at the original E4M3 range.
- Runtime validation used a 65,536-token configured context, not the full native
262,144-token context.
- Multimodal generation quality has not been evaluated on this artifact.
- Outputs may still be incorrect, overconfident, repetitive, biased, toxic, or
unsafe.
## License
Apache 2.0, following both the underlying Qwen checkpoint and NVIDIA's
quantized derivative. See the
[pinned Qwen license](https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/995ad96eacd98c81ed38be0c5b274b04031597b0/LICENSE),
the [Qwen model card](https://huggingface.co/Qwen/Qwen3.6-35B-A3B), and the
[NVIDIA ModelOpt checkpoint card](https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4).
(co-written with GPT-5.6-Sol)