Instructions to use NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp") config = load_config("NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-35B-A3B-Distill · oQ8-fp16-mtp — a Novaeon.Studio build
A fast, agentic, reasoning local worker for Apple Silicon — with a working MTP head. This is an oMLX oQ8 (near‑uniform 8‑bit, router fp16) quantization of empero-ai/Qwen3.8-35B-A3B-Distill — a distillation of the Qwen3.8 frontier teachers (including Qwen3.8 Flash Next) into the sparse Qwen3.6‑35B‑A3B MoE architecture — repacked for the Mac in oMLX‑native format. Built, tuned, and benchmarked on an Apple M5 Max (128 GB) by Novaeon.Studio.
Why this build exists: we run local agents on Apple Silicon and wanted frontier‑distilled reasoning at 8‑bit fidelity, in oMLX format, with a documented optimal serving profile. Unlike other 35B‑A3B builds we ship, this one's native MTP head actually accelerates decode under oQ8 — so it's tuned MTP‑on. Everything below is measured on our hardware, not copied.
| Base model | empero-ai/Qwen3.8-35B-A3B-Distill (Apache‑2.0) |
| Distilled from | Qwen3.8 2.4T‑A95B and Qwen3.8 Flash Next (teacher CoT traces; math/code weighted) |
| Architecture | Qwen3.5/3.6‑family MoE + vision encoder · 35B total / ~3B active · 256 experts (8 routed) · 40 layers · hybrid GDN linear + full attention |
| Quantization | oQ8, group size 64, float16 scales/non‑quant weights · ~8.6 effective bpw · ~39.5 GB on disk |
| MTP head | Native, shipped in‑checkpoint — recognized and accelerates decode under oQ8 (see below) |
| Context | 262,144 native · vision preserved |
| Engine | oMLX (Apple MLX) — VLM engine |
Best for & why it's here
Best for: the default local agent workhorse — tool-using agents, multi-step reasoning, coding, long-context (256k+), and vision, all on Apple Silicon. Fast enough to sit behind live agents and crons.
Why we published it: this is our own fleet seat — the model every Novaeon agent runs on. We wanted frontier-distilled reasoning at 8-bit fidelity in oMLX format, with a native MTP head that actually accelerates decode and a documented optimal serving profile. 6/6 on our agentic probe, effectively refusal-free, ~80–93 tok/s.
Highlights
- Frontier‑distilled reasoning. SFT on Qwen3.8 teacher chain‑of‑thought; every answer opens with a
<think>block (served asreasoning_content). Setenable_thinking: falsefor plain answers. - Working MTP → fast. The native multi‑token‑prediction head is not a decorative graft — under oQ8 it sustains real speculative acceptance (draft‑6 optimal), keeping decode at frontier‑seat speed.
- Super‑agentic. 6/6 on our controlled tool‑use probe suite (tool selection, parallel calls, sequential chains, argument fidelity, correct abstention, no hallucinated tools).
- Vision intact. Multimodal (text + image → text); the vision tower is preserved from the base.
Measured benchmarks (Apple M5 Max 128 GB · oMLX)
All numbers below were measured on our hardware, on these exact weights, with the seat isolated (no other load) using the optimal profile (thinking off).
Decode throughput vs context
| Prompt context | ~0 | ~3.9k | ~7.9k | ~15.9k | ~31.9k | ~64.7k |
|---|---|---|---|---|---|---|
| Decode (tok/s) | 93 | 81 | 82 | 69 | 57 | 41 |
| TTFT (s) | 0.30 | 1.4 | 1.4 | 2.0 | 2.2 | 3.7 |
Cold load ≈ 1 s (warm registry). Peak short‑context decode reaches ~110 tok/s with MTP draft‑6 in ideal runs.
Agentic & tool‑use
Our probe on this build — 6/6 (temp 0): tool selection, parallel tool calls, sequential tool chains, argument fidelity, abstention when no tool is needed, and no hallucinated tools. Standard probes (JSON, code, reasoning, thinking, vision) all pass.
- IFEval (ours, this build): 84.86 avg — measured locally against these quantized weights (thinking off), scored with the official IFEval verifier (541 prompts, 834 instructions).
| IFEval (this build) | prompt‑strict | prompt‑loose | inst‑strict | inst‑loose | avg |
|---|---|---|---|---|---|
| Qwen3.8‑35B‑A3B‑Distill oQ8‑fp16‑mtp | 80.22 | 83.92 | 86.33 | 88.97 | 84.86 |
Optimal oMLX settings (figured out empirically)
We swept the tuning levers on this build. Because the native MTP head genuinely accelerates decode under oQ8, the optimal profile is MTP‑on with ANE‑prefill off — the inverse of a build whose MTP is inert. (With a working MTP, ANE‑prefill's dispatch overhead is a net loss.)
| Setting | Value | Why |
|---|---|---|
mtp_enabled |
true | native MTP accelerates decode (+~40% short) — real speculative acceptance under oQ8 |
mtp_num_draft_tokens |
6 | swept 4/6/8 → 6 optimal (~110 tok/s short) |
qwen35_ane_prefill_enabled |
false | net‑regresses decode when MTP is active (dispatch overhead) |
turboquant_kv_enabled / turboquant_kv_bits |
true / 8 | holds decode up as context grows; tiny KV (hybrid attention) |
qwen35_oq_a8_enabled |
false | A8 collapsed long‑context decode in testing |
dflash_enabled |
false | net‑regressed decode in A/B |
moe_expert_offload_enabled |
false | keep experts resident (128 GB is ample) |
Drop‑in ~/.omlx/model_settings.json entry
{
"max_context_window": 262144,
"mtp_enabled": true,
"mtp_num_draft_tokens": 6,
"qwen35_ane_prefill_enabled": false,
"turboquant_kv_enabled": true,
"turboquant_kv_bits": 8.0,
"turboquant_skip_last": true,
"qwen35_oq_a8_enabled": false,
"dflash_enabled": false,
"moe_expert_offload_enabled": false
}
On the MTP head: empero ships the multi‑token‑prediction head natively in the checkpoint (the model.safetensors.index.json maps the mtp.* tensors directly). oMLX recognizes it (mtp_compatible: true) and it accelerates decode under oQ8 — our A/B shows MTP‑on ≈ +40% short‑context tok/s over MTP‑off, i.e. real draft acceptance. This is unusual for an 8‑bit MoE build and is why we serve with mtp_enabled: true.
Long context: the base uses mRoPE with a high 10M theta, so context extrapolates past the 262,144 native window without YaRN — verified coherent (needle‑in‑haystack recall) at 320k tokens. The hybrid attention (only ~10 of 40 layers carry a growing KV cache) keeps memory tiny even at very long context.
Base model benchmarks
Reported by the base‑model authors (empero-ai/Qwen3.8-35B-A3B-Distill), zero‑shot, vs the Qwen3.6-35B-A3B base it distills onto:
| Task | Metric | Qwen3.6‑35B‑A3B (base) | This distill | Δ |
|---|---|---|---|---|
| MMLU (57 subj.) | acc | 0.838 | 0.834 | −0.004 (within noise) |
| ARC‑Challenge | acc_norm | 0.548 | 0.591 | +0.044 |
| ARC‑Easy | acc_norm | 0.717 | 0.766 | +0.048 |
Reproduced from the base model card; the distillation lifts reasoning (ARC) while holding MMLU. Full teacher/trace details are on the base card.
Quickstart
# oMLX (recommended, Apple Silicon)
omlx serve NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp --port 8000
# then hit the OpenAI-compatible endpoint at http://127.0.0.1:8000/v1
# mlx-vlm
pip install -U mlx-vlm
python -m mlx_vlm.generate --model NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp \
--max-tokens 4096 --temperature 0.6 --top-p 0.95 \
--prompt "A snail climbs 3m by day, slips 2m by night, in a 10m well. How many days to escape?"
Recommended inference parameters
temperature 0.6 · top_p 0.95 · top_k 20 · reasoning model (emits <think>; set enable_thinking: false for plain answers) · parse and strip <think>…</think> for end users.
Attribution & license
- Base model:
empero-ai/Qwen3.8-35B-A3B-Distill— Apache‑2.0, by Empero. A distillation of Qwen3.8 frontier teachers into the Qwen3.6‑35B‑A3B architecture. - This quantization & serving profile: Novaeon.Studio, 2026. Released under Apache‑2.0, same as the base.
This is an independent community quantization. It is not endorsed by Empero. All credit for the model's capabilities belongs to the base‑model and Qwen authors.
@misc{novaeon2026qwen38distilloq8,
title = {Qwen3.8-35B-A3B-Distill oQ8-fp16-mtp: an oMLX build for Apple Silicon},
author = {Novaeon.Studio},
year = {2026},
note = {Quantization of empero-ai/Qwen3.8-35B-A3B-Distill},
url = {https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp}
}
novæon — digital business architecture + AI · novaeon.studio
- Downloads last month
- 2,208
8-bit
Model tree for NovaeonStudio/Qwen3.8-35B-A3B-Distill-oQ8-fp16-mtp
Base model
Qwen/Qwen3.6-35B-A3B

