Instructions to use TensorFold/Qwen3.8-Flash-Next-MLX-oQ4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TensorFold/Qwen3.8-Flash-Next-MLX-oQ4 with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("TensorFold/Qwen3.8-Flash-Next-MLX-oQ4") config = load_config("TensorFold/Qwen3.8-Flash-Next-MLX-oQ4") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use TensorFold/Qwen3.8-Flash-Next-MLX-oQ4 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Qwen3.8-Flash-Next-MLX-oQ4"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TensorFold/Qwen3.8-Flash-Next-MLX-oQ4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use TensorFold/Qwen3.8-Flash-Next-MLX-oQ4 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Qwen3.8-Flash-Next-MLX-oQ4"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TensorFold/Qwen3.8-Flash-Next-MLX-oQ4
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TensorFold/Qwen3.8-Flash-Next-MLX-oQ4 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Qwen3.8-Flash-Next-MLX-oQ4"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TensorFold/Qwen3.8-Flash-Next-MLX-oQ4" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
oQ4-MTP: degenerate repetition loop -> hits max_tokens and truncates the tool call (oQ3-MTP is fine, same settings)
Thanks for the quant. Reporting a reproducible failure I only see on oQ4-MTP — the oQ3-MTP quant of the same model, on the same machine, same server, same sampling settings, never shows it.
Symptom
On long agentic/coding turns the model falls into a repetition loop and keeps generating until it hits max_tokens (finish_reason=length, 32000 tokens). Because the run is cut off mid-generation, the tool-call envelope is never closed, so the whole turn is lost — the client gets a truncated block instead of a result.
Server-side, the truncation looks like this:
10:07:29 - omlx.patches.mlx_lm_mtp.batch_generator - MTP[0] finish=length tokens=28383 cycles=12487
tok/cycle=2.27 accept=15891/21250 (74.8%)
10:07:29 - omlx.api.tool_calling - WARNING - Unclosed tool-call envelope at end of stream;
withheld 18780 characters are available for content recovery
(start_marker='\n<function=write>\n<parameter=filePath>...')
10:07:29 - omlx.server - Chat completion: model=Qwen3.8-Flash-Next-MLX-oQ4-MTP,
32000 tokens in 916.32s, prompt: 10922, finish_reason=length, max_tokens=32000
Second occurrence, ~25 minutes later, on a different prompt:
10:33:31 - MTP[2] finish=length tokens=14684 cycles=3984 tok/cycle=3.69
accept=10698/10842 (98.7%) depth[d1=3798/3869,d2=3596/3626,d3=3304/3347]
10:33:31 - Chat completion: 32000 tokens in 1033.92s, prompt: 55878, finish_reason=length
That 98.7% MTP acceptance rate (vs. the 61–86% typical of healthy generation in my logs) is, I think, the clearest fingerprint of the problem: the draft head predicts the next tokens almost perfectly because the output has collapsed into repeating text. It is a symptom of the loop, not the cause — but it makes the loop trivially detectable from the logs.
oQ3 vs oQ4 comparison (same host, same day, same settings)
requests logged finish_reason=length unclosed tool-call envelopes
Qwen3.8-Flash-Next-MLX-oQ3-MTP 177 0 0
Qwen3.8-Flash-Next-MLX-oQ4-MTP 8 2 (25%) 1
The per-model settings for the two quants are byte-identical in my server (I diffed them): temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, repetition_penalty=1.0, presence_penalty=0.0, mtp_enabled=true, max_context_window=262144. So this does not look like a sampling-config problem on my side — the only variable is the quantization.
Secondary issue (may or may not be related)
During the same sessions I also hit a KV-cache failure that killed the stream, again only with oQ4:
09:51:18 - omlx.scheduler - WARNING - Cache corruption detected:
'QSAKVCache' object has no attribute 'extend', clearing cache and re-prefilling...
09:52:11 - omlx.server - ERROR - Error during chat streaming:
Cache corruption not recoverable after retries: 'QSAKVCache' object has no attribute 'extend'
This is plausibly an oMLX/mlx-vlm-side bug rather than a quant bug, so I'm mentioning it only for completeness — but it also never fired on oQ3 for me.
Environment
MacBook Pro, Apple M5 Max, 128 GB unified memory, macOS 26.5.1
oMLX 0.6.3, VLM batched engine, qwen4_exp compat patch, Lightning MTP enabled (language_model.mtp. checkpoint layout), PLE mode mmap
Model resolved at revision 43a82b3f0ff64fa417fd09ca046580f08d19b0d6, reported size 110.82 GB (109.94 GB text-only), ~76 GB resident once loaded
Client: OpenCode via the OpenAI-compatible endpoint, streaming, tool calling enabled, max_tokens=32000
Prompts where it triggered: agentic coding turns with 10k–56k token prompts and a write tool call producing a long file
Questions
Has the oQ4 recipe been sanity-checked for degenerate repetition on long generations (e.g. a 8–16k-token continuation test), or is the calibration set the same one used for oQ3?
Do any specific layers/tensors get more aggressive treatment in oQ4 than in oQ3 (lm_head, embeddings, attention output projections, or the MTP head itself)? A repetition collapse that shows up at 4-bit-ish and not at 3-bit-ish is unusual enough that a layer-selection difference seems more likely than raw bit width.
Is the MTP head quantized identically in the two uploads? If the oQ4 MTP head is more aggressively quantized, the near-100% acceptance runs might indicate the draft head is actively steering the model into the loop rather than just following it.
Happy to re-test any re-upload. Note I have already deleted my local copy to free the 106 GB, so I can't dump per-tensor stats from it — but I can re-download and run a fixed repetition benchmark against both quants if that's useful.
Thanks for the detailed report and logs. We’re looking into the oQ4-MTP quantisation and testing it against oQ3 on longer agentic runs; we’ll post an update here once we’ve reproduced the issue and know whether a rebuild is needed.
Thanks f
我用的这上模型稳定行还可以。一直在用这个模型用。感谢作者