oQ4-MTP: degenerate repetition loop -> hits max_tokens and truncates the tool call (oQ3-MTP is fine, same settings)

#2
by TokenAI-zer - opened

Thanks for the quant. Reporting a reproducible failure I only see on oQ4-MTP — the oQ3-MTP quant of the same model, on the same machine, same server, same sampling settings, never shows it.

Symptom
On long agentic/coding turns the model falls into a repetition loop and keeps generating until it hits max_tokens (finish_reason=length, 32000 tokens). Because the run is cut off mid-generation, the tool-call envelope is never closed, so the whole turn is lost — the client gets a truncated block instead of a result.

Server-side, the truncation looks like this:

10:07:29 - omlx.patches.mlx_lm_mtp.batch_generator - MTP[0] finish=length tokens=28383 cycles=12487
tok/cycle=2.27 accept=15891/21250 (74.8%)
10:07:29 - omlx.api.tool_calling - WARNING - Unclosed tool-call envelope at end of stream;
withheld 18780 characters are available for content recovery
(start_marker='\n<function=write>\n<parameter=filePath>...')
10:07:29 - omlx.server - Chat completion: model=Qwen3.8-Flash-Next-MLX-oQ4-MTP,
32000 tokens in 916.32s, prompt: 10922, finish_reason=length, max_tokens=32000

Second occurrence, ~25 minutes later, on a different prompt:

10:33:31 - MTP[2] finish=length tokens=14684 cycles=3984 tok/cycle=3.69
accept=10698/10842 (98.7%) depth[d1=3798/3869,d2=3596/3626,d3=3304/3347]
10:33:31 - Chat completion: 32000 tokens in 1033.92s, prompt: 55878, finish_reason=length

That 98.7% MTP acceptance rate (vs. the 61–86% typical of healthy generation in my logs) is, I think, the clearest fingerprint of the problem: the draft head predicts the next tokens almost perfectly because the output has collapsed into repeating text. It is a symptom of the loop, not the cause — but it makes the loop trivially detectable from the logs.

oQ3 vs oQ4 comparison (same host, same day, same settings)
requests logged finish_reason=length unclosed tool-call envelopes
Qwen3.8-Flash-Next-MLX-oQ3-MTP 177 0 0
Qwen3.8-Flash-Next-MLX-oQ4-MTP 8 2 (25%) 1
The per-model settings for the two quants are byte-identical in my server (I diffed them): temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, repetition_penalty=1.0, presence_penalty=0.0, mtp_enabled=true, max_context_window=262144. So this does not look like a sampling-config problem on my side — the only variable is the quantization.

Secondary issue (may or may not be related)
During the same sessions I also hit a KV-cache failure that killed the stream, again only with oQ4:

09:51:18 - omlx.scheduler - WARNING - Cache corruption detected:
'QSAKVCache' object has no attribute 'extend', clearing cache and re-prefilling...
09:52:11 - omlx.server - ERROR - Error during chat streaming:
Cache corruption not recoverable after retries: 'QSAKVCache' object has no attribute 'extend'

This is plausibly an oMLX/mlx-vlm-side bug rather than a quant bug, so I'm mentioning it only for completeness — but it also never fired on oQ3 for me.

Environment
MacBook Pro, Apple M5 Max, 128 GB unified memory, macOS 26.5.1
oMLX 0.6.3, VLM batched engine, qwen4_exp compat patch, Lightning MTP enabled (language_model.mtp. checkpoint layout), PLE mode mmap
Model resolved at revision 43a82b3f0ff64fa417fd09ca046580f08d19b0d6, reported size 110.82 GB (109.94 GB text-only), ~76 GB resident once loaded
Client: OpenCode via the OpenAI-compatible endpoint, streaming, tool calling enabled, max_tokens=32000
Prompts where it triggered: agentic coding turns with 10k–56k token prompts and a write tool call producing a long file
Questions
Has the oQ4 recipe been sanity-checked for degenerate repetition on long generations (e.g. a 8–16k-token continuation test), or is the calibration set the same one used for oQ3?
Do any specific layers/tensors get more aggressive treatment in oQ4 than in oQ3 (lm_head, embeddings, attention output projections, or the MTP head itself)? A repetition collapse that shows up at 4-bit-ish and not at 3-bit-ish is unusual enough that a layer-selection difference seems more likely than raw bit width.
Is the MTP head quantized identically in the two uploads? If the oQ4 MTP head is more aggressively quantized, the near-100% acceptance runs might indicate the draft head is actively steering the model into the loop rather than just following it.
Happy to re-test any re-upload. Note I have already deleted my local copy to free the 106 GB, so I can't dump per-tensor stats from it — but I can re-download and run a fixed repetition benchmark against both quants if that's useful.

TensorFold org
•
edited Aug 28

Thanks for the detailed report and logs. We’re looking into the oQ4-MTP quantisation and testing it against oQ3 on longer agentic runs; we’ll post an update here once we’ve reproduced the issue and know whether a rebuild is needed.

Thanks f

我用的这上模型稳定行还可以。一直在用这个模型用。感谢作者

Sign up or log in to comment