Vision requests fail: shipped chat_template.jinja is a text-only variant (no image-token emission)

#1
by JordiHugging - opened

Summary

Any vision request against GLM-5.3-Flash-oQ8e-mtp fails with:

Failed to process inputs with error: More images were provided than image tokens.

This happens on single-turn image requests too, so it is not a history-mixing issue. The vision weights themselves are fine (VLM engine, processor_config.json, image processor all present and functional).

Root cause

The checkpoint's chat_template.jinja is a text-only derivative. Its visible_text macro renders this for any media item instead of emitting image tokens:

<reminder>You are unable to process this image because you don't have multi-modal input ability. Try different methods.</reminder>

The official zai-org/GLM-5.3-Flash template instead defines emit_image() / emit_video() / emit_audio() macros that expand media items into the image-token triplet the GLM-5.3 processor requires (it replaces each token with merge_size^2-scaled placeholders, see mlx_vlm/models/glm5_next/processing.py). With zero image tokens in the rendered prompt, the processor's image_idx != len(grids) check raises the error above.

Fix (verified)

Replacing the checkpoint's chat_template.jinja with the official zai-org/GLM-5.3-Flash template fixes vision completely. The official template is a superset: identical reasoning_effort / clear_thinking handling (so the text/thinking path is unchanged), plus the media emission macros.

Verified on oMLX 0.7.0.dev4 (M3 Ultra 512GB), both cases that previously failed now return correct answers with active thinking:

  • fresh single-turn image request
  • image request preceded by plain text history

Could a future revision of this repo ship the official template (or a vision-capable variant)? Thanks for the excellent quantization otherwise β€” benchmark-class quality.

Sign up or log in to comment