Instructions to use dfp-official/GLM-5.3-Flash-oQ8e-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use dfp-official/GLM-5.3-Flash-oQ8e-mtp with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download dfp-official/GLM-5.3-Flash-oQ8e-mtp --local-dir GLM-5.3-Flash-oQ8e-mtp
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Vision requests fail: shipped chat_template.jinja is a text-only variant (no image-token emission)
Summary
Any vision request against GLM-5.3-Flash-oQ8e-mtp fails with:
Failed to process inputs with error: More images were provided than image tokens.
This happens on single-turn image requests too, so it is not a history-mixing issue. The vision weights themselves are fine (VLM engine, processor_config.json, image processor all present and functional).
Root cause
The checkpoint's chat_template.jinja is a text-only derivative. Its visible_text macro renders this for any media item instead of emitting image tokens:
<reminder>You are unable to process this image because you don't have multi-modal input ability. Try different methods.</reminder>
The official zai-org/GLM-5.3-Flash template instead defines emit_image() / emit_video() / emit_audio() macros that expand media items into the image-token triplet the GLM-5.3 processor requires (it replaces each token with merge_size^2-scaled placeholders, see mlx_vlm/models/glm5_next/processing.py). With zero image tokens in the rendered prompt, the processor's image_idx != len(grids) check raises the error above.
Fix (verified)
Replacing the checkpoint's chat_template.jinja with the official zai-org/GLM-5.3-Flash template fixes vision completely. The official template is a superset: identical reasoning_effort / clear_thinking handling (so the text/thinking path is unchanged), plus the media emission macros.
Verified on oMLX 0.7.0.dev4 (M3 Ultra 512GB), both cases that previously failed now return correct answers with active thinking:
- fresh single-turn image request
- image request preceded by plain text history
Could a future revision of this repo ship the official template (or a vision-capable variant)? Thanks for the excellent quantization otherwise β benchmark-class quality.