Seems to be OK in first short tests. The most accessible and workable of GLM 5.3 Flash deployments

#3
by iamtanmay - opened

Unlike your MiMoV2.6 Flash REAP which bugged during reasoning (max), this model seems clean and stable at first glance. Since I am getting 5-7 tok/s currently, I haven't done longer tests yet.

I will conduct longer tests

I did another day of testing... the prune itself is really good, it doesn't seem to cause problems with the underlying model

However, the model itself is rather unsuited for a homelab. I can run Deepseek v4.1 Flash at 9-10 tok/s over tens of thousands of token context on 2 32GB GPUs, Qwen Next on a single 32GB GPU at 21 tok/s

GLM starts at 7 tok/s, but by 20k context its down to 1.4 tok/s... on 3 32GB GPUs

The cloud version of GLM 5.3 flash is the most capable model on the market today, so I was really hoping to get it working locally.

Its a real shame, but I hope they learn from Qwen and Deepseek to improve their efficiency

Sign up or log in to comment