Broken model unfortunately

#1
by iamtanmay - opened

Hi Patrick, tried the Q2_K with a prompt "Write a C# 6.0 RPG like Skyrim, no graphics only backend"

It started outputting gibberish in the Think tokens, random symbols like *#$ and words/sentences that made no sense

Temp 1.0, top-p 0.95 as Xiaomi recommended on latest llama-cpp ROCm

Hi Patrick, tried the Q2_K with a prompt "Write a C# 6.0 RPG like Skyrim, no graphics only backend"

It started outputting gibberish in the Think tokens, random symbols like *#$ and words/sentences that made no sense

Temp 1.0, top-p 0.95 as Xiaomi recommended on latest llama-cpp ROCm

Hello, thanks for letting me know! I will do more evaluation this week and see which quantizations and formats work well.

I'm using a 128gb unified ram device, but you might be able to point claude or codex at these repos and the base unpruned teacher Mimo-v2.6-Flash, then prune that model 50% with a calibration set more weighted to game development. Its an option that make a model better suited towards game-oriented fine tuning.

https://github.com/patrickbdevaney/xiaomi-2.6-flash-REAP
https://github.com/patrickbdevaney/glm-5.3-reap

I need to eval the base gguf and mxfp4 more thoroughly, thanks again

The one good recent REAP of similar size that I saw was "https://huggingface.co/AnonimousA/Qwen3.8-Flash-Next-REAP-320-GGUF". This was a second attempt, because his first attempt shaved off too many experts and was unstable. Perhaps 50% is too much for MiMo V2.6 ?

One thing nice about MiMo vs your GLM 5.3 Flash prune is that I could launch it out of the box with llama.cpp. I have a unified memory rig which is slow, and a fast dedicated GPU cluster with driver issues (MI50). I spent days rebuilding your GLM 5.3 fork for Vulkan and ROCm.

Vulkan on dedicated GPUs was getting me < 0.1 tok/s, Unified memory < 3 tok/s. After a lot of debugging, I just got my ROCm setup working ~ 7tok/s and am beginning to run longer tests on your GLM 5.3 quants. So far they seem really good.

Sign up or log in to comment