Well done !

#7
by iamtanmay - opened

Q2 is running at 20-25 TPS on a Mi50 32GB with 250K context Q5 and offload a 5 layers to RAM. Incredible performance

The quality of the thinking and output is also top class. I generally ask "Write an RPG in C# 6.0 like Skyrim, no graphics only backend", and it generated an entire project of 14 scripts, well planned and consistent

Its the best LLM on such a compact size. The next level up would be a GLM 5.3, which is a far bigger model, but not so far in intelligence

Keeping the size slightly higher than your first iteration was definitely worth it. You were right about the intelligence knee

Sign up or log in to comment