Are there no options for a single Strix Halo?
It's a pity.
Try dwarfstar, there is a Q2 quant, and if you get my toolbox it's got performance improvements. https://github.com/kyuz0/ai-toolbox-cockpit is the easiest way to get ds4 up and running on strix halo.
But, I'd say that GLM 5.3 doesn't resist that level of qauntization well, so I am finding it is not as good as IQ2_XXS for DeepSeek V4 Pro.
Try dwarfstar, there is a Q2 quant, and if you get my toolbox it's got performance improvements. https://github.com/kyuz0/ai-toolbox-cockpit is the easiest way to get ds4 up and running on strix halo.
But, I'd say that GLM 5.3 doesn't resist that level of qauntization well, so I am finding it is not as good as IQ2_XXS for DeepSeek V4 Pro.
I would like to use your toolboxes with an AMD 7900 XTX and W7800 Pro eGPU on my Strix Halo, but gfx1100 is not supported. Can i build it myself with these architectures?
Try dwarfstar, there is a Q2 quant, and if you get my toolbox it's got performance improvements. https://github.com/kyuz0/ai-toolbox-cockpit is the easiest way to get ds4 up and running on strix halo.
But, I'd say that GLM 5.3 doesn't resist that level of qauntization well, so I am finding it is not as good as IQ2_XXS for DeepSeek V4 Pro.
I would like to use your toolboxes with an AMD 7900 XTX and W7800 Pro eGPU on my Strix Halo, but gfx1100 is not supported. Can i build it myself with these architectures?
Should be possible, just clone the repo and ask your favourite LLM to add support for those architectures as well, most LLMs will be able to do it and build the container I'm sure!
Try dwarfstar, there is a Q2 quant, and if you get my toolbox it's got performance improvements. https://github.com/kyuz0/ai-toolbox-cockpit is the easiest way to get ds4 up and running on strix halo.
But, I'd say that GLM 5.3 doesn't resist that level of qauntization well, so I am finding it is not as good as IQ2_XXS for DeepSeek V4 Pro.
I tried DS4, but it seemed worse to me in terms of agent coding.
There is an opinion that GLM is better even at 1 bit; I want to test that.
I am interested in long-context coding.