How much faster?

#1
by EllaPriest45 - opened

First thank you for your work.

Second, how much faster are these compared to the original one? 5-10-30-50% faster?

Also, since theres only 1 AMD gpus with more than 24gb of vram... can we get GGUF version of that because damn

Hi @EllaPriest44, You are most welcome,

It's a hard one to answer the performance varies vastly depending on mode you are using and how many inputs (images and videos you are providing it), bearing in mind unlike most gaming GPUs, the GPU I am testing on is from the Strix-halo family since GPU family is optimized for AI workloads, I will provide benchmarks within the next week but bearing in mind they are for Ryzen-AI workstations rather than gaming GPU.

You have raised an excellent point, these models were developed with the target hardware assumed to be Ryzen-AI workstations (up to ~127GB vram like AMD Strix-halo or similar), I will try my best to get it under 24GB (I cant make any promises), please know that gguf unfortunately does not yield the same performance boost, from what I understand gguf format will adversely impact performance for ROCm using triton.

Hi @EllaPriest44,

I tried to make it smaller and the smallest I can get it without pruning the model is 24.002GB which still puts it outside the reach of the average AMD gaming GPU, as I mentioned my intended audience is AMD Ryzen AI workstations with 64GB - 128GB unified RAM. I cannot in good conscience distribute a self pruned model which may ultimately prove to be unstable or yield substandard quality.

Sign up or log in to comment