Thank you!!

#1
by jc2375 - opened

How would you rate this model’s strengths on the GB10? lag/latency?

Honestly, since my efforts to make it work, I haven't used it a lot. I have an extremely tuned serve for qwen 3.5 that I usually use. Image gen, and tts don't feature as much in what I do.. that said, I had someone give me some feedback on it recently, so since there was interest in it, I've been working on honing the serve a bit. I'm trying to optimize it's tok/s so it's more useful. The big advantage of this model is going to be using it in workloads that do all the modalities it serves. Because it actually sees images or hears sound, (in their token form) then it should be able to understand it in ways that a sidecar solution vision tower or stt wont. I haven't actually tested that though?

For any one of it's modalities, there are better options if that's all you need.

It's not a snail, but it doesn't compare well against text gen models that have great speculative decoding solutions.

Sign up or log in to comment