Will this work with 16 GB VRAM?

#2
by cnsiva - opened

Will this work with the RTX 4070 Ti S (16 GB) graphics card? If it does, what is the maximum context size it can handle?

This is actually beat model for 16gb. It is moe and only 3b active. Remember to compile the turbo-tan/llama.cpp-tq3 main branch

I'm using the YTan2000 base Qwen3.6 model (the model this is based on) with a laptop 3060 - 6GB VRAM and 65k context and it's FAST - like 40 t/s with low context and 30 t/s when context is full. It's close to 50% faster than the Unsloth IQ4_XS. I am downloading this now and I'm so excited.

Owner

Thanks @bigjeff5 for leaving a feedback, not many like you come to give feedback!

Can we expect Ornith-1.5 ? @YTan2000

Yes

Sign up or log in to comment