Kudos — combined your model with DFlash/AEON for GB10

#8
by robbatt - opened

Kudos and thank you for your work! I combined this model with z-lab/Qwen3.6-27B-DFlash to create a faster version — quantized to NVFP4 and served with DFlash speculative decoding on an NVIDIA DGX Spark GB10. However I lack the hardware to train a proper DFlash drafter for the 40B target — can you help out? I added benchmark results and the full setup to the model card: https://huggingface.co/robbatt/Qwen3.6-40B-Deckard-NVFP4. Cheers, Rob

Owner

Thank you!

RE: Drafter ; not sure what you are asking here - you mean MTP version?
or?

I have not made an MTP version because there are still pipeline issues with MTP.

Sign up or log in to comment