Kudos — combined your model with DFlash/AEON for GB10

#7
by robbatt - opened

Kudos and thank you for your work! I used the AEON vLLM image to serve DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking — a 40B dense Claude 4.6 Opus fine-tune — as NVFP4 with DFlash speculative decoding on an NVIDIA DGX Spark GB10, giving 2.5× throughput boost even with a mismatched drafter. Benchmark results here: https://huggingface.co/robbatt/Qwen3.6-40B-Deckard-NVFP4. Cheers, Rob

Owner

Awesome feedback thanks for the data!

AEON-7 changed discussion status to closed

Sign up or log in to comment