VERSION: PrismAura quantized version (DGX Spark optimized)

#70
by trithemius - opened

Hi all. In case you need, here my PrismaQuant Aura quantization, optimized for DGX Spark: https://huggingface.co/trithemius/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-MTP-PrismAura-5.5bit

DavidAU changed discussion title from PrismAura quantized version (DGX Spark optimized) to VERSION: PrismAura quantized version (DGX Spark optimized)

Awesome quants, was using prisma for vanilla qwen as well. Am able to run this in 5090 @525W using vllm with two configs :

  1. 215k fp8 kv context with mtp (100+tok/sec)
  2. full 262k fp8 kv context without mtp (50+ tok/sec).

Sign up or log in to comment