Qwen3.8-27B mixed-precision EXL3 (measured)
Four EXL3/Trellis quants of Qwen3.8-27B optimizing fidelity, flexibility, capacity, or 262k context. Every KLD measured vs BF16.
Image-Text-to-Text • 11B • Updated • 12.1k • 11Note Recommended default. 21.61 GB download, 20.31 GiB resident, KLD 0.007406 body-only / 0.007532 as served, 178 s cold start, attention fixed at K6.
malaiwah/Qwen3.8-27B-EXL3-K5K6
Image-Text-to-Text • 15B • Updated • 1.13k • 1Note Flexible edition. 30.57 GB download because attention ships BF16 and is encoded at load (957 s cold, cache required), but the width is a launch-time knob: K6 0.008157, K5 0.012135 (~206k context on 32 GB), K4 0.027530.
malaiwah/Qwen3.8-27B-K4
Image-Text-to-Text • 14B • Updated • 949Note Capacity edition. 17.89 GiB resident is the only build in this family that fits native 262,144 context on a 32 GB card (289,577 KV tokens measured), at KLD 0.030736.
malaiwah/Qwen3.8-27B-EXL3-K5K6-context
Image-Text-to-Text • 10B • Updated • 678 • 6Note Context edition. Native 262,144 context on a 32 GB card with int8 embeddings (18.13 GiB resident), retrieval verified at 227,334 tokens, KLD 0.009738 - 26 % below official FP8.