ThinkingCap Qwen 3.8 Quants
Quantized ThinkingCap-Qwen3.8-27B: NVFP4 and NVFP4A4-AWQ (vLLM), GGUF (llama.cpp), MLX (Apple silicon).
Image-Text-to-Text • 17B • Updated • 1.63k • 23Note 🍀 NVFP4 weight-only, tied GDN scales, vLLM, 20.6 GB. Hopper (Marlin) and Blackwell; needs a 32 GB GPU (RTX 5090). Evals vs bf16: accuracy −3.0 to +0.9 pp per benchmark. Mean stronger shortening, median kept.
bottlecapai/ThinkingCap-Qwen3.8-27B-NVFP4A4-AWQ
Image-Text-to-Text • 20B • Updated • 622 • 11Note 🖤 NVFP4 weights and activations (AWQ), vLLM, 23.4 GB. BLACKWELL ONLY (CUTLASS FP4); 32 GB GPU. Meant for batched serving over NVFP4. Evals vs bf16: accuracy −1.3 to +1.0 pp, shortening kept.
bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF
Image-Text-to-Text • 27B • Updated • 163k • 58Note 🦙 GGUF, llama.cpp/LM Studio/Ollama; CUDA, Metal, Vulkan or CPU. Weights plus ~2 GB per 32k context (+0.9 GB mmproj): 🦅 IQ4_XS 15.5 GB, 🐆 Q4_K_M 17.4 GB (24 GB GPU, 32 GB Mac); 🐎 Q6_K 23.9 (32 GB GPU), 🐅 Q8_0 29.0 (48 GB GPU); 🐻 f16 54.7 (80 GB). Q4_K_M evals vs bf16: accuracy −3.0 to +0.8 pp per benchmark, none significant. Shortening kept.
bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ
Image-Text-to-Text • 27B • Updated • 1.25k • 10Note 🍎 MLX for Apple silicon, mixed 4/8-bit DWQ, 22.5 GB; runs on a 32 GB Mac. Serve with mlx-vlm or oMLX: vision input and MTP self-speculative decoding work there (mlx-lm drops both).