Not Ninfer version

#1
by s1lverkin - opened

Hello, would you be able to push your version before modifying it to Ninfer format?

I would like to use json schema output, which in Ninfer it doesn't work.

Thank you

It looks beautiful, but is there any way to squeeze this so it can be usefull at 16GB VRAM Q_Q ?

This is my first model, I'll see what I can do with it. Thank you.

Owner
•
This comment has been hidden
Owner
•
This comment has been hidden

Thanks!
Pushed the pre-NInfer version: https://huggingface.co/dragoy/Swift-Qwen3.8-27B-abliterated-NVFP4 - standard HF safetensors (NVFP4+FP8, vLLM/transformers compatible).
JSON-schema structured output works through vLLM guided decoding; there is a short usage snippet in the README.
@s1lverkin

For 16 GB cards: the NVFP4 weights are 21.8 GiB, so they don't fit on a single 16 GB card.

Working options:
(1) GGUF IQ4_XS with imatrix and the MTP head, ~14.5 GiB - https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF (abliterated Swift, text-only, llama.cpp/LM Studio);
(2) 2x16 GB with TP2 - documented working at https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4/discussions/16 (54-120 tok/s with MTP);
(3) the NVFP4 repo https://huggingface.co/dragoy/Swift-Qwen3.8-27B-abliterated-NVFP4 with CPU offload - full quality, lower speed. The mamba/attention hybrid keeps the KV cache tiny, so on small cards the weight size, not the context, is the binding limit.

@xor7589

Sign up or log in to comment