Instructions to use LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
Use Docker
docker model run hf.co/LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
- Ollama
How to use LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF with Ollama:
ollama run hf.co/LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF with Docker Model Runner:
docker model run hf.co/LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
- Lemonade
How to use LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.6-27B-NVFP4-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "LibertAIDAI/Qwen3.6-27B-NVFP4-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Why q4 model is so small?
This Q4 Model: 15G
Other Q4 Models: 18G->21G
Im curious about the difference between models and want to know how to reach this result.
Im thinking about to reimplement this(+ remove mtp, vision and change some layer to mxfp4 to accelerate) to fit 27B to my 4090.
If anyone knows how to do this, please tell me. Thanks
Why it's 15 GB
It's not the FFN-NVFP4 win that does most of the work β NVFP4 is ~4.5 bits/elem packed (4-bit weight + FP8 block scale
per 16 elems), versus Q4_K_M at ~4.84 effective bits/elem. That's only ~0.3 bits/weight cheaper. The bigger savings
come from not boosting embed/lm_head/etc. to higher precision, which most "Q4" GGUFs (unsloth UD, etc.) routinely do.
The actual tensor inventory of Qwen3.6-27B-NVFP4-Q4_K_M.gguf:
βββββββββββββ¬ββββββββ¬βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β # tensors β dtype β what β
βββββββββββββΌββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β 192 β NVFP4 β every ffn_gate / ffn_up / ffn_down weight (64 layers Γ 3) β
βββββββββββββΌββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β 33 β Q6_K β output.weight + the biggest projections β
βββββββββββββΌββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β 273 β Q4_K β token_embd, attn_gate, attn q/k/v on the 16 full-attn layers, ssm_out, β¦ β
βββββββββββββΌββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β 737 β F32 β norms, biases, NVFP4 per-tensor input scales β
βββββββββββββ΄ββββββββ΄βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
MTP and vision are already gone. The base GGUF was built without the MTP head; the vision tower is in the separate
mmproj-Qwen3.6-27B-F16.gguf (889 MB) and you only need it for image input. So for your "text-only on a 4090" goal you
can ignore the mmproj β that's already what you want.
How to reproduce / shrink further
- Source: mmangkad/Qwen3.6-27B-NVFP4 β NVIDIA ModelOpt v0.42 activation-aware NVFP4 calibration, repacked. You cannot
add NVFP4 to more tensors without re-running ModelOpt β NVFP4 needs activation calibration, llama.cpp will not
synthesize it from BF16. (MXFP4 can be assigned at quantize-time, no calibration needed.) - llama.cpp at master (β₯ #22196 for Blackwell NVFP4 MMA, plus #20505/#20506/#22611 for the Qwen3.5/3.6 convert path).
- python convert_hf_to_gguf.py β BF16 GGUF with the 192 FFN tensors already typed as NVFP4 (the convert script
picks up hf_quant_config.json). - llama-quantize --tensor-type '=' input.gguf out.gguf Q4_K_M to retype anything you want. To push lower
than Q4_K_M:
- --tensor-type 'attn_q.weight=MXFP4' --tensor-type 'attn_k.weight=MXFP4' --tensor-type 'attn_v.weight=MXFP4'
--tensor-type 'attn_output.weight=MXFP4'
- Or flip the SSM in_proj (attn_qkv.weight in this arch) β that's the 48Γ Q6_K block that dominates the non-FFN
footprint. - For 4090 specifically: NVFP4 and MXFP4 both fall back to the dp4a / MMQ paths on Ada (sm_89) β no hardware
tensor-core speedup there. The size shrinks, but it won't go faster than well-tuned Q4_K_M. Your real constraint on a
24 GB 4090 won't be weights, it'll be KV cache (64 layers Γ 4 KV heads Γ 256 dim is heavy at long context). Use -ctk
q4_0 -ctv q4_0 and you'll fit way more context than weight quant tweaks will buy you.