Instructions to use vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt") model = AutoModelForMultimodalLM.from_pretrained("vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt
- SGLang
How to use vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt with Docker Model Runner:
docker model run hf.co/vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt
Download README.md from vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt: direct link, hf CLI and curl.
- Browser
- Download file 4.75 kB
-
https://huggingface.co/vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt/resolve/5da9eb4005b4fad2b2d9dd04168215196f8af85f/README.md
- Command line
-
hf download hf://vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt@5da9eb4005b4fad2b2d9dd04168215196f8af85f/README.md
-
curl -L -o README.md https://huggingface.co/vroomfondel/Qwen3.8-27B-NVFP4-ModelOpt/resolve/5da9eb4005b4fad2b2d9dd04168215196f8af85f/README.md
library_name: transformers
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
pipeline_tag: text-generation
tags:
- modelopt
- nvfp4
- fp4
qwen3.8-27b-nvfp4-modelopt
NVFP4 quantization of Qwen/Qwen3.8-27B via NVIDIA ModelOpt.
Quantization details (auto-generated)
- source model: Qwen/Qwen3.8-27B
- qformat:
nvfp4kv_cache:fp8 - calibration:
?samples from? - producer: NVIDIA ModelOpt
? - generated: ?
Before/after sample generation was skipped for this run (SKIP_GENERATE=1).
Notes
Validation status -- structurally verified and smoke-tested coherent
Quantized and served on a single DGX Spark (GB10/sm121) on 2026-08-15. The export passed every structural check: the vision tower is BF16 (333 tensors, zero scale tensors, dtypes and shapes identical to the source), quant_algo is NVFP4 rather than MIXED_PRECISION, the FP8 KV scales sit on exactly the 16 full-attention layers (indices 3, 7, ..., 63), in_proj_qkv is NVFP4-packed on all 48 Gated-DeltaNet layers while conv1d/in_proj_a/in_proj_b stayed BF16, and all 15 mtp.* tensors are present and unquantized. Serving on SGLang 0.5.17 then produced coherent output on all four probe types (German two-sentence explanation, a multi-step train word problem, a five-sentence historical paragraph, a translation), each finishing with finish_reason=stop. The word problem was solved correctly (17:00, with the right derivation), which is the more informative signal: a broken Gated-DeltaNet path degrades into word salad rather than into arithmetic mistakes. This is a four-prompt smoke test, not an evaluation -- GSM8K or a comparable suite is still owed before any quality claim.
Uniform W4A4 on a Gated-DeltaNet hybrid
Unlike NVIDIA's Qwen3.5/3.6 NVFP4 releases, which quantize the FFN/expert path and leave attention in BF16, this build quantizes attention as well, including the Gated-DeltaNet linear-attention path that carries 48 of the model's 64 layers. That is the deliberate point of the profile and it is also where the risk sits: an inadequately loaded scale on the fused linear_attn.in_proj_qkv degrades this architecture into complete word salad rather than into a measurable accuracy drop. Judge the build on generated output, not on the fact that it loads.
Serving requires the qwen3_5 attention-quant and KV-scale loader changes
Because attention is quantized and FP8 KV scales are baked into the checkpoint, SGLang needs the qwen3_5 attention-quant override and baked-KV-scale loader changes (sgl-project/sglang PR #31220) plus the NVFP4 scalar-scale fix for merged and fused linears (PR #29151, merged upstream 2026-07-13) that the fused Gated-DeltaNet in_proj_qkv depends on. Use the flashinfer attention backend: the triton backend hits a forward-time crash in RadixLinearAttention.forward for this exact configuration (sgl-project/sglang#29577, still open). A reliable check that the KV scales actually loaded is that the server logs "Using FP8 KV cache but no scaling factors provided" zero times.
Multimodal, vision tower kept BF16
Calibration is text-only, so the 27-layer vision tower, its merger and the embeddings are excluded from quantization and stay BF16, avoiding the amax=0 degenerate-quant failure mode. The language-side FFNs that consume projected image tokens ARE quantized, and they were calibrated on text alone, so the image path is the least-validated surface of this build and should be checked against the BF16 source before being relied on.
BF16 MTP head, speculative decoding off by default
The NEXTN/MTP draft head (mtp_num_hidden_layers=1) ships unquantized: transformers drops mtp.* at load for every Qwen3.5 architecture, so ModelOpt never sees it and passes the tensors through as BF16. Speculative decoding is therefore available but is left disabled in the shipped serving profile until bare serving is verified coherent, so that a quality regression can never be confused with a drafter problem.
Quantize on a single Spark with offload, not force-on-GPU
Quantize with SEQ_DEVICE_MAP=0 (offload / device_map=auto). GPU and CPU share the same ~121 GB unified memory pool on a DGX Spark, so offloading the 55.6 GB source costs nothing here. Do NOT use SEQ_DEVICE_MAP=1 with a high GPU_MAX_MEM_PCT on a single Spark: that reserves most of the shared pool as GPU and OOM-kills weight loading regardless of the percentage. Force-on-GPU is a multi-GPU (4x H200) setting only.