Image-Text-to-Text
Transformers
Safetensors
qwen3_5
qwen3.8
nvfp4
fp4
compressed-tensors
mtp
vision
thinking
v100
conversational
8-bit precision
Instructions to use philbert440/Qwen3.8-27B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use philbert440/Qwen3.8-27B-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="philbert440/Qwen3.8-27B-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("philbert440/Qwen3.8-27B-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("philbert440/Qwen3.8-27B-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use philbert440/Qwen3.8-27B-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "philbert440/Qwen3.8-27B-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "philbert440/Qwen3.8-27B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/philbert440/Qwen3.8-27B-NVFP4
- SGLang
How to use philbert440/Qwen3.8-27B-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "philbert440/Qwen3.8-27B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "philbert440/Qwen3.8-27B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "philbert440/Qwen3.8-27B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "philbert440/Qwen3.8-27B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use philbert440/Qwen3.8-27B-NVFP4 with Docker Model Runner:
docker model run hf.co/philbert440/Qwen3.8-27B-NVFP4
Model card: V100 measured performance (warm multi-pass), draft-mode + cudagraph sizing guidance, chart
Browse files- README.md +34 -4
- images/throughput-v100.png +0 -0
README.md
CHANGED
|
@@ -53,9 +53,38 @@ ModelOpt-exported NVFP4 checkpoints require capability 7.5+ and reject Volta. On
|
|
| 53 |
|
| 54 |
## Measured performance
|
| 55 |
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 59 |
|
| 60 |
## The base model
|
| 61 |
|
|
@@ -122,7 +151,8 @@ python -m vllm.entrypoints.openai.api_server \
|
|
| 122 |
--tensor-parallel-size 2 --gpu-memory-utilization 0.78 \
|
| 123 |
--max-model-len 32768 --kv-cache-dtype fp8_e5m2 \
|
| 124 |
--enable-prefix-caching --reasoning-parser qwen3 \
|
| 125 |
-
--
|
|
|
|
| 126 |
```
|
| 127 |
|
| 128 |
SM70 notes, learned the hard way:
|
|
|
|
| 53 |
|
| 54 |
## Measured performance
|
| 55 |
|
| 56 |
+

|
| 57 |
+
|
| 58 |
+
**Methodology:** warm serve (2 discarded warmup generations), fixed-length generations via
|
| 59 |
+
`ignore_eos` so every run produces exactly the stated token count, official Qwen3.8 sampling per
|
| 60 |
+
mode (thinking `1.0/0.95/20`, instruct `0.7/0.80/20 + presence 1.5`), varied prompts. Reported
|
| 61 |
+
as mean ± sd tokens/s. Rig: 2×V100-32GB, 1Cat-vLLM 1.2.2, TP2, `fp8_e5m2` KV,
|
| 62 |
+
`max_num_seqs 4`, MTP K=2.
|
| 63 |
+
|
| 64 |
+
| Regime | greedy draft | probabilistic draft |
|
| 65 |
+
|---|---|---|
|
| 66 |
+
| 512-tok, thinking (n=10) | 53.0 ± 2.5 | **54.4 ± 1.3** |
|
| 67 |
+
| 2048-tok, thinking (n=3) | 51.1 ± 2.5 | **54.1 ± 1.1** |
|
| 68 |
+
| 512-tok, instruct (n=6) | 50.5 ± 1.8 | 50.0 ± 1.6 |
|
| 69 |
+
| Mean acceptance length, whole workload | 2.25 | 2.46 |
|
| 70 |
+
|
| 71 |
+
**Concurrency** (4-way, 512-tok, aggregate): **~170 tok/s** with
|
| 72 |
+
`{"cudagraph_mode":"piecewise"}` (auto capture sizes), ~165 with `full_and_piecewise` — on par
|
| 73 |
+
with the [W4A16 sibling](https://huggingface.co/philbert440/Qwen3.8-27B-W4A16-AWQ). One
|
| 74 |
+
sizing rule matters: with MTP, each sequence schedules `K+1` tokens per step, so **never set
|
| 75 |
+
explicit `cudagraph_capture_sizes` below `max_num_seqs × (K+1)`** — a cap of `[1,2,4,8]` at
|
| 76 |
+
batch 4 pushes concurrent decode off CUDA graphs and collapses aggregate throughput ~3×
|
| 77 |
+
(measured 55–72 tok/s; reproduces identically on the W4A16 sibling, so it's a config trap, not
|
| 78 |
+
a format property). Engine-default auto sizing is correct.
|
| 79 |
+
|
| 80 |
+
**Pick the draft mode by workload:** verification rejection-samples against the target model,
|
| 81 |
+
so output quality is identical either way. At the official temp-1.0 thinking sampling,
|
| 82 |
+
**probabilistic** matches the verified distribution and wins (+2–6%); on low-temperature
|
| 83 |
+
workloads the two converge (see instruct row); at `temperature 0` greedy is the natural choice.
|
| 84 |
+
|
| 85 |
+
Quality validation (passed on this rig): factual coherence, think-tag discipline (zero
|
| 86 |
+
`<think>` leakage with thinking disabled), vision (image understanding through the VLM path),
|
| 87 |
+
GSM8K sample 3/3, and long-form generation with no repetition/degeneration.
|
| 88 |
|
| 89 |
## The base model
|
| 90 |
|
|
|
|
| 151 |
--tensor-parallel-size 2 --gpu-memory-utilization 0.78 \
|
| 152 |
--max-model-len 32768 --kv-cache-dtype fp8_e5m2 \
|
| 153 |
--enable-prefix-caching --reasoning-parser qwen3 \
|
| 154 |
+
--compilation-config '{"cudagraph_mode":"piecewise"}' \
|
| 155 |
+
--speculative-config '{"method":"mtp","num_speculative_tokens":2,"attention_backend":"FLASH_ATTN_V100","draft_sample_method":"probabilistic"}'
|
| 156 |
```
|
| 157 |
|
| 158 |
SM70 notes, learned the hard way:
|
images/throughput-v100.png
ADDED
|