Instructions to use Qwen/Qwen3.8-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen3.8-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Qwen/Qwen3.8-27B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Qwen/Qwen3.8-27B") model = AutoModelForMultimodalLM.from_pretrained("Qwen/Qwen3.8-27B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Qwen/Qwen3.8-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen3.8-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Qwen/Qwen3.8-27B
- SGLang
How to use Qwen/Qwen3.8-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Qwen/Qwen3.8-27B with Docker Model Runner:
docker model run hf.co/Qwen/Qwen3.8-27B
KV Cache Precision Benchmarks
I used Qwen3.8-27B to test and measure how different KV cache configurations affect quality.
Comparing the following configuration knobs:
- KV cache data type (BF16, FP8)
- TurboQuant KV cache compression presets (k8v4, 4bit, k3v4, 3bit)
- Context scaling by using YaRN (native 256k, extended 512k, 1M tokens max. context)
- Weights data type (BF16, FP8)
Setup
- Model: Qwen3.8-27B (BF16 and FP8)
- Inference Engine: vLLM nightly snapshot
- Benchmark: OpenAI MRCR v2 (post-12/5/2025 bugfix), 2/4/8-needle
- Protocol: 15 samples/run (5 per needle bucket n=2/4/8), temp=0, seed=42, concurrency=1, max-tokens=2048, thinking off. Score < 0.5 = hard fail.
- Sample sets: in-range ~212–227k tokens (all configs below unless noted); out-of-range ~525–547k (section 4 only).
- Hardware: single DGX Spark
Decisions
I picked OpenAI's MRCR benchmark* for this task, as I quickly realized that a primitive one-needle search, buried in a large repitition of the same sentence was too simple for even aggressive KV cache configurations to fail.
The MRCR benchmark OTOH challenges the LLM much more by using multi-needle search, packed into sentences of high similarity. So it makes the results much more sensitive to KV value precision.
Basically this benchmark tests for information retrieval only though, i.e. whether the right needle was found in a large context, and whether the nearby context of that needle was returned by the LLM exactly verbatim in its response.
What this benchmark does not tell is the impact on complex tasks like agentic software engineering. I can imagine that those complex tasks are much more sensitive to V-precision than this benchmark does, which is probably more K-precision sensitive.
NOTE: I restricted the benchmark to only 15 samples per configuration, so there is some noticeable noise in the results. I had to restrict it simply because of the very long prefill time on my single DGX Spark machine. Even with 15 samples it took several days to run all these benchmarks, so I had to make a tradeoff.
1. KV dtype × Max. Context (in-range samples)
Achieved score as primary value, hard-fails as secondary information in round brackets:
| KV dtype | 256k (no YaRN) | 512k (YaRN f=2) | 1M (YaRN f=4) |
|---|---|---|---|
| BF16 | 0.833 (3/15) | 0.741 (5/15) | 0.742 (5/15) |
| FP8 | 0.717 (5/15) | 0.737 (5/15) | 0.742 (5/15) |
| TQ-k8v4 | 0.833 (3/15) | 0.833 (3/15) | 0.741 (5/15) |
| TQ-4bit_nc | 0.777 (4/15) | 0.833 (3/15) | 0.692 (6/15) |
| TQ-k3v4_nc | 0.777 (4/15) | 0.686 (5/15) | 0.686 (6/15) |
| TQ-3bit_nc | 0.673 (6/15) | 0.777 (4/15) | 0.638 (7/15) |
- The scores show that extending max. context by using YaRN is not free.
- While BF16 KV cache dtype shines at native context, it appears to fall down to FP8 level on YaRN extended context.
- FP8 remains quite stable over all context lengths.
- TQ-k8v4 appearing to be better than FP8 is most probably just noise due to the low amount of samples used. The key data type of this TurboQuant preset is plain FP8 type (no K-compression, only V-compression).
2. Per-needle results (KV dtypes × Max. Context)
| KV dtype | Max. Context | n2 | n4 | n8 |
|---|---|---|---|---|
| BF16 | 256k (no YaRN) | 0.998 (0/5) | 0.832 (1/5) | 0.669 (2/5) |
| BF16 | 512k (YaRN f=2) | 0.998 (0/5) | 0.832 (1/5) | 0.394 (4/5) |
| BF16 | 1M (YaRN f=4) | 0.998 (0/5) | 0.704 (2/5) | 0.523 (3/5) |
| FP8 | 256k (no YaRN) | 0.998 (0/5) | 0.798 (1/5) | 0.355 (4/5) |
| FP8 | 512k (YaRN f=2) | 0.998 (0/5) | 0.832 (1/5) | 0.380 (4/5) |
| FP8 | 1M (YaRN f=4) | 0.998 (0/5) | 0.704 (2/5) | 0.523 (3/5) |
| TQ-k8v4 | 256k (no YaRN) | 0.998 (0/5) | 0.832 (1/5) | 0.669 (2/5) |
| TQ-k8v4 | 512k (YaRN f=2) | 0.998 (0/5) | 0.832 (1/5) | 0.669 (2/5) |
| TQ-k8v4 | 1M (YaRN f=4) | 0.998 (0/5) | 0.555 (3/5) | 0.669 (2/5) |
| TQ-3bit_nc | 256k (no YaRN) | 0.844 (1/5) | 0.821 (1/5) | 0.355 (4/5) |
| TQ-3bit_nc | 512k (YaRN f=2) | 0.998 (0/5) | 0.812 (1/5) | 0.523 (3/5) |
| TQ-3bit_nc | 1M (YaRN f=4) | 0.997 (0/5) | 0.538 (3/5) | 0.380 (4/5) |
- Here you can see that needle count matters, almost no configuration failed on only 2 needles.
- Only the most aggressive TurboQuant presets failed on small needle count.
3. TurboQuant compression vs. quality (no-YaRN)
Note that I have not measured the perplexity (PPL) values in the following table, they were taken directly from the comments attached to vLLM's TQ presets.
| preset | compression | PPL | score | fails |
|---|---|---|---|---|
| TQ-k8v4 | 2.6× | +1.17% | 0.833 | 3/15 |
| TQ-4bit_nc | 3.8× | +2.71% | 0.777 | 4/15 |
| TQ-k3v4_nc | 3.5× | +10.63% | 0.777 | 4/15 |
| TQ-3bit_nc | 4.9× | +20.59% | 0.673 | 6/15 |
- TQ-k8v4 appears to be almost free, while providing good compression ratio.
- TQ-4bit_nc appears to be a good trade-off between compression and quality.
- The other two TurboQuant presets are probably too aggressive.
4. Out-of-range ~547k samples (YaRN f=4, 1M max. context)
While the previous benchmark results were all taken by using the same samples <256k tokens, I also ran some benchmarks on samples that were about ~547k tokens in length, so true long context prompts. However I had no more time for an extensive exploration.
| KV dtype | score | fails |
|---|---|---|
| BF16 | 0.789 | 4/15 |
| FP8 | 0.696 | 6/15 |
- While BF16 fell off before on smaller prompts on YaRN extended context, on true large prompts like here though it keeps its dominance over FP8.
- Maybe both suffer more hard fails on true long prompts than on short prompts, however hard to compare, as these are completely different sample sets than before.
Weight quantizations
| KV | BF16 weights | FP8 weights |
|---|---|---|
| BF16KV | 0.827 (3/15) | 0.833 (3/15) |
| FP8KV | 0.782 (4/15) | 0.717 (5/15) |
I have only performed few benchmarks on this aspect, but they suggested that BF16 vs. FP8 weights had no real impact on the benchmarks' quality results. They were either bit-identical or within noise range. Therefore I used FP8 weights exclusively for all previous benchmarks instead.