Qwen3.8-27B for NInfer β€” RTX 5080 16 GB, true 128K + Vision

Project-maintained NInfer artifact for Qwen3.8-27B on a single NVIDIA RTX 5080 16 GB.

Current production runtime: v1.5

The model artifact itself remains byte-identical across the later runtime releases; v1.5 is the current production-qualified NInfer runtime/profile for this artifact.

At a glance

Field Current validated value
Artifact qwen3_8_27b.ninfer
Artifact size 16,461,267,456 bytes
SHA-256 c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21
Runtime release v1.5
Primary hardware RTX 5080 16 GB
Context 131,072 tokens
KV capacity 131,072 tokens
KV dtype Q4 group64
Prefill chunk 1792
Speculation MTP-3
CUDA Graph enabled
Host-mapped embeddings enabled
Vision profile 2048 tokens
Main text-model quantization mixed Q3/Q4/Q5, ~3.953 BPW
Canonical v1.5 prefill 1,374.383 tok/s
Canonical v1.5 sustained decode 112.215 tok/s

Canonical project source and validation records:

https://github.com/toddballinger/ninfer-5080

Download

Using the Hugging Face CLI:

hf download ninfer-5080/Qwen3.8-27B-RTX5080 \
  qwen3_8_27b.ninfer \
  qwen3_8_27b.ninfer.conversion.json \
  --local-dir .

Verify the artifact:

sha256sum qwen3_8_27b.ninfer

Expected:

c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21  qwen3_8_27b.ninfer

A matching SHA-256 identifies the exact validated project artifact regardless of filename or download machine.

Artifact identity vs runtime version

The model artifact and NInfer runtime are versioned independently.

The canonical artifact identity is:

SIZE=16461267456
SHA256=c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21

The same artifact has remained valid through later runtime optimization work. A newer NInfer runtime does not imply that a new .ninfer file is required.

Current production release β€” v1.5

Item Value
Release tag qwen3.8-27b-rtx5080-128k-vision-v1.5
Release commit e4edd6d5c5f9f7996de0f3d9f6c311e883452580
Validated source tree e4353f061bf0e378c83472cb2bcaf99e65681f4e
ninfer-serve SHA-256 928e5615ef453786f47f79b6af2152d2f8f8d61307656f23c47fa45b5ed41167
CUDA 13.4.92
NVIDIA driver 615.71.09

v1.5 standardizes release performance reporting on ninfer_bench and is the current RTX 5080 production-qualified profile.

Canonical v1.5 benchmark

pp118001+tg2048 on a single RTX 5080 16 GB:

Metric Result
Prefill 1,374.383 tok/s
Sustained decode 112.215 tok/s
Context / KV capacity 131,072 / 131,072
KV dtype Q4 group64
Prefill chunk 1792
Speculation MTP-3
CUDA Graph enabled
Host-mapped embeddings enabled
Vision profile 2048 tokens

Benchmark contract:

fixture=bench/fixtures/workflow-118k-v1/ninfer_bench_118001.ids
corpus_sha256=5b08da2c7b7ea5cafad2fab5699dccbcbce86040d8a37219b8c21f094d1d1eb7
prompt_tokens=118001
test=pp118001+tg2048
warmup=1
measured_repetitions=2
prefill=1374.383 tok/s
sustained_decode=112.215 tok/s

Historical v1.3/v1.4 measurements remain in the repository validation records.

Recommended v1.5 serving profile

./build/apps/ninfer-serve /path/to/qwen3_8_27b.ninfer \
  --host 0.0.0.0 \
  --port 8080 \
  --model-id local-model \
  --max-context 131072 \
  --kv-capacity 131072 \
  --prefill-chunk 1792 \
  --kv-dtype q4 \
  --spec mtp \
  --draft-tokens 3 \
  --max-concurrency 1 \
  --max-pending-requests 16 \
  --pending-timeout-ms 180000 \
  --embedding-host \
  --vision \
  --vision-max-tokens 2048 \
  --default-thinking-budget 2048 \
  --prefix-checkpoint-policy rolling-tool

CUDA Graph is enabled by default.

Validated startup envelope

free after startup  885.94 MiB
planned slack       806.92 MiB
vision workspace    132.3142 MiB

True 128K qualification

The production profile uses both a 131,072-token context and a 131,072-token KV capacity. Qualification uses an actual 118,001-token prompt, not merely a configured maximum context value.

This is the workload used for the canonical v1.5 benchmark above.

Multimodal validation

The HostMapped Vision path has been validated for:

  • deterministic image understanding
  • deterministic video understanding
  • multi-image conversation history
  • cached historical-media accounting
  • coexistence with the full 131,072-token text context/KV allocation

Details:

https://github.com/toddballinger/ninfer-5080/blob/main/docs/VISION_128K.md

Quantization profile

Main text-core distribution:

Format Share
Q3G64_F16S 42.42%
Q4G64_F16S 45.92%
Q5G64_F16S 11.57%
BF16 / FP32 ~0.10%

Effective main-model quantization: ~3.953 BPW.

Model sources

Qwen3.8-27B

  • repo: Qwen/Qwen3.8-27B
  • revision: 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0

DFlash2

  • repo: z-lab/Qwen3.8-27B-DFlash2
  • revision: 50307d4c4cde6860d4eee73e2547cd786fe8e8a4

Both upstream Hugging Face repositories currently declare Apache-2.0 licensing.

Reproducibility

The canonical groupwise artifact is built by a CPU-only GitHub Actions workflow on an explicitly provisioned runner β€” no local GPU is required.

The integrated workflow:

  • uses true row-sliced Safetensors reads;
  • streams artifact payload assembly;
  • pins conversion-critical source revisions and dependencies;
  • performs CPU/RAM/disk runner-capacity checks before model downloads;
  • structurally inspects the generated artifact;
  • requires the exact canonical byte count and SHA-256 before publication;
  • publishes qwen3_8_27b.ninfer.conversion.json alongside the artifact.

The production workflow has reproduced the exact canonical identity:

artifact_bytes=16461267456
artifact_sha256=c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21
canonical_groupwise_identity=PASS

Standard ubuntu-latest is not treated as sufficient for the complete 27B conversion. The NVFP4 workflow profile is separate and does not inherit the groupwise artifact's expected size or SHA-256.

Documentation

Credits

This artifact and its RTX 5080 production integration are maintained by Todd Ballinger / ninfer-5080.

starskyzheng made a significant contribution through PR #1, including the original low-memory Qwen3.8-27B conversion path and initial GitHub Actions automation/reproducibility work.

The integrated implementation was subsequently extended and hardened in ninfer-5080, including true row-sliced Safetensors reads, streamed payload assembly, regression coverage, pinned conversion dependencies, exact artifact size/SHA validation, provenance-report publishing, runner-capacity safeguards, production-runtime integration and final validation.

This project also builds on:

  • NInfer upstream β€” Neroued and contributors
  • Qwen3.8-27B β€” Qwen team
  • DFlash2 β€” z-lab and project contributors

Please preserve applicable upstream copyright, attribution and license notices.

Downloads last month
3,260
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ninfer-5080/Qwen3.8-27B-RTX5080

Base model

Qwen/Qwen3.8-27B
Quantized
(1321)
this model

Space using ninfer-5080/Qwen3.8-27B-RTX5080 1