tmballin commited on
Commit
8a56ba1
·
verified ·
1 Parent(s): a448b20

docs: polish official RTX 5080 model card

Browse files
Files changed (1) hide show
  1. README.md +217 -45
README.md CHANGED
@@ -10,87 +10,259 @@ tags:
10
  - qwen3.8
11
  - cuda
12
  - rtx-5080
 
 
13
  - multimodal
14
  - vision
15
- - long-context
16
  ---
17
 
18
- # Qwen3.8-27B for NInfer — RTX 5080
 
 
 
 
 
 
 
 
 
 
 
19
 
20
- Official validated NInfer artifact for Qwen3.8-27B on the NVIDIA RTX 5080 16 GB profile.
21
 
22
- ## Artifact
 
 
23
 
24
  | Field | Value |
25
  |---|---|
26
- | Filename | `qwen3_8_27b.ninfer` |
27
  | Size | `16,461,267,456 bytes` |
28
  | SHA-256 | `c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21` |
29
- | NInfer model | Qwen3.8-27B |
30
- | Target GPU profile | RTX 5080 16 GB |
31
- | Effective main-model quantization | ~3.95 BPW |
32
- | Context | 131,072 |
33
  | KV capacity | 131,072 |
34
  | KV dtype | Q4 group64 |
35
  | Speculation | MTP-3 |
36
- | Vision | enabled |
37
- | Maximum validated Vision tokens | 2048 |
38
 
39
- ## Validated runtime
 
40
 
41
- The validated v1.3 production runtime is:
42
 
43
- `ceb32f7d002edab224a83a2e2609f45fca4f8919`
44
 
45
- Validated `ninfer-serve` SHA-256:
 
 
 
46
 
47
- `3179bfbcb88a72c04b983f28c25c62db468fbc8ef267fe043899de30a4281c56`
48
 
49
- Validated runtime profile:
50
 
51
- ```text
52
- --max-context 131072
53
- --kv-capacity 131072
54
- --prefill-chunk 896
55
- --kv-dtype q4
56
- --spec mtp
57
- --draft-tokens 3
58
- --no-cuda-graph
59
- --max-concurrency 1
60
- --vision
61
- --vision-max-tokens 2048
62
- --default-thinking-budget 2048
63
- --prefix-checkpoint-policy rolling-tool
64
- Verification
65
 
66
- After downloading:
67
 
68
- sha256sum qwen3_8_27b.ninfer
 
69
 
70
- Expected:
 
 
 
 
 
 
 
 
 
 
 
71
 
 
72
  c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21
73
- Source and provenance
74
 
75
- Canonical source repository:
76
 
77
- https://github.com/toddballinger/ninfer-5080
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78
 
79
- Validated model sources:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
80
 
81
  Qwen/Qwen3.8-27B
82
- revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
 
 
 
 
83
  z-lab/Qwen3.8-27B-DFlash2
84
- revision 50307d4c4cde6860d4eee73e2547cd786fe8e8a4
 
 
 
 
 
 
 
 
 
 
 
 
 
85
 
86
- The artifact above remains the canonical validated model artifact across subsequent NInfer runtime optimisations.
87
 
88
- Automated CPU-only reproduction and publication through GitHub Actions is being integrated separately. Until that workflow is merged and proves byte-identical reproduction, this repository should be treated as the official validated reference artifact.
 
89
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
90
  Credits
91
 
92
- NInfer upstream development: Neroued
 
93
 
94
- RTX 5080 optimisation, validation and release profile: Todd Ballinger
 
 
 
 
95
 
96
- Low-memory automated conversion work: starskyzheng
 
10
  - qwen3.8
11
  - cuda
12
  - rtx-5080
13
+ - long-context
14
+ - speculative-decoding
15
  - multimodal
16
  - vision
 
17
  ---
18
 
19
+ # Qwen3.8-27B for NInfer — RTX 5080 16 GB, true 128K + Vision
20
+
21
+ Project-maintained NInfer artifact for **Qwen3.8-27B** on a single
22
+ **NVIDIA RTX 5080 16 GB**, validated with:
23
+
24
+ - 131,072-token context
25
+ - 131,072-token KV capacity
26
+ - Q4 group64 KV
27
+ - MTP-3 speculative decoding
28
+ - Vision input
29
+ - mixed Q3/Q4/Q5 model quantization
30
+ - approximately 3.953 effective BPW for the main text model
31
 
32
+ Canonical source and validation records:
33
 
34
+ https://github.com/toddballinger/ninfer-5080
35
+
36
+ ## Official artifact
37
 
38
  | Field | Value |
39
  |---|---|
40
+ | File | `qwen3_8_27b.ninfer` |
41
  | Size | `16,461,267,456 bytes` |
42
  | SHA-256 | `c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21` |
43
+ | Format | NInfer native `.ninfer` |
44
+ | Target | Qwen3.8-27B |
45
+ | Primary hardware profile | RTX 5080 16 GB |
46
+ | Max context | 131,072 |
47
  | KV capacity | 131,072 |
48
  | KV dtype | Q4 group64 |
49
  | Speculation | MTP-3 |
50
+ | Vision | validated |
 
51
 
52
+ This file is intended for **NInfer**. It is not a Transformers checkpoint,
53
+ Safetensors distribution, or GGUF file.
54
 
55
+ ## Download
56
 
57
+ Using the Hugging Face CLI:
58
 
59
+ ```bash
60
+ hf download ninfer-5080/Qwen3.8-27B-RTX5080 \
61
+ qwen3_8_27b.ninfer \
62
+ --local-dir .
63
 
64
+ Verify the artifact:
65
 
66
+ sha256sum qwen3_8_27b.ninfer
67
 
68
+ Expected:
 
 
 
 
 
 
 
 
 
 
 
 
 
69
 
70
+ c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21 qwen3_8_27b.ninfer
71
 
72
+ A matching SHA-256 identifies the exact validated project artifact regardless
73
+ of the filename or the machine from which it was downloaded.
74
 
75
+ Model artifact vs runtime version
76
+
77
+ The model artifact and the NInfer runtime are versioned independently.
78
+
79
+ The artifact currently published here has remained byte-identical across
80
+ multiple later runtime optimizations. A newer NInfer runtime therefore does
81
+ not imply that a new .ninfer model file is required.
82
+
83
+ The canonical artifact identity is:
84
+
85
+ bytes:
86
+ 16461267456
87
 
88
+ SHA256:
89
  c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21
90
+ Validated v1.3 production runtime
91
 
92
+ Validated source commit:
93
 
94
+ ceb32f7d002edab224a83a2e2609f45fca4f8919
95
+
96
+ Validated ninfer-serve SHA-256:
97
+
98
+ 3179bfbcb88a72c04b983f28c25c62db468fbc8ef267fe043899de30a4281c56
99
+
100
+ v1.3 adds:
101
+
102
+ corrected Q4/Q4 strided attention output handling
103
+ server-wide default thinking-budget support
104
+ rolling tool checkpoints for agent/tool-loop workloads
105
+ preservation of the validated 131K/Q4-KV/MTP-3/Vision profile
106
+
107
+ Full release record:
108
+
109
+ https://github.com/toddballinger/ninfer-5080/blob/main/docs/RELEASE_QWEN3.8_27B_RTX5080_V1.3.md
110
+
111
+ Recommended serving profile
112
+
113
+ For more GPU-memory headroom, Vision 1792 is the recommended general profile:
114
+
115
+ ./ninfer-serve qwen3_8_27b.ninfer \
116
+ --host 0.0.0.0 \
117
+ --port 8080 \
118
+ --model-id qwen3.8-27b \
119
+ --max-context 131072 \
120
+ --kv-capacity 131072 \
121
+ --prefill-chunk 896 \
122
+ --kv-dtype q4 \
123
+ --spec mtp \
124
+ --draft-tokens 3 \
125
+ --no-cuda-graph \
126
+ --max-concurrency 1 \
127
+ --default-thinking-budget 2048 \
128
+ --prefix-checkpoint-policy rolling-tool \
129
+ --vision \
130
+ --vision-max-tokens 1792
131
+
132
+ Measured startup margin:
133
+
134
+ Vision workspace 115.7751 MiB
135
+ Free after startup 26.56 MiB
136
+ Planned slack 28.88 MiB
137
+ Maximum validated Vision profile
138
+
139
+ Vision 2048 is also validated and is the profile used by the v1.3 production
140
+ OpenClaw deployment:
141
+
142
+ --vision-max-tokens 2048
143
 
144
+ Measured startup margin:
145
+
146
+ Vision workspace 132.3142 MiB
147
+ Free after startup 8.56 MiB
148
+ Planned slack 10.08 MiB
149
+
150
+ This profile is intentionally tight. Use a clean GPU.
151
+
152
+ True-128K validation
153
+
154
+ The project does not describe a configuration as "true 128K" merely because
155
+ the configured maximum is 131,072.
156
+
157
+ The qualification workload contains an actual 118,001-token prompt while
158
+ retaining a full 131,072-token KV allocation.
159
+
160
+ A qualified feature-complete runtime produced:
161
+
162
+ Metric Result
163
+ Prompt tokens 118,001
164
+ Max context 131,072
165
+ KV capacity 131,072
166
+ Prefill 1378.85 tok/s
167
+ Decode 71.44 tok/s
168
+ MTP acceptance 44.74%
169
+ MTP acceptance length 2.31 tok/round
170
+
171
+ A later Q5 A16 LinearAdd semantic-port qualification produced 1376.30 tok/s
172
+ prefill and 71.53 tok/s decode on the same workload while retaining the same
173
+ model SHA.
174
+
175
+ Benchmark results are commit-scoped; see the GitHub validation ledger rather
176
+ than treating any one result as a floating "current" benchmark.
177
+
178
+ Multimodal validation
179
+
180
+ The final HostMapped Vision path has been validated for:
181
+
182
+ deterministic image understanding
183
+ deterministic video understanding
184
+ multi-image conversation history
185
+ cached historical-media accounting
186
+ coexistence with the full 131,072 text context/KV allocation
187
+
188
+ A synthetic red/blue image was correctly identified by side, and a deterministic
189
+ red → green → blue video was returned in the correct chronological order.
190
+
191
+ Details:
192
+
193
+ https://github.com/toddballinger/ninfer-5080/blob/main/docs/VISION_128K.md
194
+
195
+ Model sources
196
+
197
+ Target model:
198
 
199
  Qwen/Qwen3.8-27B
200
+ revision:
201
+ 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
202
+
203
+ DFlash2 source:
204
+
205
  z-lab/Qwen3.8-27B-DFlash2
206
+ revision:
207
+ 50307d4c4cde6860d4eee73e2547cd786fe8e8a4
208
+
209
+ Both upstream Hugging Face repositories currently declare Apache-2.0 licensing.
210
+
211
+ Quantization profile
212
+
213
+ Main text-core distribution:
214
+
215
+ Format Share
216
+ Q3G64_F16S 42.42%
217
+ Q4G64_F16S 45.92%
218
+ Q5G64_F16S 11.57%
219
+ BF16 / FP32 ~0.10%
220
 
221
+ Effective main-model quantization:
222
 
223
+ ~3.953 BPW
224
+ Reproducibility
225
 
226
+ This Hugging Face repository currently contains the validated reference
227
+ artifact.
228
+
229
+ A CPU-only GitHub Actions conversion/publishing workflow is being integrated
230
+ separately. Before any automated build is allowed to replace this artifact while
231
+ claiming byte-identical reproduction, it should reproduce both:
232
+
233
+ SIZE:
234
+ 16461267456
235
+
236
+ SHA256:
237
+ c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21
238
+
239
+ This prevents build automation from silently replacing a known-good model with
240
+ a different artifact.
241
+
242
+ Documentation
243
+ Project overview:
244
+ https://github.com/toddballinger/ninfer-5080
245
+ Validated manifest:
246
+ https://github.com/toddballinger/ninfer-5080/blob/main/docs/VALIDATED_MANIFEST.md
247
+ v1.3 release:
248
+ https://github.com/toddballinger/ninfer-5080/blob/main/docs/RELEASE_QWEN3.8_27B_RTX5080_V1.3.md
249
+ Vision:
250
+ https://github.com/toddballinger/ninfer-5080/blob/main/docs/VISION_128K.md
251
+ Reproducibility:
252
+ https://github.com/toddballinger/ninfer-5080/blob/main/docs/REPRODUCIBILITY.md
253
+ Benchmarks:
254
+ https://github.com/toddballinger/ninfer-5080/blob/main/docs/BENCHMARKS.md
255
+ Memory profile:
256
+ https://github.com/toddballinger/ninfer-5080/blob/main/docs/MEMORY_PROFILE.md
257
  Credits
258
 
259
+ This project builds on the work of the NInfer project and the Qwen/DFlash2
260
+ ecosystem.
261
 
262
+ NInfer upstream: Neroued and contributors
263
+ Qwen3.8-27B: Qwen team
264
+ DFlash2: z-lab / project contributors
265
+ RTX 5080 optimization, validation and release profile: Todd Ballinger
266
+ Low-memory automated conversion / CI work in progress: starskyzheng
267
 
268
+ Please preserve applicable upstream copyright, attribution and license notices.