File size: 12,985 Bytes
27939b8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a67c2ba
27939b8
a67c2ba
27939b8
a67c2ba
27939b8
a67c2ba
27939b8
a67c2ba
27939b8
a67c2ba
 
 
 
 
 
 
 
27939b8
a67c2ba
27939b8
a67c2ba
 
 
 
070155a
a67c2ba
070155a
a67c2ba
070155a
a67c2ba
070155a
a67c2ba
27939b8
a67c2ba
27939b8
a67c2ba
27939b8
a67c2ba
 
 
 
 
27939b8
a67c2ba
27939b8
a67c2ba
27939b8
a67c2ba
27939b8
a67c2ba
9327d8c
a67c2ba
27939b8
a67c2ba
9327d8c
 
a67c2ba
 
bb64bae
 
 
 
a67c2ba
 
 
 
 
bb64bae
a67c2ba
070155a
27939b8
 
a67c2ba
681923f
a67c2ba
88fd261
a67c2ba
88fd261
a67c2ba
88fd261
a67c2ba
88fd261
681923f
 
a67c2ba
681923f
 
a67c2ba
 
 
 
681923f
 
 
a67c2ba
4f881d4
a67c2ba
4f881d4
a67c2ba
2f0470e
a67c2ba
27939b8
a67c2ba
27939b8
a67c2ba
27939b8
bb64bae
27939b8
a67c2ba
27939b8
a67c2ba
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
---
license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:
  - Qwen/Qwen3.8-Flash-Next
  - Qwen/Qwen3.8-Flash-Next-FP8
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
inference: false
tags:
  - qwen
  - qwen3.8
  - qwen3.8-flash-next
  - gguf
  - llama.cpp
  - amd
  - rocm
  - gfx1151
  - ryzen-ai-max-395
  - strix-halo
  - mixture-of-experts
  - iu4
  - mtp
  - speculative-decoding
  - nvme
  - ple
  - long-context
  - local-inference
---

[![Qwen3.8 Flash CIRU Strix IU4](assets/qwen38-flash-ciru-strix-iu4.jpg)](https://llm.ciru.ai/research)

# Qwen3.8-Flash-CIRU-STRIX-IU4 · runtime v3.0.0

**V3 brings faster long-context serving and the qualified QSA conversation-isolation fixes to the Strix-only runner. Model weights are unchanged.** It adds parallel attention-cell selection, indexed decode attention, cached derived history, guarded PLE lookups and wider prefill. The general profile retains maximum MTP depth 6 and batch/microbatch 1024.

Use the [v3 source archive](runtime/v3.0.0/ciru-runtime-v3.0.0-source.tar.gz) or [matching GitHub tag](https://github.com/ciru-ai/Qwen3.8-Flash-CIRU-STRIX-IU4/tree/v3.0.0). This text-only package requires the custom CIRU runtime, the target GGUF and all three `ple/` files. The `mtp/` head enables speculative decoding. Stock llama.cpp and Hugging Face hosted inference do not support this package.

## V3 serving comparison

Measured on one Ryzen AI Max+ 395 / gfx1151 / 128 GB shared-memory NixOS machine, with one model workload at a time. Inputs are identical token IDs, cold prompt cache, 128 generated tokens, one slot and 262,144-token configured capacity. Nonthinking sampler: temperature 0.7, top-p 0.8, top-k 20, min-p 0, presence penalty 1.5, repeat penalty 1, frequency penalty 0 and seed 123. EOS is honored.

| Input tokens | Profile | Prompt tok/s | Generation tok/s | Whole request (s) |
| ---: | --- | ---: | ---: | ---: |
| 4,096 | Previous CIRU | 392.00 | 22.52 | 16.34 |
| 4,096 | CIRU v3 | 455.65 | 24.60 | 14.41 |
| 4,096 | Halo | 381.49 | 35.30 | 14.69 |
| 65,536 | Previous CIRU | 284.49 | 13.33 | 239.99 |
| 65,536 | CIRU v3 | 369.81 | 24.22 | 182.57 |
| 65,536 | Halo | 263.42 | 23.28 | 254.37 |

MTP 2 is a useful optional setting for the tested lower-acceptance long requests, but it was not promoted as the general default. Across 20 short HumanEval requests, v3 MTP 6 measured 53.24 generation tok/s versus 39.63 with MTP 2. Both passed 20/20 base and extended tests. Keeping MTP 6 avoids that short-coding regression; target verification retains the full vocabulary at either depth.

| Input tokens | Optional v3 MTP 2 prompt tok/s | Generation tok/s | Whole request (s) |
| ---: | ---: | ---: | ---: |
| 4,096 | 453.09 | 29.39 | 13.62 |
| 65,536 | 373.08 | 24.88 | 180.87 |

The previous CIRU arm is the locally qualified v2.0.1 runner under its original MTP 6, b2048/u512 profile. The v3 arm uses the same weights and MTP 6, with b1024/u1024. Halo is the unmodified current fork at commit `5f851647fe5ed795dfd6c0a3fba543114879e874`, using its recommended Vulkan backend, Unsloth UD-Q4_K_XL target and published EasiiX Strix Q8 MTP head. Native KV, batch, thread, fitting and cache defaults are retained.

Halo maximum depths 2, 3, 4, 6 and native adaptive 6 were screened. Depth 3 won its short-context screen at **35.37 tok/s**, versus **31.4** at depth 2, **30.04** at depth 4, **25.91** at depth 6 and **29.31** with adaptive 6. The final comparison above uses depth 3. Halo source and weights were not modified.

CIRU and Halo have different quantizations and execution profiles: this is a serving-package comparison. Generation throughput, prompt processing and whole-request latency are separate metrics. At 4K, the general MTP 6 profile is close to Halo in total time; the optional MTP 2 setting provides the clearer latency benefit on this fixture. Long-context prompt processing shows the larger gain. The tables retain Halo's generation advantage where present; a CIRU request-time win is not a claim of winning every metric. Three clean v3 loads, two selected Halo loads and one previous-CIRU load are a bounded experiment, not a confidence interval or general ranking.

The 2.79 GB Unsloth shared Q8 head intentionally omits tensors a supporting loader borrows from the main model. The pinned Halo loader fails for missing `token_embd.weight`; Unsloth's self-contained Q8 head also fails for missing `output_hc_norm.weight`. Both attempts are recorded. The compatible [EasiiX Strix Q8 head](https://huggingface.co/EasiiX/Qwen3.8-Flash-Next-MTP-Strix-Halo-GGUF/tree/6f7900648b1c6b14f067a182c640e47971e9ab35) is used as published.

[Full report, first-piece latency and memory](benchmarks/v3.0.0/COMPARISON.md) · [Structured results](benchmarks/v3.0.0/comparison.json) · [Raw evidence archive](benchmarks/v3.0.0/strix-v3.0.0-evidence.tar.gz)

## Quality and capacity checks

| Profile | HumanEval base | EvalPlus extended tests | Recall at about 8K and 64K |
| --- | ---: | ---: | --- |
| Previous CIRU | 20/20 | 20/20 | Both keys and exact cached replay |
| CIRU v3 | 20/20 | 20/20 | Both keys and exact cached replay |
| Halo | 20/20 | 20/20 | Both keys and exact cached replay |

These are canonical HumanEval tasks 0–19, EvalPlus v0.1.10, one first sample per task, no retries and a 4096-token cap; truncations fail. Generated code runs inside a filesystem/network sandbox. This small nonthinking coding and recall panel is a regression check. It does not establish broad model equality, thinking-mode quality, tool reliability or leaderboard standing.

V3 also completed **261,888 input tokens plus 128 generated tokens** at **257.44 prompt tok/s and 18.00 generation tok/s**, with a **1024.44 s** whole request. This is a CIRU-only serving-capacity check, not a filled-256K Halo comparison or full-context accuracy result.

The final inference source passes 69 QSA mapping/state/guard cases with flags on and off, 33 actual ROCm operator reference cases and the existing 30 batch allocator tests. A separate short four-prefix diagnostic matches 15,892,480 F32 logits byte-for-byte against the previous runner. Radix selection can change threshold-tie membership and selected-list order; general bitwise equivalence is not claimed.

## Download, build and run

The unchanged model files total **135,962,881,135 bytes (126.625 GiB)**, excluding runtime and reports. The target is 73.945 GiB, the PLE payload 48.828 GiB and the Q8 MTP head 3.852 GiB. Allow at least 160 GiB for model storage/verification plus SDK and build space. Use fast NVMe and a Strix Halo machine with 128 GiB unified memory.

On Ubuntu/Debian, install Git and a Python virtual environment, then download the model and build the source:

```bash
sudo apt-get update
sudo apt-get install -y git python3-venv
python3 -m venv .venv-hf
.venv-hf/bin/python -m pip install -U huggingface_hub
. .venv-hf/bin/activate
hf download jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 \
  --revision v3.0.0 --local-dir ./model
(cd model && sha256sum -c checksums.sha256 && sha256sum -c v3.0.0-checksums.sha256)
git clone --branch v3.0.0 --single-branch \
  https://github.com/ciru-ai/Qwen3.8-Flash-CIRU-STRIX-IU4.git ciru-runtime-v3.0.0
cd ciru-runtime-v3.0.0
./scripts/ciru/setup-linux-amd.sh --install-host-deps
BUILD_DIR="$PWD/build-gfx1151-sdk" MODEL_DIR="$(realpath ../model)" \
  ./scripts/ciru/run-server.sh
```

The helper installs a private complete ROCm 10.0.0 SDK with gfx1151 libraries. Keep the SDK directory used by the build; the host needs a compatible AMD driver and access to `/dev/kfd` and its render node. New v3 GPU qualification is NixOS/ROCm 10/gfx1151. The prior clean Ubuntu qualification belongs to v2.0.1, so the Ubuntu instructions are retained build guidance, not a new v3 Ubuntu test result. See [platform/build instructions](https://github.com/ciru-ai/Qwen3.8-Flash-CIRU-STRIX-IU4/blob/v3.0.0/docs/BUILD_LINUX.md).

Existing users can keep their model directory and clone/build only the new runtime. The source archive is equivalent to the Git tag including executable modes and symlinks. The optional [tested NixOS binary payload](runtime/v3.0.0/ciru-runtime-v3.0.0-nixos-gfx1151.tar.gz) requires the recorded Nix store and ROCm SDK paths; use the source build for another installation.

The launcher enables 262,144 context capacity, F16 target KV, Q8 draft KV, the 32,768-row draft shortlist, maximum MTP depth 6, b1024/u1024, eight CPU threads and prefix caching. MTP uses `--parallel 1`; multi-slot MTP is rejected before model load. For target-only parallel serving, set `ENABLE_MTP=0` and follow the [parallel instructions](https://github.com/ciru-ai/Qwen3.8-Flash-CIRU-STRIX-IU4/blob/v3.0.0/docs/RUNNING.md#parallel-requests-and-unified-kv-cache).

Confirm `CIRU MTP shortlist enabled: 32768 / 248320 vocabulary rows; full target verification retained` in the startup log. The shortlist limits draft projection; target verification retains the full vocabulary. `MTP_DEPTH=2` selects the tested option for low-acceptance long requests; the screen does not establish the optimum for every prompt. New v3 optimization switches accept literal `0`. The optional draft attention window remains off and is unqualified when enabled.

Thinking mode remains the model default: temperature 1.0, top-p 0.95, top-k 20, min-p 0. For the nonthinking mode evaluated here:

```bash
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.8-Flash-CIRU-STRIX-IU4",
    "messages": [{"role": "user", "content": "Write a Python CSV validator with tests."}],
    "chat_template_kwargs": {"enable_thinking": false},
    "temperature": 0.7, "top_p": 0.8, "top_k": 20,
    "min_p": 0, "presence_penalty": 1.5, "cache_prompt": true
  }'
```

These sampling defaults follow the [Qwen model card](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/f5d08274bafd880402bd16f5e3e6c514136ec06c/README.md#best-practices). Ordinary chat can retain prefix caching; the benchmark's cold requests, fixed seed and output cap are measurement controls.

## What ships in v3

V3 combines the previously qualified QSA sequence isolation/indexer-copy fixes with the portable retained hybrid work: wider aligned QSA prefill, J32 expert tiling for supported prefill shapes, radix selection, indexed F16 attention, direct guarded PLE lookup, derived-history caching keyed by sequence and packed tiny F32 gathers. The F32 derived cache adds about 192 MiB at 256K. Hybrid external-GPU ownership, peer transfers, expert splitting, scheduler overlap and experimental compact masks are excluded.

The original READY package, source, evidence and checksum trees are included in the [prior-package archive](benchmarks/v3.0.0/qsa-v2.0.1-prior-package.tar.gz). Its v2.0.1 fixes ship as part of v3; no separate public v2.0.1 tag is claimed. Existing public v2.0 tags and weight identities remain unchanged. [Historical v2.0 results](https://huggingface.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4/blob/v2.0/README.md) remain available; their 42.3 tok/s short greedy probe and broader older quality suites are different workloads and were not rerun as full suites for v3.

The IU4 model name is retained. The target has 1,223 tensors: 144 Q4_1 routed-expert tensors, 328 Q5_K, 290 Q8_0, 48 Q5_1, 25 BF16 and 388 F32. The standard launcher uses ordinary GGUF types; its Q4_1 matrix path expands packed values into byte lanes for IU8 WMMA. It does not activate the separate native IU4/E3 bank path. The PLE payload preserves exact FP8 E4M3 weights with one BF16 scale.

[Model file tree](https://huggingface.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4/tree/v3.0.0) · [Weight checksums](checksums.sha256) · [V3 runtime/report checksums](v3.0.0-checksums.sha256) · [Source identity](benchmarks/v3.0.0/git-source.json) · [GitHub release](https://github.com/ciru-ai/Qwen3.8-Flash-CIRU-STRIX-IU4/releases/tag/v3.0.0)

## Lineage, license and credit

Text lineage is [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/tree/f5d08274bafd880402bd16f5e3e6c514136ec06c); PLE lineage is [Qwen3.8-Flash-Next-FP8](https://huggingface.co/Qwen/Qwen3.8-Flash-Next-FP8/tree/bcd9f01ddc9cff2316eb84281bebcd5b058bddce). Runtime lineage starts from [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp/commit/f5e85d43a048f3d5adefb4c5e29867d8077fba62). Model artifacts use the included Qwen Community License 1.0; runtime code retains MIT and component notices.

Credits include Qwen, ggml-org, Ryan Monsurate's MTP integration, AMD's ROCm ecosystem and the contributors in the source notices. Daniel Han's authorship of the imported QSA repair is preserved. Thanks to halo-box, EasiiX, Unsloth, Daniel Han Chen, Laurent Zuijdwijk and Agention AI for public models/runtimes, and the HumanEval and EvalPlus authors for evaluation tools. CIRU is an independent community research project; AMD and Qwen marks do not imply sponsorship.