jcbtc's picture
Release CIRU Strix runtime v3.0.0 with qualified QSA and serving improvements
a67c2ba verified
|
Raw
History Blame
13 kB
metadata
license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:
  - Qwen/Qwen3.8-Flash-Next
  - Qwen/Qwen3.8-Flash-Next-FP8
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
inference: false
tags:
  - qwen
  - qwen3.8
  - qwen3.8-flash-next
  - gguf
  - llama.cpp
  - amd
  - rocm
  - gfx1151
  - ryzen-ai-max-395
  - strix-halo
  - mixture-of-experts
  - iu4
  - mtp
  - speculative-decoding
  - nvme
  - ple
  - long-context
  - local-inference

Qwen3.8 Flash CIRU Strix IU4

Qwen3.8-Flash-CIRU-STRIX-IU4 · runtime v3.0.0

V3 brings faster long-context serving and the qualified QSA conversation-isolation fixes to the Strix-only runner. Model weights are unchanged. It adds parallel attention-cell selection, indexed decode attention, cached derived history, guarded PLE lookups and wider prefill. The general profile retains maximum MTP depth 6 and batch/microbatch 1024.

Use the v3 source archive or matching GitHub tag. This text-only package requires the custom CIRU runtime, the target GGUF and all three ple/ files. The mtp/ head enables speculative decoding. Stock llama.cpp and Hugging Face hosted inference do not support this package.

V3 serving comparison

Measured on one Ryzen AI Max+ 395 / gfx1151 / 128 GB shared-memory NixOS machine, with one model workload at a time. Inputs are identical token IDs, cold prompt cache, 128 generated tokens, one slot and 262,144-token configured capacity. Nonthinking sampler: temperature 0.7, top-p 0.8, top-k 20, min-p 0, presence penalty 1.5, repeat penalty 1, frequency penalty 0 and seed 123. EOS is honored.

Input tokens Profile Prompt tok/s Generation tok/s Whole request (s)
4,096 Previous CIRU 392.00 22.52 16.34
4,096 CIRU v3 455.65 24.60 14.41
4,096 Halo 381.49 35.30 14.69
65,536 Previous CIRU 284.49 13.33 239.99
65,536 CIRU v3 369.81 24.22 182.57
65,536 Halo 263.42 23.28 254.37

MTP 2 is a useful optional setting for the tested lower-acceptance long requests, but it was not promoted as the general default. Across 20 short HumanEval requests, v3 MTP 6 measured 53.24 generation tok/s versus 39.63 with MTP 2. Both passed 20/20 base and extended tests. Keeping MTP 6 avoids that short-coding regression; target verification retains the full vocabulary at either depth.

Input tokens Optional v3 MTP 2 prompt tok/s Generation tok/s Whole request (s)
4,096 453.09 29.39 13.62
65,536 373.08 24.88 180.87

The previous CIRU arm is the locally qualified v2.0.1 runner under its original MTP 6, b2048/u512 profile. The v3 arm uses the same weights and MTP 6, with b1024/u1024. Halo is the unmodified current fork at commit 5f851647fe5ed795dfd6c0a3fba543114879e874, using its recommended Vulkan backend, Unsloth UD-Q4_K_XL target and published EasiiX Strix Q8 MTP head. Native KV, batch, thread, fitting and cache defaults are retained.

Halo maximum depths 2, 3, 4, 6 and native adaptive 6 were screened. Depth 3 won its short-context screen at 35.37 tok/s, versus 31.4 at depth 2, 30.04 at depth 4, 25.91 at depth 6 and 29.31 with adaptive 6. The final comparison above uses depth 3. Halo source and weights were not modified.

CIRU and Halo have different quantizations and execution profiles: this is a serving-package comparison. Generation throughput, prompt processing and whole-request latency are separate metrics. At 4K, the general MTP 6 profile is close to Halo in total time; the optional MTP 2 setting provides the clearer latency benefit on this fixture. Long-context prompt processing shows the larger gain. The tables retain Halo's generation advantage where present; a CIRU request-time win is not a claim of winning every metric. Three clean v3 loads, two selected Halo loads and one previous-CIRU load are a bounded experiment, not a confidence interval or general ranking.

The 2.79 GB Unsloth shared Q8 head intentionally omits tensors a supporting loader borrows from the main model. The pinned Halo loader fails for missing token_embd.weight; Unsloth's self-contained Q8 head also fails for missing output_hc_norm.weight. Both attempts are recorded. The compatible EasiiX Strix Q8 head is used as published.

Full report, first-piece latency and memory · Structured results · Raw evidence archive

Quality and capacity checks

Profile HumanEval base EvalPlus extended tests Recall at about 8K and 64K
Previous CIRU 20/20 20/20 Both keys and exact cached replay
CIRU v3 20/20 20/20 Both keys and exact cached replay
Halo 20/20 20/20 Both keys and exact cached replay

These are canonical HumanEval tasks 0–19, EvalPlus v0.1.10, one first sample per task, no retries and a 4096-token cap; truncations fail. Generated code runs inside a filesystem/network sandbox. This small nonthinking coding and recall panel is a regression check. It does not establish broad model equality, thinking-mode quality, tool reliability or leaderboard standing.

V3 also completed 261,888 input tokens plus 128 generated tokens at 257.44 prompt tok/s and 18.00 generation tok/s, with a 1024.44 s whole request. This is a CIRU-only serving-capacity check, not a filled-256K Halo comparison or full-context accuracy result.

The final inference source passes 69 QSA mapping/state/guard cases with flags on and off, 33 actual ROCm operator reference cases and the existing 30 batch allocator tests. A separate short four-prefix diagnostic matches 15,892,480 F32 logits byte-for-byte against the previous runner. Radix selection can change threshold-tie membership and selected-list order; general bitwise equivalence is not claimed.

Download, build and run

The unchanged model files total 135,962,881,135 bytes (126.625 GiB), excluding runtime and reports. The target is 73.945 GiB, the PLE payload 48.828 GiB and the Q8 MTP head 3.852 GiB. Allow at least 160 GiB for model storage/verification plus SDK and build space. Use fast NVMe and a Strix Halo machine with 128 GiB unified memory.

On Ubuntu/Debian, install Git and a Python virtual environment, then download the model and build the source:

sudo apt-get update
sudo apt-get install -y git python3-venv
python3 -m venv .venv-hf
.venv-hf/bin/python -m pip install -U huggingface_hub
. .venv-hf/bin/activate
hf download jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 \
  --revision v3.0.0 --local-dir ./model
(cd model && sha256sum -c checksums.sha256 && sha256sum -c v3.0.0-checksums.sha256)
git clone --branch v3.0.0 --single-branch \
  https://github.com/ciru-ai/Qwen3.8-Flash-CIRU-STRIX-IU4.git ciru-runtime-v3.0.0
cd ciru-runtime-v3.0.0
./scripts/ciru/setup-linux-amd.sh --install-host-deps
BUILD_DIR="$PWD/build-gfx1151-sdk" MODEL_DIR="$(realpath ../model)" \
  ./scripts/ciru/run-server.sh

The helper installs a private complete ROCm 10.0.0 SDK with gfx1151 libraries. Keep the SDK directory used by the build; the host needs a compatible AMD driver and access to /dev/kfd and its render node. New v3 GPU qualification is NixOS/ROCm 10/gfx1151. The prior clean Ubuntu qualification belongs to v2.0.1, so the Ubuntu instructions are retained build guidance, not a new v3 Ubuntu test result. See platform/build instructions.

Existing users can keep their model directory and clone/build only the new runtime. The source archive is equivalent to the Git tag including executable modes and symlinks. The optional tested NixOS binary payload requires the recorded Nix store and ROCm SDK paths; use the source build for another installation.

The launcher enables 262,144 context capacity, F16 target KV, Q8 draft KV, the 32,768-row draft shortlist, maximum MTP depth 6, b1024/u1024, eight CPU threads and prefix caching. MTP uses --parallel 1; multi-slot MTP is rejected before model load. For target-only parallel serving, set ENABLE_MTP=0 and follow the parallel instructions.

Confirm CIRU MTP shortlist enabled: 32768 / 248320 vocabulary rows; full target verification retained in the startup log. The shortlist limits draft projection; target verification retains the full vocabulary. MTP_DEPTH=2 selects the tested option for low-acceptance long requests; the screen does not establish the optimum for every prompt. New v3 optimization switches accept literal 0. The optional draft attention window remains off and is unqualified when enabled.

Thinking mode remains the model default: temperature 1.0, top-p 0.95, top-k 20, min-p 0. For the nonthinking mode evaluated here:

curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.8-Flash-CIRU-STRIX-IU4",
    "messages": [{"role": "user", "content": "Write a Python CSV validator with tests."}],
    "chat_template_kwargs": {"enable_thinking": false},
    "temperature": 0.7, "top_p": 0.8, "top_k": 20,
    "min_p": 0, "presence_penalty": 1.5, "cache_prompt": true
  }'

These sampling defaults follow the Qwen model card. Ordinary chat can retain prefix caching; the benchmark's cold requests, fixed seed and output cap are measurement controls.

What ships in v3

V3 combines the previously qualified QSA sequence isolation/indexer-copy fixes with the portable retained hybrid work: wider aligned QSA prefill, J32 expert tiling for supported prefill shapes, radix selection, indexed F16 attention, direct guarded PLE lookup, derived-history caching keyed by sequence and packed tiny F32 gathers. The F32 derived cache adds about 192 MiB at 256K. Hybrid external-GPU ownership, peer transfers, expert splitting, scheduler overlap and experimental compact masks are excluded.

The original READY package, source, evidence and checksum trees are included in the prior-package archive. Its v2.0.1 fixes ship as part of v3; no separate public v2.0.1 tag is claimed. Existing public v2.0 tags and weight identities remain unchanged. Historical v2.0 results remain available; their 42.3 tok/s short greedy probe and broader older quality suites are different workloads and were not rerun as full suites for v3.

The IU4 model name is retained. The target has 1,223 tensors: 144 Q4_1 routed-expert tensors, 328 Q5_K, 290 Q8_0, 48 Q5_1, 25 BF16 and 388 F32. The standard launcher uses ordinary GGUF types; its Q4_1 matrix path expands packed values into byte lanes for IU8 WMMA. It does not activate the separate native IU4/E3 bank path. The PLE payload preserves exact FP8 E4M3 weights with one BF16 scale.

Model file tree · Weight checksums · V3 runtime/report checksums · Source identity · GitHub release

Lineage, license and credit

Text lineage is Qwen/Qwen3.8-Flash-Next; PLE lineage is Qwen3.8-Flash-Next-FP8. Runtime lineage starts from ggml-org/llama.cpp. Model artifacts use the included Qwen Community License 1.0; runtime code retains MIT and component notices.

Credits include Qwen, ggml-org, Ryan Monsurate's MTP integration, AMD's ROCm ecosystem and the contributors in the source notices. Daniel Han's authorship of the imported QSA repair is preserved. Thanks to halo-box, EasiiX, Unsloth, Daniel Han Chen, Laurent Zuijdwijk and Agention AI for public models/runtimes, and the HumanEval and EvalPlus authors for evaluation tools. CIRU is an independent community research project; AMD and Qwen marks do not imply sponsorship.