How to use from
OpenClaw
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:" \
  --custom-provider-id llama-cpp \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored

A ~3 bpw GGUF quantization of the OrcaRouter Qwen3.8-27B, carrying an embedded MTP draft head and a custom OrcaRouter-native imatrix. Built to run on a single 16 GB GPU.

This is a quantization of the OrcaRouter Qwen3.8-27B checkpoint (a build on Qwen/Qwen3.8-27B) — not a new fine-tune of the trunk. The trunk weights are quantized. Quantization follows the GSQ/RCO methodology from ISTA-DASLab, with an imatrix calibrated on an OrcaRouter-native corpus. The distributed artifact also contains an embedded MTP draft head.

Release status: v2.1 is the recommended release candidate. Two bounded paired runs on unchanged build2 completed (§12): v2.0 and v2.1 each passed the 66,316-token centered-needle probe twice. These runs did not reproduce the earlier v2.0 collapse on that configuration and does not establish a general fix or broad release qualification.

Measurement discipline (read first)

Every measured row is stamped with context, flags, engine (build-kv / build3), and date. Rows from different configurations are never merged into one number. In particular: short-context gate numbers and full-history long-context numbers are different workloads and are shown in separate tables.

There are two distinct long-context regimes, and they are not interchangeable:

  • Full-history, correctness-qualified (the honest baseline): every prior token resident, no eviction. This is the row we recommend for long-context work because retrieval is intact.
  • KVMem window (speed-only): evicts early/mid history. Faster, but retrieval fails (see §4C). It is documented for transparency, not as a recommended configuration.

1. TL;DR

Recommended file Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf
Size / SHA-256 10,382,717,184 bytes (9.67 GiB) · AB955B5083D9CDF0BF55C37ACDCAE359B78756C4544D97960C23D8FCA98FEB9B
Trunk IQ3_XXS, 3.04 bpw
Draft head Embedded MTP, blk.64, 15 tensors: seven weight tensors Q6_K, attn_output.weight IQ4_XS, seven F32 norm tensors
Base OrcaRouter Qwen3.8-27B checkpoint (build of Qwen/Qwen3.8-27B)
Template froggeric-qwen3.8-tool-use.jinja v22.5, baked into the GGUF (--jinja alone selects it)
Modality text + tools in this file; vision needs the separate mmproj/ projector
Long context, correctness-qualified full-history 196K resident, q4_0 K / q4_0 V, spec OFF, KVMem OFF → decode 33.7 t/s, prefill 662 t/s, NIAH 3/3 depths PASS ({build-kv fa5994f5c, 2026-10-02})
Short-context gate (build3) short-context decode probe 106.4 t/s (input-token count not retained — not a decode-after-16K-ingest figure) · prefill 1473 t/s · separate ~16,053-token NIAH 6/6 · tool 13/14 · coherence 4/4 · MTP accept 0.91

The two headline rows above are different workloads. 33.7 t/s is the correctness-qualified full-history 196K number. 106.4 t/s is a short-context decode probe (the gate JSON retains three speed samples but not the probe's input-token count), and the 6/6 NIAH is a separate ~16,053-token haystack. Do not read 106.4 as decoding after ingesting 16K, and do not quote one row as the other.

2. What this is

We quantized the OrcaRouter checkpoint to a single-file ~3 bpw GGUF a 16 GB consumer card can serve with useful context and speculative decoding. Two design choices matter:

  • GSQ/RCO allocation. We reproduce ISTA-DASLab's GSQ/RCO per-tensor allocation map and apply it to the OrcaRouter base. The full 866-row map ships as REF-IQ3_XXS-mtp.rco-allocation.txt.
  • An embedded MTP draft head. Its 15 tensors include seven Q6_K weight tensors, attn_output.weight in IQ4_XS, and seven F32 norm tensors, so --spec-type draft-mtp gives built-in speculation with no second file.

Hardware (stated once): every number was measured on a single RTX 5070 Ti 16 GB (GB203, sm_120, 70 SM), Windows, CUDA 13.3. Metrics are single-GPU, single-build.

2b. Architecture (verified from the GGUF)

Property Value
Architecture qwen35 (Qwen3_5ForConditionalGeneration)
Layers 64 transformer = 16 full-attention + 48 GDN/linear-attention; +1 MTP block (block_count 65)
Attention head_count 24, head_count_kv 4 → GQA 6; head_dim 256; ctx 262144
GDN state_size 128, conv_kernel 4, key-heads (group_count) 16, value-heads (time_step_rank) 48 → GVA group 3, inner_size 6144
Hidden / FFN 5120 / 17408; vocab 248320
MTP qwen35.nextn_predict_layers = 1 → blk.64.nextn.* (Q6_K) + enorm/hnorm F32

Note: the GDN state dimension is 128, not the attention head dim (256). They are different subsystems.

3. Changelog

  • v2.1 (current, recommended). Targets the previously observed >64K generation collapse. Root cause: two FFN tensors in the per-tensor GSQ/RCO allocation were too aggressive — ffn_gate.weight = IQ1_S and ffn_up.weight = IQ1_M. Raising exactly those two addresses that observed failure while landing at 3.04 bpw (half the map unchanged). Evidence: the gates-only path collapsed; the full 6-sensitive-tensor fix at 3.78 bpw still collapsed; a plain IQ3_XXS mix passed; only the FFN change to IQ3_XXS passed. Anchored true tokenizer name: Staged_Tmpl. In the paired frozen-build2 smoke, both v2.0 and v2.1 passed one identical 66,316-token probe; that smoke does not distinguish versions or establish a general fix.
  • v2.0. OrcaRouter base + S1 MTP head + custom imatrix + RCO allocation. Superseded — earlier >64K collapse observation was not reproduced in the paired build2 smoke. Retained for reproducibility.
  • v1.x (Huihui line). Frozen / deprecated.

4. Benchmarks

4A. Correctness-qualified full-history 196K baseline (the recommended long-context config)

Metric Value Context Flags Engine Date
Decode 33.7 t/s (33.6 / 33.91 / 33.7 / 33.8) 196,044–196,068 tokens resident -ctk q4_0 -ctv q4_0, no spec, KVMEM_ENABLE=0, -c 262144 -fa on -ngl 999 -b 512 -ub 512 build-kv fa5994f5c 2026-10-02
Prefill 662–664 t/s same same build-kv fa5994f5c 2026-10-02
NIAH recall 3/3 PASS (depths 0.1 / 0.5 / 0.9) full 196K same build-kv fa5994f5c 2026-10-02

Profile: full-history 196K, build-kv fa5994f5c, 2026-10-02. Depth ID: each answer quotes TEAL-FALCON-4402 from the full history; no empty output, no collapse.

MTP on this exact baseline (A1 finding, separate rows): --spec-type draft-mtp --spec-draft-n-max 2 @full-KV 196K → decode 27.0–28.2 t/s on open prose (accept ~0.47) — i.e. a regression vs 33.7; on the recall prompts it rose to 39.1–39.6 t/s (accept 0.82–0.92) and prefill fell to ~166 t/s. MTP value is workload-dependent on this baseline. (A1_MTP_FINDING_20261002.md.)

4B. Short-context gate (build3) — decode speed probe + separate ~16K NIAH

# Metric Value Context Flags Engine Date
1 Decode (short-context probe; prior recorded profile) 106.4 t/s short-context; input-token count not retained -fa on -ctk q8_0 -ctv q4_0 --spec-type draft-mtp --spec-draft-n-max 2 build3 19fff3e38 2026-10-01
2 Prefill 1473.4 t/s 16K same build3 2026-10-01
3 Needle (NIAH) 6/6 ~16K haystack, depths 0.1/0.5/0.9 same build3 2026-10-01
4 Tool-call JSON 13/14 14-case suite same build3 2026-10-01
5 Coherence 4/4 4-case rubric same build3 2026-10-01
6 MTP accept 0.91 16K embedded MTP, -n-max 2 build3 2026-10-01

4C. KVMem-window config — documented for transparency, NOT recommended

Metric Value Context Flags Recall Engine
Decode 74.5 t/s (71.7–94.5) 196,068 resident draft-mtp,ngram-mod, KVMem on, budget 2048 2/6 FAIL (depths 0.1, 0.5 missed) build3

The historic 74.5 t/s @196K and the recall failure are the same configuration: the KVMem budget (2048) evicts early/mid KV, so the deep-history needles are lost. Against the 33.7 baseline two variables differ at once (spec and KVMem) — the whole gap must not be attributed to either. If you need 196K retrieval, use 4A; if you need speed on a short effective history, a window is a conscious trade. No faster full-history 196K configuration is qualified in these results (stated without causal attribution: both variables — KVMem budget 2048 and MTP/ngram — differ from the baseline).

5. Quality gate (16K short-context, build3)

Check Result Floor Verdict
Decode 106.4 t/s 94.0 PASS
Prefill 1473.4 t/s 1377.9 PASS
Needle (~16K, 6 cases) 6/6 6 PASS
Tool-call (14 cases) 13/14 12 PASS
Coherence (4 cases) 4/4 4 PASS
MTP accept 0.91 0.897 PASS

The single tool-call failure is tc-07; reported rather than rounded up.

Quality suites — historical v2.0 values; no v2.1 scores

The v2.1 fix touched FFN gates, so these are carried forward for reference only:

Suite Score
GPQA-Diamond 75.25% (198 q, harness-conditional)
WikiText-2 perplexity 6.17
IFEval 70.24 / 76.26
TruthfulQA MC1 77.60 / MC2 81.98
XSTest-safe over-refusal 0% (0/250)

6. MTP

Embedded mixed-type head (blk.64, seven Q6_K weight tensors, attn_output.weight IQ4_XS, seven F32 norm tensors): --spec-type draft-mtp --spec-draft-n-max 2. No second file. Two acceptance regimes, both stated: 0.91 at the 16K short-context gate; ~0.47–0.67 at long context (workload-dependent, see 4A). MTP output may differ from serial generation on a quantized target — fix the seed, and disable speculation when exact serial behavior matters.

7. Long context — honest scope

  • Full-history 196K (correctness-qualified): decode 33.7 t/s, prefill 662 t/s, recall 3/3 PASS (§4A). This is the number to quote.
  • KVMem window: faster (74.5 t/s) but retrieval fails (2/6) — not recommended (§4C).
  • 66K paired smoke: both v2.0 and v2.1 returned the embedded needle with multi-token output from the same 66,316-token prompt on frozen build2. The earlier v2.0 collapse was not reproduced in this paired run; one centered-depth probe does not establish a general fix. The separate 16K and 196K recall rows are documented with their own engine/configuration below.
  • Coverage: one qualified full-history point at ~196K with three tested needle depths. Other context lengths through 256K, and repeated long-context reliability, remain untested. The paired 66K smoke passed for both versions; neither result establishes broad long-context reliability.
  • The 6/6 needle gate is a ~16K haystack test, not a 196K/256K test.

8. TTS co-residency (separate configuration — do not merge with §4)

The 27B at 256K and a Qwen3-TTS-1.7B custom-voice talker fit together on the same 16 GB card: the 27B at 256K measured 77.2 t/s with the talker, in 14,751 MiB used / 1,245 MiB free. Voice synthesis ~0.7–1.2 s for ~5–9 s of audio. This 77.2 t/s is a distinct configuration (shared card, talker resident) and must not be read as the standalone long-context decode number (§4A is). Reproducibility caveat: exact LLM flags, effective input length, KVMem settings, the TTS binary/backend, and a recorded CPU-vs-CUDA launcher discrepancy are not fully resolved in the artifacts — treat this row as not-yet-reproducible until they are recovered. Do not guess them.

# 1) 27B text server
llama-server -m Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf \
  --jinja --ctx-size 262144 -b 512 -ub 512 -fa on -ngl 99 -ctk q8_0 -ctv q4_0 \
  --spec-type draft-mtp --spec-draft-n-max 2 --port 8080
# 2) Qwen3-TTS-1.7B talker
qwentts.cpp tts-server --talker Qwen3-TTS-1.7B --voice serena --port 8081

9. Quickstart

A. Default — MTP, 32K:

llama-server -m Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf \
  --jinja --ctx-size 32768 -b 512 -ub 512 -fa on -ngl 99 -ctk q8_0 -ctv q4_0 \
  --spec-type draft-mtp --spec-draft-n-max 2 --port 8080

B. Full-history long context (recall intact; the qualified 196K config):

llama-server -m Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf \
  --port 8080 -np 1 --jinja -ngl 999 -c 262144 -b 512 -ub 512 -fa on \
  -ctk q4_0 -ctv q4_0 --cache-ram 0 -t 8
# no --spec-type, KVMEM_ENABLE=0. Headroom note: reduce -c to <=216193 for >=1536 MiB margin.

The Froggeric template is baked in (tokenizer.chat_template, v22.5) — --jinja alone selects it.

Vision (optional): text-only works without it; add mmproj/mmproj-Qwen3.8-27B-BF16.gguf.

10. Known issues

  1. Historical >64K collapse observation — v2.1 raises two FFN tensors to IQ3_XXS; however, the paired frozen-build2 66,316-token smoke passed for both v2.0 and v2.1. This run does not reproduce the failure or establish a version-specific fix; see §12.
  2. Two long-context regimes; do not merge — 33.7 t/s is full-history/recall-3/3; 74.5 t/s is KVMem/recall-2/6.
  3. MTP has two regimes — 0.91 (16K gate) vs ~0.47–0.67 (long context); and on the full-KV 196K baseline MTP regressed open prose (27 vs 33.7). Not a universal speedup.
  4. Long-context retrieval not gated across 66K–256K; the 6/6 gate is ~16K.
  5. Aggressive ~3-bit quant — knowledge recall is weaker than higher-bit variants; the model leans on retrieval/tools.
  6. Deferred v2.1 suites: GPQA-Diamond (198 questions), IFEval (541 items), WikiText-2 PPL, TruthfulQA, XSTest-safe, and the large-target-length matrix were NOT RUN in this release review. Datasets were acquired; scores were not produced. Historical values below belong to v2.0 only.
  7. Engine-sensitive KV (scoped to the tested binaries only) — on the tested build2, the q8_0-K path fell back to a slow kernel; on the tested build3 19fff3e38, it is full speed. This is not a claim about every build bearing those names. Use the tested build3 with -ctk q8_0 -ctv q4_0, or q4_0/q4_0 on the tested build2.
  8. RCO allocation reproduced from ISTA-DASLab's published map (not re-searched), applied to the uncensored base. The imatrix is custom OrcaRouter-native.

Uncensored-use notice

Abliterated model: refusal is substantially reduced by design. XSTest-safe measures over-refusal (0/250) — a compliance-sensitivity check, not a safety or adversarial-robustness evaluation. The deployer is responsible for compliant, lawful use.

11. Reproduction appendix

  • Qualified baseline: Den build-kv, rev fa5994f5c (build 10821); ggml-cuda.dll SHA256[:16] 86CEFBB7197436AC. Command in §9B.
  • 16K gate: build3 commit 19fff3e38; flags in §4B.
  • SHAs: primary file AB955B5083D9CDF0BF55C37ACDCAE359B78756C4544D97960C23D8FCA98FEB9B (10,382,717,184 bytes). Per-file hashes in SHA256SUMS.txt.
  • Quantization (abridged):
    llama-quantize --imatrix imatrix.dat \
      --tensor-type-file REF-IQ3_XXS-mtp.rco-allocation.txt \
      model-f16.gguf Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf IQ3_XXS
    
  • Seeds: use fixed seeds for comparisons; the paired build2 profile used seed 12345.
File What
Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf recommended release candidate, 3.04 bpw; paired 66K functional smoke passed twice; full quality qualification deferred
mmproj/mmproj-Qwen3.8-27B-BF16.gguf optional vision projector
REF-IQ3_XXS-mtp.rco-allocation.txt 866-row per-tensor allocation map
imatrix.dat custom OrcaRouter-native calibration matrix
SHA256SUMS.txt per-file hashes

12. Paired frozen-build2 smoke (bounded; 2026-10-02)

The full base profile 173016 and repeat 182907 ran both versions serially on unchanged build2 with identical baked template, prompts, sampling, and settings. Both passed short recall (3,491 prompt tokens), exact one-case weather JSON (75 prompt tokens), the expanded coherence harness (1,024-token cap), and the same centered-depth 66,316-token needle probe. All returned nonempty final answers with finish_reason=stop; the long responses contained the needle and multi-token output. Each paired run passed at 66K; neither reproduced the earlier v2.0 collapse in this configuration or establishes a version-specific or general fix.

Profile settings: q4_0/q4_0, speculation OFF, KVMem OFF, --cache-ram 0, -np 1, -b 512 -ub 512, flash attention on, 8 threads, seed 12345, temperature 0.2, top-p 0.95, top-k 40, repeat penalty 1.05; -c 8192 for short cases and -c 70000 for the long probe. At 66,316 prompt tokens, v2.0 returned the needle in 280 tokens (39.257 decode t/s; 1125.164 prefill t/s); v2.1 returned it in 150 tokens (39.537 decode t/s; 1141.281 prefill t/s). In repeat 182907, v2.0 generated 280 tokens (38.933 decode; 1126.218 prefill t/s) and v2.1 150 (39.237 decode; 1141.007 prefill t/s). Unequal single-prompt output lengths make these descriptive values, not a speed comparison.

Earlier profiles are retained separately: profile 165659 had a 256-token coherence cap and ended with empty content / finish_reason=length for both; profile 170932 was a short-only rerun with -Skip66K -CoherenceMaxTokens 1024. Do not merge them with the latest combined profile.

The separate optional MTP profile 173756 reached its 256-token output cap for both versions with empty assistant content and finish_reason=length; reasoning was present but the final answer was incomplete. Its throughput is not a qualified response or coherence result and is not reported as a performance win. A later one-turn MTP coherence profile 180503 completed four-sentence answers for both versions at the 1,024-token cap with finish_reason=stop. This supports only that single short response, not multi-turn robustness, general MTP quality, long-context behavior, or a speed gain. Broad quality and release qualification remain deferred. Artifact SHA-256 values and tested profile settings are listed above.

Credits

  • Qwen — architecture and pretrained weights.
  • OrcaRouter — uncensored checkpoint (orcarouter/Qwen3.8-27B-Uncensored).
  • GSQ/RCO — GSQ (Dadgarnia et al.) and RCO (Helcig & Alistarh), ISTA-DASLab; allocation source ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF.
  • Froggeric — chat template (froggeric-qwen3.8-tool-use.jinja, v22.5).
  • llama.cpp / GGML — runtime and GGUF format.

Community quantization; not affiliated with Qwen, OrcaRouter, or ISTA-DASLab.

Downloads last month
41,048
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored

Base model

Qwen/Qwen3.8-27B
Quantized
(59)
this model
Quantizations
1 model

Space using RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored 1