Instructions to use RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS # Run inference directly in the terminal: llama cli -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS # Run inference directly in the terminal: llama cli -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS # Run inference directly in the terminal: ./llama-cli -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS # Run inference directly in the terminal: ./build/bin/llama-cli -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
Use Docker
docker model run hf.co/RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
- LM Studio
- Jan
- vLLM
How to use RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
- Ollama
How to use RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored with Ollama:
ollama run hf.co/RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
- Unsloth Desktop
- Pi
How to use RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored with Docker Model Runner:
docker model run hf.co/RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
- Lemonade
How to use RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
Run and chat with the model
lemonade run user.Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored-IQ3_XXS
List all available models
lemonade list
- Hermes Agent
How to use RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:IQ3_XXS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
--auth-choice custom-api-key \
--custom-base-url http://127.0.0.1:8080/v1 \
--custom-model-id "RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored:" \
--custom-provider-id llama-cpp \
--custom-compatibility openai \
--custom-text-input \
--accept-risk \
--skip-healthRun OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"- Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored
- Measurement discipline (read first)
- 1. TL;DR
- 2. What this is
- 2b. Architecture (verified from the GGUF)
- 3. Changelog
- 4. Benchmarks
- 5. Quality gate (16K short-context, build3)
- 6. MTP
- 7. Long context — honest scope
- 8. TTS co-residency (separate configuration — do not merge with §4)
- 9. Quickstart
- 10. Known issues
- 11. Reproduction appendix
- 12. Paired frozen-build2 smoke (bounded; 2026-10-02)
- Credits
- Measurement discipline (read first)
Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored
A ~3 bpw GGUF quantization of the OrcaRouter Qwen3.8-27B, carrying an embedded MTP draft head and a custom OrcaRouter-native imatrix. Built to run on a single 16 GB GPU.
This is a quantization of the OrcaRouter Qwen3.8-27B checkpoint (a build on Qwen/Qwen3.8-27B) — not a new
fine-tune of the trunk. The trunk weights are quantized.
Quantization follows the GSQ/RCO methodology from ISTA-DASLab, with an imatrix calibrated on an
OrcaRouter-native corpus. The distributed artifact also contains an embedded MTP draft head.
Release status: v2.1 is the recommended release candidate. Two bounded paired runs on unchanged build2 completed (§12): v2.0 and v2.1 each passed the 66,316-token centered-needle probe twice. These runs did not reproduce the earlier v2.0 collapse on that configuration and does not establish a general fix or broad release qualification.
Measurement discipline (read first)
Every measured row is stamped with context, flags, engine (build-kv / build3), and date.
Rows from different configurations are never merged into one number. In particular: short-context gate
numbers and full-history long-context numbers are different workloads and are shown in separate tables.
There are two distinct long-context regimes, and they are not interchangeable:
- Full-history, correctness-qualified (the honest baseline): every prior token resident, no eviction. This is the row we recommend for long-context work because retrieval is intact.
- KVMem window (speed-only): evicts early/mid history. Faster, but retrieval fails (see §4C). It is documented for transparency, not as a recommended configuration.
1. TL;DR
| Recommended file | Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf |
| Size / SHA-256 | 10,382,717,184 bytes (9.67 GiB) · AB955B5083D9CDF0BF55C37ACDCAE359B78756C4544D97960C23D8FCA98FEB9B |
| Trunk | IQ3_XXS, 3.04 bpw |
| Draft head | Embedded MTP, blk.64, 15 tensors: seven weight tensors Q6_K, attn_output.weight IQ4_XS, seven F32 norm tensors |
| Base | OrcaRouter Qwen3.8-27B checkpoint (build of Qwen/Qwen3.8-27B) |
| Template | froggeric-qwen3.8-tool-use.jinja v22.5, baked into the GGUF (--jinja alone selects it) |
| Modality | text + tools in this file; vision needs the separate mmproj/ projector |
| Long context, correctness-qualified | full-history 196K resident, q4_0 K / q4_0 V, spec OFF, KVMem OFF → decode 33.7 t/s, prefill 662 t/s, NIAH 3/3 depths PASS ({build-kv fa5994f5c, 2026-10-02}) |
| Short-context gate (build3) | short-context decode probe 106.4 t/s (input-token count not retained — not a decode-after-16K-ingest figure) · prefill 1473 t/s · separate ~16,053-token NIAH 6/6 · tool 13/14 · coherence 4/4 · MTP accept 0.91 |
The two headline rows above are different workloads. 33.7 t/s is the correctness-qualified full-history 196K number. 106.4 t/s is a short-context decode probe (the gate JSON retains three speed samples but not the probe's input-token count), and the 6/6 NIAH is a separate ~16,053-token haystack. Do not read 106.4 as decoding after ingesting 16K, and do not quote one row as the other.
2. What this is
We quantized the OrcaRouter checkpoint to a single-file ~3 bpw GGUF a 16 GB consumer card can serve with useful context and speculative decoding. Two design choices matter:
- GSQ/RCO allocation. We reproduce ISTA-DASLab's GSQ/RCO per-tensor allocation map and apply it to the
OrcaRouter base. The full 866-row map ships as
REF-IQ3_XXS-mtp.rco-allocation.txt. - An embedded MTP draft head. Its 15 tensors include seven Q6_K weight tensors,
attn_output.weightin IQ4_XS, and seven F32 norm tensors, so--spec-type draft-mtpgives built-in speculation with no second file.
Hardware (stated once): every number was measured on a single RTX 5070 Ti 16 GB (GB203, sm_120, 70 SM), Windows, CUDA 13.3. Metrics are single-GPU, single-build.
2b. Architecture (verified from the GGUF)
| Property | Value |
|---|---|
| Architecture | qwen35 (Qwen3_5ForConditionalGeneration) |
| Layers | 64 transformer = 16 full-attention + 48 GDN/linear-attention; +1 MTP block (block_count 65) |
| Attention | head_count 24, head_count_kv 4 → GQA 6; head_dim 256; ctx 262144 |
| GDN | state_size 128, conv_kernel 4, key-heads (group_count) 16, value-heads (time_step_rank) 48 → GVA group 3, inner_size 6144 |
| Hidden / FFN | 5120 / 17408; vocab 248320 |
| MTP | qwen35.nextn_predict_layers = 1 → blk.64.nextn.* (Q6_K) + enorm/hnorm F32 |
Note: the GDN state dimension is 128, not the attention head dim (256). They are different subsystems.
3. Changelog
- v2.1 (current, recommended). Targets the previously observed >64K generation collapse. Root cause: two FFN tensors in the
per-tensor GSQ/RCO allocation were too aggressive —
ffn_gate.weight= IQ1_S andffn_up.weight= IQ1_M. Raising exactly those two addresses that observed failure while landing at 3.04 bpw (half the map unchanged). Evidence: the gates-only path collapsed; the full 6-sensitive-tensor fix at 3.78 bpw still collapsed; a plain IQ3_XXS mix passed; only the FFN change to IQ3_XXS passed. Anchored true tokenizer name:Staged_Tmpl. In the paired frozen-build2 smoke, both v2.0 and v2.1 passed one identical 66,316-token probe; that smoke does not distinguish versions or establish a general fix. - v2.0. OrcaRouter base + S1 MTP head + custom imatrix + RCO allocation. Superseded — earlier >64K collapse observation was not reproduced in the paired build2 smoke. Retained for reproducibility.
- v1.x (Huihui line). Frozen / deprecated.
4. Benchmarks
4A. Correctness-qualified full-history 196K baseline (the recommended long-context config)
| Metric | Value | Context | Flags | Engine | Date |
|---|---|---|---|---|---|
| Decode | 33.7 t/s (33.6 / 33.91 / 33.7 / 33.8) | 196,044–196,068 tokens resident | -ctk q4_0 -ctv q4_0, no spec, KVMEM_ENABLE=0, -c 262144 -fa on -ngl 999 -b 512 -ub 512 |
build-kv fa5994f5c |
2026-10-02 |
| Prefill | 662–664 t/s | same | same | build-kv fa5994f5c |
2026-10-02 |
| NIAH recall | 3/3 PASS (depths 0.1 / 0.5 / 0.9) | full 196K | same | build-kv fa5994f5c |
2026-10-02 |
Profile: full-history 196K, build-kv fa5994f5c, 2026-10-02.
Depth ID: each answer quotes TEAL-FALCON-4402 from the full history; no empty output, no collapse.
MTP on this exact baseline (A1 finding, separate rows): --spec-type draft-mtp --spec-draft-n-max 2
@full-KV 196K → decode 27.0–28.2 t/s on open prose (accept ~0.47) — i.e. a regression vs 33.7; on the
recall prompts it rose to 39.1–39.6 t/s (accept 0.82–0.92) and prefill fell to ~166 t/s. MTP value is
workload-dependent on this baseline. (A1_MTP_FINDING_20261002.md.)
4B. Short-context gate (build3) — decode speed probe + separate ~16K NIAH
| # | Metric | Value | Context | Flags | Engine | Date |
|---|---|---|---|---|---|---|
| 1 | Decode (short-context probe; prior recorded profile) | 106.4 t/s | short-context; input-token count not retained | -fa on -ctk q8_0 -ctv q4_0 --spec-type draft-mtp --spec-draft-n-max 2 |
build3 19fff3e38 |
2026-10-01 |
| 2 | Prefill | 1473.4 t/s | 16K | same | build3 | 2026-10-01 |
| 3 | Needle (NIAH) | 6/6 | ~16K haystack, depths 0.1/0.5/0.9 | same | build3 | 2026-10-01 |
| 4 | Tool-call JSON | 13/14 | 14-case suite | same | build3 | 2026-10-01 |
| 5 | Coherence | 4/4 | 4-case rubric | same | build3 | 2026-10-01 |
| 6 | MTP accept | 0.91 | 16K | embedded MTP, -n-max 2 |
build3 | 2026-10-01 |
4C. KVMem-window config — documented for transparency, NOT recommended
| Metric | Value | Context | Flags | Recall | Engine |
|---|---|---|---|---|---|
| Decode | 74.5 t/s (71.7–94.5) | 196,068 resident | draft-mtp,ngram-mod, KVMem on, budget 2048 |
2/6 FAIL (depths 0.1, 0.5 missed) | build3 |
The historic 74.5 t/s @196K and the recall failure are the same configuration: the KVMem budget (2048) evicts early/mid KV, so the deep-history needles are lost. Against the 33.7 baseline two variables differ at once (spec and KVMem) — the whole gap must not be attributed to either. If you need 196K retrieval, use 4A; if you need speed on a short effective history, a window is a conscious trade. No faster full-history 196K configuration is qualified in these results (stated without causal attribution: both variables — KVMem budget 2048 and MTP/ngram — differ from the baseline).
5. Quality gate (16K short-context, build3)
| Check | Result | Floor | Verdict |
|---|---|---|---|
| Decode | 106.4 t/s | 94.0 | PASS |
| Prefill | 1473.4 t/s | 1377.9 | PASS |
| Needle (~16K, 6 cases) | 6/6 | 6 | PASS |
| Tool-call (14 cases) | 13/14 | 12 | PASS |
| Coherence (4 cases) | 4/4 | 4 | PASS |
| MTP accept | 0.91 | 0.897 | PASS |
The single tool-call failure is tc-07; reported rather than rounded up.
Quality suites — historical v2.0 values; no v2.1 scores
The v2.1 fix touched FFN gates, so these are carried forward for reference only:
| Suite | Score |
|---|---|
| GPQA-Diamond | 75.25% (198 q, harness-conditional) |
| WikiText-2 perplexity | 6.17 |
| IFEval | 70.24 / 76.26 |
| TruthfulQA | MC1 77.60 / MC2 81.98 |
| XSTest-safe over-refusal | 0% (0/250) |
6. MTP
Embedded mixed-type head (blk.64, seven Q6_K weight tensors, attn_output.weight IQ4_XS, seven F32 norm tensors): --spec-type draft-mtp --spec-draft-n-max 2. No second file.
Two acceptance regimes, both stated: 0.91 at the 16K short-context gate; ~0.47–0.67 at long context
(workload-dependent, see 4A). MTP output may differ from serial generation on a quantized target — fix the
seed, and disable speculation when exact serial behavior matters.
7. Long context — honest scope
- Full-history 196K (correctness-qualified): decode 33.7 t/s, prefill 662 t/s, recall 3/3 PASS (§4A). This is the number to quote.
- KVMem window: faster (74.5 t/s) but retrieval fails (2/6) — not recommended (§4C).
- 66K paired smoke: both v2.0 and v2.1 returned the embedded needle with multi-token output from the same 66,316-token prompt on frozen build2. The earlier v2.0 collapse was not reproduced in this paired run; one centered-depth probe does not establish a general fix. The separate 16K and 196K recall rows are documented with their own engine/configuration below.
- Coverage: one qualified full-history point at ~196K with three tested needle depths. Other context lengths through 256K, and repeated long-context reliability, remain untested. The paired 66K smoke passed for both versions; neither result establishes broad long-context reliability.
- The 6/6 needle gate is a ~16K haystack test, not a 196K/256K test.
8. TTS co-residency (separate configuration — do not merge with §4)
The 27B at 256K and a Qwen3-TTS-1.7B custom-voice talker fit together on the same 16 GB card: the 27B at 256K measured 77.2 t/s with the talker, in 14,751 MiB used / 1,245 MiB free. Voice synthesis ~0.7–1.2 s for ~5–9 s of audio. This 77.2 t/s is a distinct configuration (shared card, talker resident) and must not be read as the standalone long-context decode number (§4A is). Reproducibility caveat: exact LLM flags, effective input length, KVMem settings, the TTS binary/backend, and a recorded CPU-vs-CUDA launcher discrepancy are not fully resolved in the artifacts — treat this row as not-yet-reproducible until they are recovered. Do not guess them.
# 1) 27B text server
llama-server -m Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf \
--jinja --ctx-size 262144 -b 512 -ub 512 -fa on -ngl 99 -ctk q8_0 -ctv q4_0 \
--spec-type draft-mtp --spec-draft-n-max 2 --port 8080
# 2) Qwen3-TTS-1.7B talker
qwentts.cpp tts-server --talker Qwen3-TTS-1.7B --voice serena --port 8081
9. Quickstart
A. Default — MTP, 32K:
llama-server -m Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf \
--jinja --ctx-size 32768 -b 512 -ub 512 -fa on -ngl 99 -ctk q8_0 -ctv q4_0 \
--spec-type draft-mtp --spec-draft-n-max 2 --port 8080
B. Full-history long context (recall intact; the qualified 196K config):
llama-server -m Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf \
--port 8080 -np 1 --jinja -ngl 999 -c 262144 -b 512 -ub 512 -fa on \
-ctk q4_0 -ctv q4_0 --cache-ram 0 -t 8
# no --spec-type, KVMEM_ENABLE=0. Headroom note: reduce -c to <=216193 for >=1536 MiB margin.
The Froggeric template is baked in (tokenizer.chat_template, v22.5) — --jinja alone selects it.
Vision (optional): text-only works without it; add mmproj/mmproj-Qwen3.8-27B-BF16.gguf.
10. Known issues
- Historical >64K collapse observation — v2.1 raises two FFN tensors to IQ3_XXS; however, the paired frozen-build2 66,316-token smoke passed for both v2.0 and v2.1. This run does not reproduce the failure or establish a version-specific fix; see §12.
- Two long-context regimes; do not merge — 33.7 t/s is full-history/recall-3/3; 74.5 t/s is KVMem/recall-2/6.
- MTP has two regimes — 0.91 (16K gate) vs ~0.47–0.67 (long context); and on the full-KV 196K baseline MTP regressed open prose (27 vs 33.7). Not a universal speedup.
- Long-context retrieval not gated across 66K–256K; the 6/6 gate is ~16K.
- Aggressive ~3-bit quant — knowledge recall is weaker than higher-bit variants; the model leans on retrieval/tools.
- Deferred v2.1 suites: GPQA-Diamond (198 questions), IFEval (541 items), WikiText-2 PPL, TruthfulQA, XSTest-safe, and the large-target-length matrix were NOT RUN in this release review. Datasets were acquired; scores were not produced. Historical values below belong to v2.0 only.
- Engine-sensitive KV (scoped to the tested binaries only) — on the tested
build2, theq8_0-K path fell back to a slow kernel; on the testedbuild319fff3e38, it is full speed. This is not a claim about every build bearing those names. Use the testedbuild3with-ctk q8_0 -ctv q4_0, orq4_0/q4_0on the testedbuild2. - RCO allocation reproduced from ISTA-DASLab's published map (not re-searched), applied to the uncensored base. The imatrix is custom OrcaRouter-native.
Uncensored-use notice
Abliterated model: refusal is substantially reduced by design. XSTest-safe measures over-refusal (0/250) — a compliance-sensitivity check, not a safety or adversarial-robustness evaluation. The deployer is responsible for compliant, lawful use.
11. Reproduction appendix
- Qualified baseline: Den build-kv, rev
fa5994f5c(build 10821);ggml-cuda.dllSHA256[:16]86CEFBB7197436AC. Command in §9B. - 16K gate: build3 commit
19fff3e38; flags in §4B. - SHAs: primary file
AB955B5083D9CDF0BF55C37ACDCAE359B78756C4544D97960C23D8FCA98FEB9B(10,382,717,184 bytes). Per-file hashes inSHA256SUMS.txt. - Quantization (abridged):
llama-quantize --imatrix imatrix.dat \ --tensor-type-file REF-IQ3_XXS-mtp.rco-allocation.txt \ model-f16.gguf Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf IQ3_XXS - Seeds: use fixed seeds for comparisons; the paired build2 profile used seed 12345.
| File | What |
|---|---|
Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.1.gguf |
recommended release candidate, 3.04 bpw; paired 66K functional smoke passed twice; full quality qualification deferred |
mmproj/mmproj-Qwen3.8-27B-BF16.gguf |
optional vision projector |
REF-IQ3_XXS-mtp.rco-allocation.txt |
866-row per-tensor allocation map |
imatrix.dat |
custom OrcaRouter-native calibration matrix |
SHA256SUMS.txt |
per-file hashes |
12. Paired frozen-build2 smoke (bounded; 2026-10-02)
The full base profile 173016 and repeat 182907 ran both versions serially on unchanged
build2 with identical baked template, prompts, sampling, and settings. Both passed short recall (3,491 prompt
tokens), exact one-case weather JSON (75 prompt tokens), the expanded coherence harness (1,024-token cap), and
the same centered-depth 66,316-token needle probe. All returned nonempty final answers with
finish_reason=stop; the long responses contained the needle and multi-token output. Each paired run passed at 66K; neither reproduced the earlier v2.0 collapse in this configuration or establishes a version-specific or general fix.
Profile settings: q4_0/q4_0, speculation OFF, KVMem OFF, --cache-ram 0, -np 1, -b 512 -ub 512,
flash attention on, 8 threads, seed 12345, temperature 0.2, top-p 0.95, top-k 40, repeat penalty 1.05;
-c 8192 for short cases and -c 70000 for the long probe. At 66,316 prompt tokens, v2.0 returned the
needle in 280 tokens (39.257 decode t/s; 1125.164 prefill t/s); v2.1 returned it in 150 tokens (39.537 decode t/s; 1141.281 prefill t/s). In repeat 182907, v2.0 generated 280 tokens (38.933 decode; 1126.218 prefill t/s) and v2.1 150 (39.237 decode; 1141.007 prefill t/s). Unequal single-prompt output lengths make these descriptive values, not a speed comparison.
Earlier profiles are retained separately: profile 165659 had a 256-token coherence
cap and ended with empty content / finish_reason=length for both; profile 170932 was a
short-only rerun with -Skip66K -CoherenceMaxTokens 1024. Do not merge them with the latest combined profile.
The separate optional MTP profile 173756 reached its 256-token output cap for
both versions with empty assistant content and finish_reason=length; reasoning was present but the final
answer was incomplete. Its throughput is not a qualified response or coherence result and is not reported as a
performance win. A later one-turn MTP coherence profile 180503 completed
four-sentence answers for both versions at the 1,024-token cap with finish_reason=stop. This supports only that
single short response, not multi-turn robustness, general MTP quality, long-context behavior, or a speed gain.
Broad quality and release qualification remain deferred. Artifact SHA-256 values and tested profile settings are listed above.
Credits
- Qwen — architecture and pretrained weights.
- OrcaRouter — uncensored checkpoint (
orcarouter/Qwen3.8-27B-Uncensored). - GSQ/RCO — GSQ (Dadgarnia et al.) and RCO (Helcig & Alistarh), ISTA-DASLab; allocation source
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF. - Froggeric — chat template (
froggeric-qwen3.8-tool-use.jinja, v22.5). - llama.cpp / GGML — runtime and GGUF format.
Community quantization; not affiliated with Qwen, OrcaRouter, or ISTA-DASLab.
- Downloads last month
- 41,048
3-bit
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp# Start a local OpenAI-compatible server: llama serve -hf RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored: