OrcaSAQ2 Cyber 27B β€” 99 tok/s at 128K context on RTX 5070 Ti + RTX 4070

#6
by QbitX - opened

I have been testing OrcaSAQ-2-Cyber-27B-Uncensored on my local setup, and the performance has been better than I expected.

Config and result: https://docs.qbitz.foo/docs/local-llm/orcasaq-and-dflash
OrcaSAQ2 first analysis: https://docs.qbitz.foo/docs/malware-analysis/orca-first-test

Hardware

  • RTX 5070 Ti 16 GB
  • RTX 4070 12 GB
  • i7-13700K
  • 32 GB RAM
  • Linux + llama.cpp

The GPUs give me roughly 28 GB of usable VRAM in total.

Test method

I warmed each configuration up with two throwaway requests first, then took the median decode speed across real coding prompts.

All tests used:

  • temperature: 1.0
  • cpu_kv = 0
  • fully GPU-resident KV
  • same machine / general workload

Results

Configuration Context KV Speculative decoding Decode Acceptance
OrcaSAQ + DFlash2 131K q8_0 DFlash2 n5 99.2 tok/s 60%
OrcaSAQ + MTP 131K q8_0 MTP n3 70.8 tok/s 57%
OrcaSAQ max-context 262K q4_0 MTP n3 ~59 tok/s 56%
Qwen3.8 27B Q4 + DFlash2 131K q8_0 DFlash2 n5 91.6 tok/s 82%

What surprised me most

The result that was most suprising for me was 99.2 tok/s at 131K context with OrcaSAQ + DFlash2.

At the same 131K context, switching from the built-in MTP drafter to DFlash2 takes me from:

70.8 β†’ 99.2 tok/s

So roughly a 40% increase in decode speed.

That is a pretty massive difference just from changing the drafter.

OrcaSAQ vs my normal Qwen setup

What I also found interesting is the comparison with my normal Qwen3.8 27B Q4 setup.

The base model actually gets much higher DFlash2 draft acceptance:

82% vs 60%

But OrcaSAQ still ends up faster overall:

OrcaSAQ: 99.2 tok/s
Qwen3.8 Q4: 91.6 tok/s

I am thinking its because of the smaller/lighter OrcaSAQ weights that are making enough of a difference during normal decoding to outweigh the lower speculative acceptance rate.

So even though DFlash2 predicts fewer accepted tokens with OrcaSAQ, the model itself is fast enough that the total result still comes out ahead.

One important llama.cpp detail

I currently need two different llama.cpp builds depending on what I want.

My patched fork supports the q4_0 CUDA KV-cache path I need for 262K context, but it cannot load the current DFlash2 drafter.

Upstream llama.cpp loads DFlash2 correctly, so that is what I use for the 131K speed configuration.

So in practice I ended up with two presets:

Speed preset

131K context + q8_0 KV + DFlash2 n5

β†’ 99.2 tok/s

Max-context preset

262K native context + q4_0 KV + MTP n3

β†’ ~59 tok/s

Honestly, ~59 tok/s while keeping the full 262K native context is already very usable.

But ~100 tok/s at 128K makes the smaller-context configuration ridiculously responsive for a 27B model running locally.

So depending on what I am doing, I can basically choose between:

  • ~100 tok/s with 131K context
  • ~59 tok/s with the full 262K context

From a speed / VRAM / context perspective, OrcaSAQ has been extremely impressive on this hardware.

Sign up or log in to comment