OrcaSAQ2 Cyber 27B reached 68 tok/s and 240K context on 2× RTX 3060 12GB

#4
by SpaceMo0 - opened

I tested orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF on two RTX 3060 12GB cards using llamAmpere (so good).

Final configuration:

Context: 245,760 tokens (240K)
KV cache: Q4_0
MTP depth: 4
Batch/ubatch: 384 / 96
Tensor split: 1.25,1
Backend: llamAmpere, CUDA SM86
Driver: 580.178.04

Results:

Short code generation: 68.10 tok/s
12K-token prefill: 490 tok/s
Long-context test: 219,155 input tokens
Long-context prefill: 221.6 tok/s
Long-context total time: 990 seconds
Sentinel retrieval: Correct
Truncation: None
Remaining VRAM: 361 / 377 MiB

For comparison, my Qwen3.8 27B UD-Q5_K_M profile reached 49.61 tok/s with llamAmpere at 96K context.

Qwen Q5 size: 19.76 GB
OrcaSAQ2 size: 15.68 GB
Qwen MTP acceptance: 87.9%
Orca MTP acceptance: 93.0%

That makes Orca roughly 37% faster in my short-prompt tests while using a much larger configured context. The same two deterministic coding prompts were used, although output lengths differed, so this is not a perfect scientific comparison.

256K also loaded after reducing the batch to 128/32, but prefill dropped to about 320 tok/s and minimum free VRAM
fell to 255 MiB. I therefore kept 240K as the best speed/context compromise.

This is the best local 27B result I have measured so far.

OrcaRouter org

Thanks for your details test result!

was meant to comment but i created a new discussion if its of interest to read SpaceMo0 https://huggingface.co/orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF/discussions/6

Sign up or log in to comment