Instructions to use orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF # Run inference directly in the terminal: llama cli -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF # Run inference directly in the terminal: llama cli -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF # Run inference directly in the terminal: ./llama-cli -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
Use Docker
docker model run hf.co/orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
- LM Studio
- Jan
- vLLM
How to use orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
- Ollama
How to use orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF with Ollama:
ollama run hf.co/orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
- Unsloth Desktop
- Pi
How to use orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF with Docker Model Runner:
docker model run hf.co/orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
- Lemonade
How to use orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
Run and chat with the model
lemonade run user.OrcaSAQ-2-Cyber-27B-Uncensored-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
OrcaSAQ2 Cyber 27B β 99 tok/s at 128K context on RTX 5070 Ti + RTX 4070
I have been testing OrcaSAQ-2-Cyber-27B-Uncensored on my local setup, and the performance has been better than I expected.
Config and result: https://docs.qbitz.foo/docs/local-llm/orcasaq-and-dflash
OrcaSAQ2 first analysis: https://docs.qbitz.foo/docs/malware-analysis/orca-first-test
Hardware
- RTX 5070 Ti 16 GB
- RTX 4070 12 GB
- i7-13700K
- 32 GB RAM
- Linux + llama.cpp
The GPUs give me roughly 28 GB of usable VRAM in total.
Test method
I warmed each configuration up with two throwaway requests first, then took the median decode speed across real coding prompts.
All tests used:
- temperature:
1.0 cpu_kv = 0- fully GPU-resident KV
- same machine / general workload
Results
| Configuration | Context | KV | Speculative decoding | Decode | Acceptance |
|---|---|---|---|---|---|
| OrcaSAQ + DFlash2 | 131K | q8_0 | DFlash2 n5 | 99.2 tok/s | 60% |
| OrcaSAQ + MTP | 131K | q8_0 | MTP n3 | 70.8 tok/s | 57% |
| OrcaSAQ max-context | 262K | q4_0 | MTP n3 | ~59 tok/s | 56% |
| Qwen3.8 27B Q4 + DFlash2 | 131K | q8_0 | DFlash2 n5 | 91.6 tok/s | 82% |
What surprised me most
The result that was most suprising for me was 99.2 tok/s at 131K context with OrcaSAQ + DFlash2.
At the same 131K context, switching from the built-in MTP drafter to DFlash2 takes me from:
70.8 β 99.2 tok/s
So roughly a 40% increase in decode speed.
That is a pretty massive difference just from changing the drafter.
OrcaSAQ vs my normal Qwen setup
What I also found interesting is the comparison with my normal Qwen3.8 27B Q4 setup.
The base model actually gets much higher DFlash2 draft acceptance:
82% vs 60%
But OrcaSAQ still ends up faster overall:
OrcaSAQ: 99.2 tok/s
Qwen3.8 Q4: 91.6 tok/s
I am thinking its because of the smaller/lighter OrcaSAQ weights that are making enough of a difference during normal decoding to outweigh the lower speculative acceptance rate.
So even though DFlash2 predicts fewer accepted tokens with OrcaSAQ, the model itself is fast enough that the total result still comes out ahead.
One important llama.cpp detail
I currently need two different llama.cpp builds depending on what I want.
My patched fork supports the q4_0 CUDA KV-cache path I need for 262K context, but it cannot load the current DFlash2 drafter.
Upstream llama.cpp loads DFlash2 correctly, so that is what I use for the 131K speed configuration.
So in practice I ended up with two presets:
Speed preset
131K context + q8_0 KV + DFlash2 n5
β 99.2 tok/s
Max-context preset
262K native context + q4_0 KV + MTP n3
β ~59 tok/s
Honestly, ~59 tok/s while keeping the full 262K native context is already very usable.
But ~100 tok/s at 128K makes the smaller-context configuration ridiculously responsive for a 27B model running locally.
So depending on what I am doing, I can basically choose between:
- ~100 tok/s with 131K context
- ~59 tok/s with the full 262K context
From a speed / VRAM / context perspective, OrcaSAQ has been extremely impressive on this hardware.