Text Generation
GGUF
refusal-ablation
capability-preserving
saber
rys
layer-surgery
qwen3.5
multimodal
27b
conversational
Instructions to use DJLougen/Ornstein-27B-SABER-RYS-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use DJLougen/Ornstein-27B-SABER-RYS-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
Use Docker
docker model run hf.co/DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use DJLougen/Ornstein-27B-SABER-RYS-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "DJLougen/Ornstein-27B-SABER-RYS-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DJLougen/Ornstein-27B-SABER-RYS-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
- Ollama
How to use DJLougen/Ornstein-27B-SABER-RYS-GGUF with Ollama:
ollama run hf.co/DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use DJLougen/Ornstein-27B-SABER-RYS-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use DJLougen/Ornstein-27B-SABER-RYS-GGUF with Docker Model Runner:
docker model run hf.co/DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
- Lemonade
How to use DJLougen/Ornstein-27B-SABER-RYS-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Ornstein-27B-SABER-RYS-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use DJLougen/Ornstein-27B-SABER-RYS-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use DJLougen/Ornstein-27B-SABER-RYS-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "DJLougen/Ornstein-27B-SABER-RYS-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -214,9 +214,21 @@ This model is released for research purposes. It demonstrates that safety refusa
|
|
| 214 |
|
| 215 |
---
|
| 216 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 217 |
## References
|
| 218 |
|
| 219 |
-
- [LLM Neuroanatomy
|
| 220 |
-
- [LLM Neuroanatomy
|
| 221 |
-
- [alainnothere/llm-circuit-finder](https://github.com/alainnothere/llm-circuit-finder) — GGUF surgery tools
|
| 222 |
- [XpressAI/Qwen3.5-27B-RYS-UD-Q4_K_XL-GGUF](https://huggingface.co/XpressAI/Qwen3.5-27B-RYS-UD-Q4_K_XL-GGUF) — Reference RYS model with BFCLv4 benchmarks
|
|
|
|
| 214 |
|
| 215 |
---
|
| 216 |
|
| 217 |
+
## Acknowledgments
|
| 218 |
+
|
| 219 |
+
The **RYS (Repeat Your Self)** layer-duplication method was discovered and developed by **David Noel Ng** ([@dnhkng](https://github.com/dnhkng)). The Pareto-optimal configurations for Qwen3.5-27B, the Math/EQ probes, the XGBoost surrogate pipeline, and the beam search methodology are all from his work. The GGUF surgery tools used to create these models are from [alainnothere/llm-circuit-finder](https://github.com/alainnothere/llm-circuit-finder), an open-source (MIT) implementation of the RYS technique for llama.cpp.
|
| 220 |
+
|
| 221 |
+
**If you use these models, please cite David Noel Ng's work:**
|
| 222 |
+
|
| 223 |
+
> Ng, David Noel. "LLM Neuroanatomy: How I Topped the Leaderboard Without Changing a Single Weight." [dnhkng.github.io/posts/rys](https://dnhkng.github.io/posts/rys/)
|
| 224 |
+
>
|
| 225 |
+
> Ng, David Noel. "LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language." [dnhkng.github.io/posts/rys-ii](https://dnhkng.github.io/posts/rys-ii/)
|
| 226 |
+
|
| 227 |
+
The **SABER** refusal-ablation method is original to this model.
|
| 228 |
+
|
| 229 |
## References
|
| 230 |
|
| 231 |
+
- [LLM Neuroanatomy — Part I](https://dnhkng.github.io/posts/rys/) — David Noel Ng
|
| 232 |
+
- [LLM Neuroanatomy — Part II](https://dnhkng.github.io/posts/rys-ii/) — David Noel Ng (Qwen3.5-27B sweep)
|
| 233 |
+
- [alainnothere/llm-circuit-finder](https://github.com/alainnothere/llm-circuit-finder) — GGUF surgery tools (MIT)
|
| 234 |
- [XpressAI/Qwen3.5-27B-RYS-UD-Q4_K_XL-GGUF](https://huggingface.co/XpressAI/Qwen3.5-27B-RYS-UD-Q4_K_XL-GGUF) — Reference RYS model with BFCLv4 benchmarks
|