Instructions to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M # Run inference directly in the terminal: llama cli -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M # Run inference directly in the terminal: llama cli -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M # Run inference directly in the terminal: ./llama-cli -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
Use Docker
docker model run hf.co/morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
- LM Studio
- Jan
- vLLM
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
- SGLang
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with Ollama:
ollama run hf.co/morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
- Unsloth Desktop
- Pi
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with Docker Model Runner:
docker model run hf.co/morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
- Lemonade
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
Run and chat with the model
lemonade run user.Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF-IQ2_M
List all available models
lemonade list
- Hermes Agent
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF:IQ2_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF
This repository contains the GGUF quantized files for OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated.
- Original Model: OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated
- Architecture: Qwen3.5-122B-A10B
- License: Apache 2.0
- Vision: works as expected (image / video → text).
- MTP: the head is present and shape-compatible, but in our testing it produced no measurable speedup or quality gain on this checkpoint. It is shipped intact for completeness and forward-compatibility, but would need to be retrained to be useful — happy to do so if there is interest in the model.
| Quant Type | Size | Description |
|---|---|---|
| IQ2_M | 55-59 GB | Mixed Precision for Better Quality |
| IQ3_M | 62-67 GB | Mixed Precision for Better Quality |
| IQ4_NL | 76-81 GB | Mixed Precision for Better Quality |
Overview
The pipeline:
- Refusal Ablation — Residual-stream refusal directions (one per decoder layer, layers 19–45) were extracted via diff-in-means on a labeled prompt set and baked into the weights as a per-matrix delta — see the abliterix framework for the methodology.
- Healing — Stage A: Constrained-LoRA SFT on Opus reasoning data — Supervised finetuned on a curated set of Claude Opus reasoning traces (single-turn, ~8k rows). To keep the abliteration mathematically intact during training, a custom orthogonality projection is applied to every LoRA
B-matrix on residual-write modules after each optimizer step (B := B − r·(rᵀB)), so the LoRA update is forbidden from re-introducing the refusal direction. LoRA rank 32, α 64, 54 protected modules across 27 decoder layers. Verified residual after training:max ‖rᵀB‖₂ = 8.5 × 10⁻¹⁰. - Healing — Stage B: Unconstrained SFT on chosen completions — A second short SFT pass (LoRA r=16, α 32, no orthogonality constraint) on the chosen answers (including reasoning chains) from an internal preference dataset, to tighten on the deployment distribution and remove the last bits of drift introduced by Stage A.
- Kimi K2.6 Reasoning DPO — A targeted preference-optimization pass distilled from Kimi K2.6 to improve reasoning verbosity and eliminate degenerate looping. See the dedicated section below.
- Vision + MTP Restoration — The original Qwen3.5 vision tower (333 tensors, depth 27, hidden 1152) and MTP head (785 tensors, 1 hidden layer) were grafted back from the upstream
Qwen/Qwen3.5-122B-A10Bshards. Tensor names, shapes, andconfig.jsonschema (Qwen3_5MoeForConditionalGeneration,model_type: qwen3_5_moe) match the base model exactly — so this checkpoint loads anywhere the original loads.
Key Properties:
- Uncensored across the standard refusal axes
- Reasoning preserved and improved (Opus-style think-then-answer + Kimi K2.6 reasoning DPO)
- Fewer looping / repetition failures on long conversations
- Multimodal: vision (image / video) and MTP heads carried forward
- Drop-in shape compatibility with
Qwen/Qwen3.5-122B-A10B
Kimi K2.6 Reasoning DPO
On top of the base abliteration + Opus healing, this release adds a focused healing pass built from Kimi K2.6:
- ~3,000 samples distilled from Kimi K2.6 were used for DPO (Direct Preference Optimization), alongside synthetic datasets also generated from Kimi K2.6.
- Improved reasoning verbosity — the model now produces more complete, better-structured reasoning on the ~12% of requests where the previous release tended to under-explain or cut its chain-of-thought short.
- Fixed looping / repetition — degenerate loops that appeared on 2–6% of long-tail conversations (long context, multi-turn) were largely eliminated.
The DPO pass targets the language model's reasoning behavior only; the abliteration, vision tower, and MTP head are unchanged by this step.
Evaluation
This model family outperforms the full-precision (BF16) Qwen/Qwen3.5-122B-A10B baseline across reasoning, coding, and tool-use benchmarks:
| Benchmark | Qwen3.5-122B-A10B (BF16, baseline) | Qwopus3.5-122B-A10B |
|---|---|---|
| CTI | 64.8 | 71.5 |
| LiveCodeBench | 78.9 | 79.9 |
| BFCL | 72.2 | 85.6 |
BFCL is the Berkeley Function-Calling Leaderboard (tool use); LiveCodeBench is contamination-controlled code generation.
Notes
- License: Other (inherits from the Qwen3.5 base license)
- Base Model: Qwen/Qwen3.5-122B-A10B
- Healing: Opus reasoning SFT + Kimi K2.6 reasoning DPO (≈3,000 distilled samples + synthetic data)
- Modality: Text + Vision (image / video) + MTP
- Architecture: Qwen3 MoE (~10B active / 122B total) + Qwen3-VL vision tower + MTP head
Thanks
- Jackrong — for the idea of Qwopus merges (Opus distillations on Qwen models).
- wangzhang — for the wonderful abliterix framework, which was customized to do this abliteration.
Disclaimer
Use is the responsibility of the user. Ensure your usage complies with applicable laws, platform rules, and deployment requirements.
How to Use
These GGUF files are fully compatible with llama.cpp and popular graphical interfaces like LM Studio.
using llama.cpp CLI:
./llama-cli -m /path/to/model/Qwopus3.5-122B-Distill-Kimi-IQ3_M.gguf \
-p "Hello, how are you?" \
-sys "You are a helpful AI" \
-n 4096 \
-c 8192
using llama-server :
./llama-cli -m /path/to/model/Qwopus3.5-122B-Distill-Kimi-IQ3_M.gguf \
--host 0.0.0.0 \
--port 8080
- Downloads last month
- 748
2-bit
3-bit
4-bit
Model tree for morikomorizz/Qwopus3.5-122B-Distill-Kimi-Uncensored-GGUF
Base model
Qwen/Qwen3.5-122B-A10B