---
license: apache-2.0
base_model: Qwen/Qwen3.5-9B
library_name: gguf
pipeline_tag: text-generation
tags: [qwen3, coding, code, math, cybersecurity, reasoning, thinking, gguf, llama.cpp, local-llm]
---
# ⥠Qwen3.5-9B-Claude-4.8 (GGUF) â âĻ
### ð§ More accuracy, **20% less rambling** â a sharp local model for **code, math & general tasks**
> **Runs on modest hardware.** With **~8-9 GB** of VRAM *or* unified memory free, you get a private, offline
> reasoning assistant that **thinks less and lands more**. ð
> This is the **v1 edition** â tuned to reason efficiently, cut redundant chain-of-thought, and still hit the
> correct final answer. All local, all yours. ð
### ðŊ What it is
A focused fine-tune of **Qwen3.5-9B** with trace-inversed **CoT from Opus4.8** (Dataset not published anywhere) specialized for **coding, mathematics, and cybersecurity reasoning**.
The headline trait: it reaches the goal with roughly **20% shorter thinking traces** than the base model while
**maintaining final-answer accuracy** â less wandering, faster tokens-to-solution, lower latency. ð§ âĄ
---
## âĻ Highlights
- ðŠķ **~20% shorter reasoning** vs. base model, with accuracy held â tighter CoT, faster answers.
- ðŧ **Improved coding** â cleaner, more runnable solutions across common languages.
- ð§Ū **Strong math** â multi-step problems with the reasoning shown, then a clear final answer.
- ðĄïļ **Security-aware** â geared toward defensive concepts, code review, and CTF-style learning.
- ðĶ **One quant, well-tuned:** ships as **Q4_K_M** â the size/quality sweet spot.
---
## ðĶ Download (GGUF quant)
| Quant | Size | Vibe |
|------|------|------|
| ðĩ **Q4_K_M** | **~5.6 GB** | the sweet spot ð (recommended â balanced quality & footprint) |
> ðĄ Only Q4_K_M is published for now. Want another quant (Q5_K_M, Q8_0, f16)? Open a discussion and I'll consider it.
---
## ð§Ū "Will it fit?" â context cheat-sheet
Rough estimates ðĪ (assumes `q8_0` KV cache + ~1.5 GB overhead; switch to `q4_0` KV cache for â2à more context).
| Your VRAM / unified mem | ðĩ Q4_K_M (~8-9G) |
|---|---|
| **8 GB** | ~24K ctx |
| **12 GB** | ~64K |
| **16 GB** | ~100K |
| **24 GB** | comfortable headroom |
> ðĄ Apple Silicon / integrated GPUs with **unified memory** work too â same idea, just slower than a dGPU.
---
## ð How to run it
### Option A â llama.cpp (recommended) ðĶ
1. Download `âĶ-Q4_K_M.gguf` and `llama-server` from [llama.cpp](https://github.com/ggml-org/llama.cpp).
> â ïļ Use a **recent llama.cpp** build for current Qwen3 architectures.
2. Run a server:
\```bash
llama-server \
-m ./qwen3.5-9b-reasoner-Q4_K_M.gguf \
--ctx-size 16384 \
--n-gpu-layers 99 \
-fa on \
--cache-type-k q8_0 --cache-type-v q8_0 \
--temp 0.7 --top-p 0.8 --top-k 20 \
--host 0.0.0.0 --port 8080
\```
3. Connect to your agent and chat. ð
### Option B â one-click apps ðąïļ
Works in **LM Studio**, **Jan**, **Ollama**, etc. â import the GGUF, pick the quant, go. ðū
### ð§ Thinking mode
This model reasons before answering. Keep thinking enabled (the default chat template handles it).
Suggested sampling: `temp 0.7, top_p 0.95, top_k 20`. For deterministic code/math, try greedy (`temp 0`).
---
## â ïļ Good to know
- **Not safety-aligned for production.** This is a specialized reasoning fine-tune â add your own guardrails,
input/output filtering, and review before any production or user-facing deployment. Use responsibly. ð [Not uncensored]
- Strongest in **code, math, and security reasoning**; double-check general-knowledge facts and figures.
- English-centric.
---
## ð Acknowledgements
Special thanks to:
- The Qwen team for the strong Qwen3.5 base model.
- Unsloth for efficient fine-tuning frameworks.
---
## ð Base & License
- **License:** Apache 2.0
- **Base model:** [`Qwen/Qwen3.5-9B`](https://huggingface.co/Qwen/Qwen3.5-9B)