Instructions to use deucebucket/Granite-4.1-30B-Cerebellum-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use deucebucket/Granite-4.1-30B-Cerebellum-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M # Run inference directly in the terminal: llama cli -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M # Run inference directly in the terminal: llama cli -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M # Run inference directly in the terminal: ./llama-cli -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
Use Docker
docker model run hf.co/deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
- LM Studio
- Jan
- Ollama
How to use deucebucket/Granite-4.1-30B-Cerebellum-GGUF with Ollama:
ollama run hf.co/deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
- Unsloth Desktop
- Pi
How to use deucebucket/Granite-4.1-30B-Cerebellum-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use deucebucket/Granite-4.1-30B-Cerebellum-GGUF with Docker Model Runner:
docker model run hf.co/deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
- Lemonade
How to use deucebucket/Granite-4.1-30B-Cerebellum-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
Run and chat with the model
lemonade run user.Granite-4.1-30B-Cerebellum-GGUF-Q3_K_M
List all available models
lemonade list
- Hermes Agent
How to use deucebucket/Granite-4.1-30B-Cerebellum-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use deucebucket/Granite-4.1-30B-Cerebellum-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "deucebucket/Granite-4.1-30B-Cerebellum-GGUF:Q3_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Granite 4.1-30B — Cerebellum GGUF
Ablation-guided mixed-precision quantization of ibm-granite/granite-4.1-30b. 30B parameters, dense architecture with GQA, 64 layers.
What is Cerebellum?
Instead of uniform quantization, we measure which weight groups survive aggressive compression and which don't. Groups that tolerate Q2_K get demoted; groups that don't stay at Q3_K_M or higher. The result: smaller files with less quality loss than uniform quants of the same size.
Files
| File | Size | Description |
|---|---|---|
Granite-4.1-30B-Cerebellum-v2-Q3_K_M.gguf |
13 GB | Optimal mix — 3 groups demoted (attn_k, attn_q, attn_output), 4 kept at Q3_K_M |
Granite-4.1-30B-Cerebellum-v1-Q3_K_M.gguf |
12 GB | Aggressive — 5 groups demoted (all attn + ffn_gate) |
Benchmarks
Evaluated using our standardized benchmark suite (ARC-Challenge, HellaSwag, MMLU, HumanEval) with temperature=0, no thinking mode.
The model-index metadata in this card's frontmatter mirrors the recommended v2 build, measured with the local llama.cpp benchmark harness on RTX 3090.
Cerebellum v2 (13 GB) — Recommended
| Benchmark | Score | Questions |
|---|---|---|
| ARC-Challenge | 91.6% | 1,172 |
| HellaSwag | 88.9% | 10,042 |
| MMLU | 73.5% | 11,643 |
| HumanEval | 82.3% | 164 |
Size vs Quality
| Model | Size | BPW | PPL (wiki) |
|---|---|---|---|
| Q3_K_M (baseline) | 14 GB | 3.94 | 8.3736 |
| Cerebellum v2 | 13 GB | 3.76 | 8.4912 |
| Cerebellum v1 | 12 GB | 3.50 | 9.1405 |
v2 saves 1 GB (7%) over Q3_K_M with only +1.4% perplexity increase — and the 3 demoted groups actually improved perplexity individually during ablation.
Methodology
- Group ablation: Demote each of 7 weight groups (attn_k, attn_q, attn_v, attn_output, ffn_gate, ffn_up, ffn_down) to Q2_K individually. Measure PPL impact.
- Identify improvers: Three groups (attn_k, attn_q, attn_output) showed lower PPL when demoted — the Q3_K_M precision was actually hurting these layers.
- Build optimal mix: v2 demotes only the 3 groups that improve; v1 additionally demotes attn_v and ffn_gate.
Ablation Results
| Group | PPL when demoted | Delta vs baseline |
|---|---|---|
| attn_k | 8.1639 | -0.2097 (improved!) |
| attn_q | 8.3144 | -0.0592 (improved!) |
| attn_output | 8.3539 | -0.0197 (improved!) |
| attn_v | 8.4713 | +0.0977 |
| ffn_gate | 8.5967 | +0.2231 |
| ffn_up | 9.0587 | +0.6851 |
| ffn_down | 9.0190 | +0.6454 |
v2 Override Map
Demoted (Q2_K): attn_k, attn_q, attn_output (all 64 layers)
Sacred (kept at Q3_K_M): attn_v, ffn_gate, ffn_up, ffn_down
Usage
Works with any llama.cpp-compatible tool:
# llama.cpp
./llama-server --model Granite-4.1-30B-Cerebellum-v2-Q3_K_M.gguf -ngl 99 --ctx-size 4096
# Ollama (create Modelfile pointing to the GGUF)
# LM Studio (drag and drop)
# koboldcpp, text-generation-webui, etc.
Hardware Requirements
- v2 (13 GB): Fits in 16 GB VRAM with room for context. RTX 4060 Ti 16GB, RTX 3090, etc.
- v1 (12 GB): Fits in 16 GB VRAM with generous context, or tight in 12 GB.
Credits
Quantized with Cerebellum — ablation-guided mixed-precision quantization by deucebucket.
Base model by IBM Granite.
- Downloads last month
- 48
3-bit
Model tree for deucebucket/Granite-4.1-30B-Cerebellum-GGUF
Base model
ibm-granite/granite-4.1-30bEvaluation results
- normalized accuracy on AI2 Reasoning Challengetest set Local benchmark run (RTX 3090, llama.cpp)0.916
- accuracy on HellaSwagvalidation set Local benchmark run (RTX 3090, llama.cpp)0.889
- accuracy on MMLUtest set Local benchmark run (RTX 3090, llama.cpp)0.735
- pass@1 on HumanEval (pass@1)test set Local benchmark run (RTX 3090, llama.cpp)0.823
- perplexity on WikiText-2 Perplexitytest set Local benchmark run (RTX 3090, llama.cpp)8.491