Instructions to use aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
Use Docker
docker model run hf.co/aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF with Ollama:
ollama run hf.co/aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF with Docker Model Runner:
docker model run hf.co/aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
- Lemonade
How to use aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.spoomplesmaxx-whiskeyjack-12B-i1-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
spoomplesmaxx-whiskeyjack-12B-i1-GGUF
Weighted (imatrix) GGUF quants of aimeri/spoomplesmaxx-whiskeyjack-12B,
quantized on the box that trained it.
Measurements
KL divergence against this model's own bf16, on held-out text that was excluded from the calibration corpus by construction. Not against the base model, not against a benchmark — the number answers one question: how much did quantization change this model.
| file | GB | KLD mean | KLD median | KLD p99 | RMS Δp % | same top-1 % |
|---|---|---|---|---|---|---|
spoomplesmaxx-whiskeyjack-12B.i1-IQ2_M.gguf |
4.80 | 0.2262 | 0.1187 | 1.8559 | 14.50 | 81.03 |
spoomplesmaxx-whiskeyjack-12B.i1-IQ3_M.gguf |
6.17 | 0.0545 | 0.0252 | 0.4786 | 7.14 | 90.44 |
spoomplesmaxx-whiskeyjack-12B.i1-IQ4_XS.gguf |
7.17 | 0.0287 | 0.0114 | 0.2847 | 5.52 | 93.46 |
spoomplesmaxx-whiskeyjack-12B.i1-Q4_K_M.gguf |
7.95 | 0.0221 | 0.0088 | 0.1982 | 5.01 | 94.21 |
spoomplesmaxx-whiskeyjack-12B.i1-Q5_K_M.gguf |
9.24 | 0.0102 | 0.0034 | 0.1006 | 3.63 | 96.42 |
spoomplesmaxx-whiskeyjack-12B.i1-Q6_K.gguf |
10.61 | 0.0038 | 0.0012 | 0.0366 | 2.16 | 97.93 |
Calibration
| context | 4096 tokens, document-aligned |
| corpus | 11.8M tokens, 2093 documents |
| sources | 100% in-domain (the model's own training corpora) |
| separator | <eos> (a special token — see below) |
| reference imatrix | merged from unsloth/gemma-4-12b-it-GGUF |
Every document is truncated to an exact multiple of the calibration context, so
llama-imatrix's non-overlapping windows land on document boundaries rather than
straddling two unrelated scenes. The separator is a special token because
whitespace separators are BPE-mergeable: a document ending in a newline and the
next beginning with one can fuse into a single token and shift every subsequent
window.
CALIB_CTX is 4096 rather than the 8192 used for larger models in this
family, and that is measured rather than inherited: at 8192 only 0.1% of the
aviary corpus clears a single window, which would have made calibration ~95% one
sub-corpus with no tool-calling coverage at all. Gemma 4 also runs 40 of its 48
layers as sliding-window attention at 1024, so context beyond a few thousand
tokens sharpens statistics for only the 8 global layers.
Serving
The stop token is <turn|> (id 106), not <eos>. generation_config.json
carries a list and GGUF stores a single u32; this build is verified to have
picked <turn|>. Anything that waits for <eos> will run past the turn.
Thinking is selected by <|think|> at the top of the system turn. With it,
every model turn opens a <|channel>thought ... <channel|> block; without it,
there is no channel at all.
For tool use, a whole episode lives inside ONE <|turn>model: serve with
stop=["<tool_call|>"], inject <|tool_response>response:NAME{...}<tool_response|>,
and continue the same turn. A harness waiting for <turn|> after a tool call
will hang.
- Downloads last month
- 195
2-bit
3-bit
4-bit
5-bit
6-bit
Model tree for aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF
Base model
google/gemma-4-12B