spoomplesmaxx-whiskeyjack-12B-i1-GGUF

Weighted (imatrix) GGUF quants of aimeri/spoomplesmaxx-whiskeyjack-12B, quantized on the box that trained it.

Measurements

KL divergence against this model's own bf16, on held-out text that was excluded from the calibration corpus by construction. Not against the base model, not against a benchmark — the number answers one question: how much did quantization change this model.

file GB KLD mean KLD median KLD p99 RMS Δp % same top-1 %
spoomplesmaxx-whiskeyjack-12B.i1-IQ2_M.gguf 4.80 0.2262 0.1187 1.8559 14.50 81.03
spoomplesmaxx-whiskeyjack-12B.i1-IQ3_M.gguf 6.17 0.0545 0.0252 0.4786 7.14 90.44
spoomplesmaxx-whiskeyjack-12B.i1-IQ4_XS.gguf 7.17 0.0287 0.0114 0.2847 5.52 93.46
spoomplesmaxx-whiskeyjack-12B.i1-Q4_K_M.gguf 7.95 0.0221 0.0088 0.1982 5.01 94.21
spoomplesmaxx-whiskeyjack-12B.i1-Q5_K_M.gguf 9.24 0.0102 0.0034 0.1006 3.63 96.42
spoomplesmaxx-whiskeyjack-12B.i1-Q6_K.gguf 10.61 0.0038 0.0012 0.0366 2.16 97.93

Calibration

context 4096 tokens, document-aligned
corpus 11.8M tokens, 2093 documents
sources 100% in-domain (the model's own training corpora)
separator <eos> (a special token — see below)
reference imatrix merged from unsloth/gemma-4-12b-it-GGUF

Every document is truncated to an exact multiple of the calibration context, so llama-imatrix's non-overlapping windows land on document boundaries rather than straddling two unrelated scenes. The separator is a special token because whitespace separators are BPE-mergeable: a document ending in a newline and the next beginning with one can fuse into a single token and shift every subsequent window.

CALIB_CTX is 4096 rather than the 8192 used for larger models in this family, and that is measured rather than inherited: at 8192 only 0.1% of the aviary corpus clears a single window, which would have made calibration ~95% one sub-corpus with no tool-calling coverage at all. Gemma 4 also runs 40 of its 48 layers as sliding-window attention at 1024, so context beyond a few thousand tokens sharpens statistics for only the 8 global layers.

Serving

The stop token is <turn|> (id 106), not <eos>. generation_config.json carries a list and GGUF stores a single u32; this build is verified to have picked <turn|>. Anything that waits for <eos> will run past the turn.

Thinking is selected by <|think|> at the top of the system turn. With it, every model turn opens a <|channel>thought ... <channel|> block; without it, there is no channel at all.

For tool use, a whole episode lives inside ONE <|turn>model: serve with stop=["<tool_call|>"], inject <|tool_response>response:NAME{...}<tool_response|>, and continue the same turn. A harness waiting for <turn|> after a tool call will hang.

Downloads last month
195
GGUF
Model size
13B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF

Quantized
(3)
this model

Collection including aimeri/spoomplesmaxx-whiskeyjack-12B-i1-GGUF