CYBER-FROST-3.8-PS-GUFF

Quantized GGUF of Blackfrost-AI/CYBER-FROST-3.8-BF16 โ€” a cybersecurity-tuned 3.8-generation Qwen4-Exp MoE โ€” built by peasantsmith.

This is a biased-precision quant: rather than applying one global bit-width, each tensor class is placed at the precision where it does the most good, guided by imatrix-weighted reconstruction error and a floor determined from the ISTA-DASLab reference quant.

Model summary

Base model Blackfrost-AI/CYBER-FROST-3.8-BF16
Upstream Qwen/Qwen3.8-Flash-Next
Architecture qwen4exp (MoE, hybrid SSM/attention, MTP draft head)
Layers 49 (48 transformer + 1 NextN/MTP)
Experts 512 total / 10 active per token
Context length 262,144 tokens
n-gram table per_layer_token_embd, n-gram size 3
Total parameters 179.77 B
File size 112.53 GB
Bits per weight 5.01

Parameter distribution

Counted directly from the tensor table:

Component Parameters Share
Routed experts (ffn_{gate,up,down}_exps) 123.31 B 68.6%
n-gram table (per_layer_token_embd) 51.20 B 28.5%
Everything else (attention, dense, router, embeddings, norms) 5.26 B 2.9%
Total 179.77 B 100%

The routed experts and the n-gram table are 97% of the model between them โ€” those two components are where any future size/quality trade-off would be made.

Quantization comparison

Comparison against the ISTA-DASLab Qwen3.8-Flash-Next GSQ-RCO IQ3_S reference quant:

Tensor class ISTA-DASLab IQ3_S CYBER-FROST-3.8-PS-GUFF Our choice
ffn_gate_exps (routed experts) IQ3_S Q5_K 2 steps up
ffn_up_exps (routed experts) IQ3_S Q5_K 2 steps up
ffn_down_exps (routed experts) IQ3_S IQ4_NL 3 steps up
per_layer_token_embd (n-gram table) IQ4_NL IQ4_NL matched
ffn_gate_inp (router) โ€” F32 full precision
Attention / dense / norms higher precision higher precision matched

Design intent: the routed experts carry the bulk of the file and are where low-bit quantization does the most damage, so they are raised well above the reference. The n-gram table is deliberately held at the same IQ4_NL as the reference, so the PLE path is like-for-like and cannot flatter the comparison.

Nothing is quantized below the reference floor. Every tensor class is either equal to or more precise than its ISTA-DASLab counterpart.

On the Q5_K_M label. The filename carries Q5_K_M so the Hub renders a quant badge. This is a naming convention, not a literal description: the build is a per-tensor mix (see the table above), with Q5_K on ffn_gate_exps/ffn_up_exps, IQ4_NL on ffn_down_exps/per_layer_token_embd, and F32 on the router. No single label fits a mixed build; Q5_K_M reflects the dominant expert precision.

Quantization scheme

Tensor group Type
ffn_gate_exps Q5_K
ffn_up_exps Q5_K
ffn_down_exps IQ4_NL
per_layer_token_embd (n-gram table) IQ4_NL
ffn_gate_inp (router) F32

Imatrix and calibration quality

Quantization was guided by an importance matrix trained on a large, project-specific calibration corpus rather than a generic one.

Imatrix entries 902
Calibration chunks 800
Corpus project-specific cybersecurity and security-research text, plus broad general text
Corpus coverage ~99.8% of the tokenizer vocabulary observed during calibration

800 chunks over ~902 imatrix entries gives dense, near-complete coverage of the vocabulary โ€” a wider and more domain-appropriate calibration set than a small or generic default.

See the quantize.imatrix.* fields in the GGUF header for the recorded provenance.

Usage

llama-server -m CYBER-FROST-3.8-PS-GUFF-Q5_K_M.gguf --n-cpu-moe 18 --lazy-mode on -ctk q8_0 -ctv q8_0 -c 4096 -ngl 999

Or load directly from the Hub:

llama-server -hf peasantsmith/CYBER-FROST-3.8-PS-GUFF:Q5_K_M --n-cpu-moe 18 --lazy-mode on -ctk q8_0 -ctv q8_0 -c 4096 -ngl 999

Adjust --n-cpu-moe, -ngl and the context size to fit your own VRAM and RAM. The values above are a conservative starting point for a machine that cannot hold the full model in VRAM. On a machine with enough VRAM, offload more layers and raise --n-cpu-moe accordingly.

Reasoning output: this model emits chain-of-thought. Depending on the chat template and build, the answer may arrive in reasoning_content with content null. Allow a generous max_tokens โ€” a tight cap causes the model to run out of budget mid-reasoning and return no answer at all.

MTP / draft head: the model carries a NextN/MTP head. If your build supports speculative decoding, enabling it can improve throughput.

MMLU-Pro โ€” 100 questions

ISTA-DASLab IQ3_S CYBER-FROST-3.8-PS-GUFF
Accuracy (% of scored) 80.00% 82.80%
Accuracy (% of all) 76.00% 77.00%

Provenance

Built from the BF16 source with an 800-chunk imatrix. Imatrix provenance is recorded in the GGUF header (quantize.imatrix.*). Conversion and quantization were performed with a pinned llama.cpp build; no tensors were sourced from a third-party quant.

License

Inherited from the base model: qwen-community-license-1.0. See the base model repository for the full text. This quant is a derivative work and carries the same terms.

Citation

@misc{cyber-frost-3.8-ps-guff,
  title  = {CYBER-FROST-3.8-PS-GUFF},
  author = {peasantsmith},
  year   = {2026},
  note   = {Biased-precision GGUF quant of Blackfrost-AI/CYBER-FROST-3.8-BF16},
  url    = {https://huggingface.co/peasantsmith/CYBER-FROST-3.8-PS-GUFF}
}
Downloads last month
9,395
GGUF
Model size
180B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for peasantsmith/CYBER-FROST-3.8-PS-GUFF

Quantized
(11)
this model

Space using peasantsmith/CYBER-FROST-3.8-PS-GUFF 1