You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Blackfrost

KIMI-K3-DERISKED-MXFP4-GGUF

MXFP4-MOE GGUF of de-risked Kimi K3 · bit-exact experts · fits a single 8×B200 node

Built by Blackfrost · Las Vegas, NV

🔴 Licensed access — $499

Buy Kimi K3 MXFP4 GGUF for $499

Click the red button to purchase.
Enter your Hugging Face username at checkout; access to this gated repository is granted automatically after payment.

Take Notice *

Stock llama.cpp cannot load this. kimi-k3 support exists only in an unmerged pull request — ggml-org/llama.cpp#26185. Released llama.cpp builds and everything downstream of them (LM Studio, Ollama, koboldcpp, brew install llama.cpp) will fail to load these files.

Bugs have been observed. Expect sharp edges. Please report them — open a Community discussion with GPU SKU, driver, llama.cpp commit, full serve flags, prompt, sampler settings, and failure mode (load-time / loop / empty content / OOM).


Why this model exists

Kimi K3's routed experts ship from Moonshot already in MXFP4. GGML's block_mxfp4 stores exactly the same information — 4-bit codes plus E8M0 block scales, group 32 — so the experts can be carried into GGUF as a byte-neutral repack rather than a quantization.

This build (MXFP4_MOE, ftype 38) does exactly that: every routed expert tensor is bit-identical to the parent's MXFP4 packs, while the non-expert tensors — attention, norms, embeddings — are held at Q8_0. That combination lands it at 1,404.8 GiB against 1,426.8 GiB usable on 8× B200.

It is the only K3 build with bit-exact experts that fits a single node. The safetensors parent is 1,453.7 GiB and does not.


Specifications

Architecture Kimi K3 LatentMoE + KDA · general.architecture = kimi-k3
Parent BlackfrostAI/KIMI-K3-DERISKED-MXFP4 — HF safetensors
Transform MXFP4_MOE (ftype 38) · experts MXFP4 bit-exact · non-expert tensors Q8_0
Experts 896 of 896 routed retained — no pruning
On-disk 1,404.8 GiB (1,508.4 GB · 1,508,368,750,816 bytes) · 38 shards
Serve shape Single 8×B200 node · -ngl 99 -nr · ~22 GiB headroom
Status EXPERIMENTAL · NOT RELEASE-CLEARED

What is and isn't lossless here

Bit-exact: every routed expert. MXFP4 → GGML block_mxfp4 was verified on real K3 weights, not a fixture — an expert tensor pulled by ranged read from the parent, run through the converter's packer, and compared: 4-bit codes, E8M0 scale bytes and dequantized floats all bit-identical, and size-neutral. The experts are ~95% of the model.

Not bit-exact: the non-expert tensors — attention projections, norms, embeddings — are Q8_0. That is what buys the ~49 GiB which puts this under a single node's ceiling. Q8_0 is a very light touch, but it is a quantization and this card will not call the whole file lossless.

Why there is no fully-lossless GGUF: carrying non-expert tensors at full width lands above 1,426.8 GiB, i.e. off a single node, which defeats the purpose. If you need every tensor untouched, take the safetensors parent and run it multi-node.


Lineage

Base Official moonshotai/Kimi-K3
Applied Refusal-direction de-risk at the weight level (in parent) · GGUF repack
Not applied Expert pruning · re-quantization of expert packs · additional SFT/DPO
Format GGUF · llama.cpp PR #26185 KV contract

On refusal behaviour: this checkpoint inherits the parent's deliberately reduced refusal surface. It is a Blackfrost de-risked model. Do not evaluate or rate-limit it as if it were a safety-stock derivative of upstream Kimi K3.


Memory envelope — read before sizing context

Weights 1,404.8 GiB
8× B200 usable 1,426.8 GiB (8 × 178.35)
Headroom ~22 GiB

That headroom is real but tight — KV cache and the KDA recurrent state both live in it. K3's per-slot KDA state is ~443 MiB, and MLA KV runs ~27.6 KB/token across the 24 full-attention layers. -c 8192 with a small slot count fits comfortably; large contexts or many parallel slots will not. Size context deliberately rather than assuming the default is safe.

Smaller alternatives: Q2_K at 940.2 GiB leaves 486 GiB of headroom, at the cost of requantizing already-QAT experts.


Measured behaviour

Quality — not yet measured

Benchmark Upstream K3 This (de-risked, lossless) Retention
pending — — —%

Weights are bit-exact, so no quantization loss exists to measure. What is not established is the capability effect of the de-risking intervention itself. A card that ships a refusal-modified frontier checkpoint without publishing what the intervention cost is asking the reader to take it on faith. Harness, conditions and retention figures will be stated here — including any benchmark where retention is poor.

Refusal-rate measurements for the parent intervention are published on KIMI-K3-DERISKED-MXFP4.


Deployment notes

Build the PR branch — mainline will not work:

git clone https://github.com/pwilkin/llama.cpp.git && cd llama.cpp
git fetch origin kimi-k3-text && git checkout cf67f0d24511864d2d3da0769108fd6fc16d00d1
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=100 \
      -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF
cmake --build build --target llama-server -j
./build/bin/llama-server \
  -m Kimi-K3-MXFP4-MOE-00001-of-00038.gguf \
  -ngl 99 -nr -c 8192 --jinja --host 0.0.0.0 --port 8080
  • -nr / --no-repack is mandatory here. Without it llama.cpp repacks MXFP4 experts to mxfp4_8x8 single-threaded — roughly 4 s × 276 tensors ≈ 18 minutes — while allocating a full anonymous copy of the weights on top of the mapping. The symptom is RSS climbing at ~1.8 GB/s at 100% of one core with 0% GPU, apparently forever. With -nr it loads in well under a minute.
  • Point -m at shard 00001 — llama.cpp resolves the remaining 37 automatically.
  • Expert warm-up. llama.cpp does not warm experts, so first tokens are slow while mapped pages fault in.
  • Integrity. Verify shard count and byte totals after download before attributing a load failure to the weights.

Sampling

Follow Moonshot's published guidance — temperature 1.0, top_p 0.95 (1.0 for agentic), reasoning effort max. Do not use greedy decoding: at temperature 0 K3 fails to terminate on longer generations and loops. That is a sampler artifact, not a defect in the weights; generation_config.json sets eos_token_id: 163586 (<|end_of_msg|>).

Thinking is always on. The answer lands in content, the chain of thought in reasoning_content — budget max_tokens generously or content comes back empty with finish_reason: length.

Benign warning at load

load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect

Ignore it. llama-vocab.cpp emits the warning and inserts the token on the next line — self-healing. Vocab metadata here is correct and complete: [BOS] 163584 · [EOS] 163585 · <|end_of_msg|> 163586 · [EOT] 163593. Structural markers: <|open|> 163587 · <|close|> 163588 · <|sep|> 163589.


⚠️ Security: structural markers in untrusted input

K3 builds prompts in Python via encoding_k3.py::build_chat_segments, where each segment carries its own allow_special flag (default False). A Jinja template cannot express that, and llama.cpp tokenizes the rendered prompt in a single pass with parse_special=true.

Consequence: a literal <|end_of_msg|> in user content becomes control token 163586 in llama.cpp, where Moonshot's own tokenizer keeps it as ordinary text. Untrusted input can forge chat structure.

This is a general llama.cpp property, not specific to K3 or this repack. If you feed untrusted text to this model, strip or escape these four markers first: <|open|> · <|sep|> · <|close|> · <|end_of_msg|>


Disclaimer

Refusal behaviour in this checkpoint has been deliberately modified at the weight level. It is not a safety-stock model and must not be deployed, marketed, or evaluated as one.

No warranty of any kind. This checkpoint is provided "as is", without warranty express or implied, including fitness for a particular purpose. Nothing here is a guarantee that any given input will be accepted or refused, that any capability is retained, or that any category of output is unreachable.

Measurements describe what was measured. Refusal rates, throughput and retention figures reflect specific harnesses under specific conditions. They are not safety proofs and do not generalise to multimodal, tool-use, long-context or multi-turn adversarial settings.

Modification by a recipient voids this characterization. Blackfrost's obligations attach at the point of release. Any further ablation, fine-tuning, merging, quantization, or alteration by a recipient produces an artifact Blackfrost has not evaluated and does not stand behind — responsibility for that artifact transfers entirely to whoever produced it.

Operator-owned policy. Open weights mean the operator sets and enforces policy. Deploy only in controlled environments with access control, independent logging and review. Do not deploy where refusal behaviour equivalent to upstream Kimi K3 is assumed or required.


Access & licensing

Access is granted automatically on purchase — you do not wait on a manual review.

➜ Purchase access to this model — $499 — enter your Hugging Face username at checkout, and your account is granted access to this repository within moments of payment.

  • Base licence: Kimi K3 — Moonshot AI's terms apply to this derivative and travel with it.
  • Redistribution: do not redistribute weights outside your grant.
  • Evaluation recommendation: should not be evaluated by processes that assume refusal behaviour equivalent to the parent.
  • Ask us about other quant points, expert budgets, or calibration against your own threat model.

Contact Blackfrost

@Blackfrost_AI on X

DMs are open. Fastest route to a human.

Bug reports are better on the Community tab
so other users can see the fix.

Blackfrost · Las Vegas, Nevada
Frontier model engineering


KIMI-K3-DERISKED-MXFP4-GGUF · © 2026 Blackfrost Softwares Corp.
@Blackfrost_AI

Downloads last month
-
GGUF
Model size
2.8T params
Architecture
kimi-k3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF

Quantized
(2)
this model

Collection including Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF