Instructions to use Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4 # Run inference directly in the terminal: llama cli -hf Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4 # Run inference directly in the terminal: llama cli -hf Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4 # Run inference directly in the terminal: ./llama-cli -hf Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4
Use Docker
docker model run hf.co/Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4
- LM Studio
- Jan
- vLLM
How to use Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4
- Ollama
How to use Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF with Ollama:
ollama run hf.co/Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4
- Unsloth Desktop
- Docker Model Runner
How to use Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF with Docker Model Runner:
docker model run hf.co/Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4
- Lemonade
How to use Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF:MXFP4
Run and chat with the model
lemonade run user.KIMI-K3-DERISKED-MXFP4-GGUF-MXFP4
List all available models
lemonade list
- Atomic Chat
KIMI-K3-DERISKED-MXFP4-GGUF
MXFP4-MOE GGUF of de-risked Kimi K3 · bit-exact experts · fits a single 8×B200 node
Built by Blackfrost · Las Vegas, NV
🔴 Licensed access — $499
Click the red button to purchase.
Enter your Hugging Face username at checkout; access to this gated repository is granted automatically after payment.
Take Notice *
Stock llama.cpp cannot load this.
kimi-k3support exists only in an unmerged pull request — ggml-org/llama.cpp#26185. Released llama.cpp builds and everything downstream of them (LM Studio, Ollama, koboldcpp,brew install llama.cpp) will fail to load these files.Bugs have been observed. Expect sharp edges. Please report them — open a Community discussion with GPU SKU, driver, llama.cpp commit, full serve flags, prompt, sampler settings, and failure mode (load-time / loop / empty
content/ OOM).
Why this model exists
Kimi K3's routed experts ship from Moonshot already in MXFP4. GGML's block_mxfp4 stores
exactly the same information — 4-bit codes plus E8M0 block scales, group 32 — so the experts can be
carried into GGUF as a byte-neutral repack rather than a quantization.
This build (MXFP4_MOE, ftype 38) does exactly that: every routed expert tensor is bit-identical
to the parent's MXFP4 packs, while the non-expert tensors — attention, norms, embeddings — are
held at Q8_0. That combination lands it at 1,404.8 GiB against 1,426.8 GiB usable on 8× B200.
It is the only K3 build with bit-exact experts that fits a single node. The safetensors parent is 1,453.7 GiB and does not.
Specifications
| Architecture | Kimi K3 LatentMoE + KDA · general.architecture = kimi-k3 |
| Parent | BlackfrostAI/KIMI-K3-DERISKED-MXFP4 — HF safetensors |
| Transform | MXFP4_MOE (ftype 38) · experts MXFP4 bit-exact · non-expert tensors Q8_0 |
| Experts | 896 of 896 routed retained — no pruning |
| On-disk | 1,404.8 GiB (1,508.4 GB · 1,508,368,750,816 bytes) · 38 shards |
| Serve shape | Single 8×B200 node · -ngl 99 -nr · ~22 GiB headroom |
| Status | EXPERIMENTAL · NOT RELEASE-CLEARED |
What is and isn't lossless here
Bit-exact: every routed expert. MXFP4 → GGML block_mxfp4 was verified on real K3 weights,
not a fixture — an expert tensor pulled by ranged read from the parent, run through the converter's
packer, and compared: 4-bit codes, E8M0 scale bytes and dequantized floats all bit-identical,
and size-neutral. The experts are ~95% of the model.
Not bit-exact: the non-expert tensors — attention projections, norms, embeddings — are Q8_0. That is what buys the ~49 GiB which puts this under a single node's ceiling. Q8_0 is a very light touch, but it is a quantization and this card will not call the whole file lossless.
Why there is no fully-lossless GGUF: carrying non-expert tensors at full width lands above 1,426.8 GiB, i.e. off a single node, which defeats the purpose. If you need every tensor untouched, take the safetensors parent and run it multi-node.
Lineage
| Base | Official moonshotai/Kimi-K3 |
| Applied | Refusal-direction de-risk at the weight level (in parent) · GGUF repack |
| Not applied | Expert pruning · re-quantization of expert packs · additional SFT/DPO |
| Format | GGUF · llama.cpp PR #26185 KV contract |
On refusal behaviour: this checkpoint inherits the parent's deliberately reduced refusal surface. It is a Blackfrost de-risked model. Do not evaluate or rate-limit it as if it were a safety-stock derivative of upstream Kimi K3.
Memory envelope — read before sizing context
| Weights | 1,404.8 GiB |
| 8× B200 usable | 1,426.8 GiB (8 × 178.35) |
| Headroom | ~22 GiB |
That headroom is real but tight — KV cache and the KDA recurrent state both live in it. K3's
per-slot KDA state is ~443 MiB, and MLA KV runs ~27.6 KB/token across the 24 full-attention layers.
-c 8192 with a small slot count fits comfortably; large contexts or many parallel slots will not.
Size context deliberately rather than assuming the default is safe.
Smaller alternatives: Q2_K
at 940.2 GiB leaves 486 GiB of headroom, at the cost of requantizing already-QAT experts.
Measured behaviour
Quality — not yet measured
| Benchmark | Upstream K3 | This (de-risked, lossless) | Retention |
|---|---|---|---|
| pending | — | — | —% |
Weights are bit-exact, so no quantization loss exists to measure. What is not established is the capability effect of the de-risking intervention itself. A card that ships a refusal-modified frontier checkpoint without publishing what the intervention cost is asking the reader to take it on faith. Harness, conditions and retention figures will be stated here — including any benchmark where retention is poor.
Refusal-rate measurements for the parent intervention are published on KIMI-K3-DERISKED-MXFP4.
Deployment notes
Build the PR branch — mainline will not work:
git clone https://github.com/pwilkin/llama.cpp.git && cd llama.cpp
git fetch origin kimi-k3-text && git checkout cf67f0d24511864d2d3da0769108fd6fc16d00d1
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=100 \
-DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF
cmake --build build --target llama-server -j
./build/bin/llama-server \
-m Kimi-K3-MXFP4-MOE-00001-of-00038.gguf \
-ngl 99 -nr -c 8192 --jinja --host 0.0.0.0 --port 8080
-nr/--no-repackis mandatory here. Without it llama.cpp repacks MXFP4 experts tomxfp4_8x8single-threaded — roughly 4 s × 276 tensors ≈ 18 minutes — while allocating a full anonymous copy of the weights on top of the mapping. The symptom is RSS climbing at ~1.8 GB/s at 100% of one core with 0% GPU, apparently forever. With-nrit loads in well under a minute.- Point
-mat shard00001— llama.cpp resolves the remaining 37 automatically. - Expert warm-up. llama.cpp does not warm experts, so first tokens are slow while mapped pages fault in.
- Integrity. Verify shard count and byte totals after download before attributing a load failure to the weights.
Sampling
Follow Moonshot's published guidance — temperature 1.0, top_p 0.95 (1.0 for agentic), reasoning effort max. Do not use greedy decoding: at temperature 0 K3 fails to terminate on longer generations and loops. That is a sampler artifact, not a defect in the weights; generation_config.json sets eos_token_id: 163586 (<|end_of_msg|>).
Thinking is always on. The answer lands in content, the chain of thought in reasoning_content — budget max_tokens generously or content comes back empty with finish_reason: length.
Benign warning at load
load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect
Ignore it. llama-vocab.cpp emits the warning and inserts the token on the next line — self-healing. Vocab metadata here is correct and complete: [BOS] 163584 · [EOS] 163585 · <|end_of_msg|> 163586 · [EOT] 163593. Structural markers: <|open|> 163587 · <|close|> 163588 · <|sep|> 163589.
⚠️ Security: structural markers in untrusted input
K3 builds prompts in Python via encoding_k3.py::build_chat_segments, where each segment carries its own allow_special flag (default False). A Jinja template cannot express that, and llama.cpp tokenizes the rendered prompt in a single pass with parse_special=true.
Consequence: a literal <|end_of_msg|> in user content becomes control token 163586 in llama.cpp, where Moonshot's own tokenizer keeps it as ordinary text. Untrusted input can forge chat structure.
This is a general llama.cpp property, not specific to K3 or this repack. If you feed untrusted text to this model, strip or escape these four markers first: <|open|> · <|sep|> · <|close|> · <|end_of_msg|>
Disclaimer
Refusal behaviour in this checkpoint has been deliberately modified at the weight level. It is not a safety-stock model and must not be deployed, marketed, or evaluated as one.
No warranty of any kind. This checkpoint is provided "as is", without warranty express or implied, including fitness for a particular purpose. Nothing here is a guarantee that any given input will be accepted or refused, that any capability is retained, or that any category of output is unreachable.
Measurements describe what was measured. Refusal rates, throughput and retention figures reflect specific harnesses under specific conditions. They are not safety proofs and do not generalise to multimodal, tool-use, long-context or multi-turn adversarial settings.
Modification by a recipient voids this characterization. Blackfrost's obligations attach at the point of release. Any further ablation, fine-tuning, merging, quantization, or alteration by a recipient produces an artifact Blackfrost has not evaluated and does not stand behind — responsibility for that artifact transfers entirely to whoever produced it.
Operator-owned policy. Open weights mean the operator sets and enforces policy. Deploy only in controlled environments with access control, independent logging and review. Do not deploy where refusal behaviour equivalent to upstream Kimi K3 is assumed or required.
Access & licensing
Access is granted automatically on purchase — you do not wait on a manual review.
➜ Purchase access to this model — $499 — enter your Hugging Face username at checkout, and your account is granted access to this repository within moments of payment.
- Base licence: Kimi K3 — Moonshot AI's terms apply to this derivative and travel with it.
- Redistribution: do not redistribute weights outside your grant.
- Evaluation recommendation: should not be evaluated by processes that assume refusal behaviour equivalent to the parent.
- Ask us about other quant points, expert budgets, or calibration against your own threat model.
Contact Blackfrost
@Blackfrost_AI on X
DMs are open. Fastest route to a human.
Bug reports are better on the Community tab
so other users can see the fix.
Blackfrost · Las Vegas, Nevada
Frontier model engineering
KIMI-K3-DERISKED-MXFP4-GGUF · © 2026 Blackfrost Softwares Corp.
@Blackfrost_AI
- Downloads last month
- -
4-bit
Model tree for Blackfrost-Research/KIMI-K3-DERISKED-MXFP4-GGUF
Base model
Blackfrost-Research/KIMI-K3-DERISKED-MXFP4