Instructions to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Use Docker
docker model run hf.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Baekpica/MiMo-V2.6-Flash-RL-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Baekpica/MiMo-V2.6-Flash-RL-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
- Ollama
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Ollama:
ollama run hf.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Docker Model Runner:
docker model run hf.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
- Lemonade
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Run and chat with the model
lemonade run user.MiMo-V2.6-Flash-RL-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
--auth-choice custom-api-key \
--custom-base-url http://127.0.0.1:8080/v1 \
--custom-model-id "Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0" \
--custom-provider-id llama-cpp \
--custom-compatibility openai \
--custom-text-input \
--accept-risk \
--skip-healthRun OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"MiMo-V2.6-Flash-RL GGUF โ calibration reference
This repository holds intermediate artifacts actually generated while building MiMo-V2.6-Flash-RL Mixed-Quant GGUF. The final compact mixed variant belongs in that separate repository.
The source is XiaomiMiMo/MiMo-V2.6-Flash-RL, pinned to 3b38d063180c3e4aed9691fdc735f3d10b266ee4.
Reference representation
The four MXFP4-BF16 shards preserve the original routed experts through an exact MXFP4 repack and expand source FP8 dense matrices to BF16. Control tensors are stored as F32. This is not a full-BF16 source checkpoint or a Q8_0 baseline. It provides original-checkpoint values for importance-matrix collection without recalibrating from the final IQ2 weights.
| Shard | Bytes |
|---|---|
MiMo-V2.6-Flash-RL-MXFP4-BF16-00001-of-00004.gguf |
44,499,128,352 |
MiMo-V2.6-Flash-RL-MXFP4-BF16-00002-of-00004.gguf |
44,493,179,840 |
MiMo-V2.6-Flash-RL-MXFP4-BF16-00003-of-00004.gguf |
44,493,179,840 |
MiMo-V2.6-Flash-RL-MXFP4-BF16-00004-of-00004.gguf |
41,422,766,592 |
| Total | 174,908,254,624 |
Keep all shards together and open the first shard. The language artifact includes the checkpoint's three embedded MTP blocks; this is not evidence of validated speculative decoding. Multimodal encoders and the separate DFlash model are separate components and are not supplied by these four shards alone. They are included as the auxiliary files listed below.
Auxiliary artifacts
| File | Role | Validation |
|---|---|---|
mmproj-MiMo-V2.6-Flash-RL-BF16.gguf |
BF16 image/audio input encoders and projector, about 2.75 GB | One-image native/GGUF numerical comparison; full runtime media validation pending |
MiMo-V2.6-Flash-RL-DFlash-Q8_0.gguf |
Five-layer Q8_0 draft, about 1.56 GB | 63-tensor mapping/shape audit, exact F32 and sampled Q8 error checks; runtime acceptance untested |
dflash/ |
Original config, learned mask embedding, upstream example, and explicit runtime contract | Source files retained byte for byte |
The DFlash GGUF explicitly preserves 64 rotary dimensions per 128-dimensional head and attention value scale 0.612. Its config, sinks, target-layer mapping, and learned mask embedding must be consumed by a compatible runtime. The bundled upstream Python example does not implement every MiMo-specific draft setting; the runtime contract records the production reference. No stock-runtime compatibility is claimed.
The projector records the production MiMo image pixel bounds (8,192โ8,388,608),
alongside production ImageNet normalization. The pixel-bound correction preserved
all 811 tensor payloads; see projector-bounds-audit.json. Both temporal patch
weight slices match the source exactly (video-patch-weight-audit.json).
Validation status
- All 90 downloaded source repository files passed Hub checksum verification.
- Independent MXFP4 repacking and tensor-parallel QKV ordering checks passed.
- All 36,096 expert matrices were covered by a deterministic sampled-row audit: 108,257 rows matched the independently repacked source. See
reference-audit.json; this is not an exhaustive payload comparison. - The reference loaded on a B300 GPU and passed a short arithmetic decode check.
- Original-representation text imatrix collection is running as of 2026-09-22.
- Final mixed-model quality, multimodal end-to-end behavior, MTP/DFlash execution, and DGX Spark serving remain pending. No throughput or benchmark qualification is claimed here.
artifact-manifest.json records exact file sizes and SHA-256 digests. SHA256SUMS can be checked after download:
hf download Baekpica/MiMo-V2.6-Flash-RL-GGUF --local-dir ./MiMo-reference
cd ./MiMo-reference
sha256sum -c SHA256SUMS
Chat template and reproduction
chat_template.jinja is copied byte for byte from the pinned source. Use the original MiMo tokenizer and special-token mapping. Correct image, audio, and video processing additionally requires the corresponding native encoder and input protocol.
The conversion uses a local MXFP4 extension to llama.cpp revision 5836771. Reproduction scripts will accompany the mixed release and private Spark handoff. The reference's approximately 175 GB file size is not the compact DGX Spark target; see the separate mixed model card for that recipe and its current estimates.
License
MIT, inherited from the pinned upstream model.
- Downloads last month
- 561
8-bit
16-bit
Model tree for Baekpica/MiMo-V2.6-Flash-RL-GGUF
Base model
XiaomiMiMo/MiMo-V2.6-Flash-RL
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp# Start a local OpenAI-compatible server: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0