Instructions to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Use Docker
docker model run hf.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Baekpica/MiMo-V2.6-Flash-RL-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Baekpica/MiMo-V2.6-Flash-RL-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
- Ollama
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Ollama:
ollama run hf.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Docker Model Runner:
docker model run hf.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
- Lemonade
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Run and chat with the model
lemonade run user.MiMo-V2.6-Flash-RL-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Baekpica/MiMo-V2.6-Flash-RL-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Baekpica/MiMo-V2.6-Flash-RL-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download dflash-audit.json from Baekpica/MiMo-V2.6-Flash-RL-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 11.7 kB
-
https://huggingface.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF/resolve/main/dflash-audit.json
- Command line
-
hf download hf://Baekpica/MiMo-V2.6-Flash-RL-GGUF/dflash-audit.json
-
curl -L -o dflash-audit.json https://huggingface.co/Baekpica/MiMo-V2.6-Flash-RL-GGUF/resolve/main/dflash-audit.json
11.7 kB
| { | |
| "passed": true, | |
| "tensor_count": 63, | |
| "source_to_gguf_mapping": "exact set and shape match", | |
| "f32_check": "all elements equal", | |
| "q8_check": "three deterministic rows per matrix, relative L2 below 0.03", | |
| "metadata": { | |
| "dflash.block_count": 5, | |
| "dflash.context_length": 1048576, | |
| "dflash.embedding_length": 4096, | |
| "dflash.feed_forward_length": 16384, | |
| "dflash.attention.head_count": 64, | |
| "dflash.attention.head_count_kv": 8, | |
| "dflash.attention.causal": false, | |
| "dflash.rope.freq_base": 10000.0, | |
| "dflash.attention.layer_norm_rms_epsilon": 9.999999974752427e-07, | |
| "dflash.attention.key_length": 128, | |
| "dflash.attention.value_length": 128, | |
| "dflash.block_size": 8, | |
| "dflash.target_layers": [ | |
| 1, | |
| 12, | |
| 24, | |
| 36, | |
| 48 | |
| ], | |
| "dflash.attention.sliding_window": 1024, | |
| "dflash.attention.sliding_window_pattern": [ | |
| true, | |
| true, | |
| true, | |
| true, | |
| true | |
| ], | |
| "dflash.rope.dimension_count": 64, | |
| "dflash.attention.value_scale": 0.6119999885559082 | |
| }, | |
| "rows": [ | |
| { | |
| "source": "fc.weight", | |
| "gguf": "fc.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005588418338447809 | |
| }, | |
| { | |
| "source": "hidden_norm.weight", | |
| "gguf": "enc.output_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.0.input_layernorm.weight", | |
| "gguf": "blk.0.attn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.0.mlp.down_proj.weight", | |
| "gguf": "blk.0.ffn_down.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005761090200394392 | |
| }, | |
| { | |
| "source": "layers.0.mlp.gate_proj.weight", | |
| "gguf": "blk.0.ffn_gate.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005506220273673534 | |
| }, | |
| { | |
| "source": "layers.0.mlp.up_proj.weight", | |
| "gguf": "blk.0.ffn_up.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005492477677762508 | |
| }, | |
| { | |
| "source": "layers.0.post_attention_layernorm.weight", | |
| "gguf": "blk.0.ffn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.0.self_attn.attention_sink_bias", | |
| "gguf": "blk.0.attn_sinks.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.0.self_attn.k_norm.weight", | |
| "gguf": "blk.0.attn_k_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.0.self_attn.k_proj.weight", | |
| "gguf": "blk.0.attn_k.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005711852107197046 | |
| }, | |
| { | |
| "source": "layers.0.self_attn.o_proj.weight", | |
| "gguf": "blk.0.attn_output.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.00542857451364398 | |
| }, | |
| { | |
| "source": "layers.0.self_attn.q_norm.weight", | |
| "gguf": "blk.0.attn_q_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.0.self_attn.q_proj.weight", | |
| "gguf": "blk.0.attn_q.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.00566287524998188 | |
| }, | |
| { | |
| "source": "layers.0.self_attn.v_proj.weight", | |
| "gguf": "blk.0.attn_v.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005591947585344315 | |
| }, | |
| { | |
| "source": "layers.1.input_layernorm.weight", | |
| "gguf": "blk.1.attn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.1.mlp.down_proj.weight", | |
| "gguf": "blk.1.ffn_down.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005702429451048374 | |
| }, | |
| { | |
| "source": "layers.1.mlp.gate_proj.weight", | |
| "gguf": "blk.1.ffn_gate.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.0054396879859268665 | |
| }, | |
| { | |
| "source": "layers.1.mlp.up_proj.weight", | |
| "gguf": "blk.1.ffn_up.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005450685042887926 | |
| }, | |
| { | |
| "source": "layers.1.post_attention_layernorm.weight", | |
| "gguf": "blk.1.ffn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.1.self_attn.attention_sink_bias", | |
| "gguf": "blk.1.attn_sinks.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.1.self_attn.k_norm.weight", | |
| "gguf": "blk.1.attn_k_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.1.self_attn.k_proj.weight", | |
| "gguf": "blk.1.attn_k.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005590083077549934 | |
| }, | |
| { | |
| "source": "layers.1.self_attn.o_proj.weight", | |
| "gguf": "blk.1.attn_output.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005548989400267601 | |
| }, | |
| { | |
| "source": "layers.1.self_attn.q_norm.weight", | |
| "gguf": "blk.1.attn_q_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.1.self_attn.q_proj.weight", | |
| "gguf": "blk.1.attn_q.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.00569014111533761 | |
| }, | |
| { | |
| "source": "layers.1.self_attn.v_proj.weight", | |
| "gguf": "blk.1.attn_v.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005482448264956474 | |
| }, | |
| { | |
| "source": "layers.2.input_layernorm.weight", | |
| "gguf": "blk.2.attn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.2.mlp.down_proj.weight", | |
| "gguf": "blk.2.ffn_down.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005693534389138222 | |
| }, | |
| { | |
| "source": "layers.2.mlp.gate_proj.weight", | |
| "gguf": "blk.2.ffn_gate.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005409534554928541 | |
| }, | |
| { | |
| "source": "layers.2.mlp.up_proj.weight", | |
| "gguf": "blk.2.ffn_up.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005460306536406279 | |
| }, | |
| { | |
| "source": "layers.2.post_attention_layernorm.weight", | |
| "gguf": "blk.2.ffn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.2.self_attn.attention_sink_bias", | |
| "gguf": "blk.2.attn_sinks.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.2.self_attn.k_norm.weight", | |
| "gguf": "blk.2.attn_k_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.2.self_attn.k_proj.weight", | |
| "gguf": "blk.2.attn_k.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.00555426487699151 | |
| }, | |
| { | |
| "source": "layers.2.self_attn.o_proj.weight", | |
| "gguf": "blk.2.attn_output.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005506662651896477 | |
| }, | |
| { | |
| "source": "layers.2.self_attn.q_norm.weight", | |
| "gguf": "blk.2.attn_q_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.2.self_attn.q_proj.weight", | |
| "gguf": "blk.2.attn_q.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.0055401320569217205 | |
| }, | |
| { | |
| "source": "layers.2.self_attn.v_proj.weight", | |
| "gguf": "blk.2.attn_v.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005393540486693382 | |
| }, | |
| { | |
| "source": "layers.3.input_layernorm.weight", | |
| "gguf": "blk.3.attn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.3.mlp.down_proj.weight", | |
| "gguf": "blk.3.ffn_down.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005636085756123066 | |
| }, | |
| { | |
| "source": "layers.3.mlp.gate_proj.weight", | |
| "gguf": "blk.3.ffn_gate.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005447070114314556 | |
| }, | |
| { | |
| "source": "layers.3.mlp.up_proj.weight", | |
| "gguf": "blk.3.ffn_up.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005448319483548403 | |
| }, | |
| { | |
| "source": "layers.3.post_attention_layernorm.weight", | |
| "gguf": "blk.3.ffn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.3.self_attn.attention_sink_bias", | |
| "gguf": "blk.3.attn_sinks.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.3.self_attn.k_norm.weight", | |
| "gguf": "blk.3.attn_k_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.3.self_attn.k_proj.weight", | |
| "gguf": "blk.3.attn_k.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005607506725937128 | |
| }, | |
| { | |
| "source": "layers.3.self_attn.o_proj.weight", | |
| "gguf": "blk.3.attn_output.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005487577058374882 | |
| }, | |
| { | |
| "source": "layers.3.self_attn.q_norm.weight", | |
| "gguf": "blk.3.attn_q_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.3.self_attn.q_proj.weight", | |
| "gguf": "blk.3.attn_q.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005722770467400551 | |
| }, | |
| { | |
| "source": "layers.3.self_attn.v_proj.weight", | |
| "gguf": "blk.3.attn_v.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005510271061211824 | |
| }, | |
| { | |
| "source": "layers.4.input_layernorm.weight", | |
| "gguf": "blk.4.attn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.4.mlp.down_proj.weight", | |
| "gguf": "blk.4.ffn_down.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005580178461968899 | |
| }, | |
| { | |
| "source": "layers.4.mlp.gate_proj.weight", | |
| "gguf": "blk.4.ffn_gate.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005537654273211956 | |
| }, | |
| { | |
| "source": "layers.4.mlp.up_proj.weight", | |
| "gguf": "blk.4.ffn_up.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005511666182428598 | |
| }, | |
| { | |
| "source": "layers.4.post_attention_layernorm.weight", | |
| "gguf": "blk.4.ffn_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.4.self_attn.attention_sink_bias", | |
| "gguf": "blk.4.attn_sinks.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.4.self_attn.k_norm.weight", | |
| "gguf": "blk.4.attn_k_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.4.self_attn.k_proj.weight", | |
| "gguf": "blk.4.attn_k.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005514967255294323 | |
| }, | |
| { | |
| "source": "layers.4.self_attn.o_proj.weight", | |
| "gguf": "blk.4.attn_output.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005578942131251097 | |
| }, | |
| { | |
| "source": "layers.4.self_attn.q_norm.weight", | |
| "gguf": "blk.4.attn_q_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| }, | |
| { | |
| "source": "layers.4.self_attn.q_proj.weight", | |
| "gguf": "blk.4.attn_q.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005511818453669548 | |
| }, | |
| { | |
| "source": "layers.4.self_attn.v_proj.weight", | |
| "gguf": "blk.4.attn_v.weight", | |
| "type": "Q8_0", | |
| "sampled_relative_l2": 0.005557454191148281 | |
| }, | |
| { | |
| "source": "norm.weight", | |
| "gguf": "output_norm.weight", | |
| "type": "F32", | |
| "sampled_relative_l2": 0.0 | |
| } | |
| ], | |
| "scope": "Conversion integrity; speculative decoding and acceptance rate not tested." | |
| } | |