Instructions to use ggml-org/MiMo-V2.6-Flash-RL-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ggml-org/MiMo-V2.6-Flash-RL-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4 # Run inference directly in the terminal: llama cli -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4 # Run inference directly in the terminal: llama cli -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4 # Run inference directly in the terminal: ./llama-cli -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
Use Docker
docker model run hf.co/ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
- LM Studio
- Jan
- vLLM
How to use ggml-org/MiMo-V2.6-Flash-RL-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ggml-org/MiMo-V2.6-Flash-RL-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ggml-org/MiMo-V2.6-Flash-RL-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
- Ollama
How to use ggml-org/MiMo-V2.6-Flash-RL-GGUF with Ollama:
ollama run hf.co/ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
- Unsloth Desktop
- Pi
How to use ggml-org/MiMo-V2.6-Flash-RL-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ggml-org/MiMo-V2.6-Flash-RL-GGUF with Docker Model Runner:
docker model run hf.co/ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
- Lemonade
How to use ggml-org/MiMo-V2.6-Flash-RL-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
Run and chat with the model
lemonade run user.MiMo-V2.6-Flash-RL-GGUF-MXFP4
List all available models
lemonade list
- Hermes Agent
How to use ggml-org/MiMo-V2.6-Flash-RL-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ggml-org/MiMo-V2.6-Flash-RL-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload folder using huggingface_hub
Browse files- convert.log +310 -322
- mtp-MiMo-V2.6-Flash-RL-BF16.gguf +1 -1
- mtp-MiMo-V2.6-Flash-RL-MXFP4.gguf +1 -1
- mtp-MiMo-V2.6-Flash-RL-Q8_0.gguf +1 -1
convert.log
CHANGED
|
@@ -690,305 +690,293 @@ INFO:gguf.gguf_writer:upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16-00002-of
|
|
| 690 |
|
| 691 |
|
| 692 |
|
| 693 |
-
|
| 694 |
-
|
| 695 |
-
|
| 696 |
-
|
| 697 |
-
|
| 698 |
-
|
| 699 |
-
|
| 700 |
-
|
| 701 |
-
|
| 702 |
-
|
| 703 |
-
|
| 704 |
-
|
| 705 |
-
|
| 706 |
-
|
| 707 |
-
|
| 708 |
-
|
| 709 |
-
|
| 710 |
-
|
| 711 |
-
|
| 712 |
-
|
| 713 |
-
|
| 714 |
-
|
| 715 |
-
|
| 716 |
-
|
| 717 |
-
|
| 718 |
-
|
| 719 |
-
|
| 720 |
-
|
| 721 |
-
|
| 722 |
-
|
| 723 |
-
|
| 724 |
-
|
| 725 |
-
|
| 726 |
-
|
| 727 |
-
|
| 728 |
-
|
| 729 |
-
|
| 730 |
-
|
| 731 |
-
|
| 732 |
-
|
| 733 |
-
|
| 734 |
-
|
| 735 |
-
|
| 736 |
-
|
| 737 |
-
|
| 738 |
-
|
| 739 |
-
|
| 740 |
-
|
| 741 |
-
|
| 742 |
-
|
| 743 |
-
|
| 744 |
-
|
| 745 |
-
|
| 746 |
-
|
| 747 |
-
|
| 748 |
-
|
| 749 |
-
|
| 750 |
-
|
| 751 |
-
|
| 752 |
-
|
| 753 |
-
|
| 754 |
-
|
| 755 |
-
|
| 756 |
-
|
| 757 |
-
|
| 758 |
-
|
| 759 |
-
|
| 760 |
-
|
| 761 |
-
|
| 762 |
-
|
| 763 |
-
|
| 764 |
-
|
| 765 |
-
|
| 766 |
-
|
| 767 |
-
|
| 768 |
-
|
| 769 |
-
|
| 770 |
-
|
| 771 |
-
|
| 772 |
-
|
| 773 |
-
|
| 774 |
-
|
| 775 |
-
|
| 776 |
-
|
| 777 |
-
|
| 778 |
-
|
| 779 |
-
|
| 780 |
-
|
| 781 |
-
|
| 782 |
-
|
| 783 |
-
|
| 784 |
-
|
| 785 |
-
|
| 786 |
-
|
| 787 |
-
|
| 788 |
-
|
| 789 |
-
|
| 790 |
-
|
| 791 |
-
|
| 792 |
-
|
| 793 |
-
|
| 794 |
-
|
| 795 |
-
|
| 796 |
-
|
| 797 |
-
|
| 798 |
-
|
| 799 |
-
|
| 800 |
-
|
| 801 |
-
|
| 802 |
-
|
| 803 |
-
|
| 804 |
-
|
| 805 |
-
|
| 806 |
-
|
| 807 |
-
|
| 808 |
-
|
| 809 |
-
|
| 810 |
-
|
| 811 |
-
|
| 812 |
-
|
| 813 |
-
|
| 814 |
-
|
| 815 |
-
|
| 816 |
-
|
| 817 |
-
|
| 818 |
-
|
| 819 |
-
|
| 820 |
-
|
| 821 |
-
|
| 822 |
-
|
| 823 |
-
|
| 824 |
-
|
| 825 |
-
|
| 826 |
-
|
| 827 |
-
|
| 828 |
-
|
| 829 |
-
|
| 830 |
-
|
| 831 |
-
|
| 832 |
-
|
| 833 |
-
|
| 834 |
-
|
| 835 |
-
|
| 836 |
-
|
| 837 |
-
|
| 838 |
-
|
| 839 |
-
|
| 840 |
-
|
| 841 |
-
|
| 842 |
-
|
| 843 |
-
|
| 844 |
-
|
| 845 |
-
|
| 846 |
-
|
| 847 |
-
|
| 848 |
-
|
| 849 |
-
|
| 850 |
-
|
| 851 |
-
|
| 852 |
-
|
| 853 |
-
|
| 854 |
-
|
| 855 |
-
|
| 856 |
-
|
| 857 |
-
|
| 858 |
-
|
| 859 |
-
|
| 860 |
-
|
| 861 |
-
|
| 862 |
-
|
| 863 |
-
|
| 864 |
-
|
| 865 |
-
|
| 866 |
-
|
| 867 |
-
|
| 868 |
-
|
| 869 |
-
|
| 870 |
-
|
| 871 |
-
|
| 872 |
-
|
| 873 |
-
|
| 874 |
-
|
| 875 |
-
|
| 876 |
-
|
| 877 |
-
|
| 878 |
-
|
| 879 |
-
|
| 880 |
-
|
| 881 |
-
|
| 882 |
-
|
| 883 |
-
|
| 884 |
-
|
| 885 |
-
|
| 886 |
-
|
| 887 |
-
|
| 888 |
-
|
| 889 |
-
|
| 890 |
-
|
| 891 |
-
|
| 892 |
-
|
| 893 |
-
|
| 894 |
-
|
| 895 |
-
|
| 896 |
-
|
| 897 |
-
|
| 898 |
-
|
| 899 |
-
|
| 900 |
-
|
| 901 |
-
|
| 902 |
-
|
| 903 |
-
|
| 904 |
-
|
| 905 |
-
|
| 906 |
-
|
| 907 |
-
|
| 908 |
-
|
| 909 |
-
|
| 910 |
-
|
| 911 |
-
|
| 912 |
-
|
| 913 |
-
|
| 914 |
-
|
| 915 |
-
|
| 916 |
-
|
| 917 |
-
|
| 918 |
-
|
| 919 |
-
|
| 920 |
-
|
| 921 |
-
|
| 922 |
-
|
| 923 |
-
|
| 924 |
-
|
| 925 |
-
|
| 926 |
-
|
| 927 |
-
|
| 928 |
-
|
| 929 |
-
|
| 930 |
-
|
| 931 |
-
|
| 932 |
-
|
| 933 |
-
|
| 934 |
-
|
| 935 |
-
|
| 936 |
-
|
| 937 |
-
|
| 938 |
-
|
| 939 |
-
|
| 940 |
-
|
| 941 |
-
|
| 942 |
-
|
| 943 |
-
|
| 944 |
-
|
| 945 |
-
|
| 946 |
-
|
| 947 |
-
|
| 948 |
-
|
| 949 |
-
|
| 950 |
-
|
| 951 |
-
|
| 952 |
-
|
| 953 |
-
|
| 954 |
-
|
| 955 |
-
|
| 956 |
-
|
| 957 |
-
|
| 958 |
-
|
| 959 |
-
|
| 960 |
-
|
| 961 |
-
|
| 962 |
-
|
| 963 |
-
|
| 964 |
-
|
| 965 |
-
|
| 966 |
-
|
| 967 |
-
|
| 968 |
-
|
| 969 |
-
|
| 970 |
-
|
| 971 |
-
|
| 972 |
-
|
| 973 |
-
|
| 974 |
-
|
| 975 |
-
|
| 976 |
-
|
| 977 |
-
|
| 978 |
-
|
| 979 |
-
|
| 980 |
-
|
| 981 |
-
|
| 982 |
-
|
| 983 |
-
|
| 984 |
-
|
| 985 |
-
|
| 986 |
-
|
| 987 |
-
|
| 988 |
-
|
| 989 |
-
|
| 990 |
-
|
| 991 |
-
|
| 992 |
INFO:hf-to-gguf:Model successfully exported to upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16.gguf
|
| 993 |
+ python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-MiMo_V2.6_Flash_RL-PRIMARY --outtype bf16 --outfile ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf --mtp --model-name MiMo-V2.6-Flash-RL
|
| 994 |
INFO:hf-to-gguf:Loading model: model-temp-MiMo_V2.6_Flash_RL-PRIMARY
|
|
@@ -1239,12 +1227,12 @@ INFO:gguf.gguf_writer:Writing the following files:
|
|
| 1239 |
INFO:gguf.gguf_writer:upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf: n_tensors = 39, total_size = 4.5G
|
| 1240 |
|
| 1241 |
|
| 1242 |
-
|
| 1243 |
-
|
| 1244 |
-
|
| 1245 |
-
|
| 1246 |
-
|
| 1247 |
-
|
| 1248 |
INFO:hf-to-gguf:Model successfully exported to upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf
|
| 1249 |
+ python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-MiMo_V2.6_Flash_RL-PRIMARY --outtype bf16 --outfile ./upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-BF16.gguf --mmproj --model-name MiMo-V2.6-Flash-RL
|
| 1250 |
INFO:hf-to-gguf:Loading model: model-temp-MiMo_V2.6_Flash_RL-PRIMARY
|
|
@@ -2143,12 +2131,12 @@ INFO:gguf.gguf_writer:Writing the following files:
|
|
| 2143 |
INFO:gguf.gguf_writer:upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-BF16.gguf: n_tensors = 811, total_size = 2.7G
|
| 2144 |
|
| 2145 |
|
| 2146 |
-
|
| 2147 |
-
|
| 2148 |
INFO:hf-to-gguf:Model successfully exported to upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-BF16.gguf
|
| 2149 |
+ FLAGS_MXFP4=
|
| 2150 |
+ llama.cpp/build/bin/llama-quantize --keep-split ./upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16-00001-of-00002.gguf ./upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-MXFP4.gguf MXFP4_MOE
|
| 2151 |
-
version: 0.
|
| 2152 |
built with GNU 14.2.0 for Linux x86_64
|
| 2153 |
llama_quantize: quantizing './upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16-00001-of-00002.gguf' to './upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-MXFP4' as MXFP4_MOE
|
| 2154 |
llama_model_loader: additional 1 GGUFs metadata loaded.
|
|
@@ -2676,11 +2664,11 @@ llama_model_loader: - type mxfp4: 141 tensors
|
|
| 2676 |
llama_model_quantize_impl: model size = 164915.57 MiB (4.48 BPW)
|
| 2677 |
llama_model_quantize_impl: quant size = 159610.26 MiB (4.34 BPW)
|
| 2678 |
|
| 2679 |
-
llama_quantize: quantize time =
|
| 2680 |
-
llama_quantize: total time =
|
| 2681 |
+ FLAGS_Q2_K='--pure --tensor-type token_embd.weight=q8_0 --tensor-type output.weight=q6_k --tensor-type attn_=q8_0 --tensor-type ffn_down_exps=mxfp4 --tensor-type ffn_gate_exps=q2_k --tensor-type ffn_up_exps=q2_k '
|
| 2682 |
+ llama.cpp/build/bin/llama-quantize --keep-split --allow-requantize --pure --tensor-type token_embd.weight=q8_0 --tensor-type output.weight=q6_k --tensor-type attn_=q8_0 --tensor-type ffn_down_exps=mxfp4 --tensor-type ffn_gate_exps=q2_k --tensor-type ffn_up_exps=q2_k ./upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16-00001-of-00002.gguf ./upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-Q2_K.gguf MXFP4_MOE
|
| 2683 |
-
version: 0.
|
| 2684 |
built with GNU 14.2.0 for Linux x86_64
|
| 2685 |
llama_quantize: quantizing './upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16-00001-of-00002.gguf' to './upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-Q2_K' as MXFP4_MOE
|
| 2686 |
llama_model_loader: additional 1 GGUFs metadata loaded.
|
|
@@ -3400,10 +3388,10 @@ llama_tensor_get_type: blk.9.attn_qkv.weight - applying manual ov
|
|
| 3400 |
llama_model_quantize_impl: model size = 164915.57 MiB (4.48 BPW)
|
| 3401 |
llama_model_quantize_impl: quant size = 119887.91 MiB (3.26 BPW)
|
| 3402 |
|
| 3403 |
-
llama_quantize: quantize time =
|
| 3404 |
-
llama_quantize: total time =
|
| 3405 |
+ llama.cpp/build/bin/llama-quantize ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-MXFP4.gguf MXFP4_MOE
|
| 3406 |
-
version: 0.
|
| 3407 |
built with GNU 14.2.0 for Linux x86_64
|
| 3408 |
llama_quantize: quantizing './upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf' to './upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-MXFP4.gguf' as MXFP4_MOE
|
| 3409 |
llama_model_loader: loaded meta data with 42 key-value pairs and 39 tensors from ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf (version GGUF V3 (latest))
|
|
@@ -3494,10 +3482,10 @@ llama_model_loader: - type bf16: 20 tensors
|
|
| 3494 |
llama_model_quantize_impl: model size = 4268.25 MiB (16.00 BPW)
|
| 3495 |
llama_model_quantize_impl: quant size = 2267.63 MiB (8.50 BPW)
|
| 3496 |
|
| 3497 |
-
llama_quantize: quantize time =
|
| 3498 |
-
llama_quantize: total time =
|
| 3499 |
+ llama.cpp/build/bin/llama-quantize ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-Q8_0.gguf Q8_0
|
| 3500 |
-
version: 0.
|
| 3501 |
built with GNU 14.2.0 for Linux x86_64
|
| 3502 |
llama_quantize: quantizing './upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf' to './upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-Q8_0.gguf' as Q8_0
|
| 3503 |
llama_model_loader: loaded meta data with 42 key-value pairs and 39 tensors from ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf (version GGUF V3 (latest))
|
|
@@ -3588,8 +3576,8 @@ llama_model_loader: - type bf16: 20 tensors
|
|
| 3588 |
llama_model_quantize_impl: model size = 4268.25 MiB (16.00 BPW)
|
| 3589 |
llama_model_quantize_impl: quant size = 2267.63 MiB (8.50 BPW)
|
| 3590 |
|
| 3591 |
-
llama_quantize: quantize time =
|
| 3592 |
-
llama_quantize: total time =
|
| 3593 |
+ python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-MiMo_V2.6_Flash_RL-PRIMARY --outtype q8_0 --outfile ./upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-Q8_0.gguf --mmproj --model-name MiMo-V2.6-Flash-RL
|
| 3594 |
INFO:hf-to-gguf:Loading model: model-temp-MiMo_V2.6_Flash_RL-PRIMARY
|
| 3595 |
WARNING:hf-to-gguf:Failed to load model config from model-temp-MiMo_V2.6_Flash_RL-PRIMARY: The repository model-temp-MiMo_V2.6_Flash_RL-PRIMARY contains custom code which must be executed to correctly load the model. You can inspect the repository content at /tmp/convert/model-temp-MiMo_V2.6_Flash_RL-PRIMARY .
|
|
@@ -4487,9 +4475,9 @@ INFO:gguf.gguf_writer:Writing the following files:
|
|
| 4487 |
INFO:gguf.gguf_writer:upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-Q8_0.gguf: n_tensors = 811, total_size = 1.6G
|
| 4488 |
|
| 4489 |
|
| 4490 |
-
|
| 4491 |
-
|
| 4492 |
-
|
| 4493 |
INFO:hf-to-gguf:Model successfully exported to upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-Q8_0.gguf
|
| 4494 |
+ echo MiMo-V2.6-Flash-RL-MXFP4-00001-of-00002.gguf
|
| 4495 |
+ echo MiMo-V2.6-Flash-RL-MXFP4-00002-of-00002.gguf
|
|
|
|
| 690 |
|
| 691 |
|
| 692 |
|
| 693 |
+
|
| 694 |
+
|
| 695 |
+
|
| 696 |
+
|
| 697 |
+
|
| 698 |
+
|
| 699 |
+
|
| 700 |
+
|
| 701 |
+
|
| 702 |
+
|
| 703 |
+
|
| 704 |
+
|
| 705 |
+
|
| 706 |
+
|
| 707 |
+
|
| 708 |
+
|
| 709 |
+
|
| 710 |
+
|
| 711 |
+
|
| 712 |
+
|
| 713 |
+
|
| 714 |
+
|
| 715 |
+
|
| 716 |
+
|
| 717 |
+
|
| 718 |
+
|
| 719 |
+
|
| 720 |
+
|
| 721 |
+
|
| 722 |
+
|
| 723 |
+
|
| 724 |
+
|
| 725 |
+
|
| 726 |
+
|
| 727 |
+
|
| 728 |
+
|
| 729 |
+
|
| 730 |
+
|
| 731 |
+
|
| 732 |
+
|
| 733 |
+
|
| 734 |
+
|
| 735 |
+
|
| 736 |
+
|
| 737 |
+
|
| 738 |
+
|
| 739 |
+
|
| 740 |
+
|
| 741 |
+
|
| 742 |
+
|
| 743 |
+
|
| 744 |
+
|
| 745 |
+
|
| 746 |
+
|
| 747 |
+
|
| 748 |
+
|
| 749 |
+
|
| 750 |
+
|
| 751 |
+
|
| 752 |
+
|
| 753 |
+
|
| 754 |
+
|
| 755 |
+
|
| 756 |
+
|
| 757 |
+
|
| 758 |
+
|
| 759 |
+
|
| 760 |
+
|
| 761 |
+
|
| 762 |
+
|
| 763 |
+
|
| 764 |
+
|
| 765 |
+
|
| 766 |
+
|
| 767 |
+
|
| 768 |
+
|
| 769 |
+
|
| 770 |
+
|
| 771 |
+
|
| 772 |
+
|
| 773 |
+
|
| 774 |
+
|
| 775 |
+
|
| 776 |
+
|
| 777 |
+
|
| 778 |
+
|
| 779 |
+
|
| 780 |
+
|
| 781 |
+
|
| 782 |
+
|
| 783 |
+
|
| 784 |
+
|
| 785 |
+
|
| 786 |
+
|
| 787 |
+
|
| 788 |
+
|
| 789 |
+
|
| 790 |
+
|
| 791 |
+
|
| 792 |
+
|
| 793 |
+
|
| 794 |
+
|
| 795 |
+
|
| 796 |
+
|
| 797 |
+
|
| 798 |
+
|
| 799 |
+
|
| 800 |
+
|
| 801 |
+
|
| 802 |
+
|
| 803 |
+
|
| 804 |
+
|
| 805 |
+
|
| 806 |
+
|
| 807 |
+
|
| 808 |
+
|
| 809 |
+
|
| 810 |
+
|
| 811 |
+
|
| 812 |
+
|
| 813 |
+
|
| 814 |
+
|
| 815 |
+
|
| 816 |
+
|
| 817 |
+
|
| 818 |
+
|
| 819 |
+
|
| 820 |
+
|
| 821 |
+
|
| 822 |
+
|
| 823 |
+
|
| 824 |
+
|
| 825 |
+
|
| 826 |
+
|
| 827 |
+
|
| 828 |
+
|
| 829 |
+
|
| 830 |
+
|
| 831 |
+
|
| 832 |
+
|
| 833 |
+
|
| 834 |
+
|
| 835 |
+
|
| 836 |
+
|
| 837 |
+
|
| 838 |
+
|
| 839 |
+
|
| 840 |
+
|
| 841 |
+
|
| 842 |
+
|
| 843 |
+
|
| 844 |
+
|
| 845 |
+
|
| 846 |
+
|
| 847 |
+
|
| 848 |
+
|
| 849 |
+
|
| 850 |
+
|
| 851 |
+
|
| 852 |
+
|
| 853 |
+
|
| 854 |
+
|
| 855 |
+
|
| 856 |
+
|
| 857 |
+
|
| 858 |
+
|
| 859 |
+
|
| 860 |
+
|
| 861 |
+
|
| 862 |
+
|
| 863 |
+
|
| 864 |
+
|
| 865 |
+
|
| 866 |
+
|
| 867 |
+
|
| 868 |
+
|
| 869 |
+
|
| 870 |
+
|
| 871 |
+
|
| 872 |
+
|
| 873 |
+
|
| 874 |
+
|
| 875 |
+
|
| 876 |
+
|
| 877 |
+
|
| 878 |
+
|
| 879 |
+
|
| 880 |
+
|
| 881 |
+
|
| 882 |
+
|
| 883 |
+
|
| 884 |
+
|
| 885 |
+
|
| 886 |
+
|
| 887 |
+
|
| 888 |
+
|
| 889 |
+
|
| 890 |
+
|
| 891 |
+
|
| 892 |
+
|
| 893 |
+
|
| 894 |
+
|
| 895 |
+
|
| 896 |
+
|
| 897 |
+
|
| 898 |
+
|
| 899 |
+
|
| 900 |
+
|
| 901 |
+
|
| 902 |
+
|
| 903 |
+
|
| 904 |
+
|
| 905 |
+
|
| 906 |
+
|
| 907 |
+
|
| 908 |
+
|
| 909 |
+
|
| 910 |
+
|
| 911 |
+
|
| 912 |
+
|
| 913 |
+
|
| 914 |
+
|
| 915 |
+
|
| 916 |
+
|
| 917 |
+
|
| 918 |
+
|
| 919 |
+
|
| 920 |
+
|
| 921 |
+
|
| 922 |
+
|
| 923 |
+
|
| 924 |
+
|
| 925 |
+
|
| 926 |
+
|
| 927 |
+
|
| 928 |
+
|
| 929 |
+
|
| 930 |
+
|
| 931 |
+
|
| 932 |
+
|
| 933 |
+
|
| 934 |
+
|
| 935 |
+
|
| 936 |
+
|
| 937 |
+
|
| 938 |
+
|
| 939 |
+
|
| 940 |
+
|
| 941 |
+
|
| 942 |
+
|
| 943 |
+
|
| 944 |
+
|
| 945 |
+
|
| 946 |
+
|
| 947 |
+
|
| 948 |
+
|
| 949 |
+
|
| 950 |
+
|
| 951 |
+
|
| 952 |
+
|
| 953 |
+
|
| 954 |
+
|
| 955 |
+
|
| 956 |
+
|
| 957 |
+
|
| 958 |
+
|
| 959 |
+
|
| 960 |
+
|
| 961 |
+
|
| 962 |
+
|
| 963 |
+
|
| 964 |
+
|
| 965 |
+
|
| 966 |
+
|
| 967 |
+
|
| 968 |
+
|
| 969 |
+
|
| 970 |
+
|
| 971 |
+
|
| 972 |
+
|
| 973 |
+
|
| 974 |
+
|
| 975 |
+
|
| 976 |
+
|
| 977 |
+
|
| 978 |
+
|
| 979 |
+
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 980 |
INFO:hf-to-gguf:Model successfully exported to upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16.gguf
|
| 981 |
+ python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-MiMo_V2.6_Flash_RL-PRIMARY --outtype bf16 --outfile ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf --mtp --model-name MiMo-V2.6-Flash-RL
|
| 982 |
INFO:hf-to-gguf:Loading model: model-temp-MiMo_V2.6_Flash_RL-PRIMARY
|
|
|
|
| 1227 |
INFO:gguf.gguf_writer:upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf: n_tensors = 39, total_size = 4.5G
|
| 1228 |
|
| 1229 |
|
| 1230 |
+
|
| 1231 |
+
|
| 1232 |
+
|
| 1233 |
+
|
| 1234 |
+
|
| 1235 |
+
|
| 1236 |
INFO:hf-to-gguf:Model successfully exported to upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf
|
| 1237 |
+ python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-MiMo_V2.6_Flash_RL-PRIMARY --outtype bf16 --outfile ./upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-BF16.gguf --mmproj --model-name MiMo-V2.6-Flash-RL
|
| 1238 |
INFO:hf-to-gguf:Loading model: model-temp-MiMo_V2.6_Flash_RL-PRIMARY
|
|
|
|
| 2131 |
INFO:gguf.gguf_writer:upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-BF16.gguf: n_tensors = 811, total_size = 2.7G
|
| 2132 |
|
| 2133 |
|
| 2134 |
+
|
| 2135 |
+
|
| 2136 |
INFO:hf-to-gguf:Model successfully exported to upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-BF16.gguf
|
| 2137 |
+ FLAGS_MXFP4=
|
| 2138 |
+ llama.cpp/build/bin/llama-quantize --keep-split ./upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16-00001-of-00002.gguf ./upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-MXFP4.gguf MXFP4_MOE
|
| 2139 |
+
version: 0.5.0-dev (build 11161, commit a0484817c)
|
| 2140 |
built with GNU 14.2.0 for Linux x86_64
|
| 2141 |
llama_quantize: quantizing './upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16-00001-of-00002.gguf' to './upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-MXFP4' as MXFP4_MOE
|
| 2142 |
llama_model_loader: additional 1 GGUFs metadata loaded.
|
|
|
|
| 2664 |
llama_model_quantize_impl: model size = 164915.57 MiB (4.48 BPW)
|
| 2665 |
llama_model_quantize_impl: quant size = 159610.26 MiB (4.34 BPW)
|
| 2666 |
|
| 2667 |
+
llama_quantize: quantize time = 567689.54 ms
|
| 2668 |
+
llama_quantize: total time = 567689.54 ms
|
| 2669 |
+ FLAGS_Q2_K='--pure --tensor-type token_embd.weight=q8_0 --tensor-type output.weight=q6_k --tensor-type attn_=q8_0 --tensor-type ffn_down_exps=mxfp4 --tensor-type ffn_gate_exps=q2_k --tensor-type ffn_up_exps=q2_k '
|
| 2670 |
+ llama.cpp/build/bin/llama-quantize --keep-split --allow-requantize --pure --tensor-type token_embd.weight=q8_0 --tensor-type output.weight=q6_k --tensor-type attn_=q8_0 --tensor-type ffn_down_exps=mxfp4 --tensor-type ffn_gate_exps=q2_k --tensor-type ffn_up_exps=q2_k ./upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16-00001-of-00002.gguf ./upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-Q2_K.gguf MXFP4_MOE
|
| 2671 |
+
version: 0.5.0-dev (build 11161, commit a0484817c)
|
| 2672 |
built with GNU 14.2.0 for Linux x86_64
|
| 2673 |
llama_quantize: quantizing './upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-BF16-00001-of-00002.gguf' to './upload-MiMo_V2.6_Flash_RL/MiMo-V2.6-Flash-RL-Q2_K' as MXFP4_MOE
|
| 2674 |
llama_model_loader: additional 1 GGUFs metadata loaded.
|
|
|
|
| 3388 |
llama_model_quantize_impl: model size = 164915.57 MiB (4.48 BPW)
|
| 3389 |
llama_model_quantize_impl: quant size = 119887.91 MiB (3.26 BPW)
|
| 3390 |
|
| 3391 |
+
llama_quantize: quantize time = 1227725.91 ms
|
| 3392 |
+
llama_quantize: total time = 1227725.91 ms
|
| 3393 |
+ llama.cpp/build/bin/llama-quantize ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-MXFP4.gguf MXFP4_MOE
|
| 3394 |
+
version: 0.5.0-dev (build 11161, commit a0484817c)
|
| 3395 |
built with GNU 14.2.0 for Linux x86_64
|
| 3396 |
llama_quantize: quantizing './upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf' to './upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-MXFP4.gguf' as MXFP4_MOE
|
| 3397 |
llama_model_loader: loaded meta data with 42 key-value pairs and 39 tensors from ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf (version GGUF V3 (latest))
|
|
|
|
| 3482 |
llama_model_quantize_impl: model size = 4268.25 MiB (16.00 BPW)
|
| 3483 |
llama_model_quantize_impl: quant size = 2267.63 MiB (8.50 BPW)
|
| 3484 |
|
| 3485 |
+
llama_quantize: quantize time = 16883.06 ms
|
| 3486 |
+
llama_quantize: total time = 16883.06 ms
|
| 3487 |
+ llama.cpp/build/bin/llama-quantize ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-Q8_0.gguf Q8_0
|
| 3488 |
+
version: 0.5.0-dev (build 11161, commit a0484817c)
|
| 3489 |
built with GNU 14.2.0 for Linux x86_64
|
| 3490 |
llama_quantize: quantizing './upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf' to './upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-Q8_0.gguf' as Q8_0
|
| 3491 |
llama_model_loader: loaded meta data with 42 key-value pairs and 39 tensors from ./upload-MiMo_V2.6_Flash_RL/mtp-MiMo-V2.6-Flash-RL-BF16.gguf (version GGUF V3 (latest))
|
|
|
|
| 3576 |
llama_model_quantize_impl: model size = 4268.25 MiB (16.00 BPW)
|
| 3577 |
llama_model_quantize_impl: quant size = 2267.63 MiB (8.50 BPW)
|
| 3578 |
|
| 3579 |
+
llama_quantize: quantize time = 3866.49 ms
|
| 3580 |
+
llama_quantize: total time = 3866.49 ms
|
| 3581 |
+ python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-MiMo_V2.6_Flash_RL-PRIMARY --outtype q8_0 --outfile ./upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-Q8_0.gguf --mmproj --model-name MiMo-V2.6-Flash-RL
|
| 3582 |
INFO:hf-to-gguf:Loading model: model-temp-MiMo_V2.6_Flash_RL-PRIMARY
|
| 3583 |
WARNING:hf-to-gguf:Failed to load model config from model-temp-MiMo_V2.6_Flash_RL-PRIMARY: The repository model-temp-MiMo_V2.6_Flash_RL-PRIMARY contains custom code which must be executed to correctly load the model. You can inspect the repository content at /tmp/convert/model-temp-MiMo_V2.6_Flash_RL-PRIMARY .
|
|
|
|
| 4475 |
INFO:gguf.gguf_writer:upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-Q8_0.gguf: n_tensors = 811, total_size = 1.6G
|
| 4476 |
|
| 4477 |
|
| 4478 |
+
|
| 4479 |
+
|
| 4480 |
+
|
| 4481 |
INFO:hf-to-gguf:Model successfully exported to upload-MiMo_V2.6_Flash_RL/mmproj-MiMo-V2.6-Flash-RL-Q8_0.gguf
|
| 4482 |
+ echo MiMo-V2.6-Flash-RL-MXFP4-00001-of-00002.gguf
|
| 4483 |
+ echo MiMo-V2.6-Flash-RL-MXFP4-00002-of-00002.gguf
|
mtp-MiMo-V2.6-Flash-RL-BF16.gguf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 4481536192
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:a0ffc7356ce0ff8597f9d9b4e394df1fe68ee4726fa5867e7e99a9011c3c2293
|
| 3 |
size 4481536192
|
mtp-MiMo-V2.6-Flash-RL-MXFP4.gguf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 2383728832
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7be6df8a09060430b8ae4f106a127602a5ae8771c6d0e39e739ace0ad6083507
|
| 3 |
size 2383728832
|
mtp-MiMo-V2.6-Flash-RL-Q8_0.gguf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 2383728832
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:5f8f1c3a20a73d1dc68c83d21161121da0882f9e67d4c5c33fe5ef3f040233ca
|
| 3 |
size 2383728832
|