Instructions to use RockMan256/MiniCPM5-1B-enti-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use RockMan256/MiniCPM5-1B-enti-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
Use Docker
docker model run hf.co/RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
- LM Studio
- Jan
- Ollama
How to use RockMan256/MiniCPM5-1B-enti-GGUF with Ollama:
ollama run hf.co/RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use RockMan256/MiniCPM5-1B-enti-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use RockMan256/MiniCPM5-1B-enti-GGUF with Docker Model Runner:
docker model run hf.co/RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
- Lemonade
How to use RockMan256/MiniCPM5-1B-enti-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
Run and chat with the model
lemonade run user.MiniCPM5-1B-enti-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use RockMan256/MiniCPM5-1B-enti-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use RockMan256/MiniCPM5-1B-enti-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "RockMan256/MiniCPM5-1B-enti-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| base_model: openbmb/MiniCPM5-1B | |
| library_name: llama.cpp | |
| tags: | |
| - home-assistant | |
| - entity-resolver | |
| - intent-classification | |
| - gguf | |
| language: | |
| - ru | |
| - en | |
| # MiniCPM5-1B-enti-GGUF | |
| Маленькая модель-резолвер сущностей (entity / intent resolver) для Home Assistant. | |
| Превращает пользовательскую фразу на русском или английском в **домен + имя | |
| сущности** (или интент), которые затем маппятся в вызов HA-тула на стороне | |
| клиента. | |
| Обучена на `RockMan256/ha-ru-en-intents`. | |
| ## Как использовать (формат запроса и ответа) | |
| **Вход (системный промпт + запрос пользователя):** | |
| ``` | |
| <системный промпт> | |
| <user> включи свет в спальне | |
| ``` | |
| **Выход (модель отвечает JSON-объектом с интентом/доменом):** | |
| ```json | |
| { "name": "entityResolver", "arguments": { "domain": "light", "name": "спальня" } } | |
| ``` | |
| или для прямых команд: | |
| ```json | |
| { "name": "turn_on", "arguments": { "area": "Спальня" } } | |
| ``` | |
| ## Важно: отключение reasoning | |
| Модель упрямо включает режим размышления (think) через свой chat-template, | |
| даже если сервер запущен с `--reasoning off`. Чтобы выключить размышления, | |
| **добавь в начало системного промпта две пустые строки**: | |
| ``` | |
| \n\n<дальше твой системный промпт> | |
| ``` | |
| Без этого модель выдаёт лишний текст вместо чистого JSON (146 токенов вместо 65). | |
| ## Рекомендации | |
| - Не используй как чат-бота — это **классификатор**, а не генератор ответов. | |
| - Модель выдаёт **domain/intent**, а не готовый entity_id — нужен пост-процессинг | |
| (резолв entity через HA API по domain + fuzzy-поиск по name). | |
| - Примеры интентов: `light`, `gate`, `gate_open`, `turn_on`, `set_speed`. | |
| - Размер 1B — быстрая, но иногда нестабильна (путает вытяжку с воротами). | |