Instructions to use koshuro/MiniCPM5-1B-heretic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use koshuro/MiniCPM5-1B-heretic with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M # Run inference directly in the terminal: llama cli -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M # Run inference directly in the terminal: llama cli -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M
Use Docker
docker model run hf.co/koshuro/MiniCPM5-1B-heretic:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use koshuro/MiniCPM5-1B-heretic with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "koshuro/MiniCPM5-1B-heretic" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "koshuro/MiniCPM5-1B-heretic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/koshuro/MiniCPM5-1B-heretic:Q4_K_M
- Ollama
How to use koshuro/MiniCPM5-1B-heretic with Ollama:
ollama run hf.co/koshuro/MiniCPM5-1B-heretic:Q4_K_M
- Unsloth Desktop
- Pi
How to use koshuro/MiniCPM5-1B-heretic with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "koshuro/MiniCPM5-1B-heretic:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use koshuro/MiniCPM5-1B-heretic with Docker Model Runner:
docker model run hf.co/koshuro/MiniCPM5-1B-heretic:Q4_K_M
- Lemonade
How to use koshuro/MiniCPM5-1B-heretic with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull koshuro/MiniCPM5-1B-heretic:Q4_K_M
Run and chat with the model
lemonade run user.MiniCPM5-1B-heretic-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use koshuro/MiniCPM5-1B-heretic with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default koshuro/MiniCPM5-1B-heretic:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use koshuro/MiniCPM5-1B-heretic with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf koshuro/MiniCPM5-1B-heretic:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "koshuro/MiniCPM5-1B-heretic:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Reproduction guide
This directory contains the necessary information and assets to reproduce the results obtained during this Heretic run.
Models
- Base model: openbmb/MiniCPM5-1B (Commit:
4e9de7a)
Datasets
- Good prompts: mlabonne/harmless_alpaca (Commit:
02c6a92) - Bad prompts: mlabonne/harmful_behaviors (Commit:
01cead0) - Good evaluation prompts: mlabonne/harmless_alpaca (Commit:
02c6a92) - Bad evaluation prompts: mlabonne/harmful_behaviors (Commit:
01cead0)
Selected trial
- Trial number: 162
- KL divergence: 0.038307
- Refusals: 3/100
Environment
- Heretic: v1.4.0 (Origin: PyPI)
- PyTorch: 2.11.0+cu128
- Other dependencies: See
requirements.txt.
Contents of this directory
requirements.txt: The exact versions of all Python packages.config.toml: The exact configuration used, including the RNG seed.openbmb--MiniCPM5-1B.jsonl: The Optuna study journal containing the history of all trials.SHA256SUMS: Cryptographic hashes for all weight files.reproduce.json: A machine-readable file containing all reproducibility information.
How to reproduce
You can automate this process, including all verification steps, by downloading the
reproduce.jsonfile and runningheretic --reproduce reproduce.json.
- Install the exact version of Heretic indicated in the Environment section above, from its original source.
- Install the packages listed in
requirements.txt:pip install -r requirements.txt - Install the correct version of PyTorch:
pip install torch==2.11.0+cu128 --index-url https://download.pytorch.org/whl/cu128 - Place the provided
config.tomlin your working directory. - Run Heretic without any additional arguments:
heretic - Wait for the run to finish, then select trial 162 and export the model.
- Verify that the weight files have been exactly reproduced by comparing their SHA-256 hashes against those in
SHA256SUMS:sha256sum -c SHA256SUMS(or look at the hashes online if you uploaded to Hugging Face)
To use the included Optuna study journal
openbmb--MiniCPM5-1B.jsonl, place it in the checkpoints directory (usuallycheckpoints/) before running Heretic.This allows you to export other models from the Pareto front, or to run additional trials without having to re-run the stored trials.