How to use from
Pi
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "sahilchachra/MiniCPM5-2B-MXFP4"
Configure the model in Pi
# Install Pi:
npm install -g @earendil-works/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "mlx-lm": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "sahilchachra/MiniCPM5-2B-MXFP4"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

MiniCPM5-2B-MXFP4

This is openbmb/MiniCPM5-2B quantized to 4-bit MXFP4 (group size 32) for use with MLX on Apple Silicon.

  • Size on disk: ~1.3 GB
  • Architecture: LlamaForCausalLM (standard, no custom code required)
  • Quantized with: mlx_lm.convert -q --q-mode mxfp4

Other quantizations of this model

Use with mlx-lm

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("sahilchachra/MiniCPM5-2B-MXFP4")
prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Hello, who are you?"}],
    add_generation_prompt=True, tokenize=False,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=200))

Verification

Smoke-tested via mlx_lm.load + mlx_lm.generate — produces coherent, on-topic completions matching the base model's expected chat-reasoning style.

Run in LM Studio

Verified working in LM Studio (loads via the MLX engine, tested through the local OpenAI-compatible API at localhost:1234/v1/chat/completions):

  1. Download this repo, or symlink/copy the folder into ~/.lmstudio/models/<publisher>/<name>/.
  2. LM Studio's model indexer will pick it up automatically (or run "Rescan").
  3. Load it in the UI or via lms load <model-name>.

Confirmed a clean, correct response to a basic prompt with no truncation or garbled output.

Downloads last month
245
Safetensors
Model size
3B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sahilchachra/MiniCPM5-2B-MXFP4

Quantized
(102)
this model

Collection including sahilchachra/MiniCPM5-2B-MXFP4