Instructions to use GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic") model = AutoModelForCausalLM.from_pretrained("GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic
- SGLang
How to use GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic with Docker Model Runner:
docker model run hf.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic
Download README_zh.md from GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic: direct link, hf CLI and curl.
- Browser
- Download file 5.04 kB
-
https://huggingface.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic/resolve/934bcdfc20af23e97971704ab21ae3a8bdf7373c/README_zh.md
- Command line
-
hf download hf://GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic@934bcdfc20af23e97971704ab21ae3a8bdf7373c/README_zh.md
-
curl -L -o README_zh.md https://huggingface.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic/resolve/934bcdfc20af23e97971704ab21ae3a8bdf7373c/README_zh.md
MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic
GGUF 量化版本(用于本地部署):**MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF**
MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic 是一个基于 openbmb/MiniCPM5-2B 的紧凑型 2B 思维链(Thinking) 语言模型。使用 Claude 数据进行微调,重点增强了 Agent 工具调用 / 函数调用、代码生成 和 指令遵循 能力,同时保留了 MiniCPM5 原生的 Thinking 聊天模板和 XML 工具调用格式。
概览
| 项目 | 详情 |
|---|---|
| 基座模型 | openbmb/MiniCPM5-2B(2B 稠密 Llama 架构) |
| 后训练数据 | Claude 数据 |
| 核心能力 | Agent 工具调用、代码生成、指令遵循、思维链推理 |
| 聊天格式 | MiniCPM5 原生 Thinking 模板,支持可选的思维链输出 |
| 上下文长度 | 128K(max_position_embeddings = 131072) |
| 精度 | bfloat16 |
| 部署 | 单卡友好,适合边缘设备 / 本地部署 |
评测结果
ClawBench(Agent 代码能力)
| 模型 | QwenClawBench | WildClawBench |
|---|---|---|
| MiniCPM5-2B(基座,仅 RL) | 42.11 | 23.19 |
| MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic | 44.56(+2.45) | 24.32(+1.13) |
ClawBench 评测 Agent 代码能力——模型自主使用工具、浏览代码库、完成多步软件工程任务的综合能力。QwenClawBench 使用结构化代码场景;WildClawBench 测试多样化的真实任务。
更多评测成绩(BFCL、SWE-bench、Tau-Bench 等)补充中。
能力
- Agent 工具调用 — 基于 MiniCPM5 原生格式的可靠 XML / 函数调用风格工具使用,适配多步 Agent 工作流
- 代码生成 — 代码编写、调试及软件工程任务
- 指令遵循 — 可靠地遵循用户指令和结构化约束
- 思维链模式 — 通过 MiniCPM5 聊天模板实现链式推理
- 长上下文 — 最高支持 128K tokens(131,072 tokens)
快速开始
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [{"role": "user", "content": "写一个合并两个有序列表的 Python 函数。"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
工具调用示例
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "获取指定城市的当前天气。",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "城市名称"}
},
"required": ["city"]
}
}
}
]
messages = [
{"role": "user", "content": "北京今天天气怎么样?"}
]
text = tokenizer.apply_chat_template(
messages, tools=tools, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
采样建议
继承自 openbmb/MiniCPM5-2B:
| 场景 | 参数 |
|---|---|
| 默认 | temperature=1.0, top_p=0.95, min_p=0.0 |
| 如遇重复输出 | temperature=1.0, top_p=0.95, min_p=0.0, repetition_penalty=1.05 |
本模型为 纯思考模式(Thinking-only)——思维链推理始终开启。
不同推理框架对采样参数的支持有所差异,请参考所用框架的文档。
局限性
- 思维链输出 — 模型可能在最终回答前输出推理过程,下游应用可在展示前过滤
- 2B 规模 — 针对轻量级本地部署优化,非前沿级通用推理
来源与许可
基于 Apache-2.0 发布,继承自 MiniCPM5-2B。
致谢
- 基座模型:OpenBMB / MiniCPM5-2B
- GGUF 转换:llama.cpp