---
library_name: transformers
license: apache-2.0
language:
- en
- zh
base_model: openbmb/MiniCPM5-2B
base_model_relation: finetune
pipeline_tag: text-generation
tags:
- minicpm
- minicpm5
- llama
- text-generation
- thinking
- fable5
- tool-calling
- function-calling
- agentic
- coding
- instruction-following
- conversational
---
# MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic
GGUF quantizations for local deployment: **[MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF](https://huggingface.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF)**
**MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic** is a compact 2B **Thinking** language model built on [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B). Fine-tuned on **Claude** data with a strong focus on **agentic tool calling / function calling**, **coding**, and **instruction following**. It keeps MiniCPM5's native Thinking chat template and XML tool-call format.
For llama.cpp / Ollama / LM Studio deployment, see the **[GGUF repository](https://huggingface.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF)**.
---
## Overview
| Item | Detail |
|---|---|
| **Base model** | [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B) (2B dense Llama architecture) |
| **Post-training** | Claude data |
| **Key capabilities** | **Agentic tool calling**, coding, instruction following, chain-of-thought reasoning |
| **Chat format** | MiniCPM5 native Thinking template with optional chain-of-thought blocks |
| **Context length** | **128K** (`max_position_embeddings = 131072`) |
| **Precision** | bfloat16 |
| **Deployment** | Single-GPU friendly; suitable for edge / local use |
---
## Capabilities
- **Agentic tool calling** — reliable XML / function-calling style tool use on top of MiniCPM5's native format, designed for multi-step agentic workflows
- **Coding** — code generation, debugging, and software-engineering-style tasks
- **Instruction following** — reliable adherence to user prompts and structured constraints
- **Thinking mode** — chain-of-thought reasoning via the MiniCPM5 chat template
- **Long context** — up to **128K tokens** (131,072 tokens per `config.json`)
---
## Benchmark
### ClawBench (Agentic Coding)
| Model | QwenClawBench | WildClawBench |
|---|---|---|
| MiniCPM5-2B (Base, RL-only) | 42.11 | 23.19 |
| **MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic** | **44.56** (+2.45) | **24.32** (+1.13) |
> ClawBench evaluates agentic coding ability — the model's capacity to autonomously use tools, navigate codebases, and complete multi-step software engineering tasks. QwenClawBench uses structured coding scenarios; WildClawBench tests on diverse real-world tasks.
> **More benchmarks (BFCL, SWE-bench, Tau-Bench, etc.) coming soon.**
---
## Quick start
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [{"role": "user", "content": "Write a Python function to merge two sorted lists."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```
### Tool calling example
```python
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a given city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"}
},
"required": ["city"]
}
}
}
]
messages = [
{"role": "user", "content": "What's the weather like in Beijing?"}
]
text = tokenizer.apply_chat_template(
messages, tools=tools, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```
---
## Sampling recommendations
Inherited from [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B):
| Scenario | Params |
|---|---|
| **Default** | `temperature=1.0, top_p=0.95, min_p=0.0` |
| **If repetitive outputs** | `temperature=1.0, top_p=0.95, min_p=0.0, repetition_penalty=1.05` |
This model is **Thinking-only** — chain-of-thought reasoning is always active.
> Support for sampling parameters varies across inference frameworks — check your runtime's documentation.
---
## Limitations
- **Thinking outputs** — the model may emit reasoning blocks before the final answer; downstream apps can strip them before display
- **2B scale** — optimized for lightweight local deployment, not frontier-scale general reasoning
---
## Provenance & licensing
Released under **Apache-2.0**, inherited from [MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B).
## Acknowledgements
- Base model: [OpenBMB / MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B)
- GGUF conversion: [llama.cpp](https://github.com/ggml-org/llama.cpp)