File size: 5,719 Bytes
934bcdf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
---
library_name: transformers
license: apache-2.0
language:
  - en
  - zh
base_model: openbmb/MiniCPM5-2B
base_model_relation: finetune
pipeline_tag: text-generation
tags:
  - minicpm
  - minicpm5
  - llama
  - text-generation
  - thinking
  - fable5
  - tool-calling
  - function-calling
  - agentic
  - coding
  - instruction-following
  - conversational
---

<p align="center">
  <img src="assets/banner.png" alt="MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic" width="100%"/>
</p>

# MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic

GGUF quantizations for local deployment: **[MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF](https://huggingface.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF)**

**MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic** is a compact 2B **Thinking** language model built on [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B). Fine-tuned on **Claude** data with a strong focus on **agentic tool calling / function calling**, **coding**, and **instruction following**. It keeps MiniCPM5's native Thinking chat template and XML tool-call format.

For llama.cpp / Ollama / LM Studio deployment, see the **[GGUF repository](https://huggingface.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF)**.

---

## Overview

| Item | Detail |
|---|---|
| **Base model** | [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B) (2B dense Llama architecture) |
| **Post-training** | Claude data |
| **Key capabilities** | **Agentic tool calling**, coding, instruction following, chain-of-thought reasoning |
| **Chat format** | MiniCPM5 native Thinking template with optional chain-of-thought blocks |
| **Context length** | **128K** (`max_position_embeddings = 131072`) |
| **Precision** | bfloat16 |
| **Deployment** | Single-GPU friendly; suitable for edge / local use |

---

## Capabilities

- **Agentic tool calling** β€” reliable XML / function-calling style tool use on top of MiniCPM5's native format, designed for multi-step agentic workflows
- **Coding** β€” code generation, debugging, and software-engineering-style tasks
- **Instruction following** β€” reliable adherence to user prompts and structured constraints
- **Thinking mode** β€” chain-of-thought reasoning via the MiniCPM5 chat template
- **Long context** β€” up to **128K tokens** (131,072 tokens per `config.json`)

---

## Benchmark

### ClawBench (Agentic Coding)

| Model | QwenClawBench | WildClawBench |
|---|---|---|
| MiniCPM5-2B (Base, RL-only) | 42.11 | 23.19 |
| **MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic** | **44.56** (+2.45) | **24.32** (+1.13) |

> ClawBench evaluates agentic coding ability β€” the model's capacity to autonomously use tools, navigate codebases, and complete multi-step software engineering tasks. QwenClawBench uses structured coding scenarios; WildClawBench tests on diverse real-world tasks.

> **More benchmarks (BFCL, SWE-bench, Tau-Bench, etc.) coming soon.**

---

## Quick start

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Write a Python function to merge two sorted lists."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```

### Tool calling example

```python
tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a given city.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "City name"}
                },
                "required": ["city"]
            }
        }
    }
]

messages = [
    {"role": "user", "content": "What's the weather like in Beijing?"}
]

text = tokenizer.apply_chat_template(
    messages, tools=tools, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```

---

## Sampling recommendations

Inherited from [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B):

| Scenario | Params |
|---|---|
| **Default** | `temperature=1.0, top_p=0.95, min_p=0.0` |
| **If repetitive outputs** | `temperature=1.0, top_p=0.95, min_p=0.0, repetition_penalty=1.05` |

This model is **Thinking-only** β€” chain-of-thought reasoning is always active.

> Support for sampling parameters varies across inference frameworks β€” check your runtime's documentation.

---

## Limitations

- **Thinking outputs** β€” the model may emit reasoning blocks before the final answer; downstream apps can strip them before display
- **2B scale** β€” optimized for lightweight local deployment, not frontier-scale general reasoning

---

## Provenance & licensing

Released under **Apache-2.0**, inherited from [MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B).

## Acknowledgements

- Base model: [OpenBMB / MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B)
- GGUF conversion: [llama.cpp](https://github.com/ggml-org/llama.cpp)