--- library_name: transformers license: apache-2.0 language: - en - zh base_model: openbmb/MiniCPM5-2B base_model_relation: finetune pipeline_tag: text-generation tags: - minicpm - minicpm5 - llama - text-generation - thinking - fable5 - tool-calling - function-calling - agentic - coding - instruction-following - conversational ---

MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic

# MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic GGUF quantizations for local deployment: **[MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF](https://huggingface.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF)** **MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic** is a compact 2B **Thinking** language model built on [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B). Fine-tuned on **Claude** data with a strong focus on **agentic tool calling / function calling**, **coding**, and **instruction following**. It keeps MiniCPM5's native Thinking chat template and XML tool-call format. For llama.cpp / Ollama / LM Studio deployment, see the **[GGUF repository](https://huggingface.co/GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF)**. --- ## Overview | Item | Detail | |---|---| | **Base model** | [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B) (2B dense Llama architecture) | | **Post-training** | Claude data | | **Key capabilities** | **Agentic tool calling**, coding, instruction following, chain-of-thought reasoning | | **Chat format** | MiniCPM5 native Thinking template with optional chain-of-thought blocks | | **Context length** | **128K** (`max_position_embeddings = 131072`) | | **Precision** | bfloat16 | | **Deployment** | Single-GPU friendly; suitable for edge / local use | --- ## Capabilities - **Agentic tool calling** — reliable XML / function-calling style tool use on top of MiniCPM5's native format, designed for multi-step agentic workflows - **Coding** — code generation, debugging, and software-engineering-style tasks - **Instruction following** — reliable adherence to user prompts and structured constraints - **Thinking mode** — chain-of-thought reasoning via the MiniCPM5 chat template - **Long context** — up to **128K tokens** (131,072 tokens per `config.json`) --- ## Benchmark ### ClawBench (Agentic Coding) | Model | QwenClawBench | WildClawBench | |---|---|---| | MiniCPM5-2B (Base, RL-only) | 42.11 | 23.19 | | **MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic** | **44.56** (+2.45) | **24.32** (+1.13) | > ClawBench evaluates agentic coding ability — the model's capacity to autonomously use tools, navigate codebases, and complete multi-step software engineering tasks. QwenClawBench uses structured coding scenarios; WildClawBench tests on diverse real-world tasks. > **More benchmarks (BFCL, SWE-bench, Tau-Bench, etc.) coming soon.** --- ## Quick start ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch model_id = "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic" tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, torch_dtype=torch.bfloat16, device_map="auto", ) messages = [{"role": "user", "content": "Write a Python function to merge two sorted lists."}] text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tokenizer(text, return_tensors="pt").to(model.device) outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)) ``` ### Tool calling example ```python tools = [ { "type": "function", "function": { "name": "get_weather", "description": "Get the current weather for a given city.", "parameters": { "type": "object", "properties": { "city": {"type": "string", "description": "City name"} }, "required": ["city"] } } } ] messages = [ {"role": "user", "content": "What's the weather like in Beijing?"} ] text = tokenizer.apply_chat_template( messages, tools=tools, tokenize=False, add_generation_prompt=True ) inputs = tokenizer(text, return_tensors="pt").to(model.device) outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)) ``` --- ## Sampling recommendations Inherited from [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B): | Scenario | Params | |---|---| | **Default** | `temperature=1.0, top_p=0.95, min_p=0.0` | | **If repetitive outputs** | `temperature=1.0, top_p=0.95, min_p=0.0, repetition_penalty=1.05` | This model is **Thinking-only** — chain-of-thought reasoning is always active. > Support for sampling parameters varies across inference frameworks — check your runtime's documentation. --- ## Limitations - **Thinking outputs** — the model may emit reasoning blocks before the final answer; downstream apps can strip them before display - **2B scale** — optimized for lightweight local deployment, not frontier-scale general reasoning --- ## Provenance & licensing Released under **Apache-2.0**, inherited from [MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B). ## Acknowledgements - Base model: [OpenBMB / MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B) - GGUF conversion: [llama.cpp](https://github.com/ggml-org/llama.cpp)