GnLOLot's picture
Initial release: MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic
934bcdf verified
|
Raw History Blame Contribute Delete
5.72 kB
metadata
library_name: transformers
license: apache-2.0
language:
  - en
  - zh
base_model: openbmb/MiniCPM5-2B
base_model_relation: finetune
pipeline_tag: text-generation
tags:
  - minicpm
  - minicpm5
  - llama
  - text-generation
  - thinking
  - fable5
  - tool-calling
  - function-calling
  - agentic
  - coding
  - instruction-following
  - conversational

MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic

MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic

GGUF quantizations for local deployment: MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF

MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic is a compact 2B Thinking language model built on openbmb/MiniCPM5-2B. Fine-tuned on Claude data with a strong focus on agentic tool calling / function calling, coding, and instruction following. It keeps MiniCPM5's native Thinking chat template and XML tool-call format.

For llama.cpp / Ollama / LM Studio deployment, see the GGUF repository.


Overview

Item Detail
Base model openbmb/MiniCPM5-2B (2B dense Llama architecture)
Post-training Claude data
Key capabilities Agentic tool calling, coding, instruction following, chain-of-thought reasoning
Chat format MiniCPM5 native Thinking template with optional chain-of-thought blocks
Context length 128K (max_position_embeddings = 131072)
Precision bfloat16
Deployment Single-GPU friendly; suitable for edge / local use

Capabilities

  • Agentic tool calling β€” reliable XML / function-calling style tool use on top of MiniCPM5's native format, designed for multi-step agentic workflows
  • Coding β€” code generation, debugging, and software-engineering-style tasks
  • Instruction following β€” reliable adherence to user prompts and structured constraints
  • Thinking mode β€” chain-of-thought reasoning via the MiniCPM5 chat template
  • Long context β€” up to 128K tokens (131,072 tokens per config.json)

Benchmark

ClawBench (Agentic Coding)

Model QwenClawBench WildClawBench
MiniCPM5-2B (Base, RL-only) 42.11 23.19
MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic 44.56 (+2.45) 24.32 (+1.13)

ClawBench evaluates agentic coding ability β€” the model's capacity to autonomously use tools, navigate codebases, and complete multi-step software engineering tasks. QwenClawBench uses structured coding scenarios; WildClawBench tests on diverse real-world tasks.

More benchmarks (BFCL, SWE-bench, Tau-Bench, etc.) coming soon.


Quick start

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Write a Python function to merge two sorted lists."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Tool calling example

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a given city.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "City name"}
                },
                "required": ["city"]
            }
        }
    }
]

messages = [
    {"role": "user", "content": "What's the weather like in Beijing?"}
]

text = tokenizer.apply_chat_template(
    messages, tools=tools, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Sampling recommendations

Inherited from openbmb/MiniCPM5-2B:

Scenario Params
Default temperature=1.0, top_p=0.95, min_p=0.0
If repetitive outputs temperature=1.0, top_p=0.95, min_p=0.0, repetition_penalty=1.05

This model is Thinking-only β€” chain-of-thought reasoning is always active.

Support for sampling parameters varies across inference frameworks β€” check your runtime's documentation.


Limitations

  • Thinking outputs β€” the model may emit reasoning blocks before the final answer; downstream apps can strip them before display
  • 2B scale β€” optimized for lightweight local deployment, not frontier-scale general reasoning

Provenance & licensing

Released under Apache-2.0, inherited from MiniCPM5-2B.

Acknowledgements