GnLOLot's picture
Initial release: MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic
934bcdf verified
|
Raw History Blame
5.04 kB

MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic

MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic

GGUF 量化版本(用于本地部署):**MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic-GGUF**

English

MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic 是一个基于 openbmb/MiniCPM5-2B 的紧凑型 2B 思维链(Thinking) 语言模型。使用 Claude 数据进行微调,重点增强了 Agent 工具调用 / 函数调用、代码生成 和 指令遵循 能力,同时保留了 MiniCPM5 原生的 Thinking 聊天模板和 XML 工具调用格式。


概览

项目 详情
基座模型 openbmb/MiniCPM5-2B(2B 稠密 Llama 架构)
后训练数据 Claude 数据
核心能力 Agent 工具调用、代码生成、指令遵循、思维链推理
聊天格式 MiniCPM5 原生 Thinking 模板,支持可选的思维链输出
上下文长度 128K(max_position_embeddings = 131072)
精度 bfloat16
部署 单卡友好,适合边缘设备 / 本地部署

评测结果

ClawBench(Agent 代码能力)

模型 QwenClawBench WildClawBench
MiniCPM5-2B(基座,仅 RL) 42.11 23.19
MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic 44.56(+2.45) 24.32(+1.13)

ClawBench 评测 Agent 代码能力——模型自主使用工具、浏览代码库、完成多步软件工程任务的综合能力。QwenClawBench 使用结构化代码场景;WildClawBench 测试多样化的真实任务。

更多评测成绩(BFCL、SWE-bench、Tau-Bench 等)补充中。


能力

  • Agent 工具调用 — 基于 MiniCPM5 原生格式的可靠 XML / 函数调用风格工具使用,适配多步 Agent 工作流
  • 代码生成 — 代码编写、调试及软件工程任务
  • 指令遵循 — 可靠地遵循用户指令和结构化约束
  • 思维链模式 — 通过 MiniCPM5 聊天模板实现链式推理
  • 长上下文 — 最高支持 128K tokens(131,072 tokens)

快速开始

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "GnLOLot/MiniCPM5-2B-Claude-Fable5-1-Thinking-Agentic"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "写一个合并两个有序列表的 Python 函数。"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

工具调用示例

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "获取指定城市的当前天气。",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "城市名称"}
                },
                "required": ["city"]
            }
        }
    }
]

messages = [
    {"role": "user", "content": "北京今天天气怎么样?"}
]

text = tokenizer.apply_chat_template(
    messages, tools=tools, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

采样建议

继承自 openbmb/MiniCPM5-2B:

场景 参数
默认 temperature=1.0, top_p=0.95, min_p=0.0
如遇重复输出 temperature=1.0, top_p=0.95, min_p=0.0, repetition_penalty=1.05

本模型为 纯思考模式(Thinking-only)——思维链推理始终开启。

不同推理框架对采样参数的支持有所差异,请参考所用框架的文档。


局限性

  • 思维链输出 — 模型可能在最终回答前输出推理过程,下游应用可在展示前过滤
  • 2B 规模 — 针对轻量级本地部署优化,非前沿级通用推理

来源与许可

基于 Apache-2.0 发布,继承自 MiniCPM5-2B。

致谢