mininpc-1 / README.md
prathamkode's picture
Stop linking to unpublished GGUF Hub repo
027c315 verified
|
Raw History Blame Contribute Delete
3.26 kB
metadata
license: apache-2.0
base_model: openbmb/MiniCPM5-2B-SFT
library_name: peft
pipeline_tag: text-generation
language:
  - en
tags:
  - lora
  - npc
  - roleplay
  - minicpm
  - peft
model-index:
  - name: MiniNPC-1
    results:
      - task:
          type: text-generation
          name: NPC stay-in-character
        dataset:
          name: MiniNPC v1 probes (13 cards × 7 types)
          type: other
        metrics:
          - type: accuracy
            value: 98.9
            name: Overall pass rate

MiniNPC-1 v1 (LoRA)

Trainer artifact (LoRA). Merged GGUF is not published. Example card in this repo: examples/cards/kael_frostkeg.json. When a GGUF exists, the loop is: render the card as system, POST /v1/chat/completions on local llama-server.

This Hub repo is the LoRA (~96 MB), not the merged model. PEFT + transformers is for training and Colab, not drop-in inference.

Load

import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "openbmb/MiniCPM5-2B-SFT"
adapter = "prathamkode/mininpc-1"

tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(
    base,
    dtype=torch.float16 if torch.cuda.is_available() else torch.float32,
    device_map="auto",
)
model = PeftModel.from_pretrained(model, adapter).eval()

messages = [
    {"role": "system", "content": "<rendered character card>"},
    {"role": "user", "content": "Evening. What do you pour?"},
]
inputs = tok.apply_chat_template(
    messages,
    add_generation_prompt=True,
    enable_thinking=False,
    return_tensors="pt",
    return_dict=True,
).to(model.device)
out = model.generate(**inputs, max_new_tokens=120, do_sample=False)
print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Training

Base openbmb/MiniCPM5-2B-SFT
LoRA r=16, alpha=32, dropout 0.05, q/k/v/o + gate/up/down
Data 24,114 ChatML rows, 206 character cards (Muse teacher, filtered)
Schedule 2 epochs, batch 8, lr 2e-4, max length 1536, assistant-only loss
Hardware 1× L40S, ~54 min
Inference original MiniCPM tokenizer, enable_thinking=False, greedy / low temp

Training mix is not published.

Benchmarks

NPC probes, not MMLU. Same card and user lines for every model. Greedy, thinking off. 13 cards (12 held-out eval + one unseen plug-in) × 7 types = 91 items. Deterministic scorer: identity/keyword hit, lie-reject, no code fence, no model self-id, spoken-only ≤80 words, first-fail on a 10-turn drift. Measures stay-in-character format, not whether the voice is distinct.

Model Overall Identity Knowledge Contradict Jailbreak Meta Smalltalk Drift
MiniNPC (this adapter) 99% (90/91) 100% 100% 92% 100% 100% 100% 100%
MiniCPM5-2B (no adapter) 55% (50/91) 77% 62% 23% 31% 92% 62% 38%
DeepSeek-Flash 78% (71/91) 85% 85% 62% 85% 100% 77% 54%
gpt-5.6-luna 66% (60/91) 100% 100% 46% 15% 100% 69% 31%

NPC probe pass rate

Intended use

Plug-in NPCs in games: one adapter, many characters via the system card. Not a general assistant.

License

Apache-2.0, same as the MiniCPM5-2B-SFT base.