--- license: apache-2.0 base_model: openbmb/MiniCPM5-2B-SFT library_name: peft pipeline_tag: text-generation language: - en tags: - lora - npc - roleplay - minicpm - peft model-index: - name: MiniNPC-1 results: - task: type: text-generation name: NPC stay-in-character dataset: name: MiniNPC v1 probes (13 cards × 7 types) type: other metrics: - type: accuracy value: 98.9 name: Overall pass rate --- # MiniNPC-1 v1 (LoRA) Trainer artifact (LoRA). Merged GGUF is not published. Example card in this repo: `examples/cards/kael_frostkeg.json`. When a GGUF exists, the loop is: render the card as `system`, POST `/v1/chat/completions` on local llama-server. This Hub repo is the **LoRA** (~96 MB), not the merged model. PEFT + transformers is for training and Colab, not drop-in inference. ## Load ```python import torch from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer base = "openbmb/MiniCPM5-2B-SFT" adapter = "prathamkode/mininpc-1" tok = AutoTokenizer.from_pretrained(base) model = AutoModelForCausalLM.from_pretrained( base, dtype=torch.float16 if torch.cuda.is_available() else torch.float32, device_map="auto", ) model = PeftModel.from_pretrained(model, adapter).eval() messages = [ {"role": "system", "content": ""}, {"role": "user", "content": "Evening. What do you pour?"}, ] inputs = tok.apply_chat_template( messages, add_generation_prompt=True, enable_thinking=False, return_tensors="pt", return_dict=True, ).to(model.device) out = model.generate(**inputs, max_new_tokens=120, do_sample=False) print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)) ``` ## Training | | | |---|---| | Base | `openbmb/MiniCPM5-2B-SFT` | | LoRA | r=16, alpha=32, dropout 0.05, q/k/v/o + gate/up/down | | Data | 24,114 ChatML rows, 206 character cards (Muse teacher, filtered) | | Schedule | 2 epochs, batch 8, lr 2e-4, max length 1536, assistant-only loss | | Hardware | 1× L40S, ~54 min | | Inference | original MiniCPM tokenizer, `enable_thinking=False`, greedy / low temp | Training mix is not published. ## Benchmarks NPC probes, not MMLU. Same card and user lines for every model. Greedy, thinking off. 13 cards (12 held-out eval + one unseen plug-in) × 7 types = 91 items. Deterministic scorer: identity/keyword hit, lie-reject, no code fence, no model self-id, spoken-only ≤80 words, first-fail on a 10-turn drift. Measures stay-in-character format, not whether the voice is distinct. | Model | Overall | Identity | Knowledge | Contradict | Jailbreak | Meta | Smalltalk | Drift | |---|---|---|---|---|---|---|---|---| | **MiniNPC** (this adapter) | **99%** (90/91) | 100% | 100% | 92% | 100% | 100% | 100% | 100% | | MiniCPM5-2B (no adapter) | 55% (50/91) | 77% | 62% | 23% | 31% | 92% | 62% | 38% | | DeepSeek-Flash | 78% (71/91) | 85% | 85% | 62% | 85% | 100% | 77% | 54% | | gpt-5.6-luna | 66% (60/91) | 100% | 100% | 46% | 15% | 100% | 69% | 31% | ![NPC probe pass rate](chart.png) ## Intended use Plug-in NPCs in games: one adapter, many characters via the system card. Not a general assistant. ## License Apache-2.0, same as the MiniCPM5-2B-SFT base.