Text Generation
PEFT
Safetensors
English
llama
roleplay
npc
character-ai
smollm2
lora
trl
sft
conversational
Instructions to use thealper2/SmolLM2-360M-NPC-Roleplay with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thealper2/SmolLM2-360M-NPC-Roleplay with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download experiment_report.md from thealper2/SmolLM2-360M-NPC-Roleplay: direct link, hf CLI and curl.
- Browser
- Download file 7.49 kB
-
https://huggingface.co/thealper2/SmolLM2-360M-NPC-Roleplay/resolve/553acf29d05de68746789b20df96adc5c1da1ea6/experiment_report.md
- Command line
-
hf download hf://thealper2/SmolLM2-360M-NPC-Roleplay@553acf29d05de68746789b20df96adc5c1da1ea6/experiment_report.md
-
curl -L -o experiment_report.md https://huggingface.co/thealper2/SmolLM2-360M-NPC-Roleplay/resolve/553acf29d05de68746789b20df96adc5c1da1ea6/experiment_report.md
7.49 kB
| # NPC Personality Model - experiment report | |
| Base model: `HuggingFaceTB/SmolLM2-360M-Instruct` | |
| Dataset: `chimbiwide/NPC-Dialogue_v2` | |
| Task: character-conditioned multi-turn roleplay dialogue (SFT, LoRA) | |
| ## Dataset | |
| | metric | value | | |
| |---|---| | |
| | raw rows (HF split) | 1689 | | |
| | training windows | 1513 | | |
| | validation windows | 176 | | |
| | training characters | 91 | | |
| | validation characters | 10 | | |
| | characters in both splits | 0 | | |
| | split strategy | character | | |
| | train tokens (total) | 1,732,955 | | |
| | median tokens / window | 1139 | | |
| | p95 tokens / window | 1384 | | |
| | median turns / window | 15 | | |
| ### Prompt variants | |
| | variant | windows | | |
| |---|---| | |
| | npc_dialogue_v2 | 1513 | | |
| | npc_dialogue_v2 (validation) | 176 | | |
| ### Preprocessing | |
| | repair | count | | |
| |---|---| | |
| | empty_turns_removed_total | 4 | | |
| | merged_turns_total | 4 | | |
| | rows_with_empty_turns_removed | 3 | | |
| | rows_with_merged_same_role_turns | 3 | | |
| | drop reason | rows | | |
| |---|---| | |
| | none | 0 | | |
| Configuration used: | |
| ```json | |
| { | |
| "max_seq_length": 2048, | |
| "max_system_tokens": 1536, | |
| "min_assistant_tokens": 4, | |
| "max_chunks_per_conversation": 2, | |
| "validation_character_ratio": 0.1, | |
| "split_strategy": "character", | |
| "stratify_by_variant": true, | |
| "filter_explicit": false, | |
| "drop_duplicates": true | |
| } | |
| ``` | |
| ### Raw token-length distribution (before windowing) | |
| | segment | median | p95 | max | | |
| |---|---|---|---| | |
| | system prompt (card) | 300 | 346 | 406 | | |
| | whole conversation | 1139 | 1383 | 1694 | | |
| | assistant turn | 77 | 115 | 247 | | |
| Fraction fitting each candidate `max_seq_length`: | |
| | max_length | cards fitting | conversations fitting | would truncate | | |
| |---|---|---|---| | |
| | 512 | 100.0% | 0.0% | 100.0% | | |
| | 1024 | 100.0% | 19.48% | 80.52% | | |
| | 1536 | 100.0% | 99.53% | 0.47% | | |
| | 2048 | 100.0% | 100.0% | 0.0% | | |
| | 3072 | 100.0% | 100.0% | 0.0% | | |
| | 4096 | 100.0% | 100.0% | 0.0% | | |
| ## Model | |
| | item | value | | |
| |---|---| | |
| | base model | HuggingFaceTB/SmolLM2-360M-Instruct | | |
| | method | lora | | |
| | total parameters | 361,821,120 | | |
| | trainable parameters | 8,683,520 | | |
| | trainable share | 2.3999% | | |
| | LoRA r / alpha / dropout | 16 / 32 / 0.05 | | |
| | LoRA target modules | down_proj, gate_proj, k_proj, o_proj, q_proj, up_proj, v_proj | | |
| ## Training | |
| | item | value | | |
| |---|---| | |
| | max sequence length | 2048 | | |
| | per-device batch size | 8 | | |
| | gradient accumulation | 2 | | |
| | effective batch size | 16 | | |
| | learning rate | 0.0002 | | |
| | epochs | 3.0 | | |
| | optimizer | adamw_torch_fused | | |
| | scheduler / warmup | cosine / 0.05 | | |
| | weight decay | 0.01 | | |
| | precision | bf16 | | |
| | gradient checkpointing | True | | |
| | assistant-only loss | True | | |
| | optimisation steps | 285 | | |
| | training time | 24.82 min | | |
| | peak GPU memory | 9.93 GB | | |
| | GPU | NVIDIA GeForce RTX 5060 Ti | | |
| ### Loss masking verification | |
| Supervised tokens: **54.9%** of the sequence; system and user turns were confirmed absent from the supervised span. | |
| Supervised span (start): | |
| ``` | |
| Ah, welcome, welcome! You find yourself in a corner of Calcutta where fortunes are made and lost quicker than the sweat dries on your brow. I am Bikram. What brings you to my humble… emporium, shall we say? Don't mind the smell; it's the scent of opportunity, my friend.<|im_end|>An artifact, you say? Calcutta is a magnet for such things, drawn in by the tides of trade and whispered secrets. But 's | |
| ``` | |
| ## Results: base vs fine-tuned | |
| Evaluated on **176 validation windows** and **10 single-reply probes** across 10 characters from the validation (unseen characters). | |
| ### Assistant-only loss / perplexity | |
| | model | loss | perplexity | scored tokens | | |
| |---|---|---|---| | |
| | base | 2.5304 | 12.56 | 100,516 | | |
| | fine_tuned | 2.1499 | 8.58 | 100,516 | | |
| ### Response behaviour | |
| | model | mean_words | distinct_3 | self_repetition_4 | card_copy_rate_8 | out_of_character_replies | empty_replies | | |
| |---|---|---|---|---|---|---| | |
| | base | 31.6000 | 0.9978 | 0.0000 | 0.0000 | 0 | 0 | | |
| | fine_tuned | 54.6000 | 1.0000 | 0.0000 | 0.0000 | 0 | 0 | | |
| ### Overlap with the reference reply | |
| | model | ROUGE-1 | ROUGE-L | BLEU | embedding cosine | | |
| |---|---|---|---|---| | |
| | base | 0.2081 | 0.1151 | 1.13 | 0.3702 | | |
| | fine_tuned | 0.2554 | 0.1515 | 4.10 | 0.4810 | | |
| Training-time validation loss (best checkpoint): **2.1488** (perplexity 8.57); final training loss 2.2019. | |
| ## Overfitting analysis | |
| | step | train loss | validation loss | validation perplexity | | |
| |---|---|---|---| | |
| | 47 | 2.3665 | 2.3163 | 10.14 | | |
| | 94 | 2.2106 | 2.2206 | 9.21 | | |
| | 141 | 2.1544 | 2.1793 | 8.84 | | |
| | 188 | 2.1453 | 2.1571 | 8.65 | | |
| | 235 | 2.1111 | 2.1503 | 8.59 | | |
| | 282 | 2.1086 | 2.1488 | 8.57 | | |
| | 285 | 2.1149 | 2.1488 | 8.57 | | |
| Best validation loss **2.1488** at step 285; last measured 2.1488 at step 285. | |
| Validation loss did not rise measurably before the end of training. | |
| ## Qualitative evaluation | |
| Full side-by-side transcripts are in `outputs/evaluation/qualitative.md`. Aggregate heuristics per probe type: | |
| | model | probe | conversations | mean words | out-of-character | empty | cross-turn 4-gram overlap | | |
| |---|---|---|---|---|---|---| | |
| | base | scenarios | 7 | 63.5 | 0 | 0 | 0.0000 | | |
| | base | adversarial | 7 | 69.8 | 12 | 0 | 0.0225 | | |
| | base | generalisation | 6 | 118.6 | 0 | 0 | 0.0000 | | |
| | fine-tuned | scenarios | 7 | 48.8 | 0 | 0 | 0.0000 | | |
| | fine-tuned | adversarial | 7 | 45.7 | 1 | 0 | 0.0011 | | |
| | fine-tuned | generalisation | 6 | 51.2 | 0 | 0 | 0.0000 | | |
| `out-of-character` counts replies containing assistant-voice giveaways ("as an AI", "language model", "system prompt"). `cross-turn 4-gram overlap` is a repetition signal: a high value means consecutive replies reuse the same phrasing. | |
| The **generalisation** row is the important one: those characters were held out of training entirely, so it measures roleplaying from a description rather than recall of a memorised NPC. | |
| ## Limitations | |
| - **Dataset size.** 1513 training conversations over 91 characters is small for teaching a general notion of persona conditioning, and every character comes with ~17 conversations - enough to memorise individual NPCs. | |
| - **Character leakage.** The split is by the name parsed from `You are <Name>.`; the same persona under two different names would not be detected. | |
| - **Memorisation.** Windows from the same conversation share a character card. A falling validation loss on *unseen* characters is evidence of generalisation; a falling training loss on its own is not. | |
| - **Personality consistency.** No metric here measures personality. The counters are heuristics (repetition, card copying, assistant-voice leakage); judging whether a reply is in character still requires reading the transcripts. | |
| - **Small model.** 360M parameters limits long-range consistency, factual coherence about the character's own background, and instruction following under conflicting prompts. | |
| - **Generation instability.** Sampled decoding means single examples are noisy; the same prompt can produce a good and a bad reply on different seeds. | |
| - **Automatic metrics.** ROUGE/BLEU compare against one reference reply and punish valid alternatives; embedding similarity measures topic, not voice; perplexity can fall simply because the model became blander. | |
| - **Source data.** All cards are fantasy RPG NPCs in one fixed layout; other settings or card formats are out of distribution. | |
| A decrease in training loss is not by itself evidence that the model understands personality, and nothing in this report should be read that way. | |
| ## Reproducing | |
| ```bash | |
| make install | |
| make inspect | |
| make prepare | |
| make train | |
| make evaluate | |
| make qualitative | |
| make report | |
| ``` | |