victory888 commited on
Commit
857edf6
ยท
verified ยท
1 Parent(s): 9b1c73c

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +152 -0
README.md ADDED
@@ -0,0 +1,152 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - ko
5
+ base_model: Qwen/Qwen2.5-7B-Instruct
6
+ datasets:
7
+ - nvidia/Nemotron-Personas-Korea
8
+ pipeline_tag: text-generation
9
+ tags:
10
+ - korean
11
+ - persona
12
+ - role-play
13
+ - survey-simulation
14
+ - distillation
15
+ - qlora
16
+ ---
17
+
18
+ # kor-persona-survey-7b
19
+
20
+ *[ํ•œ๊ตญ์–ด ์š”์•ฝ์ด ์•„๋ž˜์— ์žˆ์Šต๋‹ˆ๋‹ค / Korean summary below]*
21
+
22
+ A 7B model distilled to answer survey and preference questions **as a specific Korean persona**, given a structured persona profile. Fine-tuned from Qwen2.5-7B-Instruct via QLoRA on synthetic (persona, question, response) triplets labeled by a Qwen2.5-72B-Instruct teacher, using persona profiles from [nvidia/Nemotron-Personas-Korea](https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea).
23
+
24
+ **Research question**: *How much of a 72B teacher's persona-conditioned survey-response capability can be transferred to a 7B student through synthetic-triplet distillation alone?*
25
+
26
+ ## What it does
27
+
28
+ Input: a Korean persona profile (demographics + narrative fields) and a survey/balance-game question.
29
+ Output: a structured JSON response **in character**:
30
+
31
+ ```json
32
+ {"choice": "์งฌ๋ฝ•", "confidence": 0.8, "reason": "์–ผํฐํ•œ ๊ตญ๋ฌผ ์—†์ด๋Š” ์‹์‚ฌ๊ฐ€ ํ—ˆ์ „ํ•ด์„œ"}
33
+ ```
34
+
35
+ Intended for: persona-conditioned response simulation research, synthetic survey data generation, Korean role-play agent studies.
36
+
37
+ ## How to use
38
+
39
+ ```python
40
+ from transformers import AutoModelForCausalLM, AutoTokenizer
41
+
42
+ model_id = "dobstudio/kor-persona-survey-7b"
43
+ tok = AutoTokenizer.from_pretrained(model_id)
44
+ model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
45
+
46
+ persona = """- ์„ฑ๋ณ„/๋‚˜์ด: ์—ฌ์„ฑ, 58์„ธ
47
+ - ์ง€์—ญ: ๋ถ€์‚ฐ๊ด‘์—ญ์‹œ
48
+ - ์ง์—…: ์‹๋‹น ์šด์˜
49
+ - ์Œ์‹ ์„ฑํ–ฅ: ๋งค์šด ์Œ์‹์„ ์ฆ๊ธฐ๋ฉฐ ์ง์ ‘ ๋‹ด๊ทผ ๊น€์น˜์— ์ž๋ถ€์‹ฌ์ด ์žˆ์Œ"""
50
+
51
+ question = "์งœ์žฅ๋ฉด vs ์งฌ๋ฝ•, ํ•˜๋‚˜๋งŒ ๊ณ ๋ฅธ๋‹ค๋ฉด?"
52
+
53
+ messages = [
54
+ {"role": "system", "content": f"๋‹น์‹ ์€ ์•„๋ž˜ ์ธ๋ฌผ์ž…๋‹ˆ๋‹ค. ์ด ์ธ๋ฌผ์˜ ์ž…์žฅ์—์„œ ์„ค๋ฌธ์— ๋‹ตํ•˜์„ธ์š”.\n{persona}\n\n๋ฐ˜๋“œ์‹œ JSON์œผ๋กœ๋งŒ ๋‹ตํ•˜์„ธ์š”: {{\"choice\": ..., \"confidence\": 0.0~1.0, \"reason\": \"ํ•œ ์ค„ ์ด์œ \"}}"},
55
+ {"role": "user", "content": question},
56
+ ]
57
+ inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
58
+ out = model.generate(inputs, max_new_tokens=128)
59
+ print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
60
+ ```
61
+
62
+ With vLLM (recommended for batch simulation):
63
+
64
+ ```python
65
+ from vllm import LLM, SamplingParams
66
+
67
+ llm = LLM(model="dobstudio/kor-persona-survey-7b", max_model_len=2048)
68
+ outputs = llm.chat([messages], SamplingParams(temperature=0.7, max_tokens=128))
69
+ print(outputs[0].outputs[0].text)
70
+ ```
71
+
72
+ ## Training
73
+
74
+ | | |
75
+ |---|---|
76
+ | Base model | Qwen/Qwen2.5-7B-Instruct (Apache-2.0) |
77
+ | Method | QLoRA SFT (4-bit NF4, LoRA r=16 / ฮฑ=32, all linear layers, loss on response only), 2 epochs, effective batch 48, cosine lr 1e-4 |
78
+ | Teacher | Qwen/Qwen2.5-72B-Instruct-AWQ โ€” generated response labels; disclosed for transparency |
79
+ | Training data | 27,999 synthetic triplets (28,571 after quality filtering, 95.2% pass rate from 30,000 raw): (persona from Nemotron-Personas-Korea, question from a curated 515-question bank, teacher-labeled response). Question topics: food, lifestyle, consumption, travel, leisure, work, values, digital โ€” **politically/socially sensitive topics excluded** |
80
+ | Data release | An 800-triplet sample (16 questions ร— 50 distinct personas each): https://huggingface.co/datasets/dobstudio/kor-persona-survey-sample. The full triplet set, question bank, and the population-aggregation pipeline are not released |
81
+ | Hardware | RTX 3090 (24GB) ร—4 workstation; training ran as 3-GPU DDP, 2.5 h |
82
+
83
+ ## Evaluation
84
+
85
+ All Model-level metrics are measured on **held-out questions never seen in training**, and are reproducible with the released model and the prompt format above.
86
+
87
+ Protocol: 51 held-out questions ร— 20 evaluation personas (disjoint from training personas), temperature 0.7 with fixed per-request seeds.
88
+
89
+ ### Model-level โ€” standalone model metrics (reproducible)
90
+
91
+ | Metric | Base Qwen2.5-7B | **This model** | Teacher 72B |
92
+ |---|---|---|---|
93
+ | Format compliance (valid JSON + exact option match) | 95.5% | **98.5%** | 97.0% |
94
+ | Teacher agreement (held-out) | 72.3% | **79.4%** | 96.1%ยน |
95
+ | Response consistency (5-run majority reproducibility) | 98.7% | 97.6% | โ€” |
96
+ | Persona sensitivity (TV distance vs. no-persona baseline) | 0.356 | **0.356** | โ€” |
97
+ | Position bias (first-option preference)ยฒ | โ€” | **+1.7%** | โ€” |
98
+ | Age-conditioning sensitivity (contrastive pairs)ยณ | โ€” | **+20.8%p** | โ€” |
99
+
100
+ ยน Teacher *self*-agreement across two independent samplings โ€” the effective ceiling for
101
+ teacher agreement under temperature 0.7. The distilled model reaches **82.6% of that
102
+ ceiling** (79.4 / 96.1), up from 75.2% before distillation.
103
+ ยฒ Estimated as `(first-option share, normal order + first-option share, flipped order โˆ’ 1) / 2`
104
+ over held-out questions; 0 = unbiased. Choice consistency under option-order flip is 90.7%,
105
+ statistically at the sampling-noise floor (91.4%).
106
+ ยณ Flipping only the age field (27โ†”67) in demographics-only profiles changes the chosen
107
+ option on 29.4% of questions vs. an 8.6% same-profile resampling noise floor โ€” net +20.8%p
108
+ of genuine age conditioning.
109
+
110
+ ### Scaling ablation
111
+
112
+ We verified that the released configuration saturates this method: doubling the training
113
+ data (56k triplets from 6,000 personas) left held-out teacher agreement unchanged (79.3%)
114
+ despite improving validation loss (0.391 โ†’ 0.357), and a third epoch moved it only within
115
+ sampling error (80.1%, nโ‰ˆ980). Closing the remaining gap to the teacher self-agreement
116
+ ceiling likely requires methodological changes (multi-sample distribution distillation,
117
+ preference optimization) rather than more of the same data.
118
+
119
+ ### System-level reference โ€” not yet measured
120
+
121
+ The authors plan to report population-level accuracy (MAE of predicted answer distributions
122
+ vs. real Korean survey results, using stratified sampling + post-stratification against KOSIS
123
+ census margins) in a future update. Those numbers will depend on a private aggregation
124
+ pipeline and will be reported for context only.
125
+
126
+ ## Limitations and ethical considerations
127
+
128
+ - **Synthetic personas are not real people.** Outputs simulate what a fictional profile *might* answer based on LLM priors โ€” they are not measurements of actual Korean public opinion and must not be presented as such.
129
+ - **Joint-distribution distortion in the source data.** Nemotron-Personas-Korea matches real Korean marginal distributions (age, sex, occupation) but its joint distributions (e.g., ageร—occupationร—region) deviate from reality (see arXiv:2606.12433). Segment-level conclusions drawn from raw persona samples are unreliable without statistical correction.
130
+ - **Stereotype risk.** The teacher may role-play personas stereotypically, and distillation inherits this. Persona-sensitivity metrics partially quantify conditioning, not fairness.
131
+ - **Domain bound.** Trained on non-sensitive preference/lifestyle questions. Behavior on political, medical, or otherwise sensitive questions is untested and out of scope.
132
+ - Outputs are in Korean; other languages are untested.
133
+
134
+ ## Attribution & license
135
+
136
+ - Model weights: **Apache-2.0**.
137
+ - Persona profiles: [nvidia/Nemotron-Personas-Korea](https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea), **CC BY 4.0** โ€” ยฉ NVIDIA, used with attribution as required.
138
+ - Base and teacher models: Qwen2.5 family (Apache-2.0), Alibaba Cloud.
139
+
140
+ ---
141
+
142
+ ## ํ•œ๊ตญ์–ด ์š”์•ฝ
143
+
144
+ ํ•œ๊ตญ์ธ ๊ฐ€์ƒ ์ธ๋ฌผ ํ”„๋กœํ•„์„ ์ฃผ๋ฉด **๊ทธ ์ธ๋ฌผ์˜ ์ž…์žฅ์—์„œ** ์„ค๋ฌธ/๋ฐธ๋Ÿฐ์Šค๊ฒŒ์ž„ ๋ฌธํ•ญ์— ๊ตฌ์กฐํ™”๋œ JSON์œผ๋กœ ๋‹ตํ•˜๋Š” 7B ๋ชจ๋ธ์ž…๋‹ˆ๋‹ค. Qwen2.5-72B ๊ต์‚ฌ ๋ชจ๋ธ์ด ์ƒ์„ฑํ•œ (ํŽ˜๋ฅด์†Œ๋‚˜, ๋ฌธํ•ญ, ์‘๋‹ต) ํ•ฉ์„ฑ ๋ฐ์ดํ„ฐ๋กœ Qwen2.5-7B๋ฅผ QLoRA ์ฆ๋ฅ˜ํ–ˆ์Šต๋‹ˆ๋‹ค.
145
+
146
+ - **์—ฐ๊ตฌ ์งˆ๋ฌธ**: 72B์˜ ํŽ˜๋ฅด์†Œ๋‚˜ ์กฐ๊ฑด๋ถ€ ์‘๋‹ต ๋Šฅ๋ ฅ์ด ์ฆ๋ฅ˜๋งŒ์œผ๋กœ 7B์— ์–ผ๋งˆ๋‚˜ ์ด์‹๋˜๋Š”๊ฐ€
147
+ - **ํ‰๊ฐ€**: ํ•™์Šต์— ์“ฐ์ง€ ์•Š์€ ๋ฌธํ•ญ(held-out)์—์„œ ํฌ๋งท ์ค€์ˆ˜์œจ, ๊ต์‚ฌ ์ผ์น˜์œจ, ์‘๋‹ต ์ผ๊ด€์„ฑ, ํŽ˜๋ฅด์†Œ๋‚˜ ๊ฐ์ˆ˜์„ฑ์„ ์ธก์ • โ€” ๊ณต๊ฐœ๋œ ๋ชจ๋ธ๋งŒ์œผ๋กœ ์žฌํ˜„ ๊ฐ€๋Šฅ
148
+ - **ํ•œ๊ณ„**: ๊ฐ€์ƒ ์ธ๋ฌผ์˜ ์‘๋‹ต์€ ์‹ค์ œ ์—ฌ๋ก ์ด ์•„๋‹Œ ์ถ”์ •์ด๋ฉฐ, ์›๋ณธ ๋ฐ์ดํ„ฐ์˜ ๊ฒฐํ•ฉ๋ถ„ํฌ ์™œ๊ณก์œผ๋กœ ์„ธ๋ถ€ ์ง‘๋‹จ ๋ถ„์„์—๋Š” ํ†ต๊ณ„ ๋ณด์ •์ด ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค. ๋ฏผ๊ฐ ์ฃผ์ œ๋Š” ํ•™์Šต์—์„œ ์ œ์™ธํ–ˆ์Šต๋‹ˆ๋‹ค.
149
+ - ๋ชจ์ง‘๋‹จ ์ง‘๊ณ„ ํŒŒ์ดํ”„๋ผ์ธ(์ธตํ™” ํ‘œ๋ณธ์ถ”์ถœยท์‚ฌํ›„์ธตํ™” ๊ฐ€์ค‘)์€ ๋น„๊ณต๊ฐœ์ด๋ฉฐ, ๋ณธ ๊ณต๊ฐœ๋ฌผ์€ ๊ฐœ์ธ ํŽ˜๋ฅด์†Œ๋‚˜ ์—ญํ• ๊ทน ๊ธฐ๋Šฅ๋งŒ ์ œ๊ณตํ•ฉ๋‹ˆ๋‹ค.
150
+
151
+ ---
152
+