victory888 commited on
Commit
83913d7
ยท
verified ยท
1 Parent(s): 857edf6

Upload REPORT.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. REPORT.md +140 -0
REPORT.md ADDED
@@ -0,0 +1,140 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Technical Report โ€” kor-persona-survey-7b
2
+
3
+ *Distilling persona-conditioned Korean survey responses from a 72B teacher into a 7B student*
4
+ *72B ๊ต์‚ฌ์˜ ํ•œ๊ตญ์–ด ํŽ˜๋ฅด์†Œ๋‚˜ ์กฐ๊ฑด๋ถ€ ์„ค๋ฌธ ์‘๋‹ต ๋Šฅ๋ ฅ์„ 7B ํ•™์ƒ์— ์ฆ๋ฅ˜ํ•˜๊ธฐ*
5
+
6
+ DOB Studio ยท 2026-08
7
+
8
+ ## 1. Motivation ยท ๋™๊ธฐ
9
+
10
+ Simulating survey responses with LLM role-play requires one inference per (persona,
11
+ question) pair โ€” population-scale simulation means thousands of inferences per question.
12
+ Large models judge personas well but are too slow for this workload; small models are fast
13
+ but role-play shallowly and break output formats. We ask:
14
+
15
+ > **RQ. How much of a 72B teacher's persona-conditioned survey-response capability can be
16
+ > transferred to a 7B student through synthetic-triplet distillation alone?**
17
+
18
+ **(ํ•œ๊ตญ์–ด)** LLM ์—ญํ• ๊ทน์œผ๋กœ ์„ค๋ฌธ ์‘๋‹ต์„ ์‹œ๋ฎฌ๋ ˆ์ด์…˜ํ•˜๋ ค๋ฉด (ํŽ˜๋ฅด์†Œ๋‚˜, ๋ฌธํ•ญ) ์Œ๋งˆ๋‹ค ์ถ”๋ก ์ด
19
+ 1ํšŒ์”ฉ ํ•„์š”ํ•˜๋‹ค โ€” ๋ชจ์ง‘๋‹จ ๊ทœ๋ชจ ์‹œ๋ฎฌ๋ ˆ์ด์…˜์€ ๋ฌธํ•ญ๋‹น ์ˆ˜์ฒœ ํšŒ ์ถ”๋ก ์„ ๋œปํ•œ๋‹ค. ๋Œ€ํ˜• ๋ชจ๋ธ์€
20
+ ํŽ˜๋ฅด์†Œ๋‚˜ ํŒ๋‹จ์€ ๊นŠ์ง€๋งŒ ์ด ์ž‘์—…๋Ÿ‰์—๋Š” ๋„ˆ๋ฌด ๋А๋ฆฌ๊ณ , ์†Œํ˜• ๋ชจ๋ธ์€ ๋น ๋ฅด์ง€๋งŒ ์—ญํ• ๊ทน์ด ์–•๊ณ 
21
+ ์ถœ๋ ฅ ํ˜•์‹์„ ์–ด๊ธด๋‹ค. ๊ทธ๋ž˜์„œ ๋ฌป๋Š”๋‹ค: **72B ๊ต์‚ฌ์˜ ํŽ˜๋ฅด์†Œ๋‚˜ ์กฐ๊ฑด๋ถ€ ์„ค๋ฌธ ์‘๋‹ต ๋Šฅ๋ ฅ์„, ํ•ฉ์„ฑ
22
+ ์‚ผ์ค‘ํ•ญ ์ฆ๋ฅ˜๋งŒ์œผ๋กœ 7B ํ•™์ƒ์— ์–ผ๋งˆ๋‚˜ ์ด์‹ํ•  ์ˆ˜ ์žˆ๋Š”๊ฐ€?**
23
+
24
+ ## 2. Data construction ยท ๋ฐ์ดํ„ฐ ๊ตฌ์ถ•
25
+
26
+ The source dataset (nvidia/Nemotron-Personas-Korea, 1M synthetic Korean personas, CC BY 4.0)
27
+ contains persona profiles only โ€” no questions, no answers. We manufactured the missing
28
+ supervision:
29
+
30
+ 1. **Question bank** (private): 515 binary preference questions across 8 non-sensitive
31
+ topics (food, lifestyle, consumption, travel, leisure, work, values, digital), authored
32
+ to span age/region/occupation-sensitive contrasts. A stratified 10% (51 questions) was
33
+ frozen as a held-out evaluation set **before any labeling**.
34
+ 2. **Teacher labeling**: Qwen2.5-72B-Instruct-AWQ answered (persona, question) pairs in
35
+ character โ€” 3,000 personas ร— 10 questions = 30,000 raw labels of the form
36
+ `{choice, confidence, reason}`.
37
+ 3. **Quality filtering** (95.2% pass): exact option match, valid JSON, confidence โˆˆ [0,1],
38
+ reason length bounds, per-question duplicate-reason cap.
39
+ 4. **Split hygiene**: evaluation personas (n=20) and held-out questions were excluded from
40
+ all training data; the validation split (572 triplets) is persona-disjoint from training.
41
+
42
+ **(ํ•œ๊ตญ์–ด)** ์›๋ณธ ๋ฐ์ดํ„ฐ์…‹(๊ฐ€์ƒ ํ•œ๊ตญ์ธ 100๋งŒ ๋ช…)์—๋Š” ์ธ๋ฌผ ํ”„๋กœํ•„๋งŒ ์žˆ๊ณ  ์งˆ๋ฌธยท๋Œ€๋‹ต์ด
43
+ ์—†๋‹ค. ๊ทธ๋ž˜์„œ ํ•™์Šต ์‹ ํ˜ธ๋ฅผ ์ง์ ‘ ์ œ์กฐํ–ˆ๋‹ค: โ‘  ๋น„๋ฏผ๊ฐ 8๊ฐœ ์ฃผ์ œ์˜ ์–‘์žํƒ์ผ ๋ฌธํ•ญ 515๊ฐœ๋ฅผ ์ž์ฒด
44
+ ์ž‘์„ฑ(๋น„๊ณต๊ฐœ)ํ•˜๊ณ , ๋ผ๋ฒจ๋ง **์ด์ „์—** 10%(51๋ฌธํ•ญ)๋ฅผ ํ‰๊ฐ€ ์ „์šฉ์œผ๋กœ ๋ด‰์ธ โ‘ก ๊ต์‚ฌ(72B)๊ฐ€
45
+ 3,000๋ช… ร— 10๋ฌธํ•ญ = 30,000๊ฑด์„ ์—ญํ• ๊ทน์œผ๋กœ ๋ผ๋ฒจ๋ง โ‘ข ํ˜•์‹ยท์„ ํƒ์ง€ ์ผ์น˜ยทํ™•์‹ ๋„ ๋ฒ”์œ„ยท์ค‘๋ณต
46
+ ์ด์œ  ๊ธฐ์ค€์œผ๋กœ ํ•„ํ„ฐ(ํ†ต๊ณผ์œจ 95.2%) โ‘ฃ ํ‰๊ฐ€ ์ธ๋ฌผ 20๋ช…๊ณผ ํ‰๊ฐ€ ๋ฌธํ•ญ์€ ํ•™์Šต์—์„œ ์™„์ „ ๋ฐฐ์ œ,
47
+ ๊ฒ€์ฆ 572๊ฑด์€ ์ธ๋ฌผ ๋‹จ์œ„๋กœ ๋ถ„๋ฆฌ.
48
+
49
+ ## 3. Training ยท ํ•™์Šต
50
+
51
+ QLoRA SFT on Qwen2.5-7B-Instruct: NF4 4-bit frozen base + bf16 LoRA adapters
52
+ (r=16, ฮฑ=32, all 7 linear projections ร— 28 layers โ‰ˆ 40M trainable params, 0.5%).
53
+ 27,999 triplets, 2 epochs, effective batch 48, cosine lr 1e-4, loss on the response span
54
+ only. 3-GPU DDP on RTX 3090s, 2.5 h. Final val loss 0.391 with no overfitting signature.
55
+
56
+ **(ํ•œ๊ตญ์–ด)** Qwen2.5-7B๋ฅผ QLoRA๋กœ ๋ฏธ์„ธ์กฐ์ •: ๋ฒ ์ด์Šค๋Š” 4bit๋กœ ๋™๊ฒฐํ•˜๊ณ  LoRA ์–ด๋Œ‘ํ„ฐ
57
+ (์•ฝ 4์ฒœ๋งŒ ๊ฐœ, ์ „์ฒด์˜ 0.5%)๋งŒ bf16์œผ๋กœ ํ•™์Šต. ์‚ผ์ค‘ํ•ญ 27,999๊ฑด ร— 2์—ํญ, ์œ ํšจ ๋ฐฐ์น˜ 48,
58
+ ์†์‹ค์€ ์‘๋‹ต ๋ถ€๋ถ„์—๋งŒ ์ ์šฉ. RTX 3090 3์žฅ ๋ณ‘๋ ฌ๋กœ 2.5์‹œ๊ฐ„. ์ตœ์ข… ๊ฒ€์ฆ ์†์‹ค 0.391,
59
+ ๊ณผ์ ํ•ฉ ์ง•ํ›„ ์—†์Œ.
60
+
61
+ ## 4. Evaluation ยท ํ‰๊ฐ€
62
+
63
+ Protocol: 51 held-out questions ร— 20 unseen personas, temperature 0.7, fixed seeds.
64
+ All model-level metrics are reproducible from the released weights.
65
+
66
+ | Metric ยท ์ง€ํ‘œ | Base 7B | **Distilled 7B (์ฆ๋ฅ˜ ํ›„)** | Teacher 72B |
67
+ |---|---|---|---|
68
+ | Format compliance (strict) ยท ํ˜•์‹ ์ค€์ˆ˜์œจ | 95.5% | **98.5%** | 97.0% |
69
+ | Teacher agreement ยท ๊ต์‚ฌ ์ผ์น˜์œจ | 72.3% | **79.4%** | 96.1% (self-agreement ceiling) |
70
+ | Consistency (5-run majority) ยท ์‘๋‹ต ์ผ๊ด€์„ฑ | 98.7% | 97.6% | โ€” |
71
+ | Persona sensitivity (TV) ยท ํŽ˜๋ฅด์†Œ๋‚˜ ๊ฐ์ˆ˜์„ฑ | 0.356 | **0.356** | โ€” |
72
+ | Position bias ยท ์„ ํƒ์ง€ ์ˆœ์„œ ํŽธํ–ฅ | โ€” | **+1.7%** | โ€” |
73
+ | Age-conditioning (net) ยท ๋‚˜์ด ์กฐ๊ฑดํ™”(์ˆœํšจ๊ณผ) | โ€” | **+20.8%p** | โ€” |
74
+
75
+ Key readings:
76
+ - **Ceiling-normalized transfer**: the teacher agrees with *itself* only 96.1% of the time
77
+ under sampling, so the student's 79.4% represents **82.6% of the effective ceiling**
78
+ (75.2% before distillation).
79
+ - **Format**: the student surpasses its own teacher (98.5% vs 97.0%) โ€” expected, since the
80
+ student rehearsed the schema 28k times while the teacher only follows instructions.
81
+ - **Conditioning preserved**: distillation did not collapse persona sensitivity (TV distance
82
+ unchanged); flipping only the age field changes 29.4% of choices vs an 8.6% resampling
83
+ noise floor.
84
+ - **Position bias absent**: first-option preference is +1.7%, and choice consistency under
85
+ option-order flips (90.7%) sits at the sampling-noise floor (91.4%).
86
+
87
+ **(ํ•œ๊ตญ์–ด)** ํ‰๊ฐ€๋Š” ํ•™์Šต์— ์“ฐ์ง€ ์•Š์€ 51๋ฌธํ•ญ ร— ๋ฏธํ•™์Šต ์ธ๋ฌผ 20๋ช…์œผ๋กœ ์ง„ํ–‰ํ–ˆ๊ณ , ๊ณต๊ฐœ๋œ
88
+ ๊ฐ€์ค‘์น˜๋งŒ์œผ๋กœ ์žฌํ˜„ ๊ฐ€๋Šฅํ•˜๋‹ค. ํ•ต์‹ฌ ํ•ด์„: โ‘  ๊ต์‚ฌ๋„ ํ™•๋ฅ ์  ์ƒ์„ฑ์ด๋ผ ์ž๊ธฐ ์ž์‹ ๊ณผ 96.1%๋งŒ
89
+ ์ผ์น˜ํ•˜๋ฏ€๋กœ, ํ•™์ƒ์˜ 79.4%๋Š” **์‹ค์งˆ ์ƒํ•œ ๋Œ€๋น„ 82.6% ๋„๋‹ฌ**(์ฆ๋ฅ˜ ์ „ 75.2%) โ‘ก ํ˜•์‹
90
+ ์ค€์ˆ˜์œจ์€ ํ•™์ƒ์ด ๊ต์‚ฌ๋ฅผ ์ถ”์›” โ€” ํ˜•์‹์„ 2.8๋งŒ ๋ฒˆ ๋ฐ˜๋ณต ํ•™์Šตํ•œ ํšจ๊ณผ โ‘ข ์ฆ๋ฅ˜ ๊ณผ์ •์—์„œ
91
+ ํŽ˜๋ฅด์†Œ๋‚˜ ์กฐ๊ฑดํ™” ๋Šฅ๋ ฅ ์†์‹ค ์—†์Œ(TV๊ฑฐ๋ฆฌ ์œ ์ง€), ๋‚˜์ด๋งŒ ๋’ค์ง‘์–ด๋„ ์‘๋‹ต์˜ 29.4%๊ฐ€ ๋ณ€ํ™”
92
+ (๋…ธ์ด์ฆˆ 8.6% ๋Œ€๋น„) โ‘ฃ LLM ์„ค๋ฌธ๊ธฐ์˜ ํ”ํ•œ ๊ฒฐํ•จ์ธ ์„ ํƒ์ง€ ์ˆœ์„œ ํŽธํ–ฅ์ด ์‚ฌ์‹ค์ƒ ์—†์Œ(+1.7%).
93
+
94
+ ## 5. Scaling ablation ยท ์Šค์ผ€์ผ๋ง ๊ฒ€์ฆ โ€” why this configuration is the release
95
+
96
+ We tested both obvious scale-ups under identical hyperparameters and evaluation:
97
+
98
+ | Run | Data | Epochs | Val loss | Held-out teacher agreement |
99
+ |---|---|---|---|---|
100
+ | **A (released ยท ๊ณต๊ฐœ๋ณธ)** | 28k | 2 | 0.391 | **79.4%** |
101
+ | B | 28k | 3 | 0.388 | 80.1% (within sampling error) |
102
+ | C | 56k (6,000 personas) | 2 | **0.357** | 79.3% |
103
+
104
+ Doubling persona data improved validation loss (unseen-persona fit) but left unseen-question
105
+ agreement unchanged โ€” **persona diversity is not the bottleneck**. A third epoch moved
106
+ agreement only within noise. We conclude single-sample SFT distillation saturates near 80%
107
+ on this task; the remaining ceiling gap likely requires methodological change:
108
+ multi-sample distribution distillation per question, borderline-label filtering, or
109
+ preference optimization (DPO). These are left as future work, alongside system-level
110
+ calibration against real Korean survey ground truth (MAE), which we plan to report in a
111
+ future update.
112
+
113
+ **(ํ•œ๊ตญ์–ด)** "๋ฐ์ดํ„ฐ๋ฅผ ๋” ์“ฐ๋ฉด ๋˜์ง€ ์•Š๋‚˜?"๋ฅผ ๋™์ผ ์กฐ๊ฑด์—์„œ ์ง์ ‘ ๊ฒ€์ฆํ–ˆ๋‹ค. ์ธ๋ฌผ ๋ฐ์ดํ„ฐ๋ฅผ
114
+ 2๋ฐฐ(56k)๋กœ ๋Š˜๋ฆฌ์ž ๊ฒ€์ฆ ์†์‹ค์€ ๊ฐœ์„ ๋์ง€๋งŒ(์ƒˆ ์ธ๋ฌผ ์ ์‘๋ ฅ) ๋ฏธํ•™์Šต ๋ฌธํ•ญ ์ผ์น˜์œจ์€ ๊ทธ๋Œ€๋กœ
115
+ โ€” **๋ณ‘๋ชฉ์€ ์ธ๋ฌผ ๋‹ค์–‘์„ฑ์ด ์•„๋‹ˆ๋‹ค**. 3์—ํญ๋„ ํ‘œ๋ณธ ์˜ค์ฐจ ๋‚ด ๋ณ€ํ™”์— ๊ทธ์ณค๋‹ค. ๊ฒฐ๋ก : ๋‹จ์ผ ์ƒ˜ํ”Œ
116
+ SFT ์ฆ๋ฅ˜๋Š” ์ด ๊ณผ์ œ์—์„œ ์•ฝ 80% ๋ถ€๊ทผ์—์„œ ํฌํ™”ํ•˜๋ฉฐ, ๋‚จ์€ ์ƒํ•œ ๊ฐ„๊ทน์€ ๋ฐ์ดํ„ฐ ์ฆ๋Ÿ‰์ด ์•„๋‹ˆ๋ผ
117
+ ๋ฐฉ๋ฒ• ๋ณ€๊ฒฝ(๋ฌธํ•ญ๋‹น ๊ต์‚ฌ ๋‹ค์ค‘ ์ƒ˜ํ”Œ๋ง, ๊ฒฝ๊ณ„ ๋ผ๋ฒจ ํ•„ํ„ฐ๋ง, DPO)์˜ ์˜์—ญ์ด๋‹ค. ์‹ค์ œ ์—ฌ๋ก ์กฐ์‚ฌ
118
+ ์‹ค์ธก ๋Œ€๋น„ ์‹œ์Šคํ…œ ์ •ํ™•๋„(MAE)์™€ ํ•จ๊ป˜ ํ–ฅํ›„ ์—…๋ฐ์ดํŠธ ๊ณผ์ œ๋กœ ๋‚จ๊ธด๋‹ค.
119
+
120
+ ## 6. Limitations ยท ํ•œ๊ณ„
121
+
122
+ 1. Model-level metrics measure *fidelity to the teacher*, not real-world accuracy;
123
+ population-level accuracy is unmeasured in this release.
124
+ 2. Personas are synthetic; the source data's joint distributions (ageร—occupationร—region)
125
+ deviate from Korean reality (arXiv:2606.12433) โ€” segment-level conclusions require
126
+ statistical correction that is out of scope for the released model.
127
+ 3. The teacher may role-play stereotypically and distillation inherits this; our sensitivity
128
+ metrics quantify conditioning, not fairness.
129
+ 4. Non-sensitive preference/lifestyle domains only; behavior elsewhere is untested.
130
+
131
+ **(ํ•œ๊ตญ์–ด)** โ‘  ๋ณธ ์ง€ํ‘œ๋Š” "๊ต์‚ฌ๋ฅผ ์–ผ๋งˆ๋‚˜ ๋‹ฎ์•˜๋Š”๊ฐ€"์ด์ง€ ์‹ค์„ธ๊ณ„ ์ •ํ™•๋„๊ฐ€ ์•„๋‹ˆ๋‹ค โ€” ๋ชจ์ง‘๋‹จ
132
+ ์ˆ˜์ค€ ์ •ํ™•๋„๋Š” ์ด๋ฒˆ ๊ณต๊ฐœ์—์„œ ๋ฏธ์ธก์ •. โ‘ก ํŽ˜๋ฅด์†Œ๋‚˜๋Š” ํ•ฉ์„ฑ ์ธ๋ฌผ์ด๋ฉฐ ์›๋ณธ ๋ฐ์ดํ„ฐ์˜ ๊ฒฐํ•ฉ๋ถ„ํฌ
133
+ (๋‚˜์ดร—์ง์—…ร—์ง€์—ญ)๊ฐ€ ์‹ค์ œ์™€ ์–ด๊ธ‹๋‚˜๋ฏ€๋กœ, ์ง‘๋‹จ๋ณ„ ๊ฒฐ๋ก ์—๋Š” ๋ณ„๋„ ํ†ต๊ณ„ ๋ณด์ •์ด ํ•„์š”ํ•˜๋‹ค.
134
+ โ‘ข ๊ต์‚ฌ์˜ ๊ณ ์ •๊ด€๋…์  ์—ญํ• ๊ทน์ด ํ•™์ƒ์— ์Šน๊ณ„๋  ์ˆ˜ ์žˆ๋‹ค โ€” ๊ฐ์ˆ˜์„ฑ ์ง€ํ‘œ๋Š” ์กฐ๊ฑดํ™” ์ž‘๋™ ์—ฌ๋ถ€๋ฅผ
135
+ ์žด ๋ฟ ๊ณต์ •์„ฑ์„ ์žฌ์ง€ ์•Š๋Š”๋‹ค. โ‘ฃ ๋น„๋ฏผ๊ฐ ์ทจํ–ฅยท๋ผ์ดํ”„์Šคํƒ€์ผ ์ฃผ์ œ๋กœ๋งŒ ํ•™์Šตยทํ‰๊ฐ€ํ–ˆ๋‹ค.
136
+
137
+ ## Attribution ยท ์ถœ์ฒ˜
138
+
139
+ Personas: nvidia/Nemotron-Personas-Korea (CC BY 4.0, ยฉ NVIDIA). Base & teacher:
140
+ Qwen2.5 family (Apache-2.0, Alibaba Cloud). Model weights: Apache-2.0.