victory888's picture
Upload REPORT.md with huggingface_hub
982ccac verified
|
Raw History Blame Contribute Delete
9.29 kB

Technical Report โ€” kor-persona-survey-7b

Distilling persona-conditioned Korean survey responses from a 72B teacher into a 7B student 72B ๊ต์‚ฌ์˜ ํ•œ๊ตญ์–ด ํŽ˜๋ฅด์†Œ๋‚˜ ์กฐ๊ฑด๋ถ€ ์„ค๋ฌธ ์‘๋‹ต ๋Šฅ๋ ฅ์„ 7B ํ•™์ƒ์— ์ฆ๋ฅ˜ํ•˜๊ธฐ

DOB Studio ยท 2026-08

1. Motivation ยท ๋™๊ธฐ

Simulating survey responses with LLM role-play requires one inference per (persona, question) pair โ€” population-scale simulation means thousands of inferences per question. Large models judge personas well but are too slow for this workload; small models are fast but role-play shallowly and break output formats. We ask:

RQ. How much of a 72B teacher's persona-conditioned survey-response capability can be transferred to a 7B student through synthetic-triplet distillation alone?

(ํ•œ๊ตญ์–ด) LLM ์—ญํ• ๊ทน์œผ๋กœ ์„ค๋ฌธ ์‘๋‹ต์„ ์‹œ๋ฎฌ๋ ˆ์ด์…˜ํ•˜๋ ค๋ฉด (ํŽ˜๋ฅด์†Œ๋‚˜, ๋ฌธํ•ญ) ์Œ๋งˆ๋‹ค ์ถ”๋ก ์ด 1ํšŒ์”ฉ ํ•„์š”ํ•˜๋‹ค โ€” ๋ชจ์ง‘๋‹จ ๊ทœ๋ชจ ์‹œ๋ฎฌ๋ ˆ์ด์…˜์€ ๋ฌธํ•ญ๋‹น ์ˆ˜์ฒœ ํšŒ ์ถ”๋ก ์„ ๋œปํ•œ๋‹ค. ๋Œ€ํ˜• ๋ชจ๋ธ์€ ํŽ˜๋ฅด์†Œ๋‚˜ ํŒ๋‹จ์€ ๊นŠ์ง€๋งŒ ์ด ์ž‘์—…๋Ÿ‰์—๋Š” ๋„ˆ๋ฌด ๋А๋ฆฌ๊ณ , ์†Œํ˜• ๋ชจ๋ธ์€ ๋น ๋ฅด์ง€๋งŒ ์—ญํ• ๊ทน์ด ์–•๊ณ  ์ถœ๋ ฅ ํ˜•์‹์„ ์–ด๊ธด๋‹ค. ๊ทธ๋ž˜์„œ ๋ฌป๋Š”๋‹ค: 72B ๊ต์‚ฌ์˜ ํŽ˜๋ฅด์†Œ๋‚˜ ์กฐ๊ฑด๋ถ€ ์„ค๋ฌธ ์‘๋‹ต ๋Šฅ๋ ฅ์„, ํ•ฉ์„ฑ ์‚ผ์ค‘ํ•ญ ์ฆ๋ฅ˜๋งŒ์œผ๋กœ 7B ํ•™์ƒ์— ์–ผ๋งˆ๋‚˜ ์ด์‹ํ•  ์ˆ˜ ์žˆ๋Š”๊ฐ€?

2. Data construction ยท ๋ฐ์ดํ„ฐ ๊ตฌ์ถ•

The source dataset (nvidia/Nemotron-Personas-Korea, 1M synthetic Korean personas, CC BY 4.0) contains persona profiles only โ€” no questions, no answers. We manufactured the missing supervision:

  1. Question bank (private): 515 binary preference questions across 8 non-sensitive topics (food, lifestyle, consumption, travel, leisure, work, values, digital), authored to span age/region/occupation-sensitive contrasts. A stratified 10% (51 questions) was frozen as a held-out evaluation set before any labeling.
  2. Teacher labeling: Qwen2.5-72B-Instruct-AWQ answered (persona, question) pairs in character โ€” 3,000 personas ร— 10 questions = 30,000 raw labels of the form {choice, confidence, reason}.
  3. Quality filtering (95.2% pass): exact option match, valid JSON, confidence โˆˆ [0,1], reason length bounds, per-question duplicate-reason cap.
  4. Split hygiene: evaluation personas (n=20) and held-out questions were excluded from all training data; the validation split (572 triplets) is persona-disjoint from training.

(ํ•œ๊ตญ์–ด) ์›๋ณธ ๋ฐ์ดํ„ฐ์…‹(๊ฐ€์ƒ ํ•œ๊ตญ์ธ 100๋งŒ ๋ช…)์—๋Š” ์ธ๋ฌผ ํ”„๋กœํ•„๋งŒ ์žˆ๊ณ  ์งˆ๋ฌธยท๋Œ€๋‹ต์ด ์—†๋‹ค. ๊ทธ๋ž˜์„œ ํ•™์Šต ์‹ ํ˜ธ๋ฅผ ์ง์ ‘ ์ œ์กฐํ–ˆ๋‹ค: โ‘  ๋น„๋ฏผ๊ฐ 8๊ฐœ ์ฃผ์ œ์˜ ์–‘์žํƒ์ผ ๋ฌธํ•ญ 515๊ฐœ๋ฅผ ์ž์ฒด ์ž‘์„ฑ(๋น„๊ณต๊ฐœ)ํ•˜๊ณ , ๋ผ๋ฒจ๋ง ์ด์ „์— 10%(51๋ฌธํ•ญ)๋ฅผ ํ‰๊ฐ€ ์ „์šฉ์œผ๋กœ ๋ด‰์ธ โ‘ก ๊ต์‚ฌ(72B)๊ฐ€ 3,000๋ช… ร— 10๋ฌธํ•ญ = 30,000๊ฑด์„ ์—ญํ• ๊ทน์œผ๋กœ ๋ผ๋ฒจ๋ง โ‘ข ํ˜•์‹ยท์„ ํƒ์ง€ ์ผ์น˜ยทํ™•์‹ ๋„ ๋ฒ”์œ„ยท์ค‘๋ณต ์ด์œ  ๊ธฐ์ค€์œผ๋กœ ํ•„ํ„ฐ(ํ†ต๊ณผ์œจ 95.2%) โ‘ฃ ํ‰๊ฐ€ ์ธ๋ฌผ 20๋ช…๊ณผ ํ‰๊ฐ€ ๋ฌธํ•ญ์€ ํ•™์Šต์—์„œ ์™„์ „ ๋ฐฐ์ œ, ๊ฒ€์ฆ 572๊ฑด์€ ์ธ๋ฌผ ๋‹จ์œ„๋กœ ๋ถ„๋ฆฌ.

3. Training ยท ํ•™์Šต

QLoRA SFT on Qwen2.5-7B-Instruct: NF4 4-bit frozen base + bf16 LoRA adapters (r=16, ฮฑ=32, all 7 linear projections ร— 28 layers โ‰ˆ 40M trainable params, 0.5%). 27,999 triplets, 2 epochs, effective batch 48, cosine lr 1e-4, loss on the response span only. Trained on a 4ร— RTX 3090 (24GB) workstation, ~4 h. Final val loss 0.391 with no overfitting signature.

(ํ•œ๊ตญ์–ด) Qwen2.5-7B๋ฅผ QLoRA๋กœ ๋ฏธ์„ธ์กฐ์ •: ๋ฒ ์ด์Šค๋Š” 4bit๋กœ ๋™๊ฒฐํ•˜๊ณ  LoRA ์–ด๋Œ‘ํ„ฐ (์•ฝ 4์ฒœ๋งŒ ๊ฐœ, ์ „์ฒด์˜ 0.5%)๋งŒ bf16์œผ๋กœ ํ•™์Šต. ์‚ผ์ค‘ํ•ญ 27,999๊ฑด ร— 2์—ํญ, ์œ ํšจ ๋ฐฐ์น˜ 48, ์†์‹ค์€ ์‘๋‹ต ๋ถ€๋ถ„์—๋งŒ ์ ์šฉ. RTX 3090 4์žฅ ์›Œํฌ์Šคํ…Œ์ด์…˜์—์„œ ์•ฝ 4์‹œ๊ฐ„. ์ตœ์ข… ๊ฒ€์ฆ ์†์‹ค 0.391, ๊ณผ์ ํ•ฉ ์ง•ํ›„ ์—†์Œ.

4. Evaluation ยท ํ‰๊ฐ€

Protocol: 51 held-out questions ร— 20 unseen personas, temperature 0.7, fixed seeds. All model-level metrics are reproducible from the released weights.

Metric ยท ์ง€ํ‘œ Base 7B Distilled 7B (์ฆ๋ฅ˜ ํ›„) Teacher 72B
Format compliance (strict) ยท ํ˜•์‹ ์ค€์ˆ˜์œจ 95.5% 98.5% 97.0%
Teacher agreement ยท ๊ต์‚ฌ ์ผ์น˜์œจ 72.3% 79.4% 96.1% (self-agreement ceiling)
Consistency (5-run majority) ยท ์‘๋‹ต ์ผ๊ด€์„ฑ 98.7% 97.6% โ€”
Persona sensitivity (TV) ยท ํŽ˜๋ฅด์†Œ๋‚˜ ๊ฐ์ˆ˜์„ฑ 0.356 0.356 โ€”
Position bias ยท ์„ ํƒ์ง€ ์ˆœ์„œ ํŽธํ–ฅ โ€” +1.7% โ€”
Age-conditioning (net) ยท ๋‚˜์ด ์กฐ๊ฑดํ™”(์ˆœํšจ๊ณผ) โ€” +20.8%p โ€”

Key readings:

  • Ceiling-normalized transfer: the teacher agrees with itself only 96.1% of the time under sampling, so the student's 79.4% represents 82.6% of the effective ceiling (75.2% before distillation).
  • Format: the student surpasses its own teacher (98.5% vs 97.0%) โ€” expected, since the student rehearsed the schema 28k times while the teacher only follows instructions.
  • Conditioning preserved: distillation did not collapse persona sensitivity (TV distance unchanged); flipping only the age field changes 29.4% of choices vs an 8.6% resampling noise floor.
  • Position bias absent: first-option preference is +1.7%, and choice consistency under option-order flips (90.7%) sits at the sampling-noise floor (91.4%).

(ํ•œ๊ตญ์–ด) ํ‰๊ฐ€๋Š” ํ•™์Šต์— ์“ฐ์ง€ ์•Š์€ 51๋ฌธํ•ญ ร— ๋ฏธํ•™์Šต ์ธ๋ฌผ 20๋ช…์œผ๋กœ ์ง„ํ–‰ํ–ˆ๊ณ , ๊ณต๊ฐœ๋œ ๊ฐ€์ค‘์น˜๋งŒ์œผ๋กœ ์žฌํ˜„ ๊ฐ€๋Šฅํ•˜๋‹ค. ํ•ต์‹ฌ ํ•ด์„: โ‘  ๊ต์‚ฌ๋„ ํ™•๋ฅ ์  ์ƒ์„ฑ์ด๋ผ ์ž๊ธฐ ์ž์‹ ๊ณผ 96.1%๋งŒ ์ผ์น˜ํ•˜๋ฏ€๋กœ, ํ•™์ƒ์˜ 79.4%๋Š” ์‹ค์งˆ ์ƒํ•œ ๋Œ€๋น„ 82.6% ๋„๋‹ฌ(์ฆ๋ฅ˜ ์ „ 75.2%) โ‘ก ํ˜•์‹ ์ค€์ˆ˜์œจ์€ ํ•™์ƒ์ด ๊ต์‚ฌ๋ฅผ ์ถ”์›” โ€” ํ˜•์‹์„ 2.8๋งŒ ๋ฒˆ ๋ฐ˜๋ณต ํ•™์Šตํ•œ ํšจ๊ณผ โ‘ข ์ฆ๋ฅ˜ ๊ณผ์ •์—์„œ ํŽ˜๋ฅด์†Œ๋‚˜ ์กฐ๊ฑดํ™” ๋Šฅ๋ ฅ ์†์‹ค ์—†์Œ(TV๊ฑฐ๋ฆฌ ์œ ์ง€), ๋‚˜์ด๋งŒ ๋’ค์ง‘์–ด๋„ ์‘๋‹ต์˜ 29.4%๊ฐ€ ๋ณ€ํ™” (๋…ธ์ด์ฆˆ 8.6% ๋Œ€๋น„) โ‘ฃ LLM ์„ค๋ฌธ๊ธฐ์˜ ํ”ํ•œ ๊ฒฐํ•จ์ธ ์„ ํƒ์ง€ ์ˆœ์„œ ํŽธํ–ฅ์ด ์‚ฌ์‹ค์ƒ ์—†์Œ(+1.7%).

5. Scaling ablation ยท ์Šค์ผ€์ผ๋ง ๊ฒ€์ฆ โ€” why this configuration is the release

We tested both obvious scale-ups under identical hyperparameters and evaluation:

Run Data Epochs Val loss Held-out teacher agreement
A (released ยท ๊ณต๊ฐœ๋ณธ) 28k 2 0.391 79.4%
B 28k 3 0.388 80.1% (within sampling error)
C 56k (6,000 personas) 2 0.357 79.3%

Doubling persona data improved validation loss (unseen-persona fit) but left unseen-question agreement unchanged โ€” persona diversity is not the bottleneck. A third epoch moved agreement only within noise. We conclude single-sample SFT distillation saturates near 80% on this task; the remaining ceiling gap likely requires methodological change: multi-sample distribution distillation per question, borderline-label filtering, or preference optimization (DPO). These directions, along with system-level calibration against real survey ground truth, are outside the scope of this release.

(ํ•œ๊ตญ์–ด) "๋ฐ์ดํ„ฐ๋ฅผ ๋” ์“ฐ๋ฉด ๋˜์ง€ ์•Š๋‚˜?"๋ฅผ ๋™์ผ ์กฐ๊ฑด์—์„œ ์ง์ ‘ ๊ฒ€์ฆํ–ˆ๋‹ค. ์ธ๋ฌผ ๋ฐ์ดํ„ฐ๋ฅผ 2๋ฐฐ(56k)๋กœ ๋Š˜๋ฆฌ์ž ๊ฒ€์ฆ ์†์‹ค์€ ๊ฐœ์„ ๋์ง€๋งŒ(์ƒˆ ์ธ๋ฌผ ์ ์‘๋ ฅ) ๋ฏธํ•™์Šต ๋ฌธํ•ญ ์ผ์น˜์œจ์€ ๊ทธ๋Œ€๋กœ โ€” ๋ณ‘๋ชฉ์€ ์ธ๋ฌผ ๋‹ค์–‘์„ฑ์ด ์•„๋‹ˆ๋‹ค. 3์—ํญ๋„ ํ‘œ๋ณธ ์˜ค์ฐจ ๋‚ด ๋ณ€ํ™”์— ๊ทธ์ณค๋‹ค. ๊ฒฐ๋ก : ๋‹จ์ผ ์ƒ˜ํ”Œ SFT ์ฆ๋ฅ˜๋Š” ์ด ๊ณผ์ œ์—์„œ ์•ฝ 80% ๋ถ€๊ทผ์—์„œ ํฌํ™”ํ•˜๋ฉฐ, ๋‚จ์€ ์ƒํ•œ ๊ฐ„๊ทน์€ ๋ฐ์ดํ„ฐ ์ฆ๋Ÿ‰์ด ์•„๋‹ˆ๋ผ ๋ฐฉ๋ฒ• ๋ณ€๊ฒฝ(๋ฌธํ•ญ๋‹น ๊ต์‚ฌ ๋‹ค์ค‘ ์ƒ˜ํ”Œ๋ง, ๊ฒฝ๊ณ„ ๋ผ๋ฒจ ํ•„ํ„ฐ๋ง, DPO)์˜ ์˜์—ญ์ด๋‹ค. ์ด๋Ÿฌํ•œ ๋ฐฉํ–ฅ๊ณผ ์‹ค์ œ ์—ฌ๋ก ์กฐ์‚ฌ ์‹ค์ธก ๋Œ€๋น„ ์‹œ์Šคํ…œ ์ •ํ™•๋„(MAE) ์ธก์ •์€ ์ด๋ฒˆ ๊ณต๊ฐœ์˜ ๋ฒ”์œ„ ๋ฐ–์ด๋‹ค.

6. Limitations ยท ํ•œ๊ณ„

  1. Model-level metrics measure fidelity to the teacher, not real-world accuracy; population-level accuracy is unmeasured in this release.
  2. Personas are synthetic; the source data's joint distributions (ageร—occupationร—region) deviate from Korean reality (arXiv:2606.12433) โ€” segment-level conclusions require statistical correction that is out of scope for the released model.
  3. The teacher may role-play stereotypically and distillation inherits this; our sensitivity metrics quantify conditioning, not fairness.
  4. Non-sensitive preference/lifestyle domains only; behavior elsewhere is untested.

(ํ•œ๊ตญ์–ด) โ‘  ๋ณธ ์ง€ํ‘œ๋Š” "๊ต์‚ฌ๋ฅผ ์–ผ๋งˆ๋‚˜ ๋‹ฎ์•˜๋Š”๊ฐ€"์ด์ง€ ์‹ค์„ธ๊ณ„ ์ •ํ™•๋„๊ฐ€ ์•„๋‹ˆ๋‹ค โ€” ๋ชจ์ง‘๋‹จ ์ˆ˜์ค€ ์ •ํ™•๋„๋Š” ์ด๋ฒˆ ๊ณต๊ฐœ์—์„œ ๋ฏธ์ธก์ •. โ‘ก ํŽ˜๋ฅด์†Œ๋‚˜๋Š” ํ•ฉ์„ฑ ์ธ๋ฌผ์ด๋ฉฐ ์›๋ณธ ๋ฐ์ดํ„ฐ์˜ ๊ฒฐํ•ฉ๋ถ„ํฌ (๋‚˜์ดร—์ง์—…ร—์ง€์—ญ)๊ฐ€ ์‹ค์ œ์™€ ์–ด๊ธ‹๋‚˜๋ฏ€๋กœ, ์ง‘๋‹จ๋ณ„ ๊ฒฐ๋ก ์—๋Š” ๋ณ„๋„ ํ†ต๊ณ„ ๋ณด์ •์ด ํ•„์š”ํ•˜๋‹ค. โ‘ข ๊ต์‚ฌ์˜ ๊ณ ์ •๊ด€๋…์  ์—ญํ• ๊ทน์ด ํ•™์ƒ์— ์Šน๊ณ„๋  ์ˆ˜ ์žˆ๋‹ค โ€” ๊ฐ์ˆ˜์„ฑ ์ง€ํ‘œ๋Š” ์กฐ๊ฑดํ™” ์ž‘๋™ ์—ฌ๋ถ€๋ฅผ ์žด ๋ฟ ๊ณต์ •์„ฑ์„ ์žฌ์ง€ ์•Š๋Š”๋‹ค. โ‘ฃ ๋น„๋ฏผ๊ฐ ์ทจํ–ฅยท๋ผ์ดํ”„์Šคํƒ€์ผ ์ฃผ์ œ๋กœ๋งŒ ํ•™์Šตยทํ‰๊ฐ€ํ–ˆ๋‹ค.

Attribution ยท ์ถœ์ฒ˜

Built with Qwen. Personas: nvidia/Nemotron-Personas-Korea (CC BY 4.0, ยฉ NVIDIA). Base model: Qwen2.5-7B-Instruct (Apache-2.0). Teacher: Qwen2.5-72B-Instruct-AWQ (Qwen LICENSE โ€” outputs used for training, attributed per license). Model weights: Apache-2.0.