File size: 9,289 Bytes
83913d7 71616cd 83913d7 71616cd 83913d7 a546b8e 83913d7 a546b8e 83913d7 982ccac | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 | # Technical Report โ kor-persona-survey-7b
*Distilling persona-conditioned Korean survey responses from a 72B teacher into a 7B student*
*72B ๊ต์ฌ์ ํ๊ตญ์ด ํ๋ฅด์๋ ์กฐ๊ฑด๋ถ ์ค๋ฌธ ์๋ต ๋ฅ๋ ฅ์ 7B ํ์์ ์ฆ๋ฅํ๊ธฐ*
DOB Studio ยท 2026-08
## 1. Motivation ยท ๋๊ธฐ
Simulating survey responses with LLM role-play requires one inference per (persona,
question) pair โ population-scale simulation means thousands of inferences per question.
Large models judge personas well but are too slow for this workload; small models are fast
but role-play shallowly and break output formats. We ask:
> **RQ. How much of a 72B teacher's persona-conditioned survey-response capability can be
> transferred to a 7B student through synthetic-triplet distillation alone?**
**(ํ๊ตญ์ด)** LLM ์ญํ ๊ทน์ผ๋ก ์ค๋ฌธ ์๋ต์ ์๋ฎฌ๋ ์ด์
ํ๋ ค๋ฉด (ํ๋ฅด์๋, ๋ฌธํญ) ์๋ง๋ค ์ถ๋ก ์ด
1ํ์ฉ ํ์ํ๋ค โ ๋ชจ์ง๋จ ๊ท๋ชจ ์๋ฎฌ๋ ์ด์
์ ๋ฌธํญ๋น ์์ฒ ํ ์ถ๋ก ์ ๋ปํ๋ค. ๋ํ ๋ชจ๋ธ์
ํ๋ฅด์๋ ํ๋จ์ ๊น์ง๋ง ์ด ์์
๋์๋ ๋๋ฌด ๋๋ฆฌ๊ณ , ์ํ ๋ชจ๋ธ์ ๋น ๋ฅด์ง๋ง ์ญํ ๊ทน์ด ์๊ณ
์ถ๋ ฅ ํ์์ ์ด๊ธด๋ค. ๊ทธ๋์ ๋ฌป๋๋ค: **72B ๊ต์ฌ์ ํ๋ฅด์๋ ์กฐ๊ฑด๋ถ ์ค๋ฌธ ์๋ต ๋ฅ๋ ฅ์, ํฉ์ฑ
์ผ์คํญ ์ฆ๋ฅ๋ง์ผ๋ก 7B ํ์์ ์ผ๋ง๋ ์ด์ํ ์ ์๋๊ฐ?**
## 2. Data construction ยท ๋ฐ์ดํฐ ๊ตฌ์ถ
The source dataset (nvidia/Nemotron-Personas-Korea, 1M synthetic Korean personas, CC BY 4.0)
contains persona profiles only โ no questions, no answers. We manufactured the missing
supervision:
1. **Question bank** (private): 515 binary preference questions across 8 non-sensitive
topics (food, lifestyle, consumption, travel, leisure, work, values, digital), authored
to span age/region/occupation-sensitive contrasts. A stratified 10% (51 questions) was
frozen as a held-out evaluation set **before any labeling**.
2. **Teacher labeling**: Qwen2.5-72B-Instruct-AWQ answered (persona, question) pairs in
character โ 3,000 personas ร 10 questions = 30,000 raw labels of the form
`{choice, confidence, reason}`.
3. **Quality filtering** (95.2% pass): exact option match, valid JSON, confidence โ [0,1],
reason length bounds, per-question duplicate-reason cap.
4. **Split hygiene**: evaluation personas (n=20) and held-out questions were excluded from
all training data; the validation split (572 triplets) is persona-disjoint from training.
**(ํ๊ตญ์ด)** ์๋ณธ ๋ฐ์ดํฐ์
(๊ฐ์ ํ๊ตญ์ธ 100๋ง ๋ช
)์๋ ์ธ๋ฌผ ํ๋กํ๋ง ์๊ณ ์ง๋ฌธยท๋๋ต์ด
์๋ค. ๊ทธ๋์ ํ์ต ์ ํธ๋ฅผ ์ง์ ์ ์กฐํ๋ค: โ ๋น๋ฏผ๊ฐ 8๊ฐ ์ฃผ์ ์ ์์ํ์ผ ๋ฌธํญ 515๊ฐ๋ฅผ ์์ฒด
์์ฑ(๋น๊ณต๊ฐ)ํ๊ณ , ๋ผ๋ฒจ๋ง **์ด์ ์** 10%(51๋ฌธํญ)๋ฅผ ํ๊ฐ ์ ์ฉ์ผ๋ก ๋ด์ธ โก ๊ต์ฌ(72B)๊ฐ
3,000๋ช
ร 10๋ฌธํญ = 30,000๊ฑด์ ์ญํ ๊ทน์ผ๋ก ๋ผ๋ฒจ๋ง โข ํ์ยท์ ํ์ง ์ผ์นยทํ์ ๋ ๋ฒ์ยท์ค๋ณต
์ด์ ๊ธฐ์ค์ผ๋ก ํํฐ(ํต๊ณผ์จ 95.2%) โฃ ํ๊ฐ ์ธ๋ฌผ 20๋ช
๊ณผ ํ๊ฐ ๋ฌธํญ์ ํ์ต์์ ์์ ๋ฐฐ์ ,
๊ฒ์ฆ 572๊ฑด์ ์ธ๋ฌผ ๋จ์๋ก ๋ถ๋ฆฌ.
## 3. Training ยท ํ์ต
QLoRA SFT on Qwen2.5-7B-Instruct: NF4 4-bit frozen base + bf16 LoRA adapters
(r=16, ฮฑ=32, all 7 linear projections ร 28 layers โ 40M trainable params, 0.5%).
27,999 triplets, 2 epochs, effective batch 48, cosine lr 1e-4, loss on the response span
only. Trained on a 4ร RTX 3090 (24GB) workstation, ~4 h. Final val loss 0.391 with no
overfitting signature.
**(ํ๊ตญ์ด)** Qwen2.5-7B๋ฅผ QLoRA๋ก ๋ฏธ์ธ์กฐ์ : ๋ฒ ์ด์ค๋ 4bit๋ก ๋๊ฒฐํ๊ณ LoRA ์ด๋ํฐ
(์ฝ 4์ฒ๋ง ๊ฐ, ์ ์ฒด์ 0.5%)๋ง bf16์ผ๋ก ํ์ต. ์ผ์คํญ 27,999๊ฑด ร 2์ํญ, ์ ํจ ๋ฐฐ์น 48,
์์ค์ ์๋ต ๋ถ๋ถ์๋ง ์ ์ฉ. RTX 3090 4์ฅ ์ํฌ์คํ
์ด์
์์ ์ฝ 4์๊ฐ. ์ต์ข
๊ฒ์ฆ ์์ค
0.391, ๊ณผ์ ํฉ ์งํ ์์.
## 4. Evaluation ยท ํ๊ฐ
Protocol: 51 held-out questions ร 20 unseen personas, temperature 0.7, fixed seeds.
All model-level metrics are reproducible from the released weights.
| Metric ยท ์งํ | Base 7B | **Distilled 7B (์ฆ๋ฅ ํ)** | Teacher 72B |
|---|---|---|---|
| Format compliance (strict) ยท ํ์ ์ค์์จ | 95.5% | **98.5%** | 97.0% |
| Teacher agreement ยท ๊ต์ฌ ์ผ์น์จ | 72.3% | **79.4%** | 96.1% (self-agreement ceiling) |
| Consistency (5-run majority) ยท ์๋ต ์ผ๊ด์ฑ | 98.7% | 97.6% | โ |
| Persona sensitivity (TV) ยท ํ๋ฅด์๋ ๊ฐ์์ฑ | 0.356 | **0.356** | โ |
| Position bias ยท ์ ํ์ง ์์ ํธํฅ | โ | **+1.7%** | โ |
| Age-conditioning (net) ยท ๋์ด ์กฐ๊ฑดํ(์ํจ๊ณผ) | โ | **+20.8%p** | โ |
Key readings:
- **Ceiling-normalized transfer**: the teacher agrees with *itself* only 96.1% of the time
under sampling, so the student's 79.4% represents **82.6% of the effective ceiling**
(75.2% before distillation).
- **Format**: the student surpasses its own teacher (98.5% vs 97.0%) โ expected, since the
student rehearsed the schema 28k times while the teacher only follows instructions.
- **Conditioning preserved**: distillation did not collapse persona sensitivity (TV distance
unchanged); flipping only the age field changes 29.4% of choices vs an 8.6% resampling
noise floor.
- **Position bias absent**: first-option preference is +1.7%, and choice consistency under
option-order flips (90.7%) sits at the sampling-noise floor (91.4%).
**(ํ๊ตญ์ด)** ํ๊ฐ๋ ํ์ต์ ์ฐ์ง ์์ 51๋ฌธํญ ร ๋ฏธํ์ต ์ธ๋ฌผ 20๋ช
์ผ๋ก ์งํํ๊ณ , ๊ณต๊ฐ๋
๊ฐ์ค์น๋ง์ผ๋ก ์ฌํ ๊ฐ๋ฅํ๋ค. ํต์ฌ ํด์: โ ๊ต์ฌ๋ ํ๋ฅ ์ ์์ฑ์ด๋ผ ์๊ธฐ ์์ ๊ณผ 96.1%๋ง
์ผ์นํ๋ฏ๋ก, ํ์์ 79.4%๋ **์ค์ง ์ํ ๋๋น 82.6% ๋๋ฌ**(์ฆ๋ฅ ์ 75.2%) โก ํ์
์ค์์จ์ ํ์์ด ๊ต์ฌ๋ฅผ ์ถ์ โ ํ์์ 2.8๋ง ๋ฒ ๋ฐ๋ณต ํ์ตํ ํจ๊ณผ โข ์ฆ๋ฅ ๊ณผ์ ์์
ํ๋ฅด์๋ ์กฐ๊ฑดํ ๋ฅ๋ ฅ ์์ค ์์(TV๊ฑฐ๋ฆฌ ์ ์ง), ๋์ด๋ง ๋ค์ง์ด๋ ์๋ต์ 29.4%๊ฐ ๋ณํ
(๋
ธ์ด์ฆ 8.6% ๋๋น) โฃ LLM ์ค๋ฌธ๊ธฐ์ ํํ ๊ฒฐํจ์ธ ์ ํ์ง ์์ ํธํฅ์ด ์ฌ์ค์ ์์(+1.7%).
## 5. Scaling ablation ยท ์ค์ผ์ผ๋ง ๊ฒ์ฆ โ why this configuration is the release
We tested both obvious scale-ups under identical hyperparameters and evaluation:
| Run | Data | Epochs | Val loss | Held-out teacher agreement |
|---|---|---|---|---|
| **A (released ยท ๊ณต๊ฐ๋ณธ)** | 28k | 2 | 0.391 | **79.4%** |
| B | 28k | 3 | 0.388 | 80.1% (within sampling error) |
| C | 56k (6,000 personas) | 2 | **0.357** | 79.3% |
Doubling persona data improved validation loss (unseen-persona fit) but left unseen-question
agreement unchanged โ **persona diversity is not the bottleneck**. A third epoch moved
agreement only within noise. We conclude single-sample SFT distillation saturates near 80%
on this task; the remaining ceiling gap likely requires methodological change:
multi-sample distribution distillation per question, borderline-label filtering, or
preference optimization (DPO). These directions, along with system-level calibration
against real survey ground truth, are outside the scope of this release.
**(ํ๊ตญ์ด)** "๋ฐ์ดํฐ๋ฅผ ๋ ์ฐ๋ฉด ๋์ง ์๋?"๋ฅผ ๋์ผ ์กฐ๊ฑด์์ ์ง์ ๊ฒ์ฆํ๋ค. ์ธ๋ฌผ ๋ฐ์ดํฐ๋ฅผ
2๋ฐฐ(56k)๋ก ๋๋ฆฌ์ ๊ฒ์ฆ ์์ค์ ๊ฐ์ ๋์ง๋ง(์ ์ธ๋ฌผ ์ ์๋ ฅ) ๋ฏธํ์ต ๋ฌธํญ ์ผ์น์จ์ ๊ทธ๋๋ก
โ **๋ณ๋ชฉ์ ์ธ๋ฌผ ๋ค์์ฑ์ด ์๋๋ค**. 3์ํญ๋ ํ๋ณธ ์ค์ฐจ ๋ด ๋ณํ์ ๊ทธ์ณค๋ค. ๊ฒฐ๋ก : ๋จ์ผ ์ํ
SFT ์ฆ๋ฅ๋ ์ด ๊ณผ์ ์์ ์ฝ 80% ๋ถ๊ทผ์์ ํฌํํ๋ฉฐ, ๋จ์ ์ํ ๊ฐ๊ทน์ ๋ฐ์ดํฐ ์ฆ๋์ด ์๋๋ผ
๋ฐฉ๋ฒ ๋ณ๊ฒฝ(๋ฌธํญ๋น ๊ต์ฌ ๋ค์ค ์ํ๋ง, ๊ฒฝ๊ณ ๋ผ๋ฒจ ํํฐ๋ง, DPO)์ ์์ญ์ด๋ค. ์ด๋ฌํ ๋ฐฉํฅ๊ณผ
์ค์ ์ฌ๋ก ์กฐ์ฌ ์ค์ธก ๋๋น ์์คํ
์ ํ๋(MAE) ์ธก์ ์ ์ด๋ฒ ๊ณต๊ฐ์ ๋ฒ์ ๋ฐ์ด๋ค.
## 6. Limitations ยท ํ๊ณ
1. Model-level metrics measure *fidelity to the teacher*, not real-world accuracy;
population-level accuracy is unmeasured in this release.
2. Personas are synthetic; the source data's joint distributions (ageรoccupationรregion)
deviate from Korean reality (arXiv:2606.12433) โ segment-level conclusions require
statistical correction that is out of scope for the released model.
3. The teacher may role-play stereotypically and distillation inherits this; our sensitivity
metrics quantify conditioning, not fairness.
4. Non-sensitive preference/lifestyle domains only; behavior elsewhere is untested.
**(ํ๊ตญ์ด)** โ ๋ณธ ์งํ๋ "๊ต์ฌ๋ฅผ ์ผ๋ง๋ ๋ฎ์๋๊ฐ"์ด์ง ์ค์ธ๊ณ ์ ํ๋๊ฐ ์๋๋ค โ ๋ชจ์ง๋จ
์์ค ์ ํ๋๋ ์ด๋ฒ ๊ณต๊ฐ์์ ๋ฏธ์ธก์ . โก ํ๋ฅด์๋๋ ํฉ์ฑ ์ธ๋ฌผ์ด๋ฉฐ ์๋ณธ ๋ฐ์ดํฐ์ ๊ฒฐํฉ๋ถํฌ
(๋์ดร์ง์
ร์ง์ญ)๊ฐ ์ค์ ์ ์ด๊ธ๋๋ฏ๋ก, ์ง๋จ๋ณ ๊ฒฐ๋ก ์๋ ๋ณ๋ ํต๊ณ ๋ณด์ ์ด ํ์ํ๋ค.
โข ๊ต์ฌ์ ๊ณ ์ ๊ด๋
์ ์ญํ ๊ทน์ด ํ์์ ์น๊ณ๋ ์ ์๋ค โ ๊ฐ์์ฑ ์งํ๋ ์กฐ๊ฑดํ ์๋ ์ฌ๋ถ๋ฅผ
์ด ๋ฟ ๊ณต์ ์ฑ์ ์ฌ์ง ์๋๋ค. โฃ ๋น๋ฏผ๊ฐ ์ทจํฅยท๋ผ์ดํ์คํ์ผ ์ฃผ์ ๋ก๋ง ํ์ตยทํ๊ฐํ๋ค.
## Attribution ยท ์ถ์ฒ
**Built with Qwen.** Personas: nvidia/Nemotron-Personas-Korea (CC BY 4.0, ยฉ NVIDIA).
Base model: Qwen2.5-7B-Instruct (Apache-2.0). Teacher: Qwen2.5-72B-Instruct-AWQ
(Qwen LICENSE โ outputs used for training, attributed per license). Model weights: Apache-2.0.
|