Upload REPORT.md with huggingface_hub
Browse files
REPORT.md
ADDED
|
@@ -0,0 +1,140 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Technical Report โ kor-persona-survey-7b
|
| 2 |
+
|
| 3 |
+
*Distilling persona-conditioned Korean survey responses from a 72B teacher into a 7B student*
|
| 4 |
+
*72B ๊ต์ฌ์ ํ๊ตญ์ด ํ๋ฅด์๋ ์กฐ๊ฑด๋ถ ์ค๋ฌธ ์๋ต ๋ฅ๋ ฅ์ 7B ํ์์ ์ฆ๋ฅํ๊ธฐ*
|
| 5 |
+
|
| 6 |
+
DOB Studio ยท 2026-08
|
| 7 |
+
|
| 8 |
+
## 1. Motivation ยท ๋๊ธฐ
|
| 9 |
+
|
| 10 |
+
Simulating survey responses with LLM role-play requires one inference per (persona,
|
| 11 |
+
question) pair โ population-scale simulation means thousands of inferences per question.
|
| 12 |
+
Large models judge personas well but are too slow for this workload; small models are fast
|
| 13 |
+
but role-play shallowly and break output formats. We ask:
|
| 14 |
+
|
| 15 |
+
> **RQ. How much of a 72B teacher's persona-conditioned survey-response capability can be
|
| 16 |
+
> transferred to a 7B student through synthetic-triplet distillation alone?**
|
| 17 |
+
|
| 18 |
+
**(ํ๊ตญ์ด)** LLM ์ญํ ๊ทน์ผ๋ก ์ค๋ฌธ ์๋ต์ ์๋ฎฌ๋ ์ด์
ํ๋ ค๋ฉด (ํ๋ฅด์๋, ๋ฌธํญ) ์๋ง๋ค ์ถ๋ก ์ด
|
| 19 |
+
1ํ์ฉ ํ์ํ๋ค โ ๋ชจ์ง๋จ ๊ท๋ชจ ์๋ฎฌ๋ ์ด์
์ ๋ฌธํญ๋น ์์ฒ ํ ์ถ๋ก ์ ๋ปํ๋ค. ๋ํ ๋ชจ๋ธ์
|
| 20 |
+
ํ๋ฅด์๋ ํ๋จ์ ๊น์ง๋ง ์ด ์์
๋์๋ ๋๋ฌด ๋๋ฆฌ๊ณ , ์ํ ๋ชจ๋ธ์ ๋น ๋ฅด์ง๋ง ์ญํ ๊ทน์ด ์๊ณ
|
| 21 |
+
์ถ๋ ฅ ํ์์ ์ด๊ธด๋ค. ๊ทธ๋์ ๋ฌป๋๋ค: **72B ๊ต์ฌ์ ํ๋ฅด์๋ ์กฐ๊ฑด๋ถ ์ค๋ฌธ ์๋ต ๋ฅ๋ ฅ์, ํฉ์ฑ
|
| 22 |
+
์ผ์คํญ ์ฆ๋ฅ๋ง์ผ๋ก 7B ํ์์ ์ผ๋ง๋ ์ด์ํ ์ ์๋๊ฐ?**
|
| 23 |
+
|
| 24 |
+
## 2. Data construction ยท ๋ฐ์ดํฐ ๊ตฌ์ถ
|
| 25 |
+
|
| 26 |
+
The source dataset (nvidia/Nemotron-Personas-Korea, 1M synthetic Korean personas, CC BY 4.0)
|
| 27 |
+
contains persona profiles only โ no questions, no answers. We manufactured the missing
|
| 28 |
+
supervision:
|
| 29 |
+
|
| 30 |
+
1. **Question bank** (private): 515 binary preference questions across 8 non-sensitive
|
| 31 |
+
topics (food, lifestyle, consumption, travel, leisure, work, values, digital), authored
|
| 32 |
+
to span age/region/occupation-sensitive contrasts. A stratified 10% (51 questions) was
|
| 33 |
+
frozen as a held-out evaluation set **before any labeling**.
|
| 34 |
+
2. **Teacher labeling**: Qwen2.5-72B-Instruct-AWQ answered (persona, question) pairs in
|
| 35 |
+
character โ 3,000 personas ร 10 questions = 30,000 raw labels of the form
|
| 36 |
+
`{choice, confidence, reason}`.
|
| 37 |
+
3. **Quality filtering** (95.2% pass): exact option match, valid JSON, confidence โ [0,1],
|
| 38 |
+
reason length bounds, per-question duplicate-reason cap.
|
| 39 |
+
4. **Split hygiene**: evaluation personas (n=20) and held-out questions were excluded from
|
| 40 |
+
all training data; the validation split (572 triplets) is persona-disjoint from training.
|
| 41 |
+
|
| 42 |
+
**(ํ๊ตญ์ด)** ์๋ณธ ๋ฐ์ดํฐ์
(๊ฐ์ ํ๊ตญ์ธ 100๋ง ๋ช
)์๋ ์ธ๋ฌผ ํ๋กํ๋ง ์๊ณ ์ง๋ฌธยท๋๋ต์ด
|
| 43 |
+
์๋ค. ๊ทธ๋์ ํ์ต ์ ํธ๋ฅผ ์ง์ ์ ์กฐํ๋ค: โ ๋น๋ฏผ๊ฐ 8๊ฐ ์ฃผ์ ์ ์์ํ์ผ ๋ฌธํญ 515๊ฐ๋ฅผ ์์ฒด
|
| 44 |
+
์์ฑ(๋น๊ณต๊ฐ)ํ๊ณ , ๋ผ๋ฒจ๋ง **์ด์ ์** 10%(51๋ฌธํญ)๋ฅผ ํ๊ฐ ์ ์ฉ์ผ๋ก ๋ด์ธ โก ๊ต์ฌ(72B)๊ฐ
|
| 45 |
+
3,000๋ช
ร 10๋ฌธํญ = 30,000๊ฑด์ ์ญํ ๊ทน์ผ๋ก ๋ผ๋ฒจ๋ง โข ํ์ยท์ ํ์ง ์ผ์นยทํ์ ๋ ๋ฒ์ยท์ค๋ณต
|
| 46 |
+
์ด์ ๊ธฐ์ค์ผ๋ก ํํฐ(ํต๊ณผ์จ 95.2%) โฃ ํ๊ฐ ์ธ๋ฌผ 20๋ช
๊ณผ ํ๊ฐ ๋ฌธํญ์ ํ์ต์์ ์์ ๋ฐฐ์ ,
|
| 47 |
+
๊ฒ์ฆ 572๊ฑด์ ์ธ๋ฌผ ๋จ์๋ก ๋ถ๋ฆฌ.
|
| 48 |
+
|
| 49 |
+
## 3. Training ยท ํ์ต
|
| 50 |
+
|
| 51 |
+
QLoRA SFT on Qwen2.5-7B-Instruct: NF4 4-bit frozen base + bf16 LoRA adapters
|
| 52 |
+
(r=16, ฮฑ=32, all 7 linear projections ร 28 layers โ 40M trainable params, 0.5%).
|
| 53 |
+
27,999 triplets, 2 epochs, effective batch 48, cosine lr 1e-4, loss on the response span
|
| 54 |
+
only. 3-GPU DDP on RTX 3090s, 2.5 h. Final val loss 0.391 with no overfitting signature.
|
| 55 |
+
|
| 56 |
+
**(ํ๊ตญ์ด)** Qwen2.5-7B๋ฅผ QLoRA๋ก ๋ฏธ์ธ์กฐ์ : ๋ฒ ์ด์ค๋ 4bit๋ก ๋๊ฒฐํ๊ณ LoRA ์ด๋ํฐ
|
| 57 |
+
(์ฝ 4์ฒ๋ง ๊ฐ, ์ ์ฒด์ 0.5%)๋ง bf16์ผ๋ก ํ์ต. ์ผ์คํญ 27,999๊ฑด ร 2์ํญ, ์ ํจ ๋ฐฐ์น 48,
|
| 58 |
+
์์ค์ ์๋ต ๋ถ๋ถ์๋ง ์ ์ฉ. RTX 3090 3์ฅ ๋ณ๋ ฌ๋ก 2.5์๊ฐ. ์ต์ข
๊ฒ์ฆ ์์ค 0.391,
|
| 59 |
+
๊ณผ์ ํฉ ์งํ ์์.
|
| 60 |
+
|
| 61 |
+
## 4. Evaluation ยท ํ๊ฐ
|
| 62 |
+
|
| 63 |
+
Protocol: 51 held-out questions ร 20 unseen personas, temperature 0.7, fixed seeds.
|
| 64 |
+
All model-level metrics are reproducible from the released weights.
|
| 65 |
+
|
| 66 |
+
| Metric ยท ์งํ | Base 7B | **Distilled 7B (์ฆ๋ฅ ํ)** | Teacher 72B |
|
| 67 |
+
|---|---|---|---|
|
| 68 |
+
| Format compliance (strict) ยท ํ์ ์ค์์จ | 95.5% | **98.5%** | 97.0% |
|
| 69 |
+
| Teacher agreement ยท ๊ต์ฌ ์ผ์น์จ | 72.3% | **79.4%** | 96.1% (self-agreement ceiling) |
|
| 70 |
+
| Consistency (5-run majority) ยท ์๋ต ์ผ๊ด์ฑ | 98.7% | 97.6% | โ |
|
| 71 |
+
| Persona sensitivity (TV) ยท ํ๋ฅด์๋ ๊ฐ์์ฑ | 0.356 | **0.356** | โ |
|
| 72 |
+
| Position bias ยท ์ ํ์ง ์์ ํธํฅ | โ | **+1.7%** | โ |
|
| 73 |
+
| Age-conditioning (net) ยท ๋์ด ์กฐ๊ฑดํ(์ํจ๊ณผ) | โ | **+20.8%p** | โ |
|
| 74 |
+
|
| 75 |
+
Key readings:
|
| 76 |
+
- **Ceiling-normalized transfer**: the teacher agrees with *itself* only 96.1% of the time
|
| 77 |
+
under sampling, so the student's 79.4% represents **82.6% of the effective ceiling**
|
| 78 |
+
(75.2% before distillation).
|
| 79 |
+
- **Format**: the student surpasses its own teacher (98.5% vs 97.0%) โ expected, since the
|
| 80 |
+
student rehearsed the schema 28k times while the teacher only follows instructions.
|
| 81 |
+
- **Conditioning preserved**: distillation did not collapse persona sensitivity (TV distance
|
| 82 |
+
unchanged); flipping only the age field changes 29.4% of choices vs an 8.6% resampling
|
| 83 |
+
noise floor.
|
| 84 |
+
- **Position bias absent**: first-option preference is +1.7%, and choice consistency under
|
| 85 |
+
option-order flips (90.7%) sits at the sampling-noise floor (91.4%).
|
| 86 |
+
|
| 87 |
+
**(ํ๊ตญ์ด)** ํ๊ฐ๋ ํ์ต์ ์ฐ์ง ์์ 51๋ฌธํญ ร ๋ฏธํ์ต ์ธ๋ฌผ 20๋ช
์ผ๋ก ์งํํ๊ณ , ๊ณต๊ฐ๋
|
| 88 |
+
๊ฐ์ค์น๋ง์ผ๋ก ์ฌํ ๊ฐ๋ฅํ๋ค. ํต์ฌ ํด์: โ ๊ต์ฌ๋ ํ๋ฅ ์ ์์ฑ์ด๋ผ ์๊ธฐ ์์ ๊ณผ 96.1%๋ง
|
| 89 |
+
์ผ์นํ๋ฏ๋ก, ํ์์ 79.4%๋ **์ค์ง ์ํ ๋๋น 82.6% ๋๋ฌ**(์ฆ๋ฅ ์ 75.2%) โก ํ์
|
| 90 |
+
์ค์์จ์ ํ์์ด ๊ต์ฌ๋ฅผ ์ถ์ โ ํ์์ 2.8๋ง ๋ฒ ๋ฐ๋ณต ํ์ตํ ํจ๊ณผ โข ์ฆ๋ฅ ๊ณผ์ ์์
|
| 91 |
+
ํ๋ฅด์๋ ์กฐ๊ฑดํ ๋ฅ๋ ฅ ์์ค ์์(TV๊ฑฐ๋ฆฌ ์ ์ง), ๋์ด๋ง ๋ค์ง์ด๋ ์๋ต์ 29.4%๊ฐ ๋ณํ
|
| 92 |
+
(๋
ธ์ด์ฆ 8.6% ๋๋น) โฃ LLM ์ค๋ฌธ๊ธฐ์ ํํ ๊ฒฐํจ์ธ ์ ํ์ง ์์ ํธํฅ์ด ์ฌ์ค์ ์์(+1.7%).
|
| 93 |
+
|
| 94 |
+
## 5. Scaling ablation ยท ์ค์ผ์ผ๋ง ๊ฒ์ฆ โ why this configuration is the release
|
| 95 |
+
|
| 96 |
+
We tested both obvious scale-ups under identical hyperparameters and evaluation:
|
| 97 |
+
|
| 98 |
+
| Run | Data | Epochs | Val loss | Held-out teacher agreement |
|
| 99 |
+
|---|---|---|---|---|
|
| 100 |
+
| **A (released ยท ๊ณต๊ฐ๋ณธ)** | 28k | 2 | 0.391 | **79.4%** |
|
| 101 |
+
| B | 28k | 3 | 0.388 | 80.1% (within sampling error) |
|
| 102 |
+
| C | 56k (6,000 personas) | 2 | **0.357** | 79.3% |
|
| 103 |
+
|
| 104 |
+
Doubling persona data improved validation loss (unseen-persona fit) but left unseen-question
|
| 105 |
+
agreement unchanged โ **persona diversity is not the bottleneck**. A third epoch moved
|
| 106 |
+
agreement only within noise. We conclude single-sample SFT distillation saturates near 80%
|
| 107 |
+
on this task; the remaining ceiling gap likely requires methodological change:
|
| 108 |
+
multi-sample distribution distillation per question, borderline-label filtering, or
|
| 109 |
+
preference optimization (DPO). These are left as future work, alongside system-level
|
| 110 |
+
calibration against real Korean survey ground truth (MAE), which we plan to report in a
|
| 111 |
+
future update.
|
| 112 |
+
|
| 113 |
+
**(ํ๊ตญ์ด)** "๋ฐ์ดํฐ๋ฅผ ๋ ์ฐ๋ฉด ๋์ง ์๋?"๋ฅผ ๋์ผ ์กฐ๊ฑด์์ ์ง์ ๊ฒ์ฆํ๋ค. ์ธ๋ฌผ ๋ฐ์ดํฐ๋ฅผ
|
| 114 |
+
2๋ฐฐ(56k)๋ก ๋๋ฆฌ์ ๊ฒ์ฆ ์์ค์ ๊ฐ์ ๋์ง๋ง(์ ์ธ๋ฌผ ์ ์๋ ฅ) ๋ฏธํ์ต ๋ฌธํญ ์ผ์น์จ์ ๊ทธ๋๋ก
|
| 115 |
+
โ **๋ณ๋ชฉ์ ์ธ๋ฌผ ๋ค์์ฑ์ด ์๋๋ค**. 3์ํญ๋ ํ๋ณธ ์ค์ฐจ ๋ด ๋ณํ์ ๊ทธ์ณค๋ค. ๊ฒฐ๋ก : ๋จ์ผ ์ํ
|
| 116 |
+
SFT ์ฆ๋ฅ๋ ์ด ๊ณผ์ ์์ ์ฝ 80% ๋ถ๊ทผ์์ ํฌํํ๋ฉฐ, ๋จ์ ์ํ ๊ฐ๊ทน์ ๋ฐ์ดํฐ ์ฆ๋์ด ์๋๋ผ
|
| 117 |
+
๋ฐฉ๋ฒ ๋ณ๊ฒฝ(๋ฌธํญ๋น ๊ต์ฌ ๋ค์ค ์ํ๋ง, ๊ฒฝ๊ณ ๋ผ๋ฒจ ํํฐ๋ง, DPO)์ ์์ญ์ด๋ค. ์ค์ ์ฌ๋ก ์กฐ์ฌ
|
| 118 |
+
์ค์ธก ๋๋น ์์คํ
์ ํ๋(MAE)์ ํจ๊ป ํฅํ ์
๋ฐ์ดํธ ๊ณผ์ ๋ก ๋จ๊ธด๋ค.
|
| 119 |
+
|
| 120 |
+
## 6. Limitations ยท ํ๊ณ
|
| 121 |
+
|
| 122 |
+
1. Model-level metrics measure *fidelity to the teacher*, not real-world accuracy;
|
| 123 |
+
population-level accuracy is unmeasured in this release.
|
| 124 |
+
2. Personas are synthetic; the source data's joint distributions (ageรoccupationรregion)
|
| 125 |
+
deviate from Korean reality (arXiv:2606.12433) โ segment-level conclusions require
|
| 126 |
+
statistical correction that is out of scope for the released model.
|
| 127 |
+
3. The teacher may role-play stereotypically and distillation inherits this; our sensitivity
|
| 128 |
+
metrics quantify conditioning, not fairness.
|
| 129 |
+
4. Non-sensitive preference/lifestyle domains only; behavior elsewhere is untested.
|
| 130 |
+
|
| 131 |
+
**(ํ๊ตญ์ด)** โ ๋ณธ ์งํ๋ "๊ต์ฌ๋ฅผ ์ผ๋ง๋ ๋ฎ์๋๊ฐ"์ด์ง ์ค์ธ๊ณ ์ ํ๋๊ฐ ์๋๋ค โ ๋ชจ์ง๋จ
|
| 132 |
+
์์ค ์ ํ๋๋ ์ด๋ฒ ๊ณต๊ฐ์์ ๋ฏธ์ธก์ . โก ํ๋ฅด์๋๋ ํฉ์ฑ ์ธ๋ฌผ์ด๋ฉฐ ์๋ณธ ๋ฐ์ดํฐ์ ๊ฒฐํฉ๋ถํฌ
|
| 133 |
+
(๋์ดร์ง์
ร์ง์ญ)๊ฐ ์ค์ ์ ์ด๊ธ๋๋ฏ๋ก, ์ง๋จ๋ณ ๊ฒฐ๋ก ์๋ ๋ณ๋ ํต๊ณ ๋ณด์ ์ด ํ์ํ๋ค.
|
| 134 |
+
โข ๊ต์ฌ์ ๊ณ ์ ๊ด๋
์ ์ญํ ๊ทน์ด ํ์์ ์น๊ณ๋ ์ ์๋ค โ ๊ฐ์์ฑ ์งํ๋ ์กฐ๊ฑดํ ์๋ ์ฌ๋ถ๋ฅผ
|
| 135 |
+
์ด ๋ฟ ๊ณต์ ์ฑ์ ์ฌ์ง ์๋๋ค. โฃ ๋น๋ฏผ๊ฐ ์ทจํฅยท๋ผ์ดํ์คํ์ผ ์ฃผ์ ๋ก๋ง ํ์ตยทํ๊ฐํ๋ค.
|
| 136 |
+
|
| 137 |
+
## Attribution ยท ์ถ์ฒ
|
| 138 |
+
|
| 139 |
+
Personas: nvidia/Nemotron-Personas-Korea (CC BY 4.0, ยฉ NVIDIA). Base & teacher:
|
| 140 |
+
Qwen2.5 family (Apache-2.0, Alibaba Cloud). Model weights: Apache-2.0.
|