|
Download REPORT.md from dobstudio/kor-persona-survey-7b: direct link, hf CLI and curl.
- Browser
- Download file 9.29 kB
-
https://huggingface.co/dobstudio/kor-persona-survey-7b/resolve/main/REPORT.md
- Command line
-
hf download hf://dobstudio/kor-persona-survey-7b/REPORT.md
-
curl -L -o REPORT.md https://huggingface.co/dobstudio/kor-persona-survey-7b/resolve/main/REPORT.md
9.29 kB
| # Technical Report โ kor-persona-survey-7b | |
| *Distilling persona-conditioned Korean survey responses from a 72B teacher into a 7B student* | |
| *72B ๊ต์ฌ์ ํ๊ตญ์ด ํ๋ฅด์๋ ์กฐ๊ฑด๋ถ ์ค๋ฌธ ์๋ต ๋ฅ๋ ฅ์ 7B ํ์์ ์ฆ๋ฅํ๊ธฐ* | |
| DOB Studio ยท 2026-08 | |
| ## 1. Motivation ยท ๋๊ธฐ | |
| Simulating survey responses with LLM role-play requires one inference per (persona, | |
| question) pair โ population-scale simulation means thousands of inferences per question. | |
| Large models judge personas well but are too slow for this workload; small models are fast | |
| but role-play shallowly and break output formats. We ask: | |
| > **RQ. How much of a 72B teacher's persona-conditioned survey-response capability can be | |
| > transferred to a 7B student through synthetic-triplet distillation alone?** | |
| **(ํ๊ตญ์ด)** LLM ์ญํ ๊ทน์ผ๋ก ์ค๋ฌธ ์๋ต์ ์๋ฎฌ๋ ์ด์ ํ๋ ค๋ฉด (ํ๋ฅด์๋, ๋ฌธํญ) ์๋ง๋ค ์ถ๋ก ์ด | |
| 1ํ์ฉ ํ์ํ๋ค โ ๋ชจ์ง๋จ ๊ท๋ชจ ์๋ฎฌ๋ ์ด์ ์ ๋ฌธํญ๋น ์์ฒ ํ ์ถ๋ก ์ ๋ปํ๋ค. ๋ํ ๋ชจ๋ธ์ | |
| ํ๋ฅด์๋ ํ๋จ์ ๊น์ง๋ง ์ด ์์ ๋์๋ ๋๋ฌด ๋๋ฆฌ๊ณ , ์ํ ๋ชจ๋ธ์ ๋น ๋ฅด์ง๋ง ์ญํ ๊ทน์ด ์๊ณ | |
| ์ถ๋ ฅ ํ์์ ์ด๊ธด๋ค. ๊ทธ๋์ ๋ฌป๋๋ค: **72B ๊ต์ฌ์ ํ๋ฅด์๋ ์กฐ๊ฑด๋ถ ์ค๋ฌธ ์๋ต ๋ฅ๋ ฅ์, ํฉ์ฑ | |
| ์ผ์คํญ ์ฆ๋ฅ๋ง์ผ๋ก 7B ํ์์ ์ผ๋ง๋ ์ด์ํ ์ ์๋๊ฐ?** | |
| ## 2. Data construction ยท ๋ฐ์ดํฐ ๊ตฌ์ถ | |
| The source dataset (nvidia/Nemotron-Personas-Korea, 1M synthetic Korean personas, CC BY 4.0) | |
| contains persona profiles only โ no questions, no answers. We manufactured the missing | |
| supervision: | |
| 1. **Question bank** (private): 515 binary preference questions across 8 non-sensitive | |
| topics (food, lifestyle, consumption, travel, leisure, work, values, digital), authored | |
| to span age/region/occupation-sensitive contrasts. A stratified 10% (51 questions) was | |
| frozen as a held-out evaluation set **before any labeling**. | |
| 2. **Teacher labeling**: Qwen2.5-72B-Instruct-AWQ answered (persona, question) pairs in | |
| character โ 3,000 personas ร 10 questions = 30,000 raw labels of the form | |
| `{choice, confidence, reason}`. | |
| 3. **Quality filtering** (95.2% pass): exact option match, valid JSON, confidence โ [0,1], | |
| reason length bounds, per-question duplicate-reason cap. | |
| 4. **Split hygiene**: evaluation personas (n=20) and held-out questions were excluded from | |
| all training data; the validation split (572 triplets) is persona-disjoint from training. | |
| **(ํ๊ตญ์ด)** ์๋ณธ ๋ฐ์ดํฐ์ (๊ฐ์ ํ๊ตญ์ธ 100๋ง ๋ช )์๋ ์ธ๋ฌผ ํ๋กํ๋ง ์๊ณ ์ง๋ฌธยท๋๋ต์ด | |
| ์๋ค. ๊ทธ๋์ ํ์ต ์ ํธ๋ฅผ ์ง์ ์ ์กฐํ๋ค: โ ๋น๋ฏผ๊ฐ 8๊ฐ ์ฃผ์ ์ ์์ํ์ผ ๋ฌธํญ 515๊ฐ๋ฅผ ์์ฒด | |
| ์์ฑ(๋น๊ณต๊ฐ)ํ๊ณ , ๋ผ๋ฒจ๋ง **์ด์ ์** 10%(51๋ฌธํญ)๋ฅผ ํ๊ฐ ์ ์ฉ์ผ๋ก ๋ด์ธ โก ๊ต์ฌ(72B)๊ฐ | |
| 3,000๋ช ร 10๋ฌธํญ = 30,000๊ฑด์ ์ญํ ๊ทน์ผ๋ก ๋ผ๋ฒจ๋ง โข ํ์ยท์ ํ์ง ์ผ์นยทํ์ ๋ ๋ฒ์ยท์ค๋ณต | |
| ์ด์ ๊ธฐ์ค์ผ๋ก ํํฐ(ํต๊ณผ์จ 95.2%) โฃ ํ๊ฐ ์ธ๋ฌผ 20๋ช ๊ณผ ํ๊ฐ ๋ฌธํญ์ ํ์ต์์ ์์ ๋ฐฐ์ , | |
| ๊ฒ์ฆ 572๊ฑด์ ์ธ๋ฌผ ๋จ์๋ก ๋ถ๋ฆฌ. | |
| ## 3. Training ยท ํ์ต | |
| QLoRA SFT on Qwen2.5-7B-Instruct: NF4 4-bit frozen base + bf16 LoRA adapters | |
| (r=16, ฮฑ=32, all 7 linear projections ร 28 layers โ 40M trainable params, 0.5%). | |
| 27,999 triplets, 2 epochs, effective batch 48, cosine lr 1e-4, loss on the response span | |
| only. Trained on a 4ร RTX 3090 (24GB) workstation, ~4 h. Final val loss 0.391 with no | |
| overfitting signature. | |
| **(ํ๊ตญ์ด)** Qwen2.5-7B๋ฅผ QLoRA๋ก ๋ฏธ์ธ์กฐ์ : ๋ฒ ์ด์ค๋ 4bit๋ก ๋๊ฒฐํ๊ณ LoRA ์ด๋ํฐ | |
| (์ฝ 4์ฒ๋ง ๊ฐ, ์ ์ฒด์ 0.5%)๋ง bf16์ผ๋ก ํ์ต. ์ผ์คํญ 27,999๊ฑด ร 2์ํญ, ์ ํจ ๋ฐฐ์น 48, | |
| ์์ค์ ์๋ต ๋ถ๋ถ์๋ง ์ ์ฉ. RTX 3090 4์ฅ ์ํฌ์คํ ์ด์ ์์ ์ฝ 4์๊ฐ. ์ต์ข ๊ฒ์ฆ ์์ค | |
| 0.391, ๊ณผ์ ํฉ ์งํ ์์. | |
| ## 4. Evaluation ยท ํ๊ฐ | |
| Protocol: 51 held-out questions ร 20 unseen personas, temperature 0.7, fixed seeds. | |
| All model-level metrics are reproducible from the released weights. | |
| | Metric ยท ์งํ | Base 7B | **Distilled 7B (์ฆ๋ฅ ํ)** | Teacher 72B | | |
| |---|---|---|---| | |
| | Format compliance (strict) ยท ํ์ ์ค์์จ | 95.5% | **98.5%** | 97.0% | | |
| | Teacher agreement ยท ๊ต์ฌ ์ผ์น์จ | 72.3% | **79.4%** | 96.1% (self-agreement ceiling) | | |
| | Consistency (5-run majority) ยท ์๋ต ์ผ๊ด์ฑ | 98.7% | 97.6% | โ | | |
| | Persona sensitivity (TV) ยท ํ๋ฅด์๋ ๊ฐ์์ฑ | 0.356 | **0.356** | โ | | |
| | Position bias ยท ์ ํ์ง ์์ ํธํฅ | โ | **+1.7%** | โ | | |
| | Age-conditioning (net) ยท ๋์ด ์กฐ๊ฑดํ(์ํจ๊ณผ) | โ | **+20.8%p** | โ | | |
| Key readings: | |
| - **Ceiling-normalized transfer**: the teacher agrees with *itself* only 96.1% of the time | |
| under sampling, so the student's 79.4% represents **82.6% of the effective ceiling** | |
| (75.2% before distillation). | |
| - **Format**: the student surpasses its own teacher (98.5% vs 97.0%) โ expected, since the | |
| student rehearsed the schema 28k times while the teacher only follows instructions. | |
| - **Conditioning preserved**: distillation did not collapse persona sensitivity (TV distance | |
| unchanged); flipping only the age field changes 29.4% of choices vs an 8.6% resampling | |
| noise floor. | |
| - **Position bias absent**: first-option preference is +1.7%, and choice consistency under | |
| option-order flips (90.7%) sits at the sampling-noise floor (91.4%). | |
| **(ํ๊ตญ์ด)** ํ๊ฐ๋ ํ์ต์ ์ฐ์ง ์์ 51๋ฌธํญ ร ๋ฏธํ์ต ์ธ๋ฌผ 20๋ช ์ผ๋ก ์งํํ๊ณ , ๊ณต๊ฐ๋ | |
| ๊ฐ์ค์น๋ง์ผ๋ก ์ฌํ ๊ฐ๋ฅํ๋ค. ํต์ฌ ํด์: โ ๊ต์ฌ๋ ํ๋ฅ ์ ์์ฑ์ด๋ผ ์๊ธฐ ์์ ๊ณผ 96.1%๋ง | |
| ์ผ์นํ๋ฏ๋ก, ํ์์ 79.4%๋ **์ค์ง ์ํ ๋๋น 82.6% ๋๋ฌ**(์ฆ๋ฅ ์ 75.2%) โก ํ์ | |
| ์ค์์จ์ ํ์์ด ๊ต์ฌ๋ฅผ ์ถ์ โ ํ์์ 2.8๋ง ๋ฒ ๋ฐ๋ณต ํ์ตํ ํจ๊ณผ โข ์ฆ๋ฅ ๊ณผ์ ์์ | |
| ํ๋ฅด์๋ ์กฐ๊ฑดํ ๋ฅ๋ ฅ ์์ค ์์(TV๊ฑฐ๋ฆฌ ์ ์ง), ๋์ด๋ง ๋ค์ง์ด๋ ์๋ต์ 29.4%๊ฐ ๋ณํ | |
| (๋ ธ์ด์ฆ 8.6% ๋๋น) โฃ LLM ์ค๋ฌธ๊ธฐ์ ํํ ๊ฒฐํจ์ธ ์ ํ์ง ์์ ํธํฅ์ด ์ฌ์ค์ ์์(+1.7%). | |
| ## 5. Scaling ablation ยท ์ค์ผ์ผ๋ง ๊ฒ์ฆ โ why this configuration is the release | |
| We tested both obvious scale-ups under identical hyperparameters and evaluation: | |
| | Run | Data | Epochs | Val loss | Held-out teacher agreement | | |
| |---|---|---|---|---| | |
| | **A (released ยท ๊ณต๊ฐ๋ณธ)** | 28k | 2 | 0.391 | **79.4%** | | |
| | B | 28k | 3 | 0.388 | 80.1% (within sampling error) | | |
| | C | 56k (6,000 personas) | 2 | **0.357** | 79.3% | | |
| Doubling persona data improved validation loss (unseen-persona fit) but left unseen-question | |
| agreement unchanged โ **persona diversity is not the bottleneck**. A third epoch moved | |
| agreement only within noise. We conclude single-sample SFT distillation saturates near 80% | |
| on this task; the remaining ceiling gap likely requires methodological change: | |
| multi-sample distribution distillation per question, borderline-label filtering, or | |
| preference optimization (DPO). These directions, along with system-level calibration | |
| against real survey ground truth, are outside the scope of this release. | |
| **(ํ๊ตญ์ด)** "๋ฐ์ดํฐ๋ฅผ ๋ ์ฐ๋ฉด ๋์ง ์๋?"๋ฅผ ๋์ผ ์กฐ๊ฑด์์ ์ง์ ๊ฒ์ฆํ๋ค. ์ธ๋ฌผ ๋ฐ์ดํฐ๋ฅผ | |
| 2๋ฐฐ(56k)๋ก ๋๋ฆฌ์ ๊ฒ์ฆ ์์ค์ ๊ฐ์ ๋์ง๋ง(์ ์ธ๋ฌผ ์ ์๋ ฅ) ๋ฏธํ์ต ๋ฌธํญ ์ผ์น์จ์ ๊ทธ๋๋ก | |
| โ **๋ณ๋ชฉ์ ์ธ๋ฌผ ๋ค์์ฑ์ด ์๋๋ค**. 3์ํญ๋ ํ๋ณธ ์ค์ฐจ ๋ด ๋ณํ์ ๊ทธ์ณค๋ค. ๊ฒฐ๋ก : ๋จ์ผ ์ํ | |
| SFT ์ฆ๋ฅ๋ ์ด ๊ณผ์ ์์ ์ฝ 80% ๋ถ๊ทผ์์ ํฌํํ๋ฉฐ, ๋จ์ ์ํ ๊ฐ๊ทน์ ๋ฐ์ดํฐ ์ฆ๋์ด ์๋๋ผ | |
| ๋ฐฉ๋ฒ ๋ณ๊ฒฝ(๋ฌธํญ๋น ๊ต์ฌ ๋ค์ค ์ํ๋ง, ๊ฒฝ๊ณ ๋ผ๋ฒจ ํํฐ๋ง, DPO)์ ์์ญ์ด๋ค. ์ด๋ฌํ ๋ฐฉํฅ๊ณผ | |
| ์ค์ ์ฌ๋ก ์กฐ์ฌ ์ค์ธก ๋๋น ์์คํ ์ ํ๋(MAE) ์ธก์ ์ ์ด๋ฒ ๊ณต๊ฐ์ ๋ฒ์ ๋ฐ์ด๋ค. | |
| ## 6. Limitations ยท ํ๊ณ | |
| 1. Model-level metrics measure *fidelity to the teacher*, not real-world accuracy; | |
| population-level accuracy is unmeasured in this release. | |
| 2. Personas are synthetic; the source data's joint distributions (ageรoccupationรregion) | |
| deviate from Korean reality (arXiv:2606.12433) โ segment-level conclusions require | |
| statistical correction that is out of scope for the released model. | |
| 3. The teacher may role-play stereotypically and distillation inherits this; our sensitivity | |
| metrics quantify conditioning, not fairness. | |
| 4. Non-sensitive preference/lifestyle domains only; behavior elsewhere is untested. | |
| **(ํ๊ตญ์ด)** โ ๋ณธ ์งํ๋ "๊ต์ฌ๋ฅผ ์ผ๋ง๋ ๋ฎ์๋๊ฐ"์ด์ง ์ค์ธ๊ณ ์ ํ๋๊ฐ ์๋๋ค โ ๋ชจ์ง๋จ | |
| ์์ค ์ ํ๋๋ ์ด๋ฒ ๊ณต๊ฐ์์ ๋ฏธ์ธก์ . โก ํ๋ฅด์๋๋ ํฉ์ฑ ์ธ๋ฌผ์ด๋ฉฐ ์๋ณธ ๋ฐ์ดํฐ์ ๊ฒฐํฉ๋ถํฌ | |
| (๋์ดร์ง์ ร์ง์ญ)๊ฐ ์ค์ ์ ์ด๊ธ๋๋ฏ๋ก, ์ง๋จ๋ณ ๊ฒฐ๋ก ์๋ ๋ณ๋ ํต๊ณ ๋ณด์ ์ด ํ์ํ๋ค. | |
| โข ๊ต์ฌ์ ๊ณ ์ ๊ด๋ ์ ์ญํ ๊ทน์ด ํ์์ ์น๊ณ๋ ์ ์๋ค โ ๊ฐ์์ฑ ์งํ๋ ์กฐ๊ฑดํ ์๋ ์ฌ๋ถ๋ฅผ | |
| ์ด ๋ฟ ๊ณต์ ์ฑ์ ์ฌ์ง ์๋๋ค. โฃ ๋น๋ฏผ๊ฐ ์ทจํฅยท๋ผ์ดํ์คํ์ผ ์ฃผ์ ๋ก๋ง ํ์ตยทํ๊ฐํ๋ค. | |
| ## Attribution ยท ์ถ์ฒ | |
| **Built with Qwen.** Personas: nvidia/Nemotron-Personas-Korea (CC BY 4.0, ยฉ NVIDIA). | |
| Base model: Qwen2.5-7B-Instruct (Apache-2.0). Teacher: Qwen2.5-72B-Instruct-AWQ | |
| (Qwen LICENSE โ outputs used for training, attributed per license). Model weights: Apache-2.0. | |