# Technical Report — kor-persona-survey-7b *Distilling persona-conditioned Korean survey responses from a 72B teacher into a 7B student* *72B 교사의 한국어 페르소나 조건부 설문 응답 능력을 7B 학생에 증류하기* DOB Studio · 2026-08 ## 1. Motivation · 동기 Simulating survey responses with LLM role-play requires one inference per (persona, question) pair — population-scale simulation means thousands of inferences per question. Large models judge personas well but are too slow for this workload; small models are fast but role-play shallowly and break output formats. We ask: > **RQ. How much of a 72B teacher's persona-conditioned survey-response capability can be > transferred to a 7B student through synthetic-triplet distillation alone?** **(한국어)** LLM 역할극으로 설문 응답을 시뮬레이션하려면 (페르소나, 문항) 쌍마다 추론이 1회씩 필요하다 — 모집단 규모 시뮬레이션은 문항당 수천 회 추론을 뜻한다. 대형 모델은 페르소나 판단은 깊지만 이 작업량에는 너무 느리고, 소형 모델은 빠르지만 역할극이 얕고 출력 형식을 어긴다. 그래서 묻는다: **72B 교사의 페르소나 조건부 설문 응답 능력을, 합성 삼중항 증류만으로 7B 학생에 얼마나 이식할 수 있는가?** ## 2. Data construction · 데이터 구축 The source dataset (nvidia/Nemotron-Personas-Korea, 1M synthetic Korean personas, CC BY 4.0) contains persona profiles only — no questions, no answers. We manufactured the missing supervision: 1. **Question bank** (private): 515 binary preference questions across 8 non-sensitive topics (food, lifestyle, consumption, travel, leisure, work, values, digital), authored to span age/region/occupation-sensitive contrasts. A stratified 10% (51 questions) was frozen as a held-out evaluation set **before any labeling**. 2. **Teacher labeling**: Qwen2.5-72B-Instruct-AWQ answered (persona, question) pairs in character — 3,000 personas × 10 questions = 30,000 raw labels of the form `{choice, confidence, reason}`. 3. **Quality filtering** (95.2% pass): exact option match, valid JSON, confidence ∈ [0,1], reason length bounds, per-question duplicate-reason cap. 4. **Split hygiene**: evaluation personas (n=20) and held-out questions were excluded from all training data; the validation split (572 triplets) is persona-disjoint from training. **(한국어)** 원본 데이터셋(가상 한국인 100만 명)에는 인물 프로필만 있고 질문·대답이 없다. 그래서 학습 신호를 직접 제조했다: ① 비민감 8개 주제의 양자택일 문항 515개를 자체 작성(비공개)하고, 라벨링 **이전에** 10%(51문항)를 평가 전용으로 봉인 ② 교사(72B)가 3,000명 × 10문항 = 30,000건을 역할극으로 라벨링 ③ 형식·선택지 일치·확신도 범위·중복 이유 기준으로 필터(통과율 95.2%) ④ 평가 인물 20명과 평가 문항은 학습에서 완전 배제, 검증 572건은 인물 단위로 분리. ## 3. Training · 학습 QLoRA SFT on Qwen2.5-7B-Instruct: NF4 4-bit frozen base + bf16 LoRA adapters (r=16, α=32, all 7 linear projections × 28 layers ≈ 40M trainable params, 0.5%). 27,999 triplets, 2 epochs, effective batch 48, cosine lr 1e-4, loss on the response span only. Trained on a 4× RTX 3090 (24GB) workstation, ~4 h. Final val loss 0.391 with no overfitting signature. **(한국어)** Qwen2.5-7B를 QLoRA로 미세조정: 베이스는 4bit로 동결하고 LoRA 어댑터 (약 4천만 개, 전체의 0.5%)만 bf16으로 학습. 삼중항 27,999건 × 2에폭, 유효 배치 48, 손실은 응답 부분에만 적용. RTX 3090 4장 워크스테이션에서 약 4시간. 최종 검증 손실 0.391, 과적합 징후 없음. ## 4. Evaluation · 평가 Protocol: 51 held-out questions × 20 unseen personas, temperature 0.7, fixed seeds. All model-level metrics are reproducible from the released weights. | Metric · 지표 | Base 7B | **Distilled 7B (증류 후)** | Teacher 72B | |---|---|---|---| | Format compliance (strict) · 형식 준수율 | 95.5% | **98.5%** | 97.0% | | Teacher agreement · 교사 일치율 | 72.3% | **79.4%** | 96.1% (self-agreement ceiling) | | Consistency (5-run majority) · 응답 일관성 | 98.7% | 97.6% | — | | Persona sensitivity (TV) · 페르소나 감수성 | 0.356 | **0.356** | — | | Position bias · 선택지 순서 편향 | — | **+1.7%** | — | | Age-conditioning (net) · 나이 조건화(순효과) | — | **+20.8%p** | — | Key readings: - **Ceiling-normalized transfer**: the teacher agrees with *itself* only 96.1% of the time under sampling, so the student's 79.4% represents **82.6% of the effective ceiling** (75.2% before distillation). - **Format**: the student surpasses its own teacher (98.5% vs 97.0%) — expected, since the student rehearsed the schema 28k times while the teacher only follows instructions. - **Conditioning preserved**: distillation did not collapse persona sensitivity (TV distance unchanged); flipping only the age field changes 29.4% of choices vs an 8.6% resampling noise floor. - **Position bias absent**: first-option preference is +1.7%, and choice consistency under option-order flips (90.7%) sits at the sampling-noise floor (91.4%). **(한국어)** 평가는 학습에 쓰지 않은 51문항 × 미학습 인물 20명으로 진행했고, 공개된 가중치만으로 재현 가능하다. 핵심 해석: ① 교사도 확률적 생성이라 자기 자신과 96.1%만 일치하므로, 학생의 79.4%는 **실질 상한 대비 82.6% 도달**(증류 전 75.2%) ② 형식 준수율은 학생이 교사를 추월 — 형식을 2.8만 번 반복 학습한 효과 ③ 증류 과정에서 페르소나 조건화 능력 손실 없음(TV거리 유지), 나이만 뒤집어도 응답의 29.4%가 변화 (노이즈 8.6% 대비) ④ LLM 설문기의 흔한 결함인 선택지 순서 편향이 사실상 없음(+1.7%). ## 5. Scaling ablation · 스케일링 검증 — why this configuration is the release We tested both obvious scale-ups under identical hyperparameters and evaluation: | Run | Data | Epochs | Val loss | Held-out teacher agreement | |---|---|---|---|---| | **A (released · 공개본)** | 28k | 2 | 0.391 | **79.4%** | | B | 28k | 3 | 0.388 | 80.1% (within sampling error) | | C | 56k (6,000 personas) | 2 | **0.357** | 79.3% | Doubling persona data improved validation loss (unseen-persona fit) but left unseen-question agreement unchanged — **persona diversity is not the bottleneck**. A third epoch moved agreement only within noise. We conclude single-sample SFT distillation saturates near 80% on this task; the remaining ceiling gap likely requires methodological change: multi-sample distribution distillation per question, borderline-label filtering, or preference optimization (DPO). These directions, along with system-level calibration against real survey ground truth, are outside the scope of this release. **(한국어)** "데이터를 더 쓰면 되지 않나?"를 동일 조건에서 직접 검증했다. 인물 데이터를 2배(56k)로 늘리자 검증 손실은 개선됐지만(새 인물 적응력) 미학습 문항 일치율은 그대로 — **병목은 인물 다양성이 아니다**. 3에폭도 표본 오차 내 변화에 그쳤다. 결론: 단일 샘플 SFT 증류는 이 과제에서 약 80% 부근에서 포화하며, 남은 상한 간극은 데이터 증량이 아니라 방법 변경(문항당 교사 다중 샘플링, 경계 라벨 필터링, DPO)의 영역이다. 이러한 방향과 실제 여론조사 실측 대비 시스템 정확도(MAE) 측정은 이번 공개의 범위 밖이다. ## 6. Limitations · 한계 1. Model-level metrics measure *fidelity to the teacher*, not real-world accuracy; population-level accuracy is unmeasured in this release. 2. Personas are synthetic; the source data's joint distributions (age×occupation×region) deviate from Korean reality (arXiv:2606.12433) — segment-level conclusions require statistical correction that is out of scope for the released model. 3. The teacher may role-play stereotypically and distillation inherits this; our sensitivity metrics quantify conditioning, not fairness. 4. Non-sensitive preference/lifestyle domains only; behavior elsewhere is untested. **(한국어)** ① 본 지표는 "교사를 얼마나 닮았는가"이지 실세계 정확도가 아니다 — 모집단 수준 정확도는 이번 공개에서 미측정. ② 페르소나는 합성 인물이며 원본 데이터의 결합분포 (나이×직업×지역)가 실제와 어긋나므로, 집단별 결론에는 별도 통계 보정이 필요하다. ③ 교사의 고정관념적 역할극이 학생에 승계될 수 있다 — 감수성 지표는 조건화 작동 여부를 잴 뿐 공정성을 재지 않는다. ④ 비민감 취향·라이프스타일 주제로만 학습·평가했다. ## Attribution · 출처 **Built with Qwen.** Personas: nvidia/Nemotron-Personas-Korea (CC BY 4.0, © NVIDIA). Base model: Qwen2.5-7B-Instruct (Apache-2.0). Teacher: Qwen2.5-72B-Instruct-AWQ (Qwen LICENSE — outputs used for training, attributed per license). Model weights: Apache-2.0.