Model Card for Model ID


Open Ko LLM Leaderboard Season 2 🏆 Rank-1 2024/11/01~2024/12/28


AI 와 빅데이터 분석 전문 기업인 Linkbricks Horizon-AI의 데이터사이언티스트인 지윤성(Saxo) 대표가 gemma-2-27b-it 베이스모델을 H100-80G 8개를 통해 SFT->DPO 파인 튜닝을 한 한글 언어 모델로 한국어-중국어-영어-일본어 교차 학습 데이터와 로지컬 데이터를 통하여 한중일영 언어 교차 증강 처리와 복잡한 한글 논리 문제 역시 대응 가능하도록 훈련한 모델이며 토크나이저는 단어 확장 없이 베이스 모델 그대로 사용. 특히 고객 리뷰나 소셜 포스팅 고차원 분석 및 코딩과 욕설, 비방, 음란물, 성적으로 노골적인 콘텐츠, 인종차별 또는 기타 유해 콘텐츠 탐지 개인정보 및 민감정보 패턴 탐지, 차단 또는 사용자 경고가 필요할 수 있는 기타 고위험 요청 탐지 강화된 모델
-Deepspeed Stage=3, rslora 및 BAdam Layer Mode 사용
Linkbricks Horizon-AI CEO and data scientist Yunsung Ji (Saxo) developed a Korean-specialized large language model based on Gemma-2-27B-IT, fine-tuned through a two-stage SFT → DPO pipeline using eight NVIDIA H100 80GB GPUs. The model was trained on cross-lingual datasets spanning Korean, Chinese, English, and Japanese, together with logical reasoning datasets. This enables cross-lingual knowledge augmentation and reasoning across all four languages, while significantly improving its ability to handle complex Korean-language logical reasoning tasks. The model retains the original Gemma-2-27B-IT tokenizer without additional vocabulary expansion, preserving compatibility with the base model architecture. It is particularly optimized for high-dimensional analysis of customer reviews and social media posts, as well as coding and content-safety applications. Its safety capabilities have been strengthened to detect potentially harmful content, including profanity, harassment and defamation, sexually explicit or obscene material, racist or discriminatory content, and other forms of harmful content. The model also provides enhanced detection of personally identifiable information (PII) and sensitive-information patterns, along with improved identification of high-risk requests that may require blocking, additional safeguards, or user warnings. -ollama run benedict/linkbricks-gemma2-korean:27b
Benchmark (Open Ko LLM Leader Board Season 2 : No. 1)
Model : Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B
Average : 51.37
Ko-GPQA : 25.25
Ko-Winogrande : 68.27
Ko-GSM8k : 70.96
Ko-EQ Bench : 50.25
Ko-IFEval : 49.84
KorNAT-CKA : 34.59
KorNAT-SVA : 48.42
Ko-Harmlessness : 65.66
Ko-Helpfulness : 49.12



CEO Yunsung Ji (Saxo), a data scientist at Linkbricks Horizon-AI, a company specializing in AI and big data analytics, fine-tuned the gemma-2-27b-it base model with SFT->DPO using four H100-80Gs. It is a Korean language model trained to handle complex Korean logic problems through Korean-Chinese-English-Japanese cross-training data and logical data, and Tokenizer uses the base model without word expansion.

www.horizonai.ai, www.linkbricks.com, www.linkbricks.vc

Downloads last month
899
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B

Quantized
(60)
this model

Datasets used to train Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B

Space using Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B 1