Instructions to use cadazar/han2han-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cadazar/han2han-it with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="cadazar/han2han-it", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("cadazar/han2han-it", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Han2Han IT
Instruction-tuned checkpoint of Han2Han
(han2han-ul2-base-1-it, step 43153). Han2Han is a 169M-parameter
encoder-decoder model that learns script-invariant representations of Korean
text: a document written in Hanja and its Hangul transcription land at the same
point in embedding space. The recipe (jamo and character-level embedding fusion,
morpheme-aware denoising, bidirectional Hanja-Hangul transcription) is described
in the paper, accepted to Findings of EMNLP 2026.
This repo holds the PyTorch weights, the tokenizer with its chat template, and
the modeling code needed to load them through the transformers Auto classes
with trust_remote_code=True. Training code, the Flax model, the Flax-to-PyTorch
converter, and the evaluation pipeline live in the GitHub repo.
Intended use
Han2Han is meant for representations: classification and fine-tuning on downstream tasks, and sentence embeddings that treat Hanja and Hangul spellings of the same text alike. The model generates text, but generation is not what it was built or evaluated for, and improving it is follow-up work:
- Prompt layout. The task prompt goes in the system message and the text alone in the user message, which is how the instruction tuning presented every task. With the instruction written into the user message instead, Hangul to Hanja restoration returns the Hangul input unchanged.
- Hangul to Hanja restoration (
ํ๊ธ์ ํ์๋ก ์ ์ฌํ์์ค:) works on article-length input. Given the first 478 characters of a 1922 newspaper article in Hangul, greedy decoding restores 195 of its 251 Hanja (78%). The article comes from a corpus used in pre-training, so this is an illustration and not a held-out measurement. On single sentences the model mostly returns Hangul: across 256 newspaper sentences it restores 24% of the Hanja, and 145 of the outputs contain no Hanja at all. - Hanja to Hangul transcription (
ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:) works on sentences and articles, in either layout, but can leave some Hanja untranscribed. - Stopping. Give
generate()a token limit: twice the input's tokens plus 8 leaves room for a correct answer. In fp32 the outputs measured here end on their own; in bf16 the article restoration ran on to the limit, repeating itself.
hanja_transcription_demo.ipynb in the GitHub repo was written with the other
layout (the instruction in the user message), so the Hangul to Hanja outputs in
its sections 4 to 9 understate what this checkpoint does. Its section 10 runs
the layout described here.
Usage
Runtime requirements: torch and transformers (tested with torch 2.14.1 and
transformers 5.18.0 on CPU). The model class ships in this repo, so loading needs
trust_remote_code=True; the tokenizer itself runs no custom code, but without
the flag AutoTokenizer stops to ask.
Sentence embeddings
output_sentence_embeddings=True returns the mean-pooled encoder states, one
640-dimensional vector per input.
import torch
import torch.nn.functional as F
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
repo = "cadazar/han2han-it"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
texts = [
"ๆๅ ด์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ๆฅๆฌไบบๅด ็ตๅฎถ๊น์ง๋ ไธไบบ๋ ็ผ่ฆํ ์๊ฐ ์์๋ค.",
"ํ์ฅ์ ์ผ์ํ๊ณ ๋์์ฌ ๋๊น์ง๋ ์ผ๋ณธ์ธ์ธก ํ๊ฐ๊น์ง๋ ์ผ์ธ๋ ๋ฐ๊ฒฌํ ์๊ฐ ์์๋ค.",
"ๅ็ต๋ ๅๅๅญ์์๋ ็ ดๅขจ์ ๅฆๆณ์ผ๋ก ็ฝ้ช์ ่ฑกๅพตํ ์ ์๋ค.",
]
inputs = tokenizer(texts, return_tensors="pt", padding=True)
with torch.no_grad():
embeddings = model(**inputs, output_sentence_embeddings=True)[0]
embeddings = F.normalize(embeddings, dim=-1)
print(embeddings @ embeddings.T)
# tensor([[1.0000, 0.9444, 0.8789],
# [0.9444, 1.0000, 0.8459],
# [0.8789, 0.8459, 1.0000]])
The first two rows are the same 1920s newspaper sentence in mixed script and in Hangul; the third is a different sentence.
Chat
AutoModelForCausalLM loads Han2HanForCausalLM, a wrapper that gives the
encoder-decoder the interface chat tooling expects: generate() takes the
rendered chat prompt as one sequence and returns the prompt followed by the
reply. pipeline("text-generation") works with it, and so does the
transformers server:
transformers serve --trust-remote-code
transformers chat cadazar/han2han-it --system-prompt "ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:"
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "cadazar/han2han-it"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).eval()
def run(task_prompt, text):
messages = [
{"role": "system", "content": task_prompt},
{"role": "user", "content": text},
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
)
budget = 2 * len(tokenizer(text).input_ids) + 8
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=budget, do_sample=False)
return tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(run(
"ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:",
"ๆๅ ด์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ๆฅๆฌไบบๅด ็ตๅฎถ๊น์ง๋ ไธไบบ๋ ็ผ่ฆํ ์๊ฐ ์์๋ค.",
))
# ํ์ฅ์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ์ผ๋ณธ์ธ์ธก ํ๊ฐ๊น์ง๋ ์ผ์ธ๋ ๋ฐ๊ฒฌํ ์๊ฐ ์์๋ค.
hangul = (
"๋ฏธ์ ์ ๋ ๊ณํ ์ด๋
๋ถ ์ฌ์์ผ๋ก ๋ช
๋
์ด์ฐจ ๊ฐ์ค์ด๋
๋ถ์์๋ ์กฐ์ ์ ์ฌ ํ ๋ฏธ์ ์ ๋ฐ๋ฌ์ ๋น๋ณดํ "
"๋ชฉ์ ์ผ๋ก ๋๊ฒฝ์ ์ ๊ตญ๋ฏธ์ ์์ ๋ํ๋ฅผ ๋ฐฉํ ์ฌ ๋งค๋
์ผ์ฐจ ๋ฏธ์ ์ ๋ํ๋ฅผ ๊ฐํ ๋ฐฉ์นจ์ ๋ด์ ํ๊ณ ์ "
"์ด์ญ์ก์ผ ์ค์ ์ญ์ ์ด๋
๋ถ ์ ์ดํ์์ค์ ๋ฐ์ํจ ํ, ๋ฏผ๋ณ์ ์, ์ํํํ์ ์ ๋์ , ๊น๋ํฌ, ์ด๋์ "
"์ธ ์ ์จ, ์ํ์ฐ๊ตฌํ์ ๊น๊ท์ง ์จ, ์ผ๋ณธ์ธ์ธก ์ํ๊ฐ๋ก ๊ณ ๋ชฉ๋ฐฐ์ ์ธ ์์จ ๊ธฐํ ์ํ๊ฐ์ ๊ด๊ณ์๋ "
"์ธ์ฌ๋ฅผ ์ด์ฒญํ๊ณ ์์ผ ์ ๋ฌด์ด๊ฐ ์ดํ ํ๋ฌด๋น๊ตญ์๊ฐ ํํฉํ์ฌ ์ฐจ์ ๊ดํ ์์ํ๋ฅผ ๊ฐํ๋ฐ"
)
print(run("ํ๊ธ์ ํ์๋ก ์ ์ฌํ์์ค:", hangul))
# ็พ่ก่ฆฝ ่จๅ ็ธฝ็ฃๅบ ไบ้
์ผ๋ก ๆๅนด ๅๆฌก ้่จญ็ธฝ็ฃๅบ์์๋ ๆ้ฎฎ์ ๅจ ํ ็พ่ก์ ็ผ้์ ๅๅ ฑํ
# ็ฒํ์ผ ๆฑไบฌ์ ๅธๅ็พ่กๅๅฑ่ฆฝๆ๋ฅผ ่จชํ ์ฌ ๆฏๅนด ไธๆฌก ็พ่ก๋ฅผ ้ํ ๆน้์ ๅ
งๅฎํ๊ณ ๆ
# ไบๅๅ
ญๆฅ ๅๅ ๅๆ ็ธฝ็ฃๅบ ็ฌฌไบๅ ่ญฐๅฎค์ ๆดๆฐธๅญ ๅพ, ้ไธ้ซ ่
, ๆธ็ตๅๆ์ ๆฟๅคง่ฃ, ้ๆฆ็,
# ๆ้่ฑ ๅค ่ซธๆฐ, ๆธ็ต็ก็ฉถๆ์ ้ๅฅ้ญ ๆฐ, ๆฅๆฌไบบๅด ๆธ็ตๅฎถ๋ก ้ซๆจๅน ๅค ์์จ ๅ
ถไป ๆธ็ตๅฎถ์
# ้ไฟ์๋ ไบบไบ๋ฅผ ๆ่ซํ๊ณ ๅฃฝ้ ๆฟๅ็ธฝ็ฃ ไปฅไธ ๅญธๅ็ถๅฑ่
๊ฐ ๆๅํ์ฌ ์ฐจ์ ้ํ ๅๆ๋ฅผ ้ํ๋ฐ
The outputs are shown as generated (each is one line). In the first, ไธๅทก
should have become ์ผ์. The second is the opening of the 1922 article; the
original reads ็พ่กๅฑ่ฆฝ ่จๅ ็ธฝ็ฃๅบ ไบๆก์ผ๋ก ๆๅนด ๅๆฌก ้่จญ็ธฝ็ฃๅบ์์๋ ๆ้ฎฎ์ ๅจ ํ ็พ่ก์ ็ผ้์ ่ฃจ่ฃํ ๋ชฉ์ ์ผ๋ก ๆฑไบฌ์ ๅธๅ็พ่ก้ขๅฑ่ฆฝๆ๋ฅผ ๅฃํ ์ฌ ..., so the
output drops or swaps some characters (็พ่ก่ฆฝ, ไบ้
, ๅธๅ็พ่กๅ) and
writes same-sounding ones for several names.
Prompt format, as used by the SFT collator in the GitHub repo and by the chat template:
<|system|>{system}<|user|>{user}<|end_of_turn|>, with<|assistant|>{reply}<|end_of_turn|><|user|>{user}<|end_of_turn|>for each further turn. The system part is optional in the format, but it is where this checkpoint expects the task prompt. Each direction was trained with three wordings:ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:,๋ค์ ํ๋ฌธ์ ํ๊ธ๋ก ์ฎ๊ธฐ์์ค:,ํ์ ํ๊ธฐ๋ฅผ ํ๊ธ ๋ ์์ผ๋ก ๋ณํํ์์ค:for Hanja to Hangul, andํ๊ธ์ ํ์๋ก ์ ์ฌํ์์ค:,๋ค์ ํ๊ธ ํ ์คํธ์ ์ ์ ํ ํ์๋ฅผ ์ถ๊ฐํ์์ค:,ํ์ ํ๊ธฐ๊ฐ ํ์ํ ๋ถ๋ถ์ ํ์๋ฅผ ๋ณ๊ธฐํ์์ค:for Hangul to Hanja.- The generation prompt is
<|assistant|>, or<|think|>withenable_thinking=True, in which case the reply is reasoning, then<|assistant|>, then the answer. A turn ends with<|end_of_turn|>(eos_token_id10).
The wrapper splits the prompt at its last <|end_of_turn|>: everything up to
and including it is the encoder input, and the generation prompt after it starts
the decoder. A prompt without a generation prompt raises. forward() is not
wrapped and stays encoder-decoder.
Encoder-decoder generation
AutoModelForSeq2SeqLM loads the model itself. generate() is the standard
transformers one, with a KV cache; greedy decoding, beam search, and sampling
all work. The encoder takes the prompt up to <|end_of_turn|>, and the decoder
starts from <|assistant|> (decoder_start_token_id 9).
from transformers import AutoModelForSeq2SeqLM
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
prompt = (
"<|system|>ํ์๋ฅผ ํ๊ธ๋ก ์ ์ฌํ์์ค:"
"<|user|>ๆๅ ด์ ไธๅทกํ๊ณ ๋์์ฌ ๋๊น์ง๋ ๆฅๆฌไบบๅด ็ตๅฎถ๊น์ง๋ ไธไบบ๋ ็ผ่ฆํ ์๊ฐ ์์๋ค.<|end_of_turn|>"
)
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Tokenizer
tokenizer.json is a tokenizers-library build of the SentencePiece model
(spiece.model, kept here as the source), made by
scripts/build_hf_tokenizer.py in the GitHub repo. Calling the tokenizer does
not add BOS or EOS tokens, and special tokens written into the text are mapped
to their ids. It encodes like the SentencePiece wrapper the training code uses,
with one known difference: where two segmentations of a span have exactly the
same score (runs of digits, mostly), the same pieces can come out in a different
order.
Files
| File | Contents |
|---|---|
model.safetensors |
fp32 weights, 169.2M parameters, plus the jbu / cbu subword bucket tables |
config.json, generation_config.json |
model and generation config, with auto_map entries for the Auto classes |
tokenizer.json, tokenizer_config.json, chat_template.jinja |
fast tokenizer (38400 pieces) and chat template |
spiece.model |
the SentencePiece model tokenizer.json was built from |
modeling_han2han.py, han2han_config.py |
modeling code from the GitHub repo at commit af1330e |
modeling_han2han.py (modeling_han2han_pytorch.py there) differs in two
places: the config import is relative, and the module-level register_han2han
import is replaced by the auto_map entries. han2han_config.py is the GitHub
copy at commit 7304c38; it has not changed since. config.json names
AutoModelForCausalLM under architectures so that transformers serve, which
looks that name up in the transformers namespace, can load the model.
Instruction tuning
Fine-tuned from the Han2Han pre-trained checkpoint on instruction following,
chain-of-thought reasoning, Hanja-Hangul article transcription, and
summarization data (configs/it-muon-stage_1.yaml in the GitHub repo).
Citation
@inproceedings{han2han2026,
title = {Han2Han: Efficient Language-Specific Character Representation
through Script-Aware Pre-Training for Historical Text Analysis},
author = {Adams, Cellik and Jo, EunKyoung and Kim, Juae},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026}
}
License
Apache License 2.0, the same as the GitHub repo.
- Downloads last month
- 695