Han2Han IT

Instruction-tuned checkpoint of Han2Han (han2han-ul2-base-1-it, step 43153). Han2Han is a 169M-parameter encoder-decoder model that learns script-invariant representations of Korean text: a document written in Hanja and its Hangul transcription land at the same point in embedding space. The recipe (jamo and character-level embedding fusion, morpheme-aware denoising, bidirectional Hanja-Hangul transcription) is described in the paper, accepted to Findings of EMNLP 2026.

This repo holds the PyTorch weights, the tokenizer with its chat template, and the modeling code needed to load them through the transformers Auto classes with trust_remote_code=True. Training code, the Flax model, the Flax-to-PyTorch converter, and the evaluation pipeline live in the GitHub repo.

Intended use

Han2Han is meant for representations: classification and fine-tuning on downstream tasks, and sentence embeddings that treat Hanja and Hangul spellings of the same text alike. The model generates text, but generation is not what it was built or evaluated for, and improving it is follow-up work:

  • Prompt layout. The task prompt goes in the system message and the text alone in the user message, which is how the instruction tuning presented every task. With the instruction written into the user message instead, Hangul to Hanja restoration returns the Hangul input unchanged.
  • Hangul to Hanja restoration (ํ•œ๊ธ€์„ ํ•œ์ž๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:) works on article-length input. Given the first 478 characters of a 1922 newspaper article in Hangul, greedy decoding restores 195 of its 251 Hanja (78%). The article comes from a corpus used in pre-training, so this is an illustration and not a held-out measurement. On single sentences the model mostly returns Hangul: across 256 newspaper sentences it restores 24% of the Hanja, and 145 of the outputs contain no Hanja at all.
  • Hanja to Hangul transcription (ํ•œ์ž๋ฅผ ํ•œ๊ธ€๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:) works on sentences and articles, in either layout, but can leave some Hanja untranscribed.
  • Stopping. Give generate() a token limit: twice the input's tokens plus 8 leaves room for a correct answer. In fp32 the outputs measured here end on their own; in bf16 the article restoration ran on to the limit, repeating itself.

hanja_transcription_demo.ipynb in the GitHub repo was written with the other layout (the instruction in the user message), so the Hangul to Hanja outputs in its sections 4 to 9 understate what this checkpoint does. Its section 10 runs the layout described here.

Usage

Runtime requirements: torch and transformers (tested with torch 2.14.1 and transformers 5.18.0 on CPU). The model class ships in this repo, so loading needs trust_remote_code=True; the tokenizer itself runs no custom code, but without the flag AutoTokenizer stops to ask.

Sentence embeddings

output_sentence_embeddings=True returns the mean-pooled encoder states, one 640-dimensional vector per input.

import torch
import torch.nn.functional as F
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

repo = "cadazar/han2han-it"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()

texts = [
    "ๆœƒๅ ด์„ ไธ€ๅทกํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ๆ—ฅๆœฌไบบๅด ็•ตๅฎถ๊นŒ์ง€๋„ ไธ€ไบบ๋„ ็™ผ่ฆ‹ํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.",
    "ํšŒ์žฅ์„ ์ผ์ˆœํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ์ผ๋ณธ์ธ์ธก ํ™”๊ฐ€๊นŒ์ง€๋„ ์ผ์ธ๋„ ๋ฐœ๊ฒฌํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.",
    "ๅ—็•ต๋‚˜ ๅ››ๅ›ๅญ์—์„œ๋Š” ็ ดๅขจ์˜ ๅฆ™ๆณ•์œผ๋กœ ็™ฝ้›ช์„ ่ฑกๅพตํ•  ์ˆ˜ ์žˆ๋‹ค.",
]
inputs = tokenizer(texts, return_tensors="pt", padding=True)
with torch.no_grad():
    embeddings = model(**inputs, output_sentence_embeddings=True)[0]
embeddings = F.normalize(embeddings, dim=-1)
print(embeddings @ embeddings.T)
# tensor([[1.0000, 0.9444, 0.8789],
#         [0.9444, 1.0000, 0.8459],
#         [0.8789, 0.8459, 1.0000]])

The first two rows are the same 1920s newspaper sentence in mixed script and in Hangul; the third is a different sentence.

Chat

AutoModelForCausalLM loads Han2HanForCausalLM, a wrapper that gives the encoder-decoder the interface chat tooling expects: generate() takes the rendered chat prompt as one sequence and returns the prompt followed by the reply. pipeline("text-generation") works with it, and so does the transformers server:

transformers serve --trust-remote-code
transformers chat cadazar/han2han-it --system-prompt "ํ•œ์ž๋ฅผ ํ•œ๊ธ€๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:"
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "cadazar/han2han-it"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).eval()


def run(task_prompt, text):
    messages = [
        {"role": "system", "content": task_prompt},
        {"role": "user", "content": text},
    ]
    inputs = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
    )
    budget = 2 * len(tokenizer(text).input_ids) + 8
    with torch.no_grad():
        output = model.generate(**inputs, max_new_tokens=budget, do_sample=False)
    return tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)


print(run(
    "ํ•œ์ž๋ฅผ ํ•œ๊ธ€๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:",
    "ๆœƒๅ ด์„ ไธ€ๅทกํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ๆ—ฅๆœฌไบบๅด ็•ตๅฎถ๊นŒ์ง€๋„ ไธ€ไบบ๋„ ็™ผ่ฆ‹ํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.",
))
# ํšŒ์žฅ์„ ไธ€ๅทกํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ์ผ๋ณธ์ธ์ธก ํ™”๊ฐ€๊นŒ์ง€๋„ ์ผ์ธ๋„ ๋ฐœ๊ฒฌํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.

hangul = (
    "๋ฏธ์ˆ ์ „๋žŒ ๊ณ„ํš ์ด๋…๋ถ€ ์‚ฌ์•ˆ์œผ๋กœ ๋ช…๋…„ ์ดˆ์ฐจ ๊ฐœ์„ค์ด๋…๋ถ€์—์„œ๋Š” ์กฐ์„ ์— ์žฌ ํ•œ ๋ฏธ์ˆ ์˜ ๋ฐœ๋‹ฌ์„ ๋น„๋ณดํ•  "
    "๋ชฉ์ ์œผ๋กœ ๋™๊ฒฝ์˜ ์ œ๊ตญ๋ฏธ์ˆ ์›์ „๋žŒํšŒ๋ฅผ ๋ฐฉํ•˜ ์—ฌ ๋งค๋…„ ์ผ์ฐจ ๋ฏธ์ˆ ์ „๋žŒํšŒ๋ฅผ ๊ฐœํ•  ๋ฐฉ์นจ์„ ๋‚ด์ •ํ•˜๊ณ  ์‹œ "
    "์ด์‹ญ์œก์ผ ์˜ค์ „ ์‹ญ์‹œ ์ด๋…๋ถ€ ์ œ์ดํšŒ์˜์‹ค์— ๋ฐ•์˜ํšจ ํ›„, ๋ฏผ๋ณ‘์„ ์ž, ์„œํ™”ํ˜‘ํšŒ์˜ ์ •๋Œ€์œ , ๊น€๋ˆํฌ, ์ด๋„์˜ "
    "์™ธ ์ œ์”จ, ์„œํ™”์—ฐ๊ตฌํšŒ์˜ ๊น€๊ทœ์ง„ ์”จ, ์ผ๋ณธ์ธ์ธก ์„œํ™”๊ฐ€๋กœ  ๊ณ ๋ชฉ๋ฐฐ์ˆ˜ ์™ธ ์ˆ˜์”จ ๊ธฐํƒ€ ์„œํ™”๊ฐ€์— ๊ด€๊ณ„์žˆ๋Š” "
    "์ธ์‚ฌ๋ฅผ ์ดˆ์ฒญํ•˜๊ณ  ์ˆ˜์•ผ ์ •๋ฌด์ด๊ฐ ์ดํ•˜ ํ•™๋ฌด๋‹น๊ตญ์ž๊ฐ€ ํšŒํ•ฉํ•˜์—ฌ ์ฐจ์— ๊ด€ํ•œ ์ƒ์˜ํšŒ๋ฅผ ๊ฐœํ•œ๋ฐ”"
)
print(run("ํ•œ๊ธ€์„ ํ•œ์ž๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:", hangul))
# ็พŽ่ก“่ฆฝ ่จˆๅŠƒ ็ธฝ็ฃๅบœ ไบ‹้ …์œผ๋กœ ๆ˜Žๅนด ๅˆๆฌก ้–‹่จญ็ธฝ็ฃๅบœ์—์„œ๋Š” ๆœ้ฎฎ์— ๅœจ ํ•œ ็พŽ่ก“์˜ ็™ผ้”์„ ๅ‚™ๅ ฑํ• 
# ็ˆฒํ•˜์•ผ ๆฑไบฌ์˜ ๅธๅœ‹็พŽ่ก“ๅœ’ๅฑ•่ฆฝๆœƒ๋ฅผ ่จชํ•˜ ์—ฌ ๆฏๅนด ไธ€ๆฌก ็พŽ่ก“๋ฅผ ้–‹ํ•  ๆ–น้‡์„ ๅ…งๅฎšํ•˜๊ณ  ๆ™‚
# ไบŒๅๅ…ญๆ—ฅ ๅˆๅ‰ ๅๆ™‚ ็ธฝ็ฃๅบœ ็ฌฌไบŒๅ›ž ่ญฐๅฎค์— ๆœดๆฐธๅญ ๅพŒ, ้–”ไธ™้Œซ ่€…, ๆ›ธ็•ตๅ”ๆœƒ์˜ ๆ”ฟๅคง่ฃ•, ้‡‘ๆ•ฆ็†™,
# ๆŽ้“่‹ฑ ๅค– ่ซธๆฐ, ๆ›ธ็•ต็ก็ฉถๆœƒ์˜ ้‡‘ๅฅŽ้Žญ ๆฐ, ๆ—ฅๆœฌไบบๅด ๆ›ธ็•ตๅฎถ๋กœ ้ซ˜ๆœจๅŸน ๅค– ์ˆ˜์”จ ๅ…ถไป– ๆ›ธ็•ตๅฎถ์—
# ้—œไฟ‚์žˆ๋Š” ไบบไบ‹๋ฅผ ๆ‹›่ซ‹ํ•˜๊ณ  ๅฃฝ้‡Ž ๆ”ฟๅ‹™็ธฝ็›ฃ ไปฅไธ‹ ๅญธๅ‹™็•ถๅฑ€่€…๊ฐ€ ๆœƒๅˆํ•˜์—ฌ ์ฐจ์— ้—œํ•œ ๅ•†ๆœƒ๋ฅผ ้–‹ํ•œ๋ฐ”

The outputs are shown as generated (each is one line). In the first, ไธ€ๅทก should have become ์ผ์ˆœ. The second is the opening of the 1922 article; the original reads ็พŽ่ก“ๅฑ•่ฆฝ ่จˆๅŠƒ ็ธฝ็ฃๅบœ ไบ‹ๆกˆ์œผ๋กœ ๆ˜Žๅนด ๅˆๆฌก ้–‹่จญ็ธฝ็ฃๅบœ์—์„œ๋Š” ๆœ้ฎฎ์— ๅœจ ํ•œ ็พŽ่ก“์˜ ็™ผ้”์„ ่ฃจ่ฃœํ•  ๋ชฉ์ ์œผ๋กœ ๆฑไบฌ์˜ ๅธๅœ‹็พŽ่ก“้™ขๅฑ•่ฆฝๆœƒ๋ฅผ ๅ€ฃํ•˜ ์—ฌ ..., so the output drops or swaps some characters (็พŽ่ก“่ฆฝ, ไบ‹้ …, ๅธๅœ‹็พŽ่ก“ๅœ’) and writes same-sounding ones for several names.

Prompt format, as used by the SFT collator in the GitHub repo and by the chat template:

  • <|system|>{system}<|user|>{user}<|end_of_turn|>, with <|assistant|>{reply}<|end_of_turn|><|user|>{user}<|end_of_turn|> for each further turn. The system part is optional in the format, but it is where this checkpoint expects the task prompt. Each direction was trained with three wordings: ํ•œ์ž๋ฅผ ํ•œ๊ธ€๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:, ๋‹ค์Œ ํ•œ๋ฌธ์„ ํ•œ๊ธ€๋กœ ์˜ฎ๊ธฐ์‹œ์˜ค:, ํ•œ์ž ํ‘œ๊ธฐ๋ฅผ ํ•œ๊ธ€ ๋…์Œ์œผ๋กœ ๋ณ€ํ™˜ํ•˜์‹œ์˜ค: for Hanja to Hangul, and ํ•œ๊ธ€์„ ํ•œ์ž๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:, ๋‹ค์Œ ํ•œ๊ธ€ ํ…์ŠคํŠธ์— ์ ์ ˆํ•œ ํ•œ์ž๋ฅผ ์ถ”๊ฐ€ํ•˜์‹œ์˜ค:, ํ•œ์ž ํ‘œ๊ธฐ๊ฐ€ ํ•„์š”ํ•œ ๋ถ€๋ถ„์— ํ•œ์ž๋ฅผ ๋ณ‘๊ธฐํ•˜์‹œ์˜ค: for Hangul to Hanja.
  • The generation prompt is <|assistant|>, or <|think|> with enable_thinking=True, in which case the reply is reasoning, then <|assistant|>, then the answer. A turn ends with <|end_of_turn|> (eos_token_id 10).

The wrapper splits the prompt at its last <|end_of_turn|>: everything up to and including it is the encoder input, and the generation prompt after it starts the decoder. A prompt without a generation prompt raises. forward() is not wrapped and stays encoder-decoder.

Encoder-decoder generation

AutoModelForSeq2SeqLM loads the model itself. generate() is the standard transformers one, with a KV cache; greedy decoding, beam search, and sampling all work. The encoder takes the prompt up to <|end_of_turn|>, and the decoder starts from <|assistant|> (decoder_start_token_id 9).

from transformers import AutoModelForSeq2SeqLM

model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
prompt = (
    "<|system|>ํ•œ์ž๋ฅผ ํ•œ๊ธ€๋กœ ์ „์‚ฌํ•˜์‹œ์˜ค:"
    "<|user|>ๆœƒๅ ด์„ ไธ€ๅทกํ•˜๊ณ  ๋Œ์•„์˜ฌ ๋•Œ๊นŒ์ง€๋„ ๆ—ฅๆœฌไบบๅด ็•ตๅฎถ๊นŒ์ง€๋„ ไธ€ไบบ๋„ ็™ผ่ฆ‹ํ•  ์ˆ˜๊ฐ€ ์—†์—ˆ๋‹ค.<|end_of_turn|>"
)
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
    output = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Tokenizer

tokenizer.json is a tokenizers-library build of the SentencePiece model (spiece.model, kept here as the source), made by scripts/build_hf_tokenizer.py in the GitHub repo. Calling the tokenizer does not add BOS or EOS tokens, and special tokens written into the text are mapped to their ids. It encodes like the SentencePiece wrapper the training code uses, with one known difference: where two segmentations of a span have exactly the same score (runs of digits, mostly), the same pieces can come out in a different order.

Files

File Contents
model.safetensors fp32 weights, 169.2M parameters, plus the jbu / cbu subword bucket tables
config.json, generation_config.json model and generation config, with auto_map entries for the Auto classes
tokenizer.json, tokenizer_config.json, chat_template.jinja fast tokenizer (38400 pieces) and chat template
spiece.model the SentencePiece model tokenizer.json was built from
modeling_han2han.py, han2han_config.py modeling code from the GitHub repo at commit af1330e

modeling_han2han.py (modeling_han2han_pytorch.py there) differs in two places: the config import is relative, and the module-level register_han2han import is replaced by the auto_map entries. han2han_config.py is the GitHub copy at commit 7304c38; it has not changed since. config.json names AutoModelForCausalLM under architectures so that transformers serve, which looks that name up in the transformers namespace, can load the model.

Instruction tuning

Fine-tuned from the Han2Han pre-trained checkpoint on instruction following, chain-of-thought reasoning, Hanja-Hangul article transcription, and summarization data (configs/it-muon-stage_1.yaml in the GitHub repo).

Citation

@inproceedings{han2han2026,
  title     = {Han2Han: Efficient Language-Specific Character Representation
               through Script-Aware Pre-Training for Historical Text Analysis},
  author    = {Adams, Cellik and Jo, EunKyoung and Kim, Juae},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026}
}

License

Apache License 2.0, the same as the GitHub repo.

Downloads last month
695
Safetensors
Model size
0.2B params
Tensor type
F32
ยท
I32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support