suhjae's picture
Upload promoted Level 2 huneum selector
faa94f5 verified
|
Raw History Blame Contribute Delete
1.65 kB
metadata
language:
  - ko
  - zh
language_bcp47:
  - zh-Hant
license: other
library_name: pytorch
tags:
  - joseon
  - hanmun
  - hanja
  - huneum
  - lexicalization
  - korean-history
datasets:
  - custom
metrics:
  - accuracy

Joseon Level 2 Huneum Selector

This is the promoted Level 2 훈음 selector for the Joseon-to-Day project. It selects Korean readings for Hanja spans using structured candidate sets.

This repository contains a project-specific PyTorch artifact rather than a standard Transformers model:

  • huneum_selector.pt: model weights
  • model_config.json: model/vocabulary configuration
  • candidate_manifest.json: candidate metadata
  • metrics.json: promoted evaluation metrics
  • huneum_error_analysis.json: diagnostic errors

Evaluation

Promoted dev result:

Metric Value
exact row accuracy 0.9953
correct rows 5054 / 5078
character accuracy 0.9981
selector accuracy 0.9949
deterministic single-candidate accuracy 1.0000

By label:

Label Exact
book/title/evidence 1.0000
person/name 0.9953
place 0.9948

Intended Use

Use this model after Level 1 span detection to choose readings for Korean historical names, places, book titles, and related Hanja/Hanmun spans.

Limitations

The artifact is tied to the Joseon-to-Day project code and candidate format. It is not a standalone general Hanja-to-Korean transliteration package.

Loading

Use the project loader/training code from the Joseon-to-Day repository. The artifact is intentionally published with its config and manifest so the project pipeline can reproduce the promoted Level 2 selector.