Spaces:
Running on Zero
Download data/jlpt/README.md from WolfDavid/japanese-learning-avatar: direct link, hf CLI and curl.
- Browser
- Download file 4.45 kB
-
https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/data/jlpt/README.md
- Command line
-
hf download hf://spaces/WolfDavid/japanese-learning-avatar@706670a3ece0683b9770f0b7db36e4daa955f91e/data/jlpt/README.md
-
curl -L -o README.md https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/data/jlpt/README.md
data/jlpt - JLPT vocabulary levels and kanji levels, pinned
Generated by scripts/build_jlpt.py. Do not hand-edit these files:
tests/test_data_assets.py compares every file with the SHA-256 below and pins the
counts, so an edit fails the quick loop until this file is regenerated.
Resolved: 2026-09-06
Pins: stephenmk/yomitan-jlpt-vocab release tag 2025.08.01.0 (vocabulary); davidluzgouveia/kanji-data commit 00fd7079c3890f430759536f91aa5e854ec0ca4f (kanji levels)
Regenerate (from the repo root; re-running is idempotent - byte-identical files):
.venv/Scripts/python.exe scripts/build_jlpt.py
Files
The CSVs are the upstream files byte-for-byte (columns jmdict_seq,kana,kanji,
waller_definition; jmdict_seq is the JMdict entry id, so word level is a join on
id, not a string match). kanji_levels.json is this script's projection of
the jlpt_new field of kanji.json: {literal: N5..N1} for every kanji that has a
level, sorted by literal. jlpt_new is the N1-N5 scale; the file's jlpt_old is
KANJIDIC2's pre-2010 1-4 scale and is NOT used (02-RESEARCH.md § Q2 - Kanji list).
Counts
| Measure | Value |
|---|---|
n5.csv data rows |
684 |
n4.csv data rows |
640 |
n3.csv data rows |
1,730 |
n2.csv data rows |
1,812 |
n1.csv data rows |
3,427 |
| Total data rows | 8,293 |
Rows with an empty jmdict_seq (all in n1.csv) |
14 |
Unique jmdict_seq ids |
7,748 |
| Ids appearing in more than one row | 505 |
| Ids appearing on more than one level | 447 |
| Ids listed twice within one level (two readings) | 71 |
| Kanji at N5 | 79 |
| Kanji at N4 | 166 |
| Kanji at N3 | 367 |
| Kanji at N2 | 367 |
| Kanji at N1 | 1,232 |
| Kanji with a level | 2,211 |
A word on two lists takes the easiest level it appears at (research § Q2 - Join
strategy); a word on no list is N1+; a kanji not in the map is unlisted and above
every level (D-02, D-10).
Licence
JLPT levels: Jonathan Waller's JLPT Resources (https://www.tanos.co.uk/jlpt/, CC BY) via stephenmk/yomitan-jlpt-vocab (CC BY-SA 4.0); kanji levels extracted from davidluzgouveia/kanji-data (MIT) which took its levels from the same Waller lists. There is no official JLPT vocabulary list; these are estimates.
LICENSE-yomitan-jlpt-vocab.txtis the repository'sLICENSE.txtat the tag (CC BY-SA 4.0). Any redistributed derivative of the CSVs stays CC BY-SA.LICENSE-kanji-data.txtis the repository'sLICENSEat the commit (MIT); the notice is kept beside the extracted file as MIT requires.- Jonathan Waller's terms (https://www.tanos.co.uk/jlpt/sharing/): use however you like,
credit the site.
LICENSES.mdat the repo root is the project-wide record and the page's credits line carries the attribution (plan 02-11).