# data/jlpt - JLPT vocabulary levels and kanji levels, pinned Generated by `scripts/build_jlpt.py`. **Do not hand-edit these files**: `tests/test_data_assets.py` compares every file with the SHA-256 below and pins the counts, so an edit fails the quick loop until this file is regenerated. **Resolved:** 2026-09-06 **Pins:** stephenmk/yomitan-jlpt-vocab release tag `2025.08.01.0` (vocabulary); davidluzgouveia/kanji-data commit `00fd7079c3890f430759536f91aa5e854ec0ca4f` (kanji levels) Regenerate (from the repo root; re-running is idempotent - byte-identical files): ``` .venv/Scripts/python.exe scripts/build_jlpt.py ``` ## Files | File | Bytes | SHA256 | Source | |---|---|---|---| | `n5.csv` | 24,010 | `c370961f92087e15d059548e1b3d01bab714306c739b3f3caa0adfa921d679ed` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n5.csv | | `n4.csv` | 24,772 | `a14dcc7fdc02259b22331a486c6ff66df74c16472c8aeb5a0ebe7e5fa8ee8eb4` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n4.csv | | `n3.csv` | 88,623 | `bd4d68c59cfee861351e1bdf9d99818e434392201eabbed4d3657225fe53a3ac` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n3.csv | | `n2.csv` | 95,356 | `42a4413326e857d701dc9659b2711637ab612304a339fe0d09f4f0f042d3a214` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n2.csv | | `n1.csv` | 199,103 | `7a58f0584e9ec2b0299ccb9b109f3bd03e08b90d129a714307c0a1cb72f07b32` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n1.csv | | `kanji_levels.json` | 28,746 | `de0761e2b218a81c2e58950fc89b56da9b924e056285befaf3d1d1becef7a4c4` | `jlpt_new` of https://raw.githubusercontent.com/davidluzgouveia/kanji-data/00fd7079c3890f430759536f91aa5e854ec0ca4f/kanji.json | | `LICENSE-yomitan-jlpt-vocab.txt` | 20,131 | `7abe19ec9bb73b36141b999b861d24ad855e808bafe0f81e84cce28556f6c297` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/LICENSE.txt | | `LICENSE-kanji-data.txt` | 1,070 | `4c4fe92a5adf9df114ca9f04e50329d4565c3ce511660e934cb01eb110e2ca6e` | https://raw.githubusercontent.com/davidluzgouveia/kanji-data/00fd7079c3890f430759536f91aa5e854ec0ca4f/LICENSE | The CSVs are the upstream files byte-for-byte (columns `jmdict_seq,kana,kanji,` `waller_definition`; `jmdict_seq` is the JMdict entry id, so word level is a join on id, not a string match). `kanji_levels.json` is this script's projection of the `jlpt_new` field of `kanji.json`: `{literal: N5..N1}` for every kanji that has a level, sorted by literal. `jlpt_new` is the N1-N5 scale; the file's `jlpt_old` is KANJIDIC2's pre-2010 1-4 scale and is NOT used (02-RESEARCH.md § Q2 - Kanji list). ## Counts | Measure | Value | |---|---| | `n5.csv` data rows | 684 | | `n4.csv` data rows | 640 | | `n3.csv` data rows | 1,730 | | `n2.csv` data rows | 1,812 | | `n1.csv` data rows | 3,427 | | Total data rows | 8,293 | | Rows with an empty `jmdict_seq` (all in `n1.csv`) | 14 | | Unique `jmdict_seq` ids | 7,748 | | Ids appearing in more than one row | 505 | | Ids appearing on more than one level | 447 | | Ids listed twice within one level (two readings) | 71 | | Kanji at N5 | 79 | | Kanji at N4 | 166 | | Kanji at N3 | 367 | | Kanji at N2 | 367 | | Kanji at N1 | 1,232 | | Kanji with a level | 2,211 | A word on two lists takes the easiest level it appears at (research § Q2 - Join strategy); a word on no list is `N1+`; a kanji not in the map is unlisted and above every level (D-02, D-10). ## Licence JLPT levels: Jonathan Waller's JLPT Resources (https://www.tanos.co.uk/jlpt/, CC BY) via stephenmk/yomitan-jlpt-vocab (CC BY-SA 4.0); kanji levels extracted from davidluzgouveia/kanji-data (MIT) which took its levels from the same Waller lists. There is no official JLPT vocabulary list; these are estimates. - `LICENSE-yomitan-jlpt-vocab.txt` is the repository's `LICENSE.txt` at the tag (CC BY-SA 4.0). Any redistributed derivative of the CSVs stays CC BY-SA. - `LICENSE-kanji-data.txt` is the repository's `LICENSE` at the commit (MIT); the notice is kept beside the extracted file as MIT requires. - Jonathan Waller's terms (https://www.tanos.co.uk/jlpt/sharing/): use however you like, credit the site. `LICENSES.md` at the repo root is the project-wide record and the page's credits line carries the attribution (plan 02-11).