WolfDavid's picture
feat(02-01): pinned JLPT vocabulary and kanji-level data with build script
ac406ae
|
Raw History Blame
4.45 kB
# data/jlpt - JLPT vocabulary levels and kanji levels, pinned
Generated by `scripts/build_jlpt.py`. **Do not hand-edit these files**:
`tests/test_data_assets.py` compares every file with the SHA-256 below and pins the
counts, so an edit fails the quick loop until this file is regenerated.
**Resolved:** 2026-09-06
**Pins:** stephenmk/yomitan-jlpt-vocab release tag `2025.08.01.0` (vocabulary); davidluzgouveia/kanji-data commit `00fd7079c3890f430759536f91aa5e854ec0ca4f` (kanji levels)
Regenerate (from the repo root; re-running is idempotent - byte-identical files):
```
.venv/Scripts/python.exe scripts/build_jlpt.py
```
## Files
| File | Bytes | SHA256 | Source |
|---|---|---|---|
| `n5.csv` | 24,010 | `c370961f92087e15d059548e1b3d01bab714306c739b3f3caa0adfa921d679ed` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n5.csv |
| `n4.csv` | 24,772 | `a14dcc7fdc02259b22331a486c6ff66df74c16472c8aeb5a0ebe7e5fa8ee8eb4` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n4.csv |
| `n3.csv` | 88,623 | `bd4d68c59cfee861351e1bdf9d99818e434392201eabbed4d3657225fe53a3ac` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n3.csv |
| `n2.csv` | 95,356 | `42a4413326e857d701dc9659b2711637ab612304a339fe0d09f4f0f042d3a214` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n2.csv |
| `n1.csv` | 199,103 | `7a58f0584e9ec2b0299ccb9b109f3bd03e08b90d129a714307c0a1cb72f07b32` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n1.csv |
| `kanji_levels.json` | 28,746 | `de0761e2b218a81c2e58950fc89b56da9b924e056285befaf3d1d1becef7a4c4` | `jlpt_new` of https://raw.githubusercontent.com/davidluzgouveia/kanji-data/00fd7079c3890f430759536f91aa5e854ec0ca4f/kanji.json |
| `LICENSE-yomitan-jlpt-vocab.txt` | 20,131 | `7abe19ec9bb73b36141b999b861d24ad855e808bafe0f81e84cce28556f6c297` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/LICENSE.txt |
| `LICENSE-kanji-data.txt` | 1,070 | `4c4fe92a5adf9df114ca9f04e50329d4565c3ce511660e934cb01eb110e2ca6e` | https://raw.githubusercontent.com/davidluzgouveia/kanji-data/00fd7079c3890f430759536f91aa5e854ec0ca4f/LICENSE |
The CSVs are the upstream files byte-for-byte (columns `jmdict_seq,kana,kanji,`
`waller_definition`; `jmdict_seq` is the JMdict entry id, so word level is a join on
id, not a string match). `kanji_levels.json` is this script's projection of
the `jlpt_new` field of `kanji.json`: `{literal: N5..N1}` for every kanji that has a
level, sorted by literal. `jlpt_new` is the N1-N5 scale; the file's `jlpt_old` is
KANJIDIC2's pre-2010 1-4 scale and is NOT used (02-RESEARCH.md § Q2 - Kanji list).
## Counts
| Measure | Value |
|---|---|
| `n5.csv` data rows | 684 |
| `n4.csv` data rows | 640 |
| `n3.csv` data rows | 1,730 |
| `n2.csv` data rows | 1,812 |
| `n1.csv` data rows | 3,427 |
| Total data rows | 8,293 |
| Rows with an empty `jmdict_seq` (all in `n1.csv`) | 14 |
| Unique `jmdict_seq` ids | 7,748 |
| Ids appearing in more than one row | 505 |
| Ids appearing on more than one level | 447 |
| Ids listed twice within one level (two readings) | 71 |
| Kanji at N5 | 79 |
| Kanji at N4 | 166 |
| Kanji at N3 | 367 |
| Kanji at N2 | 367 |
| Kanji at N1 | 1,232 |
| Kanji with a level | 2,211 |
A word on two lists takes the easiest level it appears at (research § Q2 - Join
strategy); a word on no list is `N1+`; a kanji not in the map is unlisted and above
every level (D-02, D-10).
## Licence
JLPT levels: Jonathan Waller's JLPT Resources (https://www.tanos.co.uk/jlpt/, CC BY) via stephenmk/yomitan-jlpt-vocab (CC BY-SA 4.0); kanji levels extracted from davidluzgouveia/kanji-data (MIT) which took its levels from the same Waller lists. There is no official JLPT vocabulary list; these are estimates.
- `LICENSE-yomitan-jlpt-vocab.txt` is the repository's `LICENSE.txt` at the tag (CC BY-SA
4.0). Any redistributed derivative of the CSVs stays CC BY-SA.
- `LICENSE-kanji-data.txt` is the repository's `LICENSE` at the commit (MIT); the notice
is kept beside the extracted file as MIT requires.
- Jonathan Waller's terms (https://www.tanos.co.uk/jlpt/sharing/): use however you like,
credit the site. `LICENSES.md` at the repo root is the project-wide record and the
page's credits line carries the attribution (plan 02-11).