Spaces:
Running on Zero
Running on Zero
|
Download data/jlpt/README.md from WolfDavid/japanese-learning-avatar: direct link, hf CLI and curl.
- Browser
- Download file 4.45 kB
-
https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/data/jlpt/README.md
- Command line
-
hf download hf://spaces/WolfDavid/japanese-learning-avatar@706670a3ece0683b9770f0b7db36e4daa955f91e/data/jlpt/README.md
-
curl -L -o README.md https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/data/jlpt/README.md
4.45 kB
| # data/jlpt - JLPT vocabulary levels and kanji levels, pinned | |
| Generated by `scripts/build_jlpt.py`. **Do not hand-edit these files**: | |
| `tests/test_data_assets.py` compares every file with the SHA-256 below and pins the | |
| counts, so an edit fails the quick loop until this file is regenerated. | |
| **Resolved:** 2026-09-06 | |
| **Pins:** stephenmk/yomitan-jlpt-vocab release tag `2025.08.01.0` (vocabulary); davidluzgouveia/kanji-data commit `00fd7079c3890f430759536f91aa5e854ec0ca4f` (kanji levels) | |
| Regenerate (from the repo root; re-running is idempotent - byte-identical files): | |
| ``` | |
| .venv/Scripts/python.exe scripts/build_jlpt.py | |
| ``` | |
| ## Files | |
| | File | Bytes | SHA256 | Source | | |
| |---|---|---|---| | |
| | `n5.csv` | 24,010 | `c370961f92087e15d059548e1b3d01bab714306c739b3f3caa0adfa921d679ed` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n5.csv | | |
| | `n4.csv` | 24,772 | `a14dcc7fdc02259b22331a486c6ff66df74c16472c8aeb5a0ebe7e5fa8ee8eb4` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n4.csv | | |
| | `n3.csv` | 88,623 | `bd4d68c59cfee861351e1bdf9d99818e434392201eabbed4d3657225fe53a3ac` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n3.csv | | |
| | `n2.csv` | 95,356 | `42a4413326e857d701dc9659b2711637ab612304a339fe0d09f4f0f042d3a214` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n2.csv | | |
| | `n1.csv` | 199,103 | `7a58f0584e9ec2b0299ccb9b109f3bd03e08b90d129a714307c0a1cb72f07b32` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/original_data/n1.csv | | |
| | `kanji_levels.json` | 28,746 | `de0761e2b218a81c2e58950fc89b56da9b924e056285befaf3d1d1becef7a4c4` | `jlpt_new` of https://raw.githubusercontent.com/davidluzgouveia/kanji-data/00fd7079c3890f430759536f91aa5e854ec0ca4f/kanji.json | | |
| | `LICENSE-yomitan-jlpt-vocab.txt` | 20,131 | `7abe19ec9bb73b36141b999b861d24ad855e808bafe0f81e84cce28556f6c297` | https://raw.githubusercontent.com/stephenmk/yomitan-jlpt-vocab/2025.08.01.0/LICENSE.txt | | |
| | `LICENSE-kanji-data.txt` | 1,070 | `4c4fe92a5adf9df114ca9f04e50329d4565c3ce511660e934cb01eb110e2ca6e` | https://raw.githubusercontent.com/davidluzgouveia/kanji-data/00fd7079c3890f430759536f91aa5e854ec0ca4f/LICENSE | | |
| The CSVs are the upstream files byte-for-byte (columns `jmdict_seq,kana,kanji,` | |
| `waller_definition`; `jmdict_seq` is the JMdict entry id, so word level is a join on | |
| id, not a string match). `kanji_levels.json` is this script's projection of | |
| the `jlpt_new` field of `kanji.json`: `{literal: N5..N1}` for every kanji that has a | |
| level, sorted by literal. `jlpt_new` is the N1-N5 scale; the file's `jlpt_old` is | |
| KANJIDIC2's pre-2010 1-4 scale and is NOT used (02-RESEARCH.md § Q2 - Kanji list). | |
| ## Counts | |
| | Measure | Value | | |
| |---|---| | |
| | `n5.csv` data rows | 684 | | |
| | `n4.csv` data rows | 640 | | |
| | `n3.csv` data rows | 1,730 | | |
| | `n2.csv` data rows | 1,812 | | |
| | `n1.csv` data rows | 3,427 | | |
| | Total data rows | 8,293 | | |
| | Rows with an empty `jmdict_seq` (all in `n1.csv`) | 14 | | |
| | Unique `jmdict_seq` ids | 7,748 | | |
| | Ids appearing in more than one row | 505 | | |
| | Ids appearing on more than one level | 447 | | |
| | Ids listed twice within one level (two readings) | 71 | | |
| | Kanji at N5 | 79 | | |
| | Kanji at N4 | 166 | | |
| | Kanji at N3 | 367 | | |
| | Kanji at N2 | 367 | | |
| | Kanji at N1 | 1,232 | | |
| | Kanji with a level | 2,211 | | |
| A word on two lists takes the easiest level it appears at (research § Q2 - Join | |
| strategy); a word on no list is `N1+`; a kanji not in the map is unlisted and above | |
| every level (D-02, D-10). | |
| ## Licence | |
| JLPT levels: Jonathan Waller's JLPT Resources (https://www.tanos.co.uk/jlpt/, CC BY) via stephenmk/yomitan-jlpt-vocab (CC BY-SA 4.0); kanji levels extracted from davidluzgouveia/kanji-data (MIT) which took its levels from the same Waller lists. There is no official JLPT vocabulary list; these are estimates. | |
| - `LICENSE-yomitan-jlpt-vocab.txt` is the repository's `LICENSE.txt` at the tag (CC BY-SA | |
| 4.0). Any redistributed derivative of the CSVs stays CC BY-SA. | |
| - `LICENSE-kanji-data.txt` is the repository's `LICENSE` at the commit (MIT); the notice | |
| is kept beside the extracted file as MIT requires. | |
| - Jonathan Waller's terms (https://www.tanos.co.uk/jlpt/sharing/): use however you like, | |
| credit the site. `LICENSES.md` at the repo root is the project-wide record and the | |
| page's credits line carries the attribution (plan 02-11). | |