Spaces:
Running on Zero
Download data/jmdict/README.md from WolfDavid/japanese-learning-avatar: direct link, hf CLI and curl.
- Browser
- Download file 3.01 kB
-
https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/data/jmdict/README.md
- Command line
-
hf download hf://spaces/WolfDavid/japanese-learning-avatar@706670a3ece0683b9770f0b7db36e4daa955f91e/data/jmdict/README.md
-
curl -L -o README.md https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/data/jmdict/README.md
data/jmdict - the compact JMdict projection, pinned
Generated by scripts/build_jmdict.py. Do not hand-edit or re-compress this file:
tests/test_data_assets.py compares it with the SHA-256 below and pins the entry count,
so any change fails the quick loop until this file is regenerated. The file is a Git
LFS object (*.gz in .gitattributes); if it starts with version https://git-lfs
instead of the gzip magic, run git lfs pull.
Resolved: 2026-09-06
Pins: scriptin/jmdict-simplified release 3.6.2+20260831182826 (JSON version 3.6.2, dictDate 2026-08-31), 218,672 words
Input: https://github.com/scriptin/jmdict-simplified/releases/download/3.6.2%2B20260831182826/jmdict-eng-3.6.2%2B20260831182826.json.tgz
Input tgz: 11,510,336 bytes, SHA-256 3c842741e2c4f1b780ad4aff6833df308866a8b7dfb9ee59a085e23dd201a65f
Regenerate (from the repo root; re-running is idempotent - byte-identical output, the gzip header carries no timestamp or name):
.venv/Scripts/python.exe scripts/build_jmdict.py
Files
| File | Bytes | SHA256 | Source |
|---|---|---|---|
jmdict-compact.json.gz |
7,662,785 | 073ee1b823485b1e08eb99d7e088d74c46b01470297e20ccf289e4cbac3ceeab |
projection of the tgz above |
Raw JSON 21,751,814 bytes -> gzip level 9 7,662,785 bytes. Loading is one
gzip.open + json.load (~0.6-2 s depending on the machine; the build prints the
measured time and tests/test_data_assets.py's fixture prints it again) - kept out of
this file so a re-run stays byte-identical.
Compact record schema
{"meta": {"source", "version", "jmdictVersion", "dictDate", "entries", "licence"},
"entries": [[id, kanji_texts, kana_texts, senses, common], ...]}
id int JMdict entry sequence number (the `jmdict_seq` of data/jlpt/*.csv)
kanji_texts [str] up to 3 kanji forms, file order (empty for kana-only words)
kana_texts [str] up to 3 kana forms, file order
senses [[str]] up to 3 senses x up to 3 English glosses each;
senses with no English gloss are dropped
common 0|1 1 if any kanji or kana form is marked common
Entries are in JMdict file order; the index in entries is the order the ranked lookup
(plan 02-03) tie-breaks on. Counts in this build: 218,672 entries, 22,637 common, 41,216 kana-only, 0 with no English sense.
Licence
Dictionary data: JMdict - Copyright (c) James William Breen and The Electronic Dictionary Research and Development Group, used under the Creative Commons Attribution-ShareAlike Licence (V4.0). https://www.edrdg.org/wiki/index.php/JMdict-EDICT_Dictionary_Project . https://www.edrdg.org/edrdg/licence.html
JSON conversion by scriptin/jmdict-simplified, release 3.6.2+20260831182826; its derived files carry the EDRDG licence. The compact projection in this directory is a derivative of JMdict and is itself CC BY-SA 4.0.
LICENSES.md at the repo root is the project-wide record and the page's credits line
carries the attribution (plan 02-11).