Spaces:
Running on Zero
Running on Zero
|
Download data/jmdict/README.md from WolfDavid/japanese-learning-avatar: direct link, hf CLI and curl.
- Browser
- Download file 3.01 kB
-
https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/data/jmdict/README.md
- Command line
-
hf download hf://spaces/WolfDavid/japanese-learning-avatar@706670a3ece0683b9770f0b7db36e4daa955f91e/data/jmdict/README.md
-
curl -L -o README.md https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/data/jmdict/README.md
3.01 kB
| # data/jmdict - the compact JMdict projection, pinned | |
| Generated by `scripts/build_jmdict.py`. **Do not hand-edit or re-compress this file**: | |
| `tests/test_data_assets.py` compares it with the SHA-256 below and pins the entry count, | |
| so any change fails the quick loop until this file is regenerated. The file is a Git | |
| LFS object (`*.gz` in `.gitattributes`); if it starts with `version https://git-lfs` | |
| instead of the gzip magic, run `git lfs pull`. | |
| **Resolved:** 2026-09-06 | |
| **Pins:** scriptin/jmdict-simplified release `3.6.2+20260831182826` (JSON `version` `3.6.2`, `dictDate` `2026-08-31`), 218,672 words | |
| **Input:** `https://github.com/scriptin/jmdict-simplified/releases/download/3.6.2%2B20260831182826/jmdict-eng-3.6.2%2B20260831182826.json.tgz` | |
| **Input tgz:** 11,510,336 bytes, SHA-256 `3c842741e2c4f1b780ad4aff6833df308866a8b7dfb9ee59a085e23dd201a65f` | |
| Regenerate (from the repo root; re-running is idempotent - byte-identical output, the | |
| gzip header carries no timestamp or name): | |
| ``` | |
| .venv/Scripts/python.exe scripts/build_jmdict.py | |
| ``` | |
| ## Files | |
| | File | Bytes | SHA256 | Source | | |
| |---|---|---|---| | |
| | `jmdict-compact.json.gz` | 7,662,785 | `073ee1b823485b1e08eb99d7e088d74c46b01470297e20ccf289e4cbac3ceeab` | projection of the tgz above | | |
| Raw JSON 21,751,814 bytes -> gzip level 9 7,662,785 bytes. Loading is one | |
| `gzip.open` + `json.load` (~0.6-2 s depending on the machine; the build prints the | |
| measured time and `tests/test_data_assets.py`'s fixture prints it again) - kept out of | |
| this file so a re-run stays byte-identical. | |
| ## Compact record schema | |
| ``` | |
| {"meta": {"source", "version", "jmdictVersion", "dictDate", "entries", "licence"}, | |
| "entries": [[id, kanji_texts, kana_texts, senses, common], ...]} | |
| id int JMdict entry sequence number (the `jmdict_seq` of data/jlpt/*.csv) | |
| kanji_texts [str] up to 3 kanji forms, file order (empty for kana-only words) | |
| kana_texts [str] up to 3 kana forms, file order | |
| senses [[str]] up to 3 senses x up to 3 English glosses each; | |
| senses with no English gloss are dropped | |
| common 0|1 1 if any kanji or kana form is marked common | |
| ``` | |
| Entries are in JMdict file order; the index in `entries` is the order the ranked lookup | |
| (plan 02-03) tie-breaks on. Counts in this build: 218,672 entries, 22,637 common, 41,216 kana-only, 0 with no English sense. | |
| ## Licence | |
| Dictionary data: **JMdict** - Copyright (c) James William Breen and The Electronic Dictionary Research and Development Group, used under the Creative Commons Attribution-ShareAlike Licence (V4.0). https://www.edrdg.org/wiki/index.php/JMdict-EDICT_Dictionary_Project . https://www.edrdg.org/edrdg/licence.html | |
| JSON conversion by scriptin/jmdict-simplified, release 3.6.2+20260831182826; its derived files carry the EDRDG licence. The compact projection in this directory is a derivative of JMdict and is itself CC BY-SA 4.0. | |
| `LICENSES.md` at the repo root is the project-wide record and the page's credits line | |
| carries the attribution (plan 02-11). | |