Irodori-TTS-v4.1-Small-Yomi
This checkpoint is a fine-tune of Irodori-TTS-v4.1-Small that reads kanji more accurately. It mainly reduces on'yomi/kun'yomi confusions (音訓の取り違え) and misreadings of less common words.
Only the text-encoding path was trained:
- the top 4 of the 25 layers of the shared ModernBERT encoder, and its final norm
- the text projector
text_norm
The caption projector was then refit to the updated encoder (see "Caption and emoji conditioning" below). The RF-DiT, the duration predictor, the speaker encoder and the codec are unchanged.
The architecture, the file format and the inference interface are the same as the base model. Voice cloning, voice design with captions, and emoji style control are used the same way.
Versions
| Version | Date | Commit | Aozora Bunko ruby | Wikipedia (held-out words) | Notes |
|---|---|---|---|---|---|
| v3 (current) | 2026-10-06 | – | 67.53% | 84.98% | more sentences per word; average of two checkpoints |
| v2 | 2026-10-03 | 0206045 |
65.80% | 85.12% | 60k more steps with the target-window loss weighting |
| v1 | 2026-10-02 | 9023f92 |
63.30% | 84.37% | first release |
The base model scores 62.50% and 84.04% on these two sets (see Evaluation).
✨ What changed
| Text (target word underlined) | Base model | This model |
|---|---|---|
| 我輩は屡々世界の人としての日本人の覚悟に関して述ぶるところがあった。 | ルルル | シバシバ |
| と、再びていねいに娘は礼を述べて、そして踵をめぐらした。 | キビツ | キビス |
| ひとりの男が、長い樋をつたって、だんだん下へおりてくるのです。 | ヒ | トイ |
| 爾来我が中村屋は三十余年を通じて、一回たりともコンミッションに悩まされたことはない。 | ジュライ | ジライ |
| 死後のことはいざ知らず、現世においては永劫浮かぶ瀬のない無間地獄というものはないはずです。 | ムケン | ムゲン |
| 異でもあり、妙でもあって、とても、市中の玩具屋を探して歩いてもある品でない。 | ガングヤ | オモチャヤ |
| 衣裳葛籠がある。 | ツズカゴ | ツズラ |
| 百舌鳥の声が喧しい程城内に交錯している。 | ヨモゾリ | モズ |
| 水は滾々として流れている。 | タンダン | コンコン |
| かれらは黙っていた、みんな槍をぴたりと脇につけ、足並を揃えて犇々と進んでいった。 | テエテエ | ヒシヒシ |
| 男の服装は半天に股引、顔は黒布で包んでいる。 | マトヒキ | モモヒキ |
These sentences come from Aozora Bunko works that were held out of training. The readings are transcribed from the generated audio by kana-whisper (target word only).
🚀 Usage
Use the inference code of Aratako/Irodori-TTS and pass this repository as the checkpoint:
uv run --no-sync python infer.py \
--hf-checkpoint takuma104/Irodori-TTS-v4.1-Small-Yomi \
--text "こんにちは、私はAIです。これは音声合成のテストです。" \
--ref-wav path/to/reference.wav \
--output-wav outputs/sample.wav
--no-ref, --caption, the Gradio apps and the other options work as described for
the base model.
🛠️ Training
The model was trained by kana-substitution self-distillation, using text only and no recorded speech:
- Teacher. The frozen base model reads a sentence correctly much more often when the target word is written in katakana. With katakana it generates the audio latents.
- Student. The text encoder sees the original kanji sentence. It is trained so that the frozen RF-DiT predicts the same velocity on the teacher's noised latents, and so that the duration predictor predicts the same length.
- Hard mining. The kana teacher is used only for sentences that the base model misreads in kanji. Sentences that it already reads correctly keep the base model's own output as the target.
- Preservation losses. The text states of the other tokens, and of general Wikipedia sentences, are kept close to the base encoder.
- Target-window weighting (v2). The velocity loss is averaged over all frames, which dilutes the target word in long corpus sentences. In v2, the frames around the target are weighted 5×: the target's position in the reading is mapped proportionally onto the frames, with a margin on both sides. With the same number of passes over the data, this fit the hard corpus sentences much faster.
- More sentences per word (v3). More steps on the same sentences only memorized them. For the words the base model misread, a local LLM wrote 6 new sentences each, and ruby sentences from 2,000 more Aozora Bunko works were added. For words that got new sentences, the base model misreads 576 held-out Aozora items; v3 reads 26.7% of them correctly, against 20.1% for v2. The continued run before averaging reached 34%.
- Contrast sentences (v3). Some kanji were learned as single-kanji words with context-specific readings (秋《とき》). Wikipedia sentences where the same kanji appears in common compounds with another reading (秋季) keep the base model's output, so that those compounds keep their reading.
- Checkpoint averaging (v3). Long continued training slowly shifted the readings of words that were never trained. The v3 weights are therefore the average of v2's fine-tuned weights and those of the run continued with the new data. The average keeps about half of the continued run's Aozora Bunko gain, and brings untrained words back close to v2's level.
The target words come from JMdict, JmdictFurigana and KANJIDIC2. Their example sentences come from three sources:
- sentences generated by a local LLM (Qwen3.6-27B)
- Japanese Wikipedia sentences whose reading UniDic, Sudachi and JMdict agree on, with homographs also checked by the LLM
- ruby readings from Aozora Bunko works
About 107k hard sentences carried the kana teacher, and 13k contrast sentences were used. Training used batch size 32 on one RTX 5090:
- v1: 90k steps
- v2: v1 plus 60k steps with the target-window weighting
- v3: the average of v2 and a run that continued v2 for 90k more steps on the new data
Code, data-building scripts and the stage-by-stage reports (in Japanese) are in takuma104/Irodori-TTS-Experiments.
📊 Evaluation
Settings: DiT in bf16, text encoder and codec in fp32, 40 RF steps, text CFG 3.0,
speaker CFG 5.0, reference voice jvs001 from the JVS corpus, seed 0. These differ from
the base model card (FP32, no reference, seeds 0–4), so the absolute numbers are not
comparable with it. Every item is compared with the base model under the same settings.
The numbers below are for this checkpoint file. Numerical precision alone moves them by a
few tenths of a point: the same weights evaluated before the export, with the
fine-tuned layers in fp32, gave 67.77% / 84.96% on the Aozora Bunko and Wikipedia sets.
"Fixed / broken" counts the items that changed from wrong to right and from right to
wrong; p is from an exact McNemar test.
Kanji readings
| Evaluation | Items | Base | This model | Fixed / broken | p |
|---|---|---|---|---|---|
| Aozora Bunko ruby (held-out works) | 3,000 | 62.50% | 67.53% | 200 / 49 | 8e-23 |
| Wikipedia sentences (held-out words) | 6,403 | 84.04% | 84.98% | 126 / 66 | 2e-5 |
| JKYB-Parakeet, Accuracy | 13,536 | 94.04% | 95.69% | 320 / 96 | 3e-29 |
The reading accuracy is the share of target words read correctly. Each item is
transcribed with kana-whisper and scored with jkyb-eval.
JKYB-Parakeet in more detail:
| Model | Accuracy ↑ | On'yomi ↑ | Kun'yomi ↑ | Appendix ↑ | Kana-CER ↓ | Sentence Kana-CER ↓ | Standard CER ↓ |
|---|---|---|---|---|---|---|---|
| Irodori-TTS-v4.1-Small | 94.04% | 92.98% | 95.87% | 83.87% | 5.51% | 0.90% | 3.93% |
| This model | 95.69% | 95.17% | 96.76% | 88.17% | 3.67% | 0.75% | 4.02% |
About JKYB. No JKYB-Parakeet sentence was used for training. Earlier stages of this fine-tune did use the benchmark in other ways:
- The Joyo on/kun reading table was taken from its keys.
- Half of its target words were held out of training to measure generalization.
- Some settings were chosen by its results.
The held-out half improved by +1.41pt and the other half by +1.91pt. The final stages no longer excluded those words. They used no JKYB data, and were developed and judged only on the Aozora Bunko and Wikipedia sets above.
General sentences
| Evaluation | Sentences | Base | This model |
|---|---|---|---|
| JSUT BASIC5000, sentence Kana-CER against the human kana labels ↓ | 2,236 | 1.881% | 1.713% |
| JVS sentences, Standard CER (whisper-large-v3-turbo) ↓ | 500 | 5.625% | 5.717% |
Neither set was used in training. The sentence-level change is 161 improved / 99 worse for JSUT and 13 / 18 for JVS. Most of the JVS differences are orthographic variants in the Whisper transcripts (時 / とき, 全て / すべて), not changed readings. Readings of everyday sentences did not measurably regress.
The author also listened to the fixed words and to general sentences. The pronunciation and the intonation of the fixed words sounded natural, and general sentences were indistinguishable from the base model. No MOS test was run.
🎙️ Caption and emoji conditioning
Irodori-TTS v4.1 encodes the reading text and the voice-design caption with one shared
ModernBERT and separate projectors. Training the top layers for the text therefore also
changes the encoder's output for captions. To keep voice design close to the base
model, the caption projector and caption_norm were refit on the updated encoder. The
refit only matches states: on 5.6k LLM-written captions and general sentences, the new
caption states are fitted to the base model's.
Difference of the conditioning states from the base model (relative L2, per token):
| Input | Difference |
|---|---|
| Captions (30 hand-written, not used in the refit), without the refit (v2) | 24.4% |
| Captions, with the refit (this checkpoint) | 7.6% |
| Reading text, general sentences | 3.9% |
| Reading text, emoji tokens inside sentences | 13.7% |
For scale, the mean-pooled states of two different captions differ by 40% on average. With the refit, the mean-pooled caption states differ from the base model's by 5% (27% without it, measured for v2).
Caption-only synthesis was also checked: 8 captions × 2 sentences, same seed, no reference audio. The median pitch differs from the base model by 0.36 semitones on average. As with the base model, the male captions stay at about 100–155 Hz and the female ones at about 190–280 Hz; a crying girl goes higher in both. For v1, the refit halved the pitch difference (0.96 → 0.45 semitones), and in an informal listening test the caption-only samples were indistinguishable from the base model. A caption can still give a slightly different voice from the base model with the same seed.
Emoji style control goes through the reading text. Emoji tokens shift more than other tokens (see the table), and emoji control was not evaluated systematically.
⚠️ Limitations
- Mostly learned words. Most of the gain is on words that appear in the fine-tuning data, read in new contexts. Rare words that were not trained improve little. On a held-out set of hard words that were never trained (898 sentences), v3 scores 81.96% against 82.74% for v2 and 80.73% for the base model. The difference from v2 is within noise (p=0.21). Some untrained compounds still change; for example, 前庭 (まえにわ) is sometimes misread.
- Literary vocabulary. Old literary words remain hard: 68% accuracy on the Aozora Bunko set. So do single-kanji homographs whose reading depends on context, names, and specialized terms.
- Evaluation scope. The evaluation is automatic and ASR-based, with one reference voice. Listening was informal, by the author.
- Base model limitations. All limitations of Irodori-TTS-v4.1-Small apply.
📚 Data
Training sources:
- JMdict, KANJIDIC2 and JmdictFurigana (CC BY-SA 4.0): target words and readings.
- Japanese Wikipedia (CC BY-SA,
wikimedia/wikipedia20231101.ja): example sentences and general sentences. - Aozora Bunko: public-domain works and their ruby.
- JVS corpus: its transcripts (from JSUT) are used as general sentences. Its voices were used as reference prompts when the base model generated the teacher latents. No JVS audio is included in this repository.
- Qwen3.6-27B: generated example sentences and checked homograph readings.
- MeCab + UniDic, SudachiPy: sentence readings and reading agreement.
Evaluation only:
- JKYB-Parakeet
- jsut-label (CC BY-SA 4.0)
sbintuitions/kana-whisperopenai/whisper-large-v3-turbo
📜 License & Ethical Restrictions
This model is released under MIT, the same as the base model. The ethical restrictions of the base model also apply:
- No impersonation.
- No misinformation or deepfakes intended to mislead.
- The users are responsible for complying with applicable laws.
🙏 Acknowledgments
- Aratako/Irodori-TTS and Irodori-TTS-v4.1-Small: the base model and its inference code.
- Joyo Kanji Yomi Benchmark (SB Intuitions), its Parakeet Edition (Parakeet Inc.), and kana-whisper: the reading evaluation.
- The JVS and JSUT corpora, Aozora Bunko, Wikipedia, the EDRDG dictionary files, JmdictFurigana, UniDic, Sudachi and Qwen.
🖊️ Citation
Please cite the base model:
@misc{irodori-tts-v4.1-small,
author = {Chihiro Arata},
title = {Irodori-TTS: A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face repository},
howpublished = {\url{https://huggingface.co/Aratako/Irodori-TTS-v4.1-Small}}
}
Model tree for takuma104/Irodori-TTS-v4.1-Small-Yomi
Base model
Aratako/Irodori-TTS-500M-v2