Model Performance is great!
Thank you for developing a great Japanese TTS for community. I tested and have some views
Advantages noted:
Accurately pronounces the pitch of homonyms: ๆฉ - ็ฎธ.
Changes the intonation at the end of sentences according to punctuation marks: "!", "~", "?", ".".
Automatically adds appropriate expressive SFX when input includes emoticons (laugh, sigh, etc.).
Clear audio quality, free from background noise.
Natural, native-like voice.
โ Sound quality is better than other TTS models currently in use.
Errors detected during testing:
Incorrect reading of numbers when written in numerals: ใ๏ผ๏ผๆญณใ is read as ใใซใใ ใฃใใใ instead of ใใฏใใกใ.
Incorrect reading of dates even when written in kanji: ใไบๆไธๆฅไปใงใ is misread.
Unstable voice during mass generation: reference audio is for a female voice but output is a male voice.
Sound changes are not applied when combining words: ใไธ้ใ is read as ใใใใใใ instead of ใใใใใใ.
Incorrect word reading: ใ้่กใ is read as ใใใใใใ instead of ใใใใใใ.
A single kanji character has multiple readings, but TTS does not yet support a mechanism to specify a reading. For example: ใ่งใ is read as ใใคใฎใ while the desired reading is ใใใฉใ; similarly with ใๅๅใ.
English words are automatically converted to Japanese katakana readings, but there may still be cases where the reading is incorrect.
When reading single words, odd characters, or short sentences without context, TTS may add inappropriate expressions (gasping, cries, etc.). For example: ใใใ, ใๆฏใ, ใใใใฆใใ ใใใ, ใๅ ็ใ.
Mispronouncing proper names: ใ็งๅฑฑๅ ็ใ is pronounced ใใใใใพใใใใใ instead of ใใใใใพใใใใใ.
When reading a long passage, there may be strange sound effects inserted (audio interruptions, speaker breathing sounds).