--- language: - ja license: apache-2.0 base_model: openai/whisper-base tags: - whisper - japanese - anime - speech-recognition - fine-tuned - automatic-speech-recognition datasets: - joujiboi/japanese-anime-speech-v2 metrics: - cer pipeline_tag: automatic-speech-recognition widget: - src: https://huggingface.co/sky9262/whisper-base-japanese-anime/resolve/main/samples/sample_1305.wav example_title: Cheerful Greeting - src: https://huggingface.co/sky9262/whisper-base-japanese-anime/resolve/main/samples/sample_1370.wav example_title: Over-the-top Shock - src: https://huggingface.co/sky9262/whisper-base-japanese-anime/resolve/main/samples/sample_1740.wav example_title: Desperate Urgency model-index: - name: whisper-base-japanese-anime results: - task: type: automatic-speech-recognition name: Automatic Speech Recognition dataset: name: joujiboi/japanese-anime-speech-v2 type: joujiboi/japanese-anime-speech-v2 split: sfw metrics: - type: cer value: 31.2 name: CER (final) - type: cer value: 44.3 name: CER (baseline) --- # whisper-base-japanese-anime **[Try the live demo on Spaces](https://huggingface.co/spaces/sky9262/whisper-base-japanese-anime)** — record audio or use pre-loaded anime samples to compare base vs fine-tuned model. A fine-tuned version of [openai/whisper-base](https://huggingface.co/openai/whisper-base) for **Japanese anime speech recognition**. ## Model Description This model was fine-tuned on the **SFW split** of [joujiboi/japanese-anime-speech-v2](https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2) — a dataset of Japanese anime voice clips paired with transcriptions. The goal is to improve Whisper's accuracy on anime-style Japanese speech, which includes expressive intonation, character voices, and domain-specific vocabulary. | Property | Value | |---|---| | **Base model** | `openai/whisper-base` (74M params) | | **Language** | Japanese (`ja`) | | **Task** | `transcribe` | | **Domain** | Anime speech | | **Training samples** | ~269k SFW samples | ## Training Details Training was done incrementally across **26 sequential ranges** of ~10k samples each (10k → 269k), with 800 gradient steps per range. This progressive approach allowed monitoring CER at each stage. ### Training Hyperparameters | Parameter | Value | |---|---| | Learning rate | `1e-6` | | LR scheduler | Linear | | Warmup steps | 200 | | Batch size (train) | 20 | | Batch size (eval) | 16 | | Gradient accumulation | 1 | | Max steps per range | 800 | | Precision | FP16 | | Optimizer | AdamW (fused) | | Weight decay | 0.0 | | Total training ranges | 26 | ### Training Progression | Stage | Samples | Val Loss | CER | |---|---|---|---| | Baseline (10k-20k) | 10k | 0.892 | 44.3% | | Mid (120k-130k) | 120k | 0.740 | 33.5% | | Late (210k-220k) | 210k | 0.621 | 31.2% | | **Final (260k-269k)** | **269k** | **0.563** | **31.2%** | ### Key Results - **CER reduction**: 44.3% → 31.2% (**~30% relative improvement**) - **Validation loss reduction**: 0.892 → 0.563 (**37% drop**) - Best CER achieved at 250k-260k range: **30.6%** - Validation loss steadily decreased throughout training ## Usage ```python import torch from transformers import WhisperProcessor, WhisperForConditionalGeneration # Load model and processor processor = WhisperProcessor.from_pretrained( "sky9262/whisper-base-japanese-anime", language="Japanese", task="transcribe", ) model = WhisperForConditionalGeneration.from_pretrained( "sky9262/whisper-base-japanese-anime" ) device = "cuda" if torch.cuda.is_available() else "cpu" model.to(device) # Transcribe audio import soundfile as sf audio_array, sampling_rate = sf.read("your_audio.wav") inputs = processor.feature_extractor( audio_array, sampling_rate=sampling_rate, return_tensors="pt", ).input_features.to(device) with torch.no_grad(): predicted_ids = model.generate( inputs, language="ja", task="transcribe", ) transcription = processor.tokenizer.batch_decode( predicted_ids, skip_special_tokens=True, )[0] print(transcription) ``` ## Audio Samples & Transcriptions 15 expressive anime voice samples (unseen test set) comparing **base `openai/whisper-base`** vs. **fine-tuned model**. These clips feature shouting, panic, teasing, anger, and more -- areas where the base model often hallucinates or loops. ### 1. Cheerful Greeting (sample 1305, 3.4s) | | Text | |---|---| | **Base Model** | 先生 おはろございまーす | | **Fine-tuned** | 先生、おはろーございまーす | ### 2. Panicked Scream (sample 1340, 5.5s) | | Text | |---|---| | **Base Model** | 今日はスカートをめくらないでくださいましー! | | **Fine-tuned** | 今日はわわー!スカートを目くらないでくださいましー! | ### 3. Crying for Help — JP/EN Mix (sample 1343, 5.8s) | | Text | |---|---| | **Base Model** | いやー、助けてください、ニーさん! Help you! | | **Fine-tuned** | やー、助けてください、兄さん!ヘルプユー | ### 4. Playful Teasing (sample 1369, 5.9s) | | Text | |---|---| | **Base Model** | うーん? もうしかして、ネギラってくれるのかねー? | | **Fine-tuned** | うーん、もしかして、ネギラってくれるのかねー | ### 5. Over-the-top Shock (sample 1370, 8.0s) | | Text | |---|---| | **Base Model** | な、なに?だいまおがだいにケータイに返ししたです!ですー! | | **Fine-tuned** | な、なにー!だいまおーが、だいに携帯に変身したです、ですー! | ### 6. Disappointed Sigh (sample 1416, 3.8s) | | Text | |---|---| | **Base Model** | えぇ…もう終わりなのか? | | **Fine-tuned** | えぇ~、もう終わりなのか? | ### 7. Stuttering Shock (sample 1469, 4.1s) | | Text | |---|---| | **Base Model** | ななな! 伸びゆきとのは痛いなによ! | | **Fine-tuned** | ななななな、信行きとのは一体なにを! | ### 8. Childish Insistence (sample 1502, 5.9s) | | Text | |---|---| | **Base Model** | いやぁ、行くったら行くの? あるいも一緒に遊ぶですよ! | | **Fine-tuned** | やぁ、行くったら行くの!アンリーも一緒に遊ぶですよ | ### 9. Furious Struggle (sample 1515, 6.0s) ⚠️ Base model hallucinates | | Text | |---|---| | **Base Model** | え、話せブレフモノ!おぉ、も、アドデカイスのか!へっへっへっへっへっへっへっ... **(loops 200+ times)** | | **Fine-tuned** | えへっ、話せぶれ物、おうも後で返すのかっ、ふわああっ! | ### 10. Whining / Pouting (sample 1542, 5.0s) | | Text | |---|---| | **Base Model** | だってだって!成功がかわいくないんだもん! | | **Fine-tuned** | だってだって、せっくくが可愛くないんだもん! | ### 11. Alarmed Demand (sample 1627, 2.9s) | | Text | |---|---| | **Base Model** | まんじゃ、痛い何があった? | | **Fine-tuned** | なんじゃ、一体何があった! | ### 12. Cute Excitement (sample 1661, 3.7s) ⚠️ Base model hallucinates | | Text | |---|---| | **Base Model** | もーもーもーもーもーもー... **(loops 200+ times)** | | **Fine-tuned** | やふぅ!もうふもふもふですー | ### 13. Angry Scolding (sample 1728, 4.6s) | | Text | |---|---| | **Base Model** | たかものは、そんなことができたら、笑顔が特にやっておるわ! | | **Fine-tuned** | 高者の、そんなことができたら、妾が特にやっておるわ | ### 14. Desperate Urgency (sample 1740, 4.5s) | | Text | |---|---| | **Base Model** | ハイヤト!しっかりしろ!ハイヤト! | | **Fine-tuned** | 隼斗、しっかりしろ、隼斗っ! | ### 15. Elongated Refusal (sample 1768, 3.0s) | | Text | |---|---| | **Base Model** | んだーめーですー | | **Fine-tuned** | ダメです! | ---
⚠️ NSFW Samples (Click to expand) — Contains explicit anime voice content. Viewer discretion advised.
> **Warning**: The following 15 samples are from the NSFW split and contain sexually explicit voice acting. > They are included to demonstrate the model's ability to transcribe emotionally extreme and non-standard speech > where the base model completely fails (repetition loops, hallucinations). ### NSFW 1. Apologetic Crying (nsfw #6, 6.2s) | | Text | |---|---| | **Base Model** | ううううううくれない!ごめんない! | | **Fine-tuned** | やぁ、やぁああっ、くっくりゃない、ごめんなないっ! | ### NSFW 2. Teasing / Seductive (nsfw #16, 5.7s) | | Text | |---|---| | **Base Model** | なぁ、ニーモクーン スマした顔して、エッチに興味あるの? | | **Fine-tuned** | なー、ニーもくん、済ました顔して、エッチに興味あるの | ### NSFW 3. Flustered Embarrassment (nsfw #29, 5.9s) | | Text | |---|---| | **Base Model** | いやいや、ウイキを見出しながら近づかないでください。なんだか一致です。 | | **Fine-tuned** | やややん、息を見出しながら近づかないでください。なんだかエッジーです | ### NSFW 4. Confused Protest (nsfw #33, 6.0s) | | Text | |---|---| | **Base Model** | 圧倒状況って何ですか?私エッチな動物さじゃないですよ! | | **Fine-tuned** | 熱情聞ってなんですか、私、エッチな動物さんじゃないですよー | ### NSFW 5. Exhibitionist Excitement (nsfw #46, 6.8s) | | Text | |---|---| | **Base Model** | あさあ、とんどんご覧になって下さい!もっと私を見たからさ!あ、あ、あ、あ! | | **Fine-tuned** | あさー、どんどんごらになってください。もっと私を見たからさ、あっ | ### NSFW 6. Victory Panting (nsfw #60, 4.8s) ⚠️ Base model hallucinates | | Text | |---|---| | **Base Model** | うううう... **(loops 400+ times)** | | **Fine-tuned** | はぁああああ…勝ちましたっ! | ### NSFW 7. Embarrassed Protest (nsfw #75, 5.6s) | | Text | |---|---| | **Base Model** | えぇ、おぉ、いい誘いでください!私、エッチさんじゃないですよ! | | **Fine-tuned** | やふぅぅぅぅぅぅ、やめてください。私、エッチさんじゃないですよー | ### NSFW 8. Short Shock (nsfw #84, 2.1s) ⚠️ Base model hallucinates | | Text | |---|---| | **Base Model** | えぇえ!あ、嘘!嘘!あ、あ、あ... **(loops 200+ times)** | | **Fine-tuned** | やや ややっ、そ、そ、そやだっ | ### NSFW 9. Angry Embarrassment (nsfw #89, 6.7s) ⚠️ Base model hallucinates | | Text | |---|---| | **Base Model** | うっうう... **(loops 400+ times)** | | **Fine-tuned** | ぱ、バクバクバクバクバクバクッ!エイチのことを言うでない!もーもーもーもーもーもー | ### NSFW 10. Overwhelmed Scream (nsfw #356, 2.7s) ⚠️ Base model hallucinates | | Text | |---|---| | **Base Model** | うう... **(loops 400+ times)** | | **Fine-tuned** | にゅうぅ~~~~~~~~っ | ### NSFW 11. Surprised Protest (nsfw #359, 6.9s) ⚠️ Fine-tuned model loops | | Text | |---|---| | **Base Model** | おい!いやぁ!あぁ!あぁ!いや、今そこないじゃん! | | **Fine-tuned** | んにゃっ、やっ、やっ、やっ... **(loops 140+ times, 444 chars)** | ### NSFW 12. Intense Overwhelm (nsfw #404, 9.2s) ⚠️ Both models hallucinate | | Text | |---|---| | **Base Model** | あーーー... **(loops 400+ times)** | | **Fine-tuned** | ああああああああああっ、それ、それ、ダメ、ダメッ | ### NSFW 13. Sudden Panic (nsfw #439, 3.3s) ⚠️ Base model hallucinates | | Text | |---|---| | **Base Model** | お、お、え、え、え... **(loops 200+ times)** | | **Fine-tuned** | んっ、あっ、えんっ、えんっ、えんっ、えんっ! | ### NSFW 14. Desperate Plea (nsfw #456, 7.0s) | | Text | |---|---| | **Base Model** | うえ、あわたって本当にでもやるし、お願いしましょ!このまま、このまま! | | **Fine-tuned** | あ、もうだめ、本当にダメです。お願いします。お願いします。こんなもん、こんなもんもんっ! | ### NSFW 15. Falling Panic (nsfw #484, 4.6s) ⚠️ Base model hallucinates | | Text | |---|---| | **Base Model** | あ、あ、あ... **(loops 200+ times)** | | **Fine-tuned** | あっ、あっ、あっ、あっ、あっ、ひゃあんっ、落ちる、落ちちゃいませんっ | --- #### Pure Moaning / Non-verbal Vocalization > These 5 samples contain **no intelligible words** — only moaning, panting, and gasping. > The base model completely hallucinates on all of them. The fine-tuned model also struggles but produces shorter, more contained output. ### NSFW 16. Soft Panting (nsfw #136, 8.8s) ⚠️ Both loop | | Text | |---|---| | **Base Model** | ふっふっふっふっふっ... **(loops 200+ times)** | | **Fine-tuned** | んっ、はっ、はっ、はっ、はっ、はっ... **(15 repeats, 44 chars vs base 444)** | ### NSFW 17. Intense Climax Scream (nsfw #176, 5.3s) ⚠️ Both loop | | Text | |---|---| | **Base Model** | うっうっうっうっ... **(loops 200+ times)** | | **Fine-tuned** | うぅ、うぅ、うぅ、うぅ、うぅ、うぅ、うぅ、うぅ **(8 repeats, 23 chars vs base 444)** | ### NSFW 18. Subdued Gasping (nsfw #272, 8.5s) ⚠️ Both loop | | Text | |---|---| | **Base Model** | あ!あ!あ!あ!... **(loops 200+ times)** | | **Fine-tuned** | あ、あんっ、あんっ、あんっ、あんっ... **(11 repeats, 41 chars vs base 444)** | ### NSFW 19. Exhausted Breathing (nsfw #384, 9.9s) ⚠️ Both loop | | Text | |---|---| | **Base Model** | ううう... **(loops 400+ times)** | | **Fine-tuned** | んっ、はぁ…はぁ…はぁ…はぁ…はぁ…はぁ… **(11 repeats, 35 chars vs base 444)** | ### NSFW 20. High-pitched Climax (nsfw #468, 5.5s) ⚠️ Both loop | | Text | |---|---| | **Base Model** | あ!あ!あ!あ!... **(loops 200+ times)** | | **Fine-tuned** | あ、あ、ああ、あああ... **(loops 400+ times, 883 chars — worst case, similar to base)** |
## Evaluation Evaluated on unseen test samples from the SFW split using **Character Error Rate (CER)**, the standard metric for Japanese ASR (since Japanese lacks clear word boundaries). CER is computed as: ``` CER = (substitutions + insertions + deletions) / len(reference) ``` ## Limitations - **Domain-specific**: Optimized for anime-style Japanese speech. Performance on other Japanese speech domains (news, conversation, etc.) may differ from the base model. - **Base model size**: This is whisper-base (74M params). Larger variants (small, medium, large) would likely achieve better CER. - **SFW only**: Trained only on the SFW portion of the dataset. - **Single speaker variability**: Anime speech has high variability in tone, pitch, and speaking style which can still cause errors. ## Citation If you use this model, please cite the original Whisper paper and the training dataset: ```bibtex @article{radford2022whisper, title={Robust Speech Recognition via Large-Scale Weak Supervision}, author={Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya}, journal={arXiv preprint arXiv:2212.04356}, year={2022} } ``` Dataset: [joujiboi/japanese-anime-speech-v2](https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2)