---
language:
- ja
license: apache-2.0
base_model: openai/whisper-base
tags:
- whisper
- japanese
- anime
- speech-recognition
- fine-tuned
- automatic-speech-recognition
datasets:
- joujiboi/japanese-anime-speech-v2
metrics:
- cer
pipeline_tag: automatic-speech-recognition
widget:
- src: https://huggingface.co/sky9262/whisper-base-japanese-anime/resolve/main/samples/sample_1305.wav
example_title: Cheerful Greeting
- src: https://huggingface.co/sky9262/whisper-base-japanese-anime/resolve/main/samples/sample_1370.wav
example_title: Over-the-top Shock
- src: https://huggingface.co/sky9262/whisper-base-japanese-anime/resolve/main/samples/sample_1740.wav
example_title: Desperate Urgency
model-index:
- name: whisper-base-japanese-anime
results:
- task:
type: automatic-speech-recognition
name: Automatic Speech Recognition
dataset:
name: joujiboi/japanese-anime-speech-v2
type: joujiboi/japanese-anime-speech-v2
split: sfw
metrics:
- type: cer
value: 31.2
name: CER (final)
- type: cer
value: 44.3
name: CER (baseline)
---
# whisper-base-japanese-anime
**[Try the live demo on Spaces](https://huggingface.co/spaces/sky9262/whisper-base-japanese-anime)** — record audio or use pre-loaded anime samples to compare base vs fine-tuned model.
A fine-tuned version of [openai/whisper-base](https://huggingface.co/openai/whisper-base) for **Japanese anime speech recognition**.
## Model Description
This model was fine-tuned on the **SFW split** of [joujiboi/japanese-anime-speech-v2](https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2) — a dataset of Japanese anime voice clips paired with transcriptions. The goal is to improve Whisper's accuracy on anime-style Japanese speech, which includes expressive intonation, character voices, and domain-specific vocabulary.
| Property | Value |
|---|---|
| **Base model** | `openai/whisper-base` (74M params) |
| **Language** | Japanese (`ja`) |
| **Task** | `transcribe` |
| **Domain** | Anime speech |
| **Training samples** | ~269k SFW samples |
## Training Details
Training was done incrementally across **26 sequential ranges** of ~10k samples each (10k → 269k), with 800 gradient steps per range. This progressive approach allowed monitoring CER at each stage.
### Training Hyperparameters
| Parameter | Value |
|---|---|
| Learning rate | `1e-6` |
| LR scheduler | Linear |
| Warmup steps | 200 |
| Batch size (train) | 20 |
| Batch size (eval) | 16 |
| Gradient accumulation | 1 |
| Max steps per range | 800 |
| Precision | FP16 |
| Optimizer | AdamW (fused) |
| Weight decay | 0.0 |
| Total training ranges | 26 |
### Training Progression
| Stage | Samples | Val Loss | CER |
|---|---|---|---|
| Baseline (10k-20k) | 10k | 0.892 | 44.3% |
| Mid (120k-130k) | 120k | 0.740 | 33.5% |
| Late (210k-220k) | 210k | 0.621 | 31.2% |
| **Final (260k-269k)** | **269k** | **0.563** | **31.2%** |
### Key Results
- **CER reduction**: 44.3% → 31.2% (**~30% relative improvement**)
- **Validation loss reduction**: 0.892 → 0.563 (**37% drop**)
- Best CER achieved at 250k-260k range: **30.6%**
- Validation loss steadily decreased throughout training
## Usage
```python
import torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration
# Load model and processor
processor = WhisperProcessor.from_pretrained(
"sky9262/whisper-base-japanese-anime",
language="Japanese",
task="transcribe",
)
model = WhisperForConditionalGeneration.from_pretrained(
"sky9262/whisper-base-japanese-anime"
)
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
# Transcribe audio
import soundfile as sf
audio_array, sampling_rate = sf.read("your_audio.wav")
inputs = processor.feature_extractor(
audio_array,
sampling_rate=sampling_rate,
return_tensors="pt",
).input_features.to(device)
with torch.no_grad():
predicted_ids = model.generate(
inputs,
language="ja",
task="transcribe",
)
transcription = processor.tokenizer.batch_decode(
predicted_ids,
skip_special_tokens=True,
)[0]
print(transcription)
```
## Audio Samples & Transcriptions
15 expressive anime voice samples (unseen test set) comparing **base `openai/whisper-base`** vs. **fine-tuned model**.
These clips feature shouting, panic, teasing, anger, and more -- areas where the base model often hallucinates or loops.
### 1. Cheerful Greeting (sample 1305, 3.4s)
| | Text |
|---|---|
| **Base Model** | 先生 おはろございまーす |
| **Fine-tuned** | 先生、おはろーございまーす |
### 2. Panicked Scream (sample 1340, 5.5s)
| | Text |
|---|---|
| **Base Model** | 今日はスカートをめくらないでくださいましー! |
| **Fine-tuned** | 今日はわわー!スカートを目くらないでくださいましー! |
### 3. Crying for Help — JP/EN Mix (sample 1343, 5.8s)
| | Text |
|---|---|
| **Base Model** | いやー、助けてください、ニーさん! Help you! |
| **Fine-tuned** | やー、助けてください、兄さん!ヘルプユー |
### 4. Playful Teasing (sample 1369, 5.9s)
| | Text |
|---|---|
| **Base Model** | うーん? もうしかして、ネギラってくれるのかねー? |
| **Fine-tuned** | うーん、もしかして、ネギラってくれるのかねー |
### 5. Over-the-top Shock (sample 1370, 8.0s)
| | Text |
|---|---|
| **Base Model** | な、なに?だいまおがだいにケータイに返ししたです!ですー! |
| **Fine-tuned** | な、なにー!だいまおーが、だいに携帯に変身したです、ですー! |
### 6. Disappointed Sigh (sample 1416, 3.8s)
| | Text |
|---|---|
| **Base Model** | えぇ…もう終わりなのか? |
| **Fine-tuned** | えぇ~、もう終わりなのか? |
### 7. Stuttering Shock (sample 1469, 4.1s)
| | Text |
|---|---|
| **Base Model** | ななな! 伸びゆきとのは痛いなによ! |
| **Fine-tuned** | ななななな、信行きとのは一体なにを! |
### 8. Childish Insistence (sample 1502, 5.9s)
| | Text |
|---|---|
| **Base Model** | いやぁ、行くったら行くの? あるいも一緒に遊ぶですよ! |
| **Fine-tuned** | やぁ、行くったら行くの!アンリーも一緒に遊ぶですよ |
### 9. Furious Struggle (sample 1515, 6.0s) ⚠️ Base model hallucinates
| | Text |
|---|---|
| **Base Model** | え、話せブレフモノ!おぉ、も、アドデカイスのか!へっへっへっへっへっへっへっ... **(loops 200+ times)** |
| **Fine-tuned** | えへっ、話せぶれ物、おうも後で返すのかっ、ふわああっ! |
### 10. Whining / Pouting (sample 1542, 5.0s)
| | Text |
|---|---|
| **Base Model** | だってだって!成功がかわいくないんだもん! |
| **Fine-tuned** | だってだって、せっくくが可愛くないんだもん! |
### 11. Alarmed Demand (sample 1627, 2.9s)
| | Text |
|---|---|
| **Base Model** | まんじゃ、痛い何があった? |
| **Fine-tuned** | なんじゃ、一体何があった! |
### 12. Cute Excitement (sample 1661, 3.7s) ⚠️ Base model hallucinates
| | Text |
|---|---|
| **Base Model** | もーもーもーもーもーもー... **(loops 200+ times)** |
| **Fine-tuned** | やふぅ!もうふもふもふですー |
### 13. Angry Scolding (sample 1728, 4.6s)
| | Text |
|---|---|
| **Base Model** | たかものは、そんなことができたら、笑顔が特にやっておるわ! |
| **Fine-tuned** | 高者の、そんなことができたら、妾が特にやっておるわ |
### 14. Desperate Urgency (sample 1740, 4.5s)
| | Text |
|---|---|
| **Base Model** | ハイヤト!しっかりしろ!ハイヤト! |
| **Fine-tuned** | 隼斗、しっかりしろ、隼斗っ! |
### 15. Elongated Refusal (sample 1768, 3.0s)
| | Text |
|---|---|
| **Base Model** | んだーめーですー |
| **Fine-tuned** | ダメです! |
---
⚠️ NSFW Samples (Click to expand) — Contains explicit anime voice content. Viewer discretion advised.
> **Warning**: The following 15 samples are from the NSFW split and contain sexually explicit voice acting.
> They are included to demonstrate the model's ability to transcribe emotionally extreme and non-standard speech
> where the base model completely fails (repetition loops, hallucinations).
### NSFW 1. Apologetic Crying (nsfw #6, 6.2s)
| | Text |
|---|---|
| **Base Model** | ううううううくれない!ごめんない! |
| **Fine-tuned** | やぁ、やぁああっ、くっくりゃない、ごめんなないっ! |
### NSFW 2. Teasing / Seductive (nsfw #16, 5.7s)
| | Text |
|---|---|
| **Base Model** | なぁ、ニーモクーン スマした顔して、エッチに興味あるの? |
| **Fine-tuned** | なー、ニーもくん、済ました顔して、エッチに興味あるの |
### NSFW 3. Flustered Embarrassment (nsfw #29, 5.9s)
| | Text |
|---|---|
| **Base Model** | いやいや、ウイキを見出しながら近づかないでください。なんだか一致です。 |
| **Fine-tuned** | やややん、息を見出しながら近づかないでください。なんだかエッジーです |
### NSFW 4. Confused Protest (nsfw #33, 6.0s)
| | Text |
|---|---|
| **Base Model** | 圧倒状況って何ですか?私エッチな動物さじゃないですよ! |
| **Fine-tuned** | 熱情聞ってなんですか、私、エッチな動物さんじゃないですよー |
### NSFW 5. Exhibitionist Excitement (nsfw #46, 6.8s)
| | Text |
|---|---|
| **Base Model** | あさあ、とんどんご覧になって下さい!もっと私を見たからさ!あ、あ、あ、あ! |
| **Fine-tuned** | あさー、どんどんごらになってください。もっと私を見たからさ、あっ |
### NSFW 6. Victory Panting (nsfw #60, 4.8s) ⚠️ Base model hallucinates
| | Text |
|---|---|
| **Base Model** | うううう... **(loops 400+ times)** |
| **Fine-tuned** | はぁああああ…勝ちましたっ! |
### NSFW 7. Embarrassed Protest (nsfw #75, 5.6s)
| | Text |
|---|---|
| **Base Model** | えぇ、おぉ、いい誘いでください!私、エッチさんじゃないですよ! |
| **Fine-tuned** | やふぅぅぅぅぅぅ、やめてください。私、エッチさんじゃないですよー |
### NSFW 8. Short Shock (nsfw #84, 2.1s) ⚠️ Base model hallucinates
| | Text |
|---|---|
| **Base Model** | えぇえ!あ、嘘!嘘!あ、あ、あ... **(loops 200+ times)** |
| **Fine-tuned** | やや ややっ、そ、そ、そやだっ |
### NSFW 9. Angry Embarrassment (nsfw #89, 6.7s) ⚠️ Base model hallucinates
| | Text |
|---|---|
| **Base Model** | うっうう... **(loops 400+ times)** |
| **Fine-tuned** | ぱ、バクバクバクバクバクバクッ!エイチのことを言うでない!もーもーもーもーもーもー |
### NSFW 10. Overwhelmed Scream (nsfw #356, 2.7s) ⚠️ Base model hallucinates
| | Text |
|---|---|
| **Base Model** | うう... **(loops 400+ times)** |
| **Fine-tuned** | にゅうぅ~~~~~~~~っ |
### NSFW 11. Surprised Protest (nsfw #359, 6.9s) ⚠️ Fine-tuned model loops
| | Text |
|---|---|
| **Base Model** | おい!いやぁ!あぁ!あぁ!いや、今そこないじゃん! |
| **Fine-tuned** | んにゃっ、やっ、やっ、やっ... **(loops 140+ times, 444 chars)** |
### NSFW 12. Intense Overwhelm (nsfw #404, 9.2s) ⚠️ Both models hallucinate
| | Text |
|---|---|
| **Base Model** | あーーー... **(loops 400+ times)** |
| **Fine-tuned** | ああああああああああっ、それ、それ、ダメ、ダメッ |
### NSFW 13. Sudden Panic (nsfw #439, 3.3s) ⚠️ Base model hallucinates
| | Text |
|---|---|
| **Base Model** | お、お、え、え、え... **(loops 200+ times)** |
| **Fine-tuned** | んっ、あっ、えんっ、えんっ、えんっ、えんっ! |
### NSFW 14. Desperate Plea (nsfw #456, 7.0s)
| | Text |
|---|---|
| **Base Model** | うえ、あわたって本当にでもやるし、お願いしましょ!このまま、このまま! |
| **Fine-tuned** | あ、もうだめ、本当にダメです。お願いします。お願いします。こんなもん、こんなもんもんっ! |
### NSFW 15. Falling Panic (nsfw #484, 4.6s) ⚠️ Base model hallucinates
| | Text |
|---|---|
| **Base Model** | あ、あ、あ... **(loops 200+ times)** |
| **Fine-tuned** | あっ、あっ、あっ、あっ、あっ、ひゃあんっ、落ちる、落ちちゃいませんっ |
---
#### Pure Moaning / Non-verbal Vocalization
> These 5 samples contain **no intelligible words** — only moaning, panting, and gasping.
> The base model completely hallucinates on all of them. The fine-tuned model also struggles but produces shorter, more contained output.
### NSFW 16. Soft Panting (nsfw #136, 8.8s) ⚠️ Both loop
| | Text |
|---|---|
| **Base Model** | ふっふっふっふっふっ... **(loops 200+ times)** |
| **Fine-tuned** | んっ、はっ、はっ、はっ、はっ、はっ... **(15 repeats, 44 chars vs base 444)** |
### NSFW 17. Intense Climax Scream (nsfw #176, 5.3s) ⚠️ Both loop
| | Text |
|---|---|
| **Base Model** | うっうっうっうっ... **(loops 200+ times)** |
| **Fine-tuned** | うぅ、うぅ、うぅ、うぅ、うぅ、うぅ、うぅ、うぅ **(8 repeats, 23 chars vs base 444)** |
### NSFW 18. Subdued Gasping (nsfw #272, 8.5s) ⚠️ Both loop
| | Text |
|---|---|
| **Base Model** | あ!あ!あ!あ!... **(loops 200+ times)** |
| **Fine-tuned** | あ、あんっ、あんっ、あんっ、あんっ... **(11 repeats, 41 chars vs base 444)** |
### NSFW 19. Exhausted Breathing (nsfw #384, 9.9s) ⚠️ Both loop
| | Text |
|---|---|
| **Base Model** | ううう... **(loops 400+ times)** |
| **Fine-tuned** | んっ、はぁ…はぁ…はぁ…はぁ…はぁ…はぁ… **(11 repeats, 35 chars vs base 444)** |
### NSFW 20. High-pitched Climax (nsfw #468, 5.5s) ⚠️ Both loop
| | Text |
|---|---|
| **Base Model** | あ!あ!あ!あ!... **(loops 200+ times)** |
| **Fine-tuned** | あ、あ、ああ、あああ... **(loops 400+ times, 883 chars — worst case, similar to base)** |
## Evaluation
Evaluated on unseen test samples from the SFW split using **Character Error Rate (CER)**, the standard metric for Japanese ASR (since Japanese lacks clear word boundaries).
CER is computed as:
```
CER = (substitutions + insertions + deletions) / len(reference)
```
## Limitations
- **Domain-specific**: Optimized for anime-style Japanese speech. Performance on other Japanese speech domains (news, conversation, etc.) may differ from the base model.
- **Base model size**: This is whisper-base (74M params). Larger variants (small, medium, large) would likely achieve better CER.
- **SFW only**: Trained only on the SFW portion of the dataset.
- **Single speaker variability**: Anime speech has high variability in tone, pitch, and speaking style which can still cause errors.
## Citation
If you use this model, please cite the original Whisper paper and the training dataset:
```bibtex
@article{radford2022whisper,
title={Robust Speech Recognition via Large-Scale Weak Supervision},
author={Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
journal={arXiv preprint arXiv:2212.04356},
year={2022}
}
```
Dataset: [joujiboi/japanese-anime-speech-v2](https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2)