whisper-base-japanese-anime

Try the live demo on Spaces — record audio or use pre-loaded anime samples to compare base vs fine-tuned model.

A fine-tuned version of openai/whisper-base for Japanese anime speech recognition.

Model Description

This model was fine-tuned on the SFW split of joujiboi/japanese-anime-speech-v2 — a dataset of Japanese anime voice clips paired with transcriptions. The goal is to improve Whisper's accuracy on anime-style Japanese speech, which includes expressive intonation, character voices, and domain-specific vocabulary.

Property Value
Base model openai/whisper-base (74M params)
Language Japanese (ja)
Task transcribe
Domain Anime speech
Training samples ~269k SFW samples

Training Details

Training was done incrementally across 26 sequential ranges of ~10k samples each (10k → 269k), with 800 gradient steps per range. This progressive approach allowed monitoring CER at each stage.

Training Hyperparameters

Parameter Value
Learning rate 1e-6
LR scheduler Linear
Warmup steps 200
Batch size (train) 20
Batch size (eval) 16
Gradient accumulation 1
Max steps per range 800
Precision FP16
Optimizer AdamW (fused)
Weight decay 0.0
Total training ranges 26

Training Progression

Stage Samples Val Loss CER
Baseline (10k-20k) 10k 0.892 44.3%
Mid (120k-130k) 120k 0.740 33.5%
Late (210k-220k) 210k 0.621 31.2%
Final (260k-269k) 269k 0.563 31.2%

Key Results

  • CER reduction: 44.3% → 31.2% (~30% relative improvement)
  • Validation loss reduction: 0.892 → 0.563 (37% drop)
  • Best CER achieved at 250k-260k range: 30.6%
  • Validation loss steadily decreased throughout training

Usage

import torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration

# Load model and processor
processor = WhisperProcessor.from_pretrained(
    "sky9262/whisper-base-japanese-anime",
    language="Japanese",
    task="transcribe",
)
model = WhisperForConditionalGeneration.from_pretrained(
    "sky9262/whisper-base-japanese-anime"
)

device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)

# Transcribe audio
import soundfile as sf

audio_array, sampling_rate = sf.read("your_audio.wav")

inputs = processor.feature_extractor(
    audio_array,
    sampling_rate=sampling_rate,
    return_tensors="pt",
).input_features.to(device)

with torch.no_grad():
    predicted_ids = model.generate(
        inputs,
        language="ja",
        task="transcribe",
    )

transcription = processor.tokenizer.batch_decode(
    predicted_ids,
    skip_special_tokens=True,
)[0]

print(transcription)

Audio Samples & Transcriptions

15 expressive anime voice samples (unseen test set) comparing base openai/whisper-base vs. fine-tuned model. These clips feature shouting, panic, teasing, anger, and more -- areas where the base model often hallucinates or loops.

1. Cheerful Greeting (sample 1305, 3.4s)

Text
Base Model 先生 おはろございまーす
Fine-tuned 先生、おはろーございまーす

2. Panicked Scream (sample 1340, 5.5s)

Text
Base Model 今日はスカートをめくらないでくださいましー!
Fine-tuned 今日はわわー!スカートを目くらないでくださいましー!

3. Crying for Help — JP/EN Mix (sample 1343, 5.8s)

Text
Base Model いやー、助けてください、ニーさん! Help you!
Fine-tuned やー、助けてください、兄さん!ヘルプユー

4. Playful Teasing (sample 1369, 5.9s)

Text
Base Model うーん? もうしかして、ネギラってくれるのかねー?
Fine-tuned うーん、もしかして、ネギラってくれるのかねー

5. Over-the-top Shock (sample 1370, 8.0s)

Text
Base Model な、なに?だいまおがだいにケータイに返ししたです!ですー!
Fine-tuned な、なにー!だいまおーが、だいに携帯に変身したです、ですー!

6. Disappointed Sigh (sample 1416, 3.8s)

Text
Base Model えぇ…もう終わりなのか?
Fine-tuned えぇ~、もう終わりなのか?

7. Stuttering Shock (sample 1469, 4.1s)

Text
Base Model ななな! 伸びゆきとのは痛いなによ!
Fine-tuned ななななな、信行きとのは一体なにを!

8. Childish Insistence (sample 1502, 5.9s)

Text
Base Model いやぁ、行くったら行くの? あるいも一緒に遊ぶですよ!
Fine-tuned やぁ、行くったら行くの!アンリーも一緒に遊ぶですよ

9. Furious Struggle (sample 1515, 6.0s) ⚠️ Base model hallucinates

Text
Base Model え、話せブレフモノ!おぉ、も、アドデカイスのか!へっへっへっへっへっへっへっ... (loops 200+ times)
Fine-tuned えへっ、話せぶれ物、おうも後で返すのかっ、ふわああっ!

10. Whining / Pouting (sample 1542, 5.0s)

Text
Base Model だってだって!成功がかわいくないんだもん!
Fine-tuned だってだって、せっくくが可愛くないんだもん!

11. Alarmed Demand (sample 1627, 2.9s)

Text
Base Model まんじゃ、痛い何があった?
Fine-tuned なんじゃ、一体何があった!

12. Cute Excitement (sample 1661, 3.7s) ⚠️ Base model hallucinates

Text
Base Model もーもーもーもーもーもー... (loops 200+ times)
Fine-tuned やふぅ!もうふもふもふですー

13. Angry Scolding (sample 1728, 4.6s)

Text
Base Model たかものは、そんなことができたら、笑顔が特にやっておるわ!
Fine-tuned 高者の、そんなことができたら、妾が特にやっておるわ

14. Desperate Urgency (sample 1740, 4.5s)

Text
Base Model ハイヤト!しっかりしろ!ハイヤト!
Fine-tuned 隼斗、しっかりしろ、隼斗っ!

15. Elongated Refusal (sample 1768, 3.0s)

Text
Base Model んだーめーですー
Fine-tuned ダメです!

⚠️ NSFW Samples (Click to expand) — Contains explicit anime voice content. Viewer discretion advised.

Warning: The following 15 samples are from the NSFW split and contain sexually explicit voice acting. They are included to demonstrate the model's ability to transcribe emotionally extreme and non-standard speech where the base model completely fails (repetition loops, hallucinations).

NSFW 1. Apologetic Crying (nsfw #6, 6.2s)

Text
Base Model ううううううくれない!ごめんない!
Fine-tuned やぁ、やぁああっ、くっくりゃない、ごめんなないっ!

NSFW 2. Teasing / Seductive (nsfw #16, 5.7s)

Text
Base Model なぁ、ニーモクーン スマした顔して、エッチに興味あるの?
Fine-tuned なー、ニーもくん、済ました顔して、エッチに興味あるの

NSFW 3. Flustered Embarrassment (nsfw #29, 5.9s)

Text
Base Model いやいや、ウイキを見出しながら近づかないでください。なんだか一致です。
Fine-tuned やややん、息を見出しながら近づかないでください。なんだかエッジーです

NSFW 4. Confused Protest (nsfw #33, 6.0s)

Text
Base Model 圧倒状況って何ですか?私エッチな動物さじゃないですよ!
Fine-tuned 熱情聞ってなんですか、私、エッチな動物さんじゃないですよー

NSFW 5. Exhibitionist Excitement (nsfw #46, 6.8s)

Text
Base Model あさあ、とんどんご覧になって下さい!もっと私を見たからさ!あ、あ、あ、あ!
Fine-tuned あさー、どんどんごらになってください。もっと私を見たからさ、あっ

NSFW 6. Victory Panting (nsfw #60, 4.8s) ⚠️ Base model hallucinates

Text
Base Model うううう... (loops 400+ times)
Fine-tuned はぁああああ…勝ちましたっ!

NSFW 7. Embarrassed Protest (nsfw #75, 5.6s)

Text
Base Model えぇ、おぉ、いい誘いでください!私、エッチさんじゃないですよ!
Fine-tuned やふぅぅぅぅぅぅ、やめてください。私、エッチさんじゃないですよー

NSFW 8. Short Shock (nsfw #84, 2.1s) ⚠️ Base model hallucinates

Text
Base Model えぇえ!あ、嘘!嘘!あ、あ、あ... (loops 200+ times)
Fine-tuned やや ややっ、そ、そ、そやだっ

NSFW 9. Angry Embarrassment (nsfw #89, 6.7s) ⚠️ Base model hallucinates

Text
Base Model うっうう... (loops 400+ times)
Fine-tuned ぱ、バクバクバクバクバクバクッ!エイチのことを言うでない!もーもーもーもーもーもー

NSFW 10. Overwhelmed Scream (nsfw #356, 2.7s) ⚠️ Base model hallucinates

Text
Base Model うう... (loops 400+ times)
Fine-tuned にゅうぅ~~~~~~~~っ

NSFW 11. Surprised Protest (nsfw #359, 6.9s) ⚠️ Fine-tuned model loops

Text
Base Model おい!いやぁ!あぁ!あぁ!いや、今そこないじゃん!
Fine-tuned んにゃっ、やっ、やっ、やっ... (loops 140+ times, 444 chars)

NSFW 12. Intense Overwhelm (nsfw #404, 9.2s) ⚠️ Both models hallucinate

Text
Base Model あーーー... (loops 400+ times)
Fine-tuned ああああああああああっ、それ、それ、ダメ、ダメッ

NSFW 13. Sudden Panic (nsfw #439, 3.3s) ⚠️ Base model hallucinates

Text
Base Model お、お、え、え、え... (loops 200+ times)
Fine-tuned んっ、あっ、えんっ、えんっ、えんっ、えんっ!

NSFW 14. Desperate Plea (nsfw #456, 7.0s)

Text
Base Model うえ、あわたって本当にでもやるし、お願いしましょ!このまま、このまま!
Fine-tuned あ、もうだめ、本当にダメです。お願いします。お願いします。こんなもん、こんなもんもんっ!

NSFW 15. Falling Panic (nsfw #484, 4.6s) ⚠️ Base model hallucinates

Text
Base Model あ、あ、あ... (loops 200+ times)
Fine-tuned あっ、あっ、あっ、あっ、あっ、ひゃあんっ、落ちる、落ちちゃいませんっ

Pure Moaning / Non-verbal Vocalization

These 5 samples contain no intelligible words — only moaning, panting, and gasping. The base model completely hallucinates on all of them. The fine-tuned model also struggles but produces shorter, more contained output.

NSFW 16. Soft Panting (nsfw #136, 8.8s) ⚠️ Both loop

Text
Base Model ふっふっふっふっふっ... (loops 200+ times)
Fine-tuned んっ、はっ、はっ、はっ、はっ、はっ... (15 repeats, 44 chars vs base 444)

NSFW 17. Intense Climax Scream (nsfw #176, 5.3s) ⚠️ Both loop

Text
Base Model うっうっうっうっ... (loops 200+ times)
Fine-tuned うぅ、うぅ、うぅ、うぅ、うぅ、うぅ、うぅ、うぅ (8 repeats, 23 chars vs base 444)

NSFW 18. Subdued Gasping (nsfw #272, 8.5s) ⚠️ Both loop

Text
Base Model あ!あ!あ!あ!... (loops 200+ times)
Fine-tuned あ、あんっ、あんっ、あんっ、あんっ... (11 repeats, 41 chars vs base 444)

NSFW 19. Exhausted Breathing (nsfw #384, 9.9s) ⚠️ Both loop

Text
Base Model ううう... (loops 400+ times)
Fine-tuned んっ、はぁ…はぁ…はぁ…はぁ…はぁ…はぁ… (11 repeats, 35 chars vs base 444)

NSFW 20. High-pitched Climax (nsfw #468, 5.5s) ⚠️ Both loop

Text
Base Model あ!あ!あ!あ!... (loops 200+ times)
Fine-tuned あ、あ、ああ、あああ... (loops 400+ times, 883 chars — worst case, similar to base)

Evaluation

Evaluated on unseen test samples from the SFW split using Character Error Rate (CER), the standard metric for Japanese ASR (since Japanese lacks clear word boundaries).

CER is computed as:

CER = (substitutions + insertions + deletions) / len(reference)

Limitations

  • Domain-specific: Optimized for anime-style Japanese speech. Performance on other Japanese speech domains (news, conversation, etc.) may differ from the base model.
  • Base model size: This is whisper-base (74M params). Larger variants (small, medium, large) would likely achieve better CER.
  • SFW only: Trained only on the SFW portion of the dataset.
  • Single speaker variability: Anime speech has high variability in tone, pitch, and speaking style which can still cause errors.

Citation

If you use this model, please cite the original Whisper paper and the training dataset:

@article{radford2022whisper,
  title={Robust Speech Recognition via Large-Scale Weak Supervision},
  author={Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
  journal={arXiv preprint arXiv:2212.04356},
  year={2022}
}

Dataset: joujiboi/japanese-anime-speech-v2

Downloads last month
60
Safetensors
Model size
72.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sky9262/whisper-base-japanese-anime

Finetuned
(764)
this model

Dataset used to train sky9262/whisper-base-japanese-anime

Space using sky9262/whisper-base-japanese-anime 1

Paper for sky9262/whisper-base-japanese-anime

Evaluation results

  • CER (final) on joujiboi/japanese-anime-speech-v2
    self-reported
    31.200
  • CER (baseline) on joujiboi/japanese-anime-speech-v2
    self-reported
    44.300