Try the live demo on Spaces — record audio or use pre-loaded anime samples to compare base vs fine-tuned model.
A fine-tuned version of openai/whisper-base for Japanese anime speech recognition.
Model Description
This model was fine-tuned on the SFW split of joujiboi/japanese-anime-speech-v2 — a dataset of Japanese anime voice clips paired with transcriptions. The goal is to improve Whisper's accuracy on anime-style Japanese speech, which includes expressive intonation, character voices, and domain-specific vocabulary.
Property
Value
Base model
openai/whisper-base (74M params)
Language
Japanese (ja)
Task
transcribe
Domain
Anime speech
Training samples
~269k SFW samples
Training Details
Training was done incrementally across 26 sequential ranges of ~10k samples each (10k → 269k), with 800 gradient steps per range. This progressive approach allowed monitoring CER at each stage.
Validation loss reduction: 0.892 → 0.563 (37% drop)
Best CER achieved at 250k-260k range: 30.6%
Validation loss steadily decreased throughout training
Usage
import torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration
# Load model and processor
processor = WhisperProcessor.from_pretrained(
"sky9262/whisper-base-japanese-anime",
language="Japanese",
task="transcribe",
)
model = WhisperForConditionalGeneration.from_pretrained(
"sky9262/whisper-base-japanese-anime"
)
device = "cuda"if torch.cuda.is_available() else"cpu"
model.to(device)
# Transcribe audioimport soundfile as sf
audio_array, sampling_rate = sf.read("your_audio.wav")
inputs = processor.feature_extractor(
audio_array,
sampling_rate=sampling_rate,
return_tensors="pt",
).input_features.to(device)
with torch.no_grad():
predicted_ids = model.generate(
inputs,
language="ja",
task="transcribe",
)
transcription = processor.tokenizer.batch_decode(
predicted_ids,
skip_special_tokens=True,
)[0]
print(transcription)
Audio Samples & Transcriptions
15 expressive anime voice samples (unseen test set) comparing base openai/whisper-base vs. fine-tuned model.
These clips feature shouting, panic, teasing, anger, and more -- areas where the base model often hallucinates or loops.
1. Cheerful Greeting (sample 1305, 3.4s)
Text
Base Model
先生 おはろございまーす
Fine-tuned
先生、おはろーございまーす
2. Panicked Scream (sample 1340, 5.5s)
Text
Base Model
今日はスカートをめくらないでくださいましー!
Fine-tuned
今日はわわー!スカートを目くらないでくださいましー!
3. Crying for Help — JP/EN Mix (sample 1343, 5.8s)
Text
Base Model
いやー、助けてください、ニーさん! Help you!
Fine-tuned
やー、助けてください、兄さん!ヘルプユー
4. Playful Teasing (sample 1369, 5.9s)
Text
Base Model
うーん? もうしかして、ネギラってくれるのかねー?
Fine-tuned
うーん、もしかして、ネギラってくれるのかねー
5. Over-the-top Shock (sample 1370, 8.0s)
Text
Base Model
な、なに?だいまおがだいにケータイに返ししたです!ですー!
Fine-tuned
な、なにー!だいまおーが、だいに携帯に変身したです、ですー!
6. Disappointed Sigh (sample 1416, 3.8s)
Text
Base Model
えぇ…もう終わりなのか?
Fine-tuned
えぇ~、もう終わりなのか?
7. Stuttering Shock (sample 1469, 4.1s)
Text
Base Model
ななな! 伸びゆきとのは痛いなによ!
Fine-tuned
ななななな、信行きとのは一体なにを!
8. Childish Insistence (sample 1502, 5.9s)
Text
Base Model
いやぁ、行くったら行くの? あるいも一緒に遊ぶですよ!
Fine-tuned
やぁ、行くったら行くの!アンリーも一緒に遊ぶですよ
9. Furious Struggle (sample 1515, 6.0s) ⚠️ Base model hallucinates
Warning: The following 15 samples are from the NSFW split and contain sexually explicit voice acting.
They are included to demonstrate the model's ability to transcribe emotionally extreme and non-standard speech
where the base model completely fails (repetition loops, hallucinations).
NSFW 1. Apologetic Crying (nsfw #6, 6.2s)
Text
Base Model
ううううううくれない!ごめんない!
Fine-tuned
やぁ、やぁああっ、くっくりゃない、ごめんなないっ!
NSFW 2. Teasing / Seductive (nsfw #16, 5.7s)
Text
Base Model
なぁ、ニーモクーン スマした顔して、エッチに興味あるの?
Fine-tuned
なー、ニーもくん、済ました顔して、エッチに興味あるの
NSFW 3. Flustered Embarrassment (nsfw #29, 5.9s)
Text
Base Model
いやいや、ウイキを見出しながら近づかないでください。なんだか一致です。
Fine-tuned
やややん、息を見出しながら近づかないでください。なんだかエッジーです
NSFW 4. Confused Protest (nsfw #33, 6.0s)
Text
Base Model
圧倒状況って何ですか?私エッチな動物さじゃないですよ!
Fine-tuned
熱情聞ってなんですか、私、エッチな動物さんじゃないですよー
NSFW 5. Exhibitionist Excitement (nsfw #46, 6.8s)
Text
Base Model
あさあ、とんどんご覧になって下さい!もっと私を見たからさ!あ、あ、あ、あ!
Fine-tuned
あさー、どんどんごらになってください。もっと私を見たからさ、あっ
NSFW 6. Victory Panting (nsfw #60, 4.8s) ⚠️ Base model hallucinates
Text
Base Model
うううう... (loops 400+ times)
Fine-tuned
はぁああああ…勝ちましたっ!
NSFW 7. Embarrassed Protest (nsfw #75, 5.6s)
Text
Base Model
えぇ、おぉ、いい誘いでください!私、エッチさんじゃないですよ!
Fine-tuned
やふぅぅぅぅぅぅ、やめてください。私、エッチさんじゃないですよー
NSFW 8. Short Shock (nsfw #84, 2.1s) ⚠️ Base model hallucinates
Text
Base Model
えぇえ!あ、嘘!嘘!あ、あ、あ... (loops 200+ times)
Fine-tuned
やや ややっ、そ、そ、そやだっ
NSFW 9. Angry Embarrassment (nsfw #89, 6.7s) ⚠️ Base model hallucinates
Text
Base Model
うっうう... (loops 400+ times)
Fine-tuned
ぱ、バクバクバクバクバクバクッ!エイチのことを言うでない!もーもーもーもーもーもー
NSFW 10. Overwhelmed Scream (nsfw #356, 2.7s) ⚠️ Base model hallucinates
NSFW 13. Sudden Panic (nsfw #439, 3.3s) ⚠️ Base model hallucinates
Text
Base Model
お、お、え、え、え... (loops 200+ times)
Fine-tuned
んっ、あっ、えんっ、えんっ、えんっ、えんっ!
NSFW 14. Desperate Plea (nsfw #456, 7.0s)
Text
Base Model
うえ、あわたって本当にでもやるし、お願いしましょ!このまま、このまま!
Fine-tuned
あ、もうだめ、本当にダメです。お願いします。お願いします。こんなもん、こんなもんもんっ!
NSFW 15. Falling Panic (nsfw #484, 4.6s) ⚠️ Base model hallucinates
Text
Base Model
あ、あ、あ... (loops 200+ times)
Fine-tuned
あっ、あっ、あっ、あっ、あっ、ひゃあんっ、落ちる、落ちちゃいませんっ
Pure Moaning / Non-verbal Vocalization
These 5 samples contain no intelligible words — only moaning, panting, and gasping.
The base model completely hallucinates on all of them. The fine-tuned model also struggles but produces shorter, more contained output.
NSFW 16. Soft Panting (nsfw #136, 8.8s) ⚠️ Both loop
Text
Base Model
ふっふっふっふっふっ... (loops 200+ times)
Fine-tuned
んっ、はっ、はっ、はっ、はっ、はっ... (15 repeats, 44 chars vs base 444)
うぅ、うぅ、うぅ、うぅ、うぅ、うぅ、うぅ、うぅ (8 repeats, 23 chars vs base 444)
NSFW 18. Subdued Gasping (nsfw #272, 8.5s) ⚠️ Both loop
Text
Base Model
あ!あ!あ!あ!... (loops 200+ times)
Fine-tuned
あ、あんっ、あんっ、あんっ、あんっ... (11 repeats, 41 chars vs base 444)
NSFW 19. Exhausted Breathing (nsfw #384, 9.9s) ⚠️ Both loop
Text
Base Model
ううう... (loops 400+ times)
Fine-tuned
んっ、はぁ…はぁ…はぁ…はぁ…はぁ…はぁ… (11 repeats, 35 chars vs base 444)
NSFW 20. High-pitched Climax (nsfw #468, 5.5s) ⚠️ Both loop
Text
Base Model
あ!あ!あ!あ!... (loops 200+ times)
Fine-tuned
あ、あ、ああ、あああ... (loops 400+ times, 883 chars — worst case, similar to base)
Evaluation
Evaluated on unseen test samples from the SFW split using Character Error Rate (CER), the standard metric for Japanese ASR (since Japanese lacks clear word boundaries).
Domain-specific: Optimized for anime-style Japanese speech. Performance on other Japanese speech domains (news, conversation, etc.) may differ from the base model.
Base model size: This is whisper-base (74M params). Larger variants (small, medium, large) would likely achieve better CER.
SFW only: Trained only on the SFW portion of the dataset.
Single speaker variability: Anime speech has high variability in tone, pitch, and speaking style which can still cause errors.
Citation
If you use this model, please cite the original Whisper paper and the training dataset:
@article{radford2022whisper,
title={Robust Speech Recognition via Large-Scale Weak Supervision},
author={Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
journal={arXiv preprint arXiv:2212.04356},
year={2022}
}