Try the live demo on Spaces โ record audio or use pre-loaded anime samples to compare base vs fine-tuned model.
A fine-tuned version of openai/whisper-base for Japanese anime speech recognition.
Model Description
This model was fine-tuned on the SFW split of joujiboi/japanese-anime-speech-v2 โ a dataset of Japanese anime voice clips paired with transcriptions. The goal is to improve Whisper's accuracy on anime-style Japanese speech, which includes expressive intonation, character voices, and domain-specific vocabulary.
Property
Value
Base model
openai/whisper-base (74M params)
Language
Japanese (ja)
Task
transcribe
Domain
Anime speech
Training samples
~269k SFW samples
Training Details
Training was done incrementally across 26 sequential ranges of ~10k samples each (10k โ 269k), with 800 gradient steps per range. This progressive approach allowed monitoring CER at each stage.
Validation loss reduction: 0.892 โ 0.563 (37% drop)
Best CER achieved at 250k-260k range: 30.6%
Validation loss steadily decreased throughout training
Usage
import torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration
# Load model and processor
processor = WhisperProcessor.from_pretrained(
"sky9262/whisper-base-japanese-anime",
language="Japanese",
task="transcribe",
)
model = WhisperForConditionalGeneration.from_pretrained(
"sky9262/whisper-base-japanese-anime"
)
device = "cuda"if torch.cuda.is_available() else"cpu"
model.to(device)
# Transcribe audioimport soundfile as sf
audio_array, sampling_rate = sf.read("your_audio.wav")
inputs = processor.feature_extractor(
audio_array,
sampling_rate=sampling_rate,
return_tensors="pt",
).input_features.to(device)
with torch.no_grad():
predicted_ids = model.generate(
inputs,
language="ja",
task="transcribe",
)
transcription = processor.tokenizer.batch_decode(
predicted_ids,
skip_special_tokens=True,
)[0]
print(transcription)
Audio Samples & Transcriptions
15 expressive anime voice samples (unseen test set) comparing base openai/whisper-base vs. fine-tuned model.
These clips feature shouting, panic, teasing, anger, and more -- areas where the base model often hallucinates or loops.
Warning: The following 15 samples are from the NSFW split and contain sexually explicit voice acting.
They are included to demonstrate the model's ability to transcribe emotionally extreme and non-standard speech
where the base model completely fails (repetition loops, hallucinations).
These 5 samples contain no intelligible words โ only moaning, panting, and gasping.
The base model completely hallucinates on all of them. The fine-tuned model also struggles but produces shorter, more contained output.
ใใใใใใใใใใ... (loops 400+ times, 883 chars โ worst case, similar to base)
Evaluation
Evaluated on unseen test samples from the SFW split using Character Error Rate (CER), the standard metric for Japanese ASR (since Japanese lacks clear word boundaries).
Domain-specific: Optimized for anime-style Japanese speech. Performance on other Japanese speech domains (news, conversation, etc.) may differ from the base model.
Base model size: This is whisper-base (74M params). Larger variants (small, medium, large) would likely achieve better CER.
SFW only: Trained only on the SFW portion of the dataset.
Single speaker variability: Anime speech has high variability in tone, pitch, and speaking style which can still cause errors.
Citation
If you use this model, please cite the original Whisper paper and the training dataset:
@article{radford2022whisper,
title={Robust Speech Recognition via Large-Scale Weak Supervision},
author={Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
journal={arXiv preprint arXiv:2212.04356},
year={2022}
}