sky9262's picture
Upload README.md with huggingface_hub
8149cee verified
|
Raw History Blame Contribute Delete
20.2 kB
metadata
language:
  - ja
license: apache-2.0
base_model: openai/whisper-base
tags:
  - whisper
  - japanese
  - anime
  - speech-recognition
  - fine-tuned
  - automatic-speech-recognition
datasets:
  - joujiboi/japanese-anime-speech-v2
metrics:
  - cer
pipeline_tag: automatic-speech-recognition
widget:
  - src: >-
      https://huggingface.co/sky9262/whisper-base-japanese-anime/resolve/main/samples/sample_1305.wav
    example_title: Cheerful Greeting
  - src: >-
      https://huggingface.co/sky9262/whisper-base-japanese-anime/resolve/main/samples/sample_1370.wav
    example_title: Over-the-top Shock
  - src: >-
      https://huggingface.co/sky9262/whisper-base-japanese-anime/resolve/main/samples/sample_1740.wav
    example_title: Desperate Urgency
model-index:
  - name: whisper-base-japanese-anime
    results:
      - task:
          type: automatic-speech-recognition
          name: Automatic Speech Recognition
        dataset:
          name: joujiboi/japanese-anime-speech-v2
          type: joujiboi/japanese-anime-speech-v2
          split: sfw
        metrics:
          - type: cer
            value: 31.2
            name: CER (final)
          - type: cer
            value: 44.3
            name: CER (baseline)

whisper-base-japanese-anime

Try the live demo on Spaces โ€” record audio or use pre-loaded anime samples to compare base vs fine-tuned model.

A fine-tuned version of openai/whisper-base for Japanese anime speech recognition.

Model Description

This model was fine-tuned on the SFW split of joujiboi/japanese-anime-speech-v2 โ€” a dataset of Japanese anime voice clips paired with transcriptions. The goal is to improve Whisper's accuracy on anime-style Japanese speech, which includes expressive intonation, character voices, and domain-specific vocabulary.

Property Value
Base model openai/whisper-base (74M params)
Language Japanese (ja)
Task transcribe
Domain Anime speech
Training samples ~269k SFW samples

Training Details

Training was done incrementally across 26 sequential ranges of ~10k samples each (10k โ†’ 269k), with 800 gradient steps per range. This progressive approach allowed monitoring CER at each stage.

Training Hyperparameters

Parameter Value
Learning rate 1e-6
LR scheduler Linear
Warmup steps 200
Batch size (train) 20
Batch size (eval) 16
Gradient accumulation 1
Max steps per range 800
Precision FP16
Optimizer AdamW (fused)
Weight decay 0.0
Total training ranges 26

Training Progression

Stage Samples Val Loss CER
Baseline (10k-20k) 10k 0.892 44.3%
Mid (120k-130k) 120k 0.740 33.5%
Late (210k-220k) 210k 0.621 31.2%
Final (260k-269k) 269k 0.563 31.2%

Key Results

  • CER reduction: 44.3% โ†’ 31.2% (~30% relative improvement)
  • Validation loss reduction: 0.892 โ†’ 0.563 (37% drop)
  • Best CER achieved at 250k-260k range: 30.6%
  • Validation loss steadily decreased throughout training

Usage

import torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration

# Load model and processor
processor = WhisperProcessor.from_pretrained(
    "sky9262/whisper-base-japanese-anime",
    language="Japanese",
    task="transcribe",
)
model = WhisperForConditionalGeneration.from_pretrained(
    "sky9262/whisper-base-japanese-anime"
)

device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)

# Transcribe audio
import soundfile as sf

audio_array, sampling_rate = sf.read("your_audio.wav")

inputs = processor.feature_extractor(
    audio_array,
    sampling_rate=sampling_rate,
    return_tensors="pt",
).input_features.to(device)

with torch.no_grad():
    predicted_ids = model.generate(
        inputs,
        language="ja",
        task="transcribe",
    )

transcription = processor.tokenizer.batch_decode(
    predicted_ids,
    skip_special_tokens=True,
)[0]

print(transcription)

Audio Samples & Transcriptions

15 expressive anime voice samples (unseen test set) comparing base openai/whisper-base vs. fine-tuned model. These clips feature shouting, panic, teasing, anger, and more -- areas where the base model often hallucinates or loops.

1. Cheerful Greeting (sample 1305, 3.4s)

Text
Base Model ๅ…ˆ็”Ÿ ใŠใฏใ‚ใ”ใ–ใ„ใพใƒผใ™
Fine-tuned ๅ…ˆ็”Ÿใ€ใŠใฏใ‚ใƒผใ”ใ–ใ„ใพใƒผใ™

2. Panicked Scream (sample 1340, 5.5s)

Text
Base Model ไปŠๆ—ฅใฏใ‚นใ‚ซใƒผใƒˆใ‚’ใ‚ใใ‚‰ใชใ„ใงใใ ใ•ใ„ใพใ—ใƒผ!
Fine-tuned ไปŠๆ—ฅใฏใ‚ใ‚ใƒผ๏ผใ‚นใ‚ซใƒผใƒˆใ‚’็›ฎใใ‚‰ใชใ„ใงใใ ใ•ใ„ใพใ—ใƒผ๏ผ

3. Crying for Help โ€” JP/EN Mix (sample 1343, 5.8s)

Text
Base Model ใ„ใ‚„ใƒผใ€ๅŠฉใ‘ใฆใใ ใ•ใ„ใ€ใƒ‹ใƒผใ•ใ‚“! Help you!
Fine-tuned ใ‚„ใƒผใ€ๅŠฉใ‘ใฆใใ ใ•ใ„ใ€ๅ…„ใ•ใ‚“๏ผใƒ˜ใƒซใƒ—ใƒฆใƒผ

4. Playful Teasing (sample 1369, 5.9s)

Text
Base Model ใ†ใƒผใ‚“? ใ‚‚ใ†ใ—ใ‹ใ—ใฆใ€ใƒใ‚ฎใƒฉใฃใฆใใ‚Œใ‚‹ใฎใ‹ใญใƒผ?
Fine-tuned ใ†ใƒผใ‚“ใ€ใ‚‚ใ—ใ‹ใ—ใฆใ€ใƒใ‚ฎใƒฉใฃใฆใใ‚Œใ‚‹ใฎใ‹ใญใƒผ

5. Over-the-top Shock (sample 1370, 8.0s)

Text
Base Model ใชใ€ใชใซ?ใ ใ„ใพใŠใŒใ ใ„ใซใ‚ฑใƒผใ‚ฟใ‚คใซ่ฟ”ใ—ใ—ใŸใงใ™!ใงใ™ใƒผ!
Fine-tuned ใชใ€ใชใซใƒผ๏ผใ ใ„ใพใŠใƒผใŒใ€ใ ใ„ใซๆบๅธฏใซๅค‰่บซใ—ใŸใงใ™ใ€ใงใ™ใƒผ๏ผ

6. Disappointed Sigh (sample 1416, 3.8s)

Text
Base Model ใˆใ‡โ€ฆใ‚‚ใ†็ต‚ใ‚ใ‚Šใชใฎใ‹?
Fine-tuned ใˆใ‡๏ฝžใ€ใ‚‚ใ†็ต‚ใ‚ใ‚Šใชใฎใ‹๏ผŸ

7. Stuttering Shock (sample 1469, 4.1s)

Text
Base Model ใชใชใช! ไผธใณใ‚†ใใจใฎใฏ็—›ใ„ใชใซใ‚ˆ!
Fine-tuned ใชใชใชใชใชใ€ไฟก่กŒใใจใฎใฏไธ€ไฝ“ใชใซใ‚’๏ผ

8. Childish Insistence (sample 1502, 5.9s)

Text
Base Model ใ„ใ‚„ใใ€่กŒใใฃใŸใ‚‰่กŒใใฎ? ใ‚ใ‚‹ใ„ใ‚‚ไธ€็ท’ใซ้Šใถใงใ™ใ‚ˆ!
Fine-tuned ใ‚„ใใ€่กŒใใฃใŸใ‚‰่กŒใใฎ๏ผใ‚ขใƒณใƒชใƒผใ‚‚ไธ€็ท’ใซ้Šใถใงใ™ใ‚ˆ

9. Furious Struggle (sample 1515, 6.0s) โš ๏ธ Base model hallucinates

Text
Base Model ใˆใ€่ฉฑใ›ใƒ–ใƒฌใƒ•ใƒขใƒŽ!ใŠใ‰ใ€ใ‚‚ใ€ใ‚ขใƒ‰ใƒ‡ใ‚ซใ‚คใ‚นใฎใ‹!ใธใฃใธใฃใธใฃใธใฃใธใฃใธใฃใธใฃ... (loops 200+ times)
Fine-tuned ใˆใธใฃใ€่ฉฑใ›ใถใ‚Œ็‰ฉใ€ใŠใ†ใ‚‚ๅพŒใง่ฟ”ใ™ใฎใ‹ใฃใ€ใตใ‚ใ‚ใ‚ใฃ๏ผ

10. Whining / Pouting (sample 1542, 5.0s)

Text
Base Model ใ ใฃใฆใ ใฃใฆ!ๆˆๅŠŸใŒใ‹ใ‚ใ„ใใชใ„ใ‚“ใ ใ‚‚ใ‚“!
Fine-tuned ใ ใฃใฆใ ใฃใฆใ€ใ›ใฃใใใŒๅฏๆ„›ใใชใ„ใ‚“ใ ใ‚‚ใ‚“๏ผ

11. Alarmed Demand (sample 1627, 2.9s)

Text
Base Model ใพใ‚“ใ˜ใ‚ƒใ€็—›ใ„ไฝ•ใŒใ‚ใฃใŸ?
Fine-tuned ใชใ‚“ใ˜ใ‚ƒใ€ไธ€ไฝ“ไฝ•ใŒใ‚ใฃใŸ๏ผ

12. Cute Excitement (sample 1661, 3.7s) โš ๏ธ Base model hallucinates

Text
Base Model ใ‚‚ใƒผใ‚‚ใƒผใ‚‚ใƒผใ‚‚ใƒผใ‚‚ใƒผใ‚‚ใƒผ... (loops 200+ times)
Fine-tuned ใ‚„ใตใ…๏ผใ‚‚ใ†ใตใ‚‚ใตใ‚‚ใตใงใ™ใƒผ

13. Angry Scolding (sample 1728, 4.6s)

Text
Base Model ใŸใ‹ใ‚‚ใฎใฏใ€ใใ‚“ใชใ“ใจใŒใงใใŸใ‚‰ใ€็ฌ‘้ก”ใŒ็‰นใซใ‚„ใฃใฆใŠใ‚‹ใ‚!
Fine-tuned ้ซ˜่€…ใฎใ€ใใ‚“ใชใ“ใจใŒใงใใŸใ‚‰ใ€ๅฆพใŒ็‰นใซใ‚„ใฃใฆใŠใ‚‹ใ‚

14. Desperate Urgency (sample 1740, 4.5s)

Text
Base Model ใƒใ‚คใƒคใƒˆ!ใ—ใฃใ‹ใ‚Šใ—ใ‚!ใƒใ‚คใƒคใƒˆ!
Fine-tuned ้šผๆ–—ใ€ใ—ใฃใ‹ใ‚Šใ—ใ‚ใ€้šผๆ–—ใฃ๏ผ

15. Elongated Refusal (sample 1768, 3.0s)

Text
Base Model ใ‚“ใ ใƒผใ‚ใƒผใงใ™ใƒผ
Fine-tuned ใƒ€ใƒกใงใ™๏ผ

โš ๏ธ NSFW Samples (Click to expand) โ€” Contains explicit anime voice content. Viewer discretion advised.

Warning: The following 15 samples are from the NSFW split and contain sexually explicit voice acting. They are included to demonstrate the model's ability to transcribe emotionally extreme and non-standard speech where the base model completely fails (repetition loops, hallucinations).

NSFW 1. Apologetic Crying (nsfw #6, 6.2s)

Text
Base Model ใ†ใ†ใ†ใ†ใ†ใ†ใใ‚Œใชใ„!ใ”ใ‚ใ‚“ใชใ„!
Fine-tuned ใ‚„ใใ€ใ‚„ใใ‚ใ‚ใฃใ€ใใฃใใ‚Šใ‚ƒใชใ„ใ€ใ”ใ‚ใ‚“ใชใชใ„ใฃ๏ผ

NSFW 2. Teasing / Seductive (nsfw #16, 5.7s)

Text
Base Model ใชใใ€ใƒ‹ใƒผใƒขใ‚ฏใƒผใƒณ ใ‚นใƒžใ—ใŸ้ก”ใ—ใฆใ€ใ‚จใƒƒใƒใซ่ˆˆๅ‘ณใ‚ใ‚‹ใฎ?
Fine-tuned ใชใƒผใ€ใƒ‹ใƒผใ‚‚ใใ‚“ใ€ๆธˆใพใ—ใŸ้ก”ใ—ใฆใ€ใ‚จใƒƒใƒใซ่ˆˆๅ‘ณใ‚ใ‚‹ใฎ

NSFW 3. Flustered Embarrassment (nsfw #29, 5.9s)

Text
Base Model ใ„ใ‚„ใ„ใ‚„ใ€ใ‚ฆใ‚คใ‚ญใ‚’่ฆ‹ๅ‡บใ—ใชใŒใ‚‰่ฟ‘ใฅใ‹ใชใ„ใงใใ ใ•ใ„ใ€‚ใชใ‚“ใ ใ‹ไธ€่‡ดใงใ™ใ€‚
Fine-tuned ใ‚„ใ‚„ใ‚„ใ‚“ใ€ๆฏใ‚’่ฆ‹ๅ‡บใ—ใชใŒใ‚‰่ฟ‘ใฅใ‹ใชใ„ใงใใ ใ•ใ„ใ€‚ใชใ‚“ใ ใ‹ใ‚จใƒƒใ‚ธใƒผใงใ™

NSFW 4. Confused Protest (nsfw #33, 6.0s)

Text
Base Model ๅœงๅ€’็Šถๆณใฃใฆไฝ•ใงใ™ใ‹?็งใ‚จใƒƒใƒใชๅ‹•็‰ฉใ•ใ˜ใ‚ƒใชใ„ใงใ™ใ‚ˆ!
Fine-tuned ็†ฑๆƒ…่žใฃใฆใชใ‚“ใงใ™ใ‹ใ€็งใ€ใ‚จใƒƒใƒใชๅ‹•็‰ฉใ•ใ‚“ใ˜ใ‚ƒใชใ„ใงใ™ใ‚ˆใƒผ

NSFW 5. Exhibitionist Excitement (nsfw #46, 6.8s)

Text
Base Model ใ‚ใ•ใ‚ใ€ใจใ‚“ใฉใ‚“ใ”่ฆงใซใชใฃใฆไธ‹ใ•ใ„!ใ‚‚ใฃใจ็งใ‚’่ฆ‹ใŸใ‹ใ‚‰ใ•!ใ‚ใ€ใ‚ใ€ใ‚ใ€ใ‚!
Fine-tuned ใ‚ใ•ใƒผใ€ใฉใ‚“ใฉใ‚“ใ”ใ‚‰ใซใชใฃใฆใใ ใ•ใ„ใ€‚ใ‚‚ใฃใจ็งใ‚’่ฆ‹ใŸใ‹ใ‚‰ใ•ใ€ใ‚ใฃ

NSFW 6. Victory Panting (nsfw #60, 4.8s) โš ๏ธ Base model hallucinates

Text
Base Model ใ†ใ†ใ†ใ†... (loops 400+ times)
Fine-tuned ใฏใใ‚ใ‚ใ‚ใ‚โ€ฆๅ‹ใกใพใ—ใŸใฃ๏ผ

NSFW 7. Embarrassed Protest (nsfw #75, 5.6s)

Text
Base Model ใˆใ‡ใ€ใŠใ‰ใ€ใ„ใ„่ช˜ใ„ใงใใ ใ•ใ„!็งใ€ใ‚จใƒƒใƒใ•ใ‚“ใ˜ใ‚ƒใชใ„ใงใ™ใ‚ˆ!
Fine-tuned ใ‚„ใตใ…ใ…ใ…ใ…ใ…ใ…ใ€ใ‚„ใ‚ใฆใใ ใ•ใ„ใ€‚็งใ€ใ‚จใƒƒใƒใ•ใ‚“ใ˜ใ‚ƒใชใ„ใงใ™ใ‚ˆใƒผ

NSFW 8. Short Shock (nsfw #84, 2.1s) โš ๏ธ Base model hallucinates

Text
Base Model ใˆใ‡ใˆ!ใ‚ใ€ๅ˜˜!ๅ˜˜!ใ‚ใ€ใ‚ใ€ใ‚... (loops 200+ times)
Fine-tuned ใ‚„ใ‚„ ใ‚„ใ‚„ใฃใ€ใใ€ใใ€ใใ‚„ใ ใฃ

NSFW 9. Angry Embarrassment (nsfw #89, 6.7s) โš ๏ธ Base model hallucinates

Text
Base Model ใ†ใฃใ†ใ†... (loops 400+ times)
Fine-tuned ใฑใ€ใƒใ‚ฏใƒใ‚ฏใƒใ‚ฏใƒใ‚ฏใƒใ‚ฏใƒใ‚ฏใƒƒ๏ผใ‚จใ‚คใƒใฎใ“ใจใ‚’่จ€ใ†ใงใชใ„๏ผใ‚‚ใƒผใ‚‚ใƒผใ‚‚ใƒผใ‚‚ใƒผใ‚‚ใƒผใ‚‚ใƒผ

NSFW 10. Overwhelmed Scream (nsfw #356, 2.7s) โš ๏ธ Base model hallucinates

Text
Base Model ใ†ใ†... (loops 400+ times)
Fine-tuned ใซใ‚…ใ†ใ…๏ฝž๏ฝž๏ฝž๏ฝž๏ฝž๏ฝž๏ฝž๏ฝžใฃ

NSFW 11. Surprised Protest (nsfw #359, 6.9s) โš ๏ธ Fine-tuned model loops

Text
Base Model ใŠใ„!ใ„ใ‚„ใ!ใ‚ใ!ใ‚ใ!ใ„ใ‚„ใ€ไปŠใใ“ใชใ„ใ˜ใ‚ƒใ‚“!
Fine-tuned ใ‚“ใซใ‚ƒใฃใ€ใ‚„ใฃใ€ใ‚„ใฃใ€ใ‚„ใฃ... (loops 140+ times, 444 chars)

NSFW 12. Intense Overwhelm (nsfw #404, 9.2s) โš ๏ธ Both models hallucinate

Text
Base Model ใ‚ใƒผใƒผใƒผ... (loops 400+ times)
Fine-tuned ใ‚ใ‚ใ‚ใ‚ใ‚ใ‚ใ‚ใ‚ใ‚ใ‚ใฃใ€ใใ‚Œใ€ใใ‚Œใ€ใƒ€ใƒกใ€ใƒ€ใƒกใƒƒ

NSFW 13. Sudden Panic (nsfw #439, 3.3s) โš ๏ธ Base model hallucinates

Text
Base Model ใŠใ€ใŠใ€ใˆใ€ใˆใ€ใˆ... (loops 200+ times)
Fine-tuned ใ‚“ใฃใ€ใ‚ใฃใ€ใˆใ‚“ใฃใ€ใˆใ‚“ใฃใ€ใˆใ‚“ใฃใ€ใˆใ‚“ใฃ๏ผ

NSFW 14. Desperate Plea (nsfw #456, 7.0s)

Text
Base Model ใ†ใˆใ€ใ‚ใ‚ใŸใฃใฆๆœฌๅฝ“ใซใงใ‚‚ใ‚„ใ‚‹ใ—ใ€ใŠ้ก˜ใ„ใ—ใพใ—ใ‚‡!ใ“ใฎใพใพใ€ใ“ใฎใพใพ!
Fine-tuned ใ‚ใ€ใ‚‚ใ†ใ ใ‚ใ€ๆœฌๅฝ“ใซใƒ€ใƒกใงใ™ใ€‚ใŠ้ก˜ใ„ใ—ใพใ™ใ€‚ใŠ้ก˜ใ„ใ—ใพใ™ใ€‚ใ“ใ‚“ใชใ‚‚ใ‚“ใ€ใ“ใ‚“ใชใ‚‚ใ‚“ใ‚‚ใ‚“ใฃ๏ผ

NSFW 15. Falling Panic (nsfw #484, 4.6s) โš ๏ธ Base model hallucinates

Text
Base Model ใ‚ใ€ใ‚ใ€ใ‚... (loops 200+ times)
Fine-tuned ใ‚ใฃใ€ใ‚ใฃใ€ใ‚ใฃใ€ใ‚ใฃใ€ใ‚ใฃใ€ใฒใ‚ƒใ‚ใ‚“ใฃใ€่ฝใกใ‚‹ใ€่ฝใกใกใ‚ƒใ„ใพใ›ใ‚“ใฃ

Pure Moaning / Non-verbal Vocalization

These 5 samples contain no intelligible words โ€” only moaning, panting, and gasping. The base model completely hallucinates on all of them. The fine-tuned model also struggles but produces shorter, more contained output.

NSFW 16. Soft Panting (nsfw #136, 8.8s) โš ๏ธ Both loop

Text
Base Model ใตใฃใตใฃใตใฃใตใฃใตใฃ... (loops 200+ times)
Fine-tuned ใ‚“ใฃใ€ใฏใฃใ€ใฏใฃใ€ใฏใฃใ€ใฏใฃใ€ใฏใฃ... (15 repeats, 44 chars vs base 444)

NSFW 17. Intense Climax Scream (nsfw #176, 5.3s) โš ๏ธ Both loop

Text
Base Model ใ†ใฃใ†ใฃใ†ใฃใ†ใฃ... (loops 200+ times)
Fine-tuned ใ†ใ…ใ€ใ†ใ…ใ€ใ†ใ…ใ€ใ†ใ…ใ€ใ†ใ…ใ€ใ†ใ…ใ€ใ†ใ…ใ€ใ†ใ… (8 repeats, 23 chars vs base 444)

NSFW 18. Subdued Gasping (nsfw #272, 8.5s) โš ๏ธ Both loop

Text
Base Model ใ‚!ใ‚!ใ‚!ใ‚!... (loops 200+ times)
Fine-tuned ใ‚ใ€ใ‚ใ‚“ใฃใ€ใ‚ใ‚“ใฃใ€ใ‚ใ‚“ใฃใ€ใ‚ใ‚“ใฃ... (11 repeats, 41 chars vs base 444)

NSFW 19. Exhausted Breathing (nsfw #384, 9.9s) โš ๏ธ Both loop

Text
Base Model ใ†ใ†ใ†... (loops 400+ times)
Fine-tuned ใ‚“ใฃใ€ใฏใโ€ฆใฏใโ€ฆใฏใโ€ฆใฏใโ€ฆใฏใโ€ฆใฏใโ€ฆ (11 repeats, 35 chars vs base 444)

NSFW 20. High-pitched Climax (nsfw #468, 5.5s) โš ๏ธ Both loop

Text
Base Model ใ‚!ใ‚!ใ‚!ใ‚!... (loops 200+ times)
Fine-tuned ใ‚ใ€ใ‚ใ€ใ‚ใ‚ใ€ใ‚ใ‚ใ‚... (loops 400+ times, 883 chars โ€” worst case, similar to base)

Evaluation

Evaluated on unseen test samples from the SFW split using Character Error Rate (CER), the standard metric for Japanese ASR (since Japanese lacks clear word boundaries).

CER is computed as:

CER = (substitutions + insertions + deletions) / len(reference)

Limitations

  • Domain-specific: Optimized for anime-style Japanese speech. Performance on other Japanese speech domains (news, conversation, etc.) may differ from the base model.
  • Base model size: This is whisper-base (74M params). Larger variants (small, medium, large) would likely achieve better CER.
  • SFW only: Trained only on the SFW portion of the dataset.
  • Single speaker variability: Anime speech has high variability in tone, pitch, and speaking style which can still cause errors.

Citation

If you use this model, please cite the original Whisper paper and the training dataset:

@article{radford2022whisper,
  title={Robust Speech Recognition via Large-Scale Weak Supervision},
  author={Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
  journal={arXiv preprint arXiv:2212.04356},
  year={2022}
}

Dataset: joujiboi/japanese-anime-speech-v2