yuriyvnv's picture
Add model card
dd94971 verified
|
Raw
History Blame Contribute Delete
4.45 kB
metadata
language:
  - pl
license: cc-by-4.0
library_name: nemo
tags:
  - automatic-speech-recognition
  - speech
  - nemo
  - parakeet
  - fastconformer
  - tdt
  - polish
  - nvidia
  - common-voice
  - bigos
  - fine-tuned
datasets:
  - amu-cai/pl-asr-bigos-v2
  - fixie-ai/common_voice_17_0
base_model: nvidia/parakeet-tdt-0.6b-v3
pipeline_tag: automatic-speech-recognition
model-index:
  - name: parakeet-tdt-0.6b-polish
    results:
      - task:
          type: automatic-speech-recognition
          name: Speech Recognition
        dataset:
          name: Common Voice 17.0 (pl) - Validation
          type: fixie-ai/common_voice_17_0
          config: pl
          split: validation
        metrics:
          - type: wer
            value: 6.07
            name: Val WER
      - task:
          type: automatic-speech-recognition
          name: Speech Recognition
        dataset:
          name: Common Voice 17.0 (pl) - Test
          type: fixie-ai/common_voice_17_0
          config: pl
          split: test
        metrics:
          - type: wer
            value: 11.81
            name: Test WER
          - type: cer
            value: 2.72
            name: Test CER

Parakeet-TDT-0.6B Polish

A Polish automatic speech recognition (ASR) model fine-tuned from nvidia/parakeet-tdt-0.6b-v3.

Model Details

Property Value
Base model nvidia/parakeet-tdt-0.6b-v3
Architecture FastConformer-TDT (600M params)
Language Polish (pl)
Input 16 kHz mono audio
Output Polish text with punctuation and capitalization
License CC-BY-4.0

Evaluation Results

Evaluated on Common Voice 17.0 Polish (raw text, no normalization):

Split WER CER Samples
Validation 6.07% -- --
Test 11.81% 2.72% 9,230

Training

Fine-tuned on a curated subset of the BIGOS v2 benchmark, filtered to retain only sources with proper casing and punctuation:

Validation uses the BIGOS v2 validation split (same source filtering). Test evaluation uses Common Voice 17.0 Polish (independent test set).

Training Configuration

Parameter Value
Optimizer AdamW
Learning rate 5e-5 (cosine annealing)
Warmup 10% of total steps
Batch size 32
Precision bf16-mixed
Gradient clipping 1.0
Early stopping 10 epochs patience on val WER
Best epoch 21

Usage

Installation

pip install nemo_toolkit[asr]

Transcribe Audio

import nemo.collections.asr as nemo_asr

# Load model
asr_model = nemo_asr.models.ASRModel.from_pretrained(
    model_name="yuriyvnv/parakeet-tdt-0.6b-polish"
)

# Transcribe
output = asr_model.transcribe(["audio.wav"])
print(output[0].text)

Transcribe with Timestamps

output = asr_model.transcribe(["audio.wav"], timestamps=True)

for stamp in output[0].timestamp["segment"]:
    print(f"{stamp['start']:.1f}s - {stamp['end']:.1f}s : {stamp['segment']}")

Long-Form Audio

For audio longer than 24 minutes, enable local attention:

asr_model.change_attention_model(
    self_attention_model="rel_pos_local_attn",
    att_context_size=[256, 256],
)
output = asr_model.transcribe(["long_audio.wav"])

Intended Use

This model is designed for transcribing Polish speech to text. It works best on:

  • Read speech and conversational Polish
  • Audio recorded at 16 kHz or higher
  • Segments up to 24 minutes (or longer with local attention enabled)

Limitations

  • Training data is sourced from read speech (audiobooks, Common Voice read prompts) and short banking dialogs; performance may differ on spontaneous or heavily accented speech
  • The model preserves punctuation and capitalization as seen in training data
  • Not suitable for real-time streaming without additional configuration