File size: 2,831 Bytes
831d766
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cfd5702
831d766
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
---
pipeline_tag: automatic-speech-recognition
language: amh
license: apache-2.0
tags:
  - trimmed
library_name: transformers
base_model: openai/whisper-medium
base_model_relation: quantized
datasets:
  - lbourdois/fineweb-2-trimming
---

# whisper-medium-amh-16384
This model is a **4.76% medium** version of [openai/whisper-medium](https://huggingface.co/openai/whisper-medium) optimized for **Amharic** language via vocabulary size reduction using the [trimming](https://huggingface.co/blog/lbourdois/introduction-to-trimming) method.  
This trimmed model should perform similarly to the original model with only 16,384 tokens and a much smaller memory footprint. However, it may not perform well for other languages as tokens not commonly used in the selected languages were removed from the vocabulary.

## Model Statistics
| Metric | Original | Trimmed | Reduction |
|--------|----------|---------|-----------|
| **Vocabulary size** | 51,865 tokens | 16,384 tokens | **68.41%** |
| **Model size** | 763,857,920 params | 727,525,376 params | **4.76%** |

![image](https://raw.githubusercontent.com/lbourdois/blog/refs/heads/master/assets/images/Trimming/whisper-medium-16384.png)

## Mining Dataset Statistics
- **Number of texts used for mining**: 200,000 texts  
- **Dataset**: [lbourdois/fineweb-2-trimming](https://huggingface.co/datasets/lbourdois/fineweb-2-trimming)

## Usage
```python
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline
import librosa

# Pipeline function
processor = AutoProcessor.from_pretrained("alphaedge-ai/whisper-medium-amh-16384")
pipe = pipeline(
    "automatic-speech-recognition",
    model="alphaedge-ai/whisper-medium-amh-16384",
    tokenizer=processor.tokenizer,
    feature_extractor=processor.feature_extractor,
    generate_kwargs={"language": "amharic", "task": "transcribe"},
)

# Loading and resampling at 16 kHz (required by Whisper)
audio_array, sampling_rate = librosa.load(audio_path, sr=16000)

# Result
result = pipe(audio_array)
print("Transcription :", result["text"])
```

## Citations

#### Whisper
```
@misc{radford2022whisper,
  doi = {10.48550/ARXIV.2212.04356},
  url = {https://arxiv.org/abs/2212.04356},
  author = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
  title = {Robust Speech Recognition via Large-Scale Weak Supervision},
  publisher = {arXiv},
  year = {2022},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
```

#### Trimming blog post
```
@misc{hf_blogpost_trimming,
      title={Introduction to Trimming}, 
      author={Loïck BOURDOIS and Tom AARSEN and Bram VANROY and Christopher AKIKI and Woojun JUNG and Manuel ROMERO and Prithiv SAKTHI},
      year={2026},
      url={https://huggingface.co/blog/lbourdois/introduction-to-trimming}, 
}
```