Instructions to use phaxz10/belle-whisper-large-v3-turbo-zh_timestamped with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use phaxz10/belle-whisper-large-v3-turbo-zh_timestamped with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('automatic-speech-recognition', 'phaxz10/belle-whisper-large-v3-turbo-zh_timestamped');
Belle-whisper-large-v3-turbo-zh โ ONNX (_timestamped)
ONNX weights for BELLE-2/Belle-whisper-large-v3-turbo-zh,
a Mandarin fine-tune of openai/whisper-large-v3-turbo, for use with
Transformers.js.
This is a _timestamped export: the decoder is exported with output_attentions=True so its
graph emits cross_attentions.{0..3}, and generation_config.json carries alignment_heads.
Without both, return_timestamps: 'word' throws "Model outputs must contain cross attentions".
Produced with the same recipe as
onnx-community/whisper-large-v3-turbo_timestamped
(transformers.js scripts/convert.py at tag 3.8.1, --output_attentions).
Files
| file | size |
|---|---|
onnx/encoder_model_fp16.onnx |
1274.40 MB |
onnx/decoder_model_merged_q4.onnx |
334.05 MB |
onnx/decoder_model_merged_quantized.onnx (q8) |
438.12 MB |
No .onnx_data: every weight file is self-contained and under the 2 GB protobuf limit.
Usage
import { pipeline } from '@huggingface/transformers'
const asr = await pipeline(
'automatic-speech-recognition',
'phaxz10/belle-whisper-large-v3-turbo-zh_timestamped',
{ dtype: { encoder_model: 'fp16', decoder_model_merged: 'q4' }, device: 'webgpu' },
)
const out = await asr(audio, { language: 'chinese', task: 'transcribe', return_timestamps: true })
Known limitation: word timestamps on short audio
return_timestamps: true (segment timings, read off Whisper's <|t|> tokens) is accurate at
every clip length tested, from 0.5 s up.
return_timestamps: 'word' (cross-attention DTW) is accurate only from roughly 8 s of audio
upward. Below that the DTW path latches onto the padded region of the 30 s mel window and returns
times far outside the clip (e.g. [16.42, 24.14] for a 4 s clip, [29.82, 29.92] for a 1.6 s one).
The stock onnx-community/whisper-large-v3-turbo_timestamped export is correct at all lengths on
the same clips, so this comes from the Mandarin fine-tune, not from the conversion.
alignment_heads here are the stock whisper-large-v3-turbo heads, which the base model already
ships and which scripts/extra/whisper.py also resolves to. Substituting other head sets (all of
layer 3, all of layers 2+3, all 80 heads) was tested and none fixed the short-clip case; several
were worse. Prefer return_timestamps: true unless your chunks are reliably longer than ~8 s.
Base model quality (from the base model card)
| benchmark | whisper-large-v3-turbo |
Belle fine-tune |
|---|---|---|
| AISHELL-1 (CER) | 8.64 | 3.07 |
| WenetSpeech-meeting (CER) | 20.31 | 13.36 |
- Downloads last month
- 45
Model tree for phaxz10/belle-whisper-large-v3-turbo-zh_timestamped
Base model
openai/whisper-large-v3