orukeet-ggml / README.md
npario's picture JoaoZaokk's picture
Duplicate from JoaoZaokk/orukeet-ggml
cc25fd4
|
Raw History Blame Contribute Delete
1.9 kB
metadata
license: cc-by-sa-4.0
base_model: oruk/orukeet
base_model_relation: quantized
pipeline_tag: automatic-speech-recognition
library_name: whisper.cpp
tags:
  - ggml
  - gguf
  - whisper.cpp
  - automatic-speech-recognition
  - on-device
  - quantized
  - parakeet

orukeet-ggml

GGML conversion of oruk/orukeet (.nemo) for the Parakeet TDT engine that ships inside whisper.cpp ≥ 1.9, in f16 plus q8_0 / q5_0 / q4_0.

Source checkpoint: oruk/orukeet by oruk (fine-tune of NVIDIA Parakeet TDT 0.6B v3) · License: cc-by-sa-4.0 (unchanged; this repo only re-packages the weights) Engine: Load with whisper.cpp ≥ 1.9 parakeet-cli -m <file> (the Parakeet TDT engine that ships inside whisper.cpp). Not compatible with mudler/parakeet.cpp GGUF files.

Files

File Quantization Size Note
ggml-orukeet-f16.bin f16 1256 MB
ggml-orukeet-q4_0.bin q4_0 356 MB
ggml-orukeet-q5_0.bin q5_0 434 MB
ggml-orukeet-q8_0.bin q8_0 669 MB

f16 is the lossless conversion; q8_0 is nearly identical in accuracy at ~55 % of the size; q5_0/q5_k are the phone-friendly choice; q4_* is smallest with a small accuracy cost.

How these were made

Converted from the upstream checkpoint with the engine's own converter, then quantized with the engine's quantizer. Each variant was checked by transcribing short Portuguese and English samples before upload.

Attribution

Weights are derivative works of the upstream model and keep its license. Please cite the original authors (oruk (fine-tune of NVIDIA Parakeet TDT 0.6B v3)). The conversion and hosting here are maintained by JoaoZaokk so that the download links used by the Odysseus / Open WebUI native apps stay stable. No warranty.