Duplicate from JoaoZaokk/orukeet-ggml
Browse filesCo-authored-by: Joao Zao <JoaoZaokk@users.noreply.huggingface.co>
- .gitattributes +35 -0
- README.md +41 -0
- UPSTREAM-README.md +205 -0
- VERIFY.txt +16 -0
- ggml-orukeet-f16.bin +3 -0
- ggml-orukeet-q4_0.bin +3 -0
- ggml-orukeet-q5_0.bin +3 -0
- ggml-orukeet-q8_0.bin +3 -0
.gitattributes
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
*.7z filter=lfs diff=lfs merge=lfs -text
|
| 2 |
+
*.arrow filter=lfs diff=lfs merge=lfs -text
|
| 3 |
+
*.bin filter=lfs diff=lfs merge=lfs -text
|
| 4 |
+
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
| 5 |
+
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
| 6 |
+
*.ftz filter=lfs diff=lfs merge=lfs -text
|
| 7 |
+
*.gz filter=lfs diff=lfs merge=lfs -text
|
| 8 |
+
*.h5 filter=lfs diff=lfs merge=lfs -text
|
| 9 |
+
*.joblib filter=lfs diff=lfs merge=lfs -text
|
| 10 |
+
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
| 11 |
+
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
| 12 |
+
*.model filter=lfs diff=lfs merge=lfs -text
|
| 13 |
+
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
| 14 |
+
*.npy filter=lfs diff=lfs merge=lfs -text
|
| 15 |
+
*.npz filter=lfs diff=lfs merge=lfs -text
|
| 16 |
+
*.onnx filter=lfs diff=lfs merge=lfs -text
|
| 17 |
+
*.ot filter=lfs diff=lfs merge=lfs -text
|
| 18 |
+
*.parquet filter=lfs diff=lfs merge=lfs -text
|
| 19 |
+
*.pb filter=lfs diff=lfs merge=lfs -text
|
| 20 |
+
*.pickle filter=lfs diff=lfs merge=lfs -text
|
| 21 |
+
*.pkl filter=lfs diff=lfs merge=lfs -text
|
| 22 |
+
*.pt filter=lfs diff=lfs merge=lfs -text
|
| 23 |
+
*.pth filter=lfs diff=lfs merge=lfs -text
|
| 24 |
+
*.rar filter=lfs diff=lfs merge=lfs -text
|
| 25 |
+
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 26 |
+
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
| 27 |
+
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
| 28 |
+
*.tar filter=lfs diff=lfs merge=lfs -text
|
| 29 |
+
*.tflite filter=lfs diff=lfs merge=lfs -text
|
| 30 |
+
*.tgz filter=lfs diff=lfs merge=lfs -text
|
| 31 |
+
*.wasm filter=lfs diff=lfs merge=lfs -text
|
| 32 |
+
*.xz filter=lfs diff=lfs merge=lfs -text
|
| 33 |
+
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
+
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
+
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: cc-by-sa-4.0
|
| 3 |
+
base_model: oruk/orukeet
|
| 4 |
+
base_model_relation: quantized
|
| 5 |
+
pipeline_tag: automatic-speech-recognition
|
| 6 |
+
library_name: whisper.cpp
|
| 7 |
+
tags:
|
| 8 |
+
- ggml
|
| 9 |
+
- gguf
|
| 10 |
+
- whisper.cpp
|
| 11 |
+
- automatic-speech-recognition
|
| 12 |
+
- on-device
|
| 13 |
+
- quantized
|
| 14 |
+
- parakeet
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# orukeet-ggml
|
| 18 |
+
|
| 19 |
+
GGML conversion of **oruk/orukeet** (.nemo) for the Parakeet TDT engine that ships inside whisper.cpp ≥ 1.9, in f16 plus q8_0 / q5_0 / q4_0.
|
| 20 |
+
|
| 21 |
+
**Source checkpoint:** [oruk/orukeet](https://huggingface.co/oruk/orukeet) by oruk (fine-tune of NVIDIA Parakeet TDT 0.6B v3) · **License:** cc-by-sa-4.0 (unchanged; this repo only re-packages the weights)
|
| 22 |
+
**Engine:** Load with [whisper.cpp](https://github.com/ggml-org/whisper.cpp) ≥ 1.9 `parakeet-cli -m <file>` (the Parakeet TDT engine that ships inside whisper.cpp). Not compatible with mudler/parakeet.cpp GGUF files.
|
| 23 |
+
|
| 24 |
+
## Files
|
| 25 |
+
|
| 26 |
+
| File | Quantization | Size | Note |
|
| 27 |
+
|---|---|---|---|
|
| 28 |
+
| `ggml-orukeet-f16.bin` | f16 | 1256 MB | |
|
| 29 |
+
| `ggml-orukeet-q4_0.bin` | q4_0 | 356 MB | |
|
| 30 |
+
| `ggml-orukeet-q5_0.bin` | q5_0 | 434 MB | |
|
| 31 |
+
| `ggml-orukeet-q8_0.bin` | q8_0 | 669 MB | |
|
| 32 |
+
|
| 33 |
+
`f16` is the lossless conversion; `q8_0` is nearly identical in accuracy at ~55 % of the size; `q5_0`/`q5_k` are the phone-friendly choice; `q4_*` is smallest with a small accuracy cost.
|
| 34 |
+
|
| 35 |
+
## How these were made
|
| 36 |
+
|
| 37 |
+
Converted from the upstream checkpoint with the engine's own converter, then quantized with the engine's quantizer. Each variant was checked by transcribing short Portuguese and English samples before upload.
|
| 38 |
+
|
| 39 |
+
## Attribution
|
| 40 |
+
|
| 41 |
+
Weights are derivative works of the upstream model and keep its license. Please cite the original authors (oruk (fine-tune of NVIDIA Parakeet TDT 0.6B v3)). The conversion and hosting here are maintained by [JoaoZaokk](https://huggingface.co/JoaoZaokk) so that the download links used by the Odysseus / Open WebUI native apps stay stable. No warranty.
|
UPSTREAM-README.md
ADDED
|
@@ -0,0 +1,205 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language: [bg, hr, cs, da, nl, en, et, fi, fr, de, el, hu, it, lv, lt, mt, pl, pt, ro, ru, sk, sl, es, sv, uk]
|
| 3 |
+
license: cc-by-sa-4.0
|
| 4 |
+
base_model: nvidia/parakeet-tdt-0.6b-v3
|
| 5 |
+
base_model_relation: finetune
|
| 6 |
+
pipeline_tag: automatic-speech-recognition
|
| 7 |
+
library_name: nemo
|
| 8 |
+
transcribe_cpp:
|
| 9 |
+
streaming: false
|
| 10 |
+
translate: false
|
| 11 |
+
lang_detect: true
|
| 12 |
+
timestamps: token
|
| 13 |
+
tags: [parakeet, tdt, onnx, sherpa-onnx, gguf, multilingual, speech-recognition, gabor, fastconformer]
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
<!-- orukeet-brand:start -->
|
| 17 |
+
<p><a href="https://oruk.ai"><img src="affiliations/oruk.png" alt="oruk" width="184"></a></p>
|
| 18 |
+
<!-- orukeet-brand:end -->
|
| 19 |
+
|
| 20 |
+
# Orukeet
|
| 21 |
+
|
| 22 |
+
<!-- orukeet-team:start -->
|
| 23 |
+
<p>
|
| 24 |
+
Nathan Roll<sup>1,2</sup> · Irene Yi<sup>1,2</sup> · Büşra Marşan<sup>1,2</sup><br>
|
| 25 |
+
Vianney Grenez<sup>1</sup> · Gabriel Stein<sup>4</sup> · Momcilo Mrkaic<sup>5</sup><br>
|
| 26 |
+
Pavle Padjin<sup>5</sup> · Vladimir Zeljkovic<sup>5</sup> · Calbert Graham<sup>1,3</sup>
|
| 27 |
+
</p>
|
| 28 |
+
|
| 29 |
+
<p><strong><sup>1</sup> Oruk AI</strong></p>
|
| 30 |
+
<table>
|
| 31 |
+
<tr>
|
| 32 |
+
<td align="center" valign="middle"><img src="affiliations/stanford.png" alt="Stanford University" width="144"><br><sup>2</sup> Stanford University</td>
|
| 33 |
+
<td align="center" valign="middle"><img src="affiliations/cambridge.png" alt="University of Cambridge" width="144"><br><sup>3</sup> University of Cambridge</td>
|
| 34 |
+
<td align="center" valign="middle"><img src="affiliations/openwhispr.png" alt="OpenWhispr" width="40"><br><sup>4</sup> OpenWhispr</td>
|
| 35 |
+
<td align="center" valign="middle"><img src="affiliations/hoid.png" alt="Hoid" width="76"><br><sup>5</sup> Hoid</td>
|
| 36 |
+
</tr>
|
| 37 |
+
</table>
|
| 38 |
+
<!-- orukeet-team:end -->
|
| 39 |
+
|
| 40 |
+
Orukeet is a 25-language speech recognizer built from NVIDIA Parakeet TDT 0.6B v3. It replaces half of the encoder's temporal depthwise filters with **12,288 fitted, frozen Gabor kernels** and trains the remaining parameters on multilingual and multi-accent data.
|
| 41 |
+
|
| 42 |
+
Orukeet outperforms Parakeet on **61 of 74 tested splits**, including LibriSpeech test-clean (**1.46% vs. 1.53% WER**), test-other (**2.86% vs. 3.14%**), and FLEURS English (**3.82% vs. 4.28%**). Across all 25 FLEURS languages, pooled WER is **9.85% vs. 11.01%**, a **10.6% relative reduction**. Final adaptation and checkpoint selection use LibriSpeech test-other.
|
| 43 |
+
|
| 44 |
+
Use Orukeet for recordings, media, batch transcription, server workers and interactive applications. NeMo, ONNX INT8, native Q8 and native F16 all derive from the same **r3 release checkpoint** (`031c8ddab484`).
|
| 45 |
+
|
| 46 |
+
[Code](https://github.com/Oruk-AI/orukeet) · [OpenWhispr PR](https://github.com/OpenWhispr/openwhispr/pull/2085) · [Technical report](orukeet-technical-report.pdf) · [Artifact hashes](ARTIFACTS.json)
|
| 47 |
+
|
| 48 |
+
## Run Orukeet with NeMo
|
| 49 |
+
|
| 50 |
+
Use a CUDA-enabled PyTorch environment with `nemo_toolkit[asr]==3.0.0` and `huggingface-hub`. The [recorded source environment](https://github.com/Oruk-AI/orukeet/blob/main/evidence/standard-asr-20260908/runtime.json) lists the exact package versions used for evaluation.
|
| 51 |
+
|
| 52 |
+
```python
|
| 53 |
+
from huggingface_hub import hf_hub_download
|
| 54 |
+
from nemo.collections.asr.models import ASRModel
|
| 55 |
+
|
| 56 |
+
checkpoint = hf_hub_download(
|
| 57 |
+
"oruk/orukeet", "orukeet-v0.1.0.nemo",
|
| 58 |
+
revision="555136b50265a132d4cea0d35560c26fc4f657ab",
|
| 59 |
+
)
|
| 60 |
+
asr = ASRModel.restore_from(checkpoint)
|
| 61 |
+
asr.eval()
|
| 62 |
+
print(asr.transcribe(["recording.wav"], return_hypotheses=True)[0].text)
|
| 63 |
+
```
|
| 64 |
+
|
| 65 |
+
`orukeet fetch source` retrieves the same hash-checked checkpoint. Further training attaches the supplied frozen-row parametrization before constructing the optimizer.
|
| 66 |
+
|
| 67 |
+
## Architecture
|
| 68 |
+
|
| 69 |
+
The model retains Parakeet's 627,008,134 parameters, 24-layer FastConformer encoder, token-and-duration transducer and tokenizer. Each encoder block contains 1,024 nine-tap temporal depthwise filters. A selected filter stores its own fitted Gabor function:
|
| 70 |
+
|
| 71 |
+
$$g(t)=A\exp\left[-\frac{(t-\mu)^2}{2\sigma^2}\right]\cos\left(2\pi f(t-\mu)+\phi\right),\quad t=-4,\ldots,4.$$
|
| 72 |
+
|
| 73 |
+
We fit all 24,576 filters and globally select the 12,288 lowest normalized squared errors. This selects 175–748 kernels per layer, with 6.32% median relative RMS error and a 13.30% cutoff. The 110,592 selected taps remain fixed; 626,897,542 scalar parameters remain trainable. Native exports materialize the fitted taps as ordinary F16 convolution weights.
|
| 74 |
+
|
| 75 |
+

|
| 76 |
+
|
| 77 |
+
## Construction
|
| 78 |
+
|
| 79 |
+
Gabor recovery uses transducer loss, encoder matching and token/duration distillation. A further 4,035 low-learning-rate updates produce the parent checkpoint. The final r3 pass applies 168 AdamW updates, with a 3% warmup and cosine decay from `5e-6` to `5e-7`, over three passes through 2,939 LibriSpeech test-other recordings. Targets preserve native casing and punctuation while correcting reference words. The same split supplies checkpoint selection. An export audit verifies that all 12,288 fitted kernels remain exact and all 651 other parameter tensors change.
|
| 80 |
+
|
| 81 |
+
[Fit and freeze recipe](https://github.com/Oruk-AI/orukeet/blob/main/training/gabor_half/README.md) · [Final adaptation](https://github.com/Oruk-AI/orukeet/blob/main/training/librispeech_ft/README.md) · [Training lineage](https://github.com/Oruk-AI/orukeet/blob/main/training/README.md)
|
| 82 |
+
|
| 83 |
+
## Evaluation
|
| 84 |
+
|
| 85 |
+
Both models decode identical recordings with NeMo greedy-batch TDT, FP32 weights and BF16 CUDA autocast. The pinned scoring code defines text normalization and compound alignment; pooled WER sums errors and normalized reference words. Lower is better.
|
| 86 |
+
|
| 87 |
+
| Comparison | Recordings | Parakeet WER | Orukeet WER |
|
| 88 |
+
|:--|--:|--:|--:|
|
| 89 |
+
| LibriSpeech test-clean | 2,620 | 1.53% | **1.46%** |
|
| 90 |
+
| LibriSpeech test-other | 2,939 | 3.14% | **2.86%** |
|
| 91 |
+
| FLEURS English | 647 | 4.28% | **3.82%** |
|
| 92 |
+
| FLEURS pooled, 25 languages | 20,146 | 11.01% | **9.85%** |
|
| 93 |
+
| Accents/domains pooled, 47 splits | 12,006 | 16.72% | **15.25%** |
|
| 94 |
+
| Accents/domains English, 20 splits | 5,120 | 9.51% | **8.84%** |
|
| 95 |
+
|
| 96 |
+
Orukeet improves 25 of 27 complete LibriSpeech/FLEURS splits and 36 of 47 accent/domain splits, including all 20 English accent/domain splits. The accent/domain sample contains 256 recordings per split and all 230 Lesbos recordings; the preceding adaptation includes 6,118 sampled recordings. Read speech and accents/domains have separate pooled results. Every recording contributes to the scores.
|
| 97 |
+
|
| 98 |
+
[All 74 paired WER/CER scores and edit counts](docs/current-checkpoint-benchmarks.md) · [Methods](docs/technical-report.md) · [Technical report](orukeet-technical-report.pdf)
|
| 99 |
+
|
| 100 |
+
## sherpa-onnx inference
|
| 101 |
+
|
| 102 |
+
The [ONNX INT8 archive](https://huggingface.co/oruk/orukeet/resolve/55a984d46f68323301837194ce647c702f55facc/onnx/sherpa-onnx-orukeet-v0.1.0-int8.tar.bz2) uses the standard Parakeet TDT v3 layout: `encoder.int8.onnx`, `decoder.int8.onnx`, `joiner.int8.onnx` and `tokens.txt`. It also includes the BPE vocabulary, weight license and attribution. Gabor filters are ordinary convolution weights; the model uses sherpa-onnx's existing offline transducer loader.
|
| 103 |
+
|
| 104 |
+
The optimized encoder evaluates 24 quantized depthwise convolutions with exactly equivalent FP32 arithmetic using operators already in ONNX Runtime. All 640 application-check transcripts match the previous export. On the same 160-clip timing sample, median file transcription is 390 ms versus 432 ms before optimization and 428 ms for stock Parakeet on M5 Max. [Execution details and receipts](https://github.com/Oruk-AI/orukeet/blob/main/evidence/speed20260910/README.md).
|
| 105 |
+
|
| 106 |
+
[OpenWhispr 1.10.0](https://github.com/OpenWhispr/openwhispr/releases/tag/v1.10.0) ships Orukeet as its recommended local model, using this format through its existing Parakeet worker. Choose **Local → Oruk → Orukeet**, then **Download**. Recognition runs locally after installation.
|
| 107 |
+
|
| 108 |
+
[Follow the file-upload walkthrough](https://oruk.ai/guides/orukeet-local-transcription#openwhispr) for the exact settings and a public sample with its observed transcript. Audio Upload needs its own model selection even when Orukeet is active for dictation.
|
| 109 |
+
|
| 110 |
+
Archive SHA-256: `f9191f30178cc9122ce2f023bf9fefafc822028307b0efa4caff645ba3fe8d0a`.
|
| 111 |
+
|
| 112 |
+
[Export and loader instructions](https://github.com/Oruk-AI/orukeet/blob/main/export/onnx/README.md) · [Conversion evidence](https://github.com/Oruk-AI/orukeet/tree/main/evidence/onnx-r3-20260910) · [OpenWhispr checks and paired scores](https://github.com/Oruk-AI/orukeet/blob/main/integrations/openwhispr/APP_BENCHMARKS.md)
|
| 113 |
+
|
| 114 |
+
## Native inference
|
| 115 |
+
|
| 116 |
+
Use Python 3.12+ in an activated virtual environment. The native package is
|
| 117 |
+
v0.1.1; the r3 weight filenames retain their original v0.1.0 names.
|
| 118 |
+
|
| 119 |
+
```sh
|
| 120 |
+
python -m pip install --upgrade \
|
| 121 |
+
https://github.com/Oruk-AI/orukeet/releases/download/v0.1.1/orukeet-0.1.1-py3-none-any.whl
|
| 122 |
+
orukeet install --device auto --cache ./orukeet-cache --output installation.json
|
| 123 |
+
```
|
| 124 |
+
|
| 125 |
+
```python
|
| 126 |
+
import json
|
| 127 |
+
from pathlib import Path
|
| 128 |
+
from orukeet import Orukeet
|
| 129 |
+
|
| 130 |
+
config = json.loads(Path("installation.json").read_text(encoding="utf-8-sig"))
|
| 131 |
+
with Orukeet(config["model"], config["runtime"], device=config["device"]) as asr:
|
| 132 |
+
print(asr.transcribe("recording.wav")["text"])
|
| 133 |
+
```
|
| 134 |
+
|
| 135 |
+
The installer verifies the Q8 weights and native runtime. It selects the
|
| 136 |
+
optimized Metal runtime on Apple silicon, CUDA on a detected NVIDIA device,
|
| 137 |
+
or CPU, subject to the available runtime for the platform. Keep the worker
|
| 138 |
+
alive across recordings to avoid repeated model loading.
|
| 139 |
+
|
| 140 |
+
[Run the complete local tutorial](https://oruk.ai/guides/orukeet-local-transcription)
|
| 141 |
+
for a supplied audio file, a reusable runner, actual output and verification
|
| 142 |
+
hashes. The native response contains transcription and window-level segment
|
| 143 |
+
times; it does not return emotion, speaking-style or speaker-diarization scores.
|
| 144 |
+
|
| 145 |
+
[Watch the 39-second recorded example](https://oruk.ai/guides/orukeet-local-transcription#watch) to hear the input and inspect the native Q8 / Metal output. The walkthrough is edited for readability; it is not a speed or accuracy benchmark.
|
| 146 |
+
|
| 147 |
+
[Usage and batch transcription](https://github.com/Oruk-AI/orukeet/blob/main/docs/usage.md)
|
| 148 |
+
· [Native runtime and measurements](https://github.com/Oruk-AI/orukeet/blob/main/runtime/README.md)
|
| 149 |
+
|
| 150 |
+
## transcribe.cpp and Handy-compatible GGUF
|
| 151 |
+
|
| 152 |
+
[`orukeet-transcribe-cpp-Q8_0.gguf`](orukeet-transcribe-cpp-Q8_0.gguf) is a Q8 export of the same r3 checkpoint for [transcribe.cpp](https://github.com/cjpais/transcribe.cpp). It uses the existing `parakeet` architecture and requires no Gabor-specific runtime. CPU and Apple Metal checks use the exact `transcribe-cpp` 0.2.0 dependency pinned by Handy.
|
| 153 |
+
|
| 154 |
+
This file has a different tensor layout from the native NeMo-Speech.cpp GGUFs above. Select the export for your runtime. [Conversion, checksums and validation](transcribe-cpp/README.md).
|
| 155 |
+
|
| 156 |
+
## Model files
|
| 157 |
+
|
| 158 |
+
| Format | File | Bytes |
|
| 159 |
+
|:--|:--|--:|
|
| 160 |
+
| NeMo source | `orukeet-v0.1.0.nemo` | 2,509,342,720 |
|
| 161 |
+
| Native Q8 | `orukeet-v0.1.0-q8.gguf` | 714,456,704 |
|
| 162 |
+
| transcribe.cpp Q8 | `orukeet-transcribe-cpp-Q8_0.gguf` | 739,508,608 |
|
| 163 |
+
| Native F16 | `orukeet-v0.1.0-f16.gguf` | 1,296,681,088 |
|
| 164 |
+
| ONNX INT8 archive | `onnx/sherpa-onnx-orukeet-v0.1.0-int8.tar.bz2` | 486,807,585 |
|
| 165 |
+
|
| 166 |
+
All formats derive from **r3**. NeMo and native files are pinned to revision `555136b50265a132d4cea0d35560c26fc4f657ab`; the ONNX archive is pinned to `55a984d46f68323301837194ce647c702f55facc`. The ONNX package occupies 671,619,800 bytes after extraction.
|
| 167 |
+
|
| 168 |
+
- NeMo SHA-256: `031c8ddab4845aeced904a7cde8e8aa57993b2e344716cf83a545b079c473b56`
|
| 169 |
+
- Q8 SHA-256: `93ce19c6d8244acbfea980eeaf970531d4f216171578ef8e041dcc2d070a45bd`
|
| 170 |
+
- F16 SHA-256: `de53fb8ec251fb07ade15baabe17b00774ae3f1112f8618b062337f90fb49194`
|
| 171 |
+
|
| 172 |
+
Q8 and F16 pass real transcription and protocol checks on Apple silicon with Metal and CPU. Conversion audits verify all 12,288 fitted kernels after F16 rounding. The table above reports NeMo recognition scores; native checks have their own model hashes and runtime receipts.
|
| 173 |
+
|
| 174 |
+
[Artifact catalog](https://github.com/Oruk-AI/orukeet/blob/main/src/orukeet/artifacts.json) · [Native conversion and validation](https://github.com/Oruk-AI/orukeet/tree/main/evidence/r3-promotion-20260908/)
|
| 175 |
+
|
| 176 |
+
## License and attribution
|
| 177 |
+
|
| 178 |
+
Code: MIT. Weights and fitted kernels: CC BY-SA 4.0, retaining NVIDIA's foundation attribution. Transcript-free metric records: CC BY 4.0. Dataset audio is obtained from its original providers under their terms.
|
| 179 |
+
|
| 180 |
+
[Data provenance](https://github.com/Oruk-AI/orukeet/blob/main/docs/data-and-licenses.md) · [Attribution](NOTICE.md)
|
| 181 |
+
|
| 182 |
+
<!-- orukeet-citation:start -->
|
| 183 |
+
## Citation
|
| 184 |
+
|
| 185 |
+
```bibtex
|
| 186 |
+
@techreport{roll2026orukeet,
|
| 187 |
+
title = {{Orukeet}: Multilingual {ASR} with Frozen {Gabor} Kernels},
|
| 188 |
+
author = {Roll, Nathan and
|
| 189 |
+
Yi, Irene and
|
| 190 |
+
Mar{\c{s}}an, B{\"u}{\c{s}}ra and
|
| 191 |
+
Grenez, Vianney and
|
| 192 |
+
Stein, Gabriel and
|
| 193 |
+
Mrkaic, Momcilo and
|
| 194 |
+
Padjin, Pavle and
|
| 195 |
+
Zeljkovic, Vladimir and
|
| 196 |
+
Graham, Calbert},
|
| 197 |
+
institution = {Oruk AI},
|
| 198 |
+
year = {2026},
|
| 199 |
+
type = {Technical report},
|
| 200 |
+
url = {https://github.com/Oruk-AI/orukeet/blob/main/output/pdf/orukeet-technical-report.pdf}
|
| 201 |
+
}
|
| 202 |
+
```
|
| 203 |
+
|
| 204 |
+
[Download BibTeX](CITATION.bib) · [Citation metadata](CITATION.cff)
|
| 205 |
+
<!-- orukeet-citation:end -->
|
VERIFY.txt
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
== ggml-orukeet-f16.bin
|
| 2 |
+
pt Bom dia, hoje vamos testar o reconhecimento de fala e português do Brasil com um modelo convertido para o formato do Sper.cpp.
|
| 3 |
+
en Good morning. Today we are testing speech recognition in English with a converted model.
|
| 4 |
+
jfk And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country.
|
| 5 |
+
== ggml-orukeet-q4_0.bin
|
| 6 |
+
pt Bom dia, hoje vamos testar o reconhecimento de fala e português do Brasil com um modelo convertido para o formato do SPER.cpp.
|
| 7 |
+
en Good morning. Today we are testing speech recognition in English with a converted model.
|
| 8 |
+
jfk And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country.
|
| 9 |
+
== ggml-orukeet-q5_0.bin
|
| 10 |
+
pt Bom dia, hoje vamos testar o reconhecimento de fala e português do Brasil com um modelo convertido para o formato do spare.cpp.
|
| 11 |
+
en Good morning. Today we are testing speech recognition in English with a converted model.
|
| 12 |
+
jfk And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country.
|
| 13 |
+
== ggml-orukeet-q8_0.bin
|
| 14 |
+
pt Bom dia, hoje vamos testar o reconhecimento de fala e português do Brasil com um modelo convertido para o formato do Sper.cpp.
|
| 15 |
+
en Good morning. Today we are testing speech recognition in English with a converted model.
|
| 16 |
+
jfk And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country.
|
ggml-orukeet-f16.bin
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:da0eaeb20761b9ec48a51873fd31c40370da81031e4d5caf0fd0b03b378b34e2
|
| 3 |
+
size 1255897319
|
ggml-orukeet-q4_0.bin
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4da777774cf1720442d82560df3ca4a02e8295767a43010077e46e66ab93724f
|
| 3 |
+
size 355615679
|
ggml-orukeet-q5_0.bin
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:5df30706701245d583e826f6bd734dcfca365283315f33adcce995cb832f30cd
|
| 3 |
+
size 433901039
|
ggml-orukeet-q8_0.bin
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:46ee7e4b7e5ec1e9870e4004027849f74508fb9aaacf0e22786a396c378d5c6a
|
| 3 |
+
size 668757119
|