SraVaani-1.0 MoE
SraVaani-1.0 (ARTPARK, IISc; 444M FastConformer-TDT, 65 Indian languages) with its 34 feed-forward blocks converted to a 16-expert Mixture-of-Experts with a per-frame linear router, accuracy recovered by distillation, packaged as GGUF for a patched parakeet.cpp.
Same parameter count and memory as the dense model; only compute is skipped. Half the experts per frame, no accuracy loss.
| Dense q8 | MoE top-8 q8 | MoE top-6 q8 | |
|---|---|---|---|
| WER, 50 held-out Hindi clips (Pi 5) | 12.9% | 12.0% | 12.3% |
| Raspberry Pi 5, 44 s audio, 4 thr, 2.4 GHz | 9.42 s | 9.02 s (1.04×) | 8.46 s (1.11×) |
| Pi 5, 3 s command incl. model load | 566 ms | 530 ms | 521 ms |
| Laptop i7-1360P, 44 s audio, 8 thr | 5.37 s | 4.62 s (1.16×) | 4.38 s (1.23×) |
| File size | 660 MB | 664 MB | 664 MB |
Files
| File | What |
|---|---|
sravaani-moe8-q8_0.gguf |
Recommended. 16 experts, 8 active per frame, q8_0 experts. WER at or below dense. |
sravaani-moe6-q8_0.gguf |
Same weights, 6 active (metadata only). ~10% faster, +0.3 to +1.2 WER without retraining. |
moe_ffn_best.pt |
PyTorch state of the 34 MoE FFN blocks after recovery training (fp16): {i}.{W1,b1,W2,b2,mu,rw,rb,lab}. For further training or re-export. |
expert_assign_E16.pt, routers_E16_h0.pt, ffn_means.pt |
Neuron→expert assignment, linear routers, imputation means (pre-training). |
parakeet_moe.patch |
Patch for parakeet.cpp (adds build_moe_ffn, MoE metadata). |
export_moe_gguf.py |
Builds the MoE GGUF from the dense q8 GGUF + moe_ffn_best.pt. |
REPORT.md |
Full technical report with method, ablations and benchmarks. |
Usage
The GGUF needs parakeet.cpp with the MoE patch (parakeet_moe.patch, ~80 lines, against v0.5.0). No custom kernel: it uses ggml's ggml_top_k and ggml_mul_mat_id.
git clone --depth 1 https://github.com/mudler/parakeet.cpp && cd parakeet.cpp
git submodule update --init --depth 1 third_party/ggml
git apply parakeet_moe.patch
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build -j --target parakeet-cli
./build/examples/cli/parakeet-cli transcribe --model sravaani-moe8-q8_0.gguf --input clip.wav --decoder tdt --threads 4
--decoder tdt is required. Input must be 16 kHz mono 16-bit PCM WAV. No language flag: the model detects the language and writes native script. Builds and runs on Raspberry Pi 5 (aarch64) in about a minute.
Unpatched parakeet.cpp will fail on this file (the dense FFN tensors are replaced by expert tensors).
How it was made
- Split. The 4096 hidden neurons of each FFN were partitioned into 16 balanced experts of 256 by k-means on the rows of
linear1.weight. No new parameters. - Route. A linear router (1024 → 16) per FFN, trained to predict each expert's activation mass on 8.5k frames of Hindi speech. 95% top-12 agreement with an oracle.
- Impute. Skipped neurons get their mean activation instead of zero (Swish has a negative tail); this folds into one constant per expert. Without it top-8 WER is 54.7%; with it 29.1% before training.
- Recover. Distillation from the frozen dense encoder (MSE on encoder output + 0.1 × per-layer MSE), FFN weights and imputation means trainable, everything else frozen. 8 shards of Vaani transcribed Hindi (~35 h), Colab A100, 6 epochs, 41 minutes. WER 27.4% → 11.7% after one epoch (dense 13.5%).
- Ship. Experts stored contiguously in GGUF (
moe.w1[1024,256,16],moe.w2[256,1024,16], q8_0) plus router, imputation offsets, metadataparakeet.moe.{n_expert,n_active,expert_size}.
Output formula per FFN: y = Σ_{e∈top-K} (W2_e·silu(W1_e x + b1_e) − c_e) + Σ_e c_e + b2, with c_e = W2_e·μ_e.
Honest notes
- Speed gain on CPU is 4 to 11% on Pi 5 and 16 to 23% on x86, against a 26% ceiling (FFN is 53% of encoder time). The rest is gather overhead in the generic ggml MoE kernel.
- Memory is unchanged. Quantize (q8 → q4) for memory; it stacks with this.
- Evaluated on Hindi only (50 held-out clips from a grocery-assistant domain plus Vaani regional Hindi for training). Other languages should work (nothing outside the FFN changed) but were not measured.
- The Pi 5 throttles at 85 °C without a heatsink; cooling it gave a larger speedup (1.22×) than the MoE conversion.
Credits
Base model: ARTPARK-IISc/SraVaani-1.0 (MIT). Dense GGUF: Henil1/Sravaani-1.0-GGUF. Runtime: parakeet.cpp by mudler. Training data: ARTPARK-IISc/Vaani-transcription-part. MoE conversion, routers, recovery training, runtime patch and benchmarks by Shrinkhala & Lalit Belwal with Claude (Anthropic). MIT.
- Downloads last month
- 19
8-bit
Model tree for sraivante/SraVaani-1.0-MoE
Base model
ARTPARK-IISc/SraVaani-1.0