SraVaani-1.0 MoE

SraVaani-1.0 (ARTPARK, IISc; 444M FastConformer-TDT, 65 Indian languages) with its 34 feed-forward blocks converted to a 16-expert Mixture-of-Experts with a per-frame linear router, accuracy recovered by distillation, packaged as GGUF for a patched parakeet.cpp.

Same parameter count and memory as the dense model; only compute is skipped. Half the experts per frame, no accuracy loss.

Dense q8 MoE top-8 q8 MoE top-6 q8
WER, 50 held-out Hindi clips (Pi 5) 12.9% 12.0% 12.3%
Raspberry Pi 5, 44 s audio, 4 thr, 2.4 GHz 9.42 s 9.02 s (1.04×) 8.46 s (1.11×)
Pi 5, 3 s command incl. model load 566 ms 530 ms 521 ms
Laptop i7-1360P, 44 s audio, 8 thr 5.37 s 4.62 s (1.16×) 4.38 s (1.23×)
File size 660 MB 664 MB 664 MB

Files

File What
sravaani-moe8-q8_0.gguf Recommended. 16 experts, 8 active per frame, q8_0 experts. WER at or below dense.
sravaani-moe6-q8_0.gguf Same weights, 6 active (metadata only). ~10% faster, +0.3 to +1.2 WER without retraining.
moe_ffn_best.pt PyTorch state of the 34 MoE FFN blocks after recovery training (fp16): {i}.{W1,b1,W2,b2,mu,rw,rb,lab}. For further training or re-export.
expert_assign_E16.pt, routers_E16_h0.pt, ffn_means.pt Neuron→expert assignment, linear routers, imputation means (pre-training).
parakeet_moe.patch Patch for parakeet.cpp (adds build_moe_ffn, MoE metadata).
export_moe_gguf.py Builds the MoE GGUF from the dense q8 GGUF + moe_ffn_best.pt.
REPORT.md Full technical report with method, ablations and benchmarks.

Usage

The GGUF needs parakeet.cpp with the MoE patch (parakeet_moe.patch, ~80 lines, against v0.5.0). No custom kernel: it uses ggml's ggml_top_k and ggml_mul_mat_id.

git clone --depth 1 https://github.com/mudler/parakeet.cpp && cd parakeet.cpp
git submodule update --init --depth 1 third_party/ggml
git apply parakeet_moe.patch
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build -j --target parakeet-cli
./build/examples/cli/parakeet-cli transcribe --model sravaani-moe8-q8_0.gguf --input clip.wav --decoder tdt --threads 4

--decoder tdt is required. Input must be 16 kHz mono 16-bit PCM WAV. No language flag: the model detects the language and writes native script. Builds and runs on Raspberry Pi 5 (aarch64) in about a minute.

Unpatched parakeet.cpp will fail on this file (the dense FFN tensors are replaced by expert tensors).

How it was made

  1. Split. The 4096 hidden neurons of each FFN were partitioned into 16 balanced experts of 256 by k-means on the rows of linear1.weight. No new parameters.
  2. Route. A linear router (1024 → 16) per FFN, trained to predict each expert's activation mass on 8.5k frames of Hindi speech. 95% top-12 agreement with an oracle.
  3. Impute. Skipped neurons get their mean activation instead of zero (Swish has a negative tail); this folds into one constant per expert. Without it top-8 WER is 54.7%; with it 29.1% before training.
  4. Recover. Distillation from the frozen dense encoder (MSE on encoder output + 0.1 × per-layer MSE), FFN weights and imputation means trainable, everything else frozen. 8 shards of Vaani transcribed Hindi (~35 h), Colab A100, 6 epochs, 41 minutes. WER 27.4% → 11.7% after one epoch (dense 13.5%).
  5. Ship. Experts stored contiguously in GGUF (moe.w1 [1024,256,16], moe.w2 [256,1024,16], q8_0) plus router, imputation offsets, metadata parakeet.moe.{n_expert,n_active,expert_size}.

Output formula per FFN: y = Σ_{e∈top-K} (W2_e·silu(W1_e x + b1_e) − c_e) + Σ_e c_e + b2, with c_e = W2_e·μ_e.

Honest notes

  • Speed gain on CPU is 4 to 11% on Pi 5 and 16 to 23% on x86, against a 26% ceiling (FFN is 53% of encoder time). The rest is gather overhead in the generic ggml MoE kernel.
  • Memory is unchanged. Quantize (q8 → q4) for memory; it stacks with this.
  • Evaluated on Hindi only (50 held-out clips from a grocery-assistant domain plus Vaani regional Hindi for training). Other languages should work (nothing outside the FFN changed) but were not measured.
  • The Pi 5 throttles at 85 °C without a heatsink; cooling it gave a larger speedup (1.22×) than the MoE conversion.

Credits

Base model: ARTPARK-IISc/SraVaani-1.0 (MIT). Dense GGUF: Henil1/Sravaani-1.0-GGUF. Runtime: parakeet.cpp by mudler. Training data: ARTPARK-IISc/Vaani-transcription-part. MoE conversion, routers, recovery training, runtime patch and benchmarks by Shrinkhala & Lalit Belwal with Claude (Anthropic). MIT.

Downloads last month
19
GGUF
Model size
0.4B params
Architecture
parakeet
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sraivante/SraVaani-1.0-MoE

Quantized
(4)
this model

Dataset used to train sraivante/SraVaani-1.0-MoE