Moonshine Streaming Tiny

Useful Sensors' Moonshine Streaming tiny (34M) English speech recognizer -- a sliding-window encoder over the raw waveform and an autoregressive decoder -- exported for loom.cpp.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from moonshine-ai/moonshine-streaming-tiny. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

mit, inherited from the base model above.

Language(s)

en

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom

model = loom.Model.from_pretrained("loom-ai-org/moonshine-streaming-tiny-loom")

# Audio is a mono float list at 16 kHz. This model decodes in the one language it was trained for and
# takes no `language=` argument -- passing one warns and is ignored, because nothing in its decode
# could act on it.
result = model.speech2text.infer(audio, timestamps=True)
print(result.text)

# It emits no timestamp tokens, so `segments` is one span covering the whole clip and
# `result.timestamped` is False. Check that before treating a start/end as a boundary the model chose.
for segment in result.segments:
    print(segment.start, segment.end, segment.text)

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

At most 81.9 s per call. The decoder's position table has 4096 encoder rows (50 per second); a longer clip is refused with an error rather than truncated. Split long audio at its pauses -- a VAD such as silero-vad-loom finds them -- and transcribe each part.

Decoding is the model card's own: greedy, and capped at 6.5 tokens per second of audio "to avoid hallucination loops". A clip shorter than about 0.3 s therefore returns no text. Like other encoder-decoder recognizers it can still repeat or invent words on noisy or very short audio.

The whole clip is one pass, every encoder layer attending through the sliding window the model was trained with -- what the model card's usage computes. This export does not run the encoder incrementally (live streaming); English only, mono 16 kHz.

Files

  • moonshine-streaming-tiny.gguf -- the model, exported with loom-exporter.
Downloads last month
53
GGUF
Model size
33.6M params
Architecture
loom-moonshine-streaming
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/moonshine-streaming-tiny-loom

Quantized
(8)
this model

Collection including loom-ai-org/moonshine-streaming-tiny-loom