Moonshine Streaming Tiny
Useful Sensors' Moonshine Streaming tiny (34M) English speech recognizer -- a sliding-window encoder over the raw waveform and an autoregressive decoder -- exported for loom.cpp.
This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.
Original model
Exported from moonshine-ai/moonshine-streaming-tiny. Weights are unmodified; this repo packages the same parameters into
loom.cpp's GGUF format.
License
mit, inherited from the base model above.
Language(s)
en
Usage
Run it with loom-py -- loom-py-rt on PyPI:
pip install -U "loom-py-rt[hub]"
import loom
model = loom.Model.from_pretrained("loom-ai-org/moonshine-streaming-tiny-loom")
# Audio is a mono float list at 16 kHz. This model decodes in the one language it was trained for and
# takes no `language=` argument -- passing one warns and is ignored, because nothing in its decode
# could act on it.
result = model.speech2text.infer(audio, timestamps=True)
print(result.text)
# It emits no timestamp tokens, so `segments` is one span covering the whole clip and
# `result.timestamped` is False. Check that before treating a start/end as a boundary the model chose.
for segment in result.segments:
print(segment.start, segment.end, segment.text)
The layer underneath
The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...)
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.
model.driver_source prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See loom-py for the API and
loom.cpp for what the engine does between the two.
Known limitations
At most 81.9 s per call. The decoder's position table has 4096 encoder rows (50 per second); a longer clip is refused with an error rather than truncated. Split long audio at its pauses -- a VAD such as silero-vad-loom finds them -- and transcribe each part.
Decoding is the model card's own: greedy, and capped at 6.5 tokens per second of audio "to avoid hallucination loops". A clip shorter than about 0.3 s therefore returns no text. Like other encoder-decoder recognizers it can still repeat or invent words on noisy or very short audio.
The whole clip is one pass, every encoder layer attending through the sliding window the model was trained with -- what the model card's usage computes. This export does not run the encoder incrementally (live streaming); English only, mono 16 kHz.
Files
moonshine-streaming-tiny.gguf-- the model, exported with loom-exporter.
- Downloads last month
- 53
We're not able to determine the quantization variants.
Model tree for loom-ai-org/moonshine-streaming-tiny-loom
Base model
moonshine-ai/moonshine-streaming-tiny