--- license: mit language: - en base_model: - moonshine-ai/moonshine-streaming-tiny pipeline_tag: automatic-speech-recognition tags: - loom - automatic-speech-recognition library_name: loom-py-rt --- # Moonshine Streaming Tiny Useful Sensors' Moonshine Streaming tiny (34M) English speech recognizer -- a sliding-window encoder over the raw waveform and an autoregressive decoder -- exported for loom.cpp. This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by [loom-exporter](https://github.com/loom-ai-org/loom-exporter). ## Original model Exported from [`moonshine-ai/moonshine-streaming-tiny`](https://huggingface.co/moonshine-ai/moonshine-streaming-tiny). Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format. ## License `mit`, inherited from the base model above. ## Language(s) `en` ## Usage Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI: ```sh pip install -U "loom-py-rt[hub]" ``` ```python import loom model = loom.Model.from_pretrained("loom-ai-org/moonshine-streaming-tiny-loom") # Audio is a mono float list at 16 kHz. This model decodes in the one language it was trained for and # takes no `language=` argument -- passing one warns and is ignored, because nothing in its decode # could act on it. result = model.speech2text.infer(audio, timestamps=True) print(result.text) # It emits no timestamp tokens, so `segments` is one span covering the whole clip and # `result.timestamped` is False. Check that before treating a start/end as a boundary the model chose. for segment in result.segments: print(segment.start, segment.end, segment.text) ``` ### The layer underneath The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)` passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name. `model.driver_source` prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and [loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two. ## Known limitations **At most 81.9 s per call.** The decoder's position table has 4096 encoder rows (50 per second); a longer clip is refused with an error rather than truncated. Split long audio at its pauses -- a VAD such as `silero-vad-loom` finds them -- and transcribe each part. Decoding is the model card's own: greedy, and capped at 6.5 tokens per second of audio "to avoid hallucination loops". A clip shorter than about 0.3 s therefore returns no text. Like other encoder-decoder recognizers it can still repeat or invent words on noisy or very short audio. The whole clip is one pass, every encoder layer attending through the sliding window the model was trained with -- what the model card's usage computes. This export does not run the encoder incrementally (live streaming); English only, mono 16 kHz. ## Files - `moonshine-streaming-tiny.gguf` -- the model, exported with loom-exporter.