fdemelo commited on
Commit
d6559f9
·
verified ·
1 Parent(s): e869f54

Re-export moonshine-streaming-tiny (loom-exporter 34483dc)

Browse files
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +80 -0
  3. moonshine-streaming-tiny.gguf +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ moonshine-streaming-tiny.gguf filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,80 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ language:
4
+ - en
5
+ base_model:
6
+ - moonshine-ai/moonshine-streaming-tiny
7
+ pipeline_tag: automatic-speech-recognition
8
+ tags:
9
+ - loom
10
+ - automatic-speech-recognition
11
+ library_name: loom-py-rt
12
+ ---
13
+ # Moonshine Streaming Tiny
14
+
15
+ Useful Sensors' Moonshine Streaming tiny (34M) English speech recognizer -- a sliding-window encoder over the raw waveform and an autoregressive decoder -- exported for loom.cpp.
16
+
17
+ This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF
18
+ that carries its own graph topologies, tokenizer (if any) and driver script, produced by
19
+ [loom-exporter](https://github.com/loom-ai-org/loom-exporter).
20
+
21
+ ## Original model
22
+
23
+ Exported from [`moonshine-ai/moonshine-streaming-tiny`](https://huggingface.co/moonshine-ai/moonshine-streaming-tiny). Weights are unmodified; this repo packages the same parameters into
24
+ loom.cpp's GGUF format.
25
+
26
+ ## License
27
+
28
+ `mit`, inherited from the base model above.
29
+
30
+ ## Language(s)
31
+
32
+ `en`
33
+
34
+ ## Usage
35
+
36
+ Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI:
37
+
38
+ ```sh
39
+ pip install -U "loom-py-rt[hub]"
40
+ ```
41
+
42
+ ```python
43
+ import loom
44
+
45
+ model = loom.Model.from_pretrained("loom-ai-org/moonshine-streaming-tiny-loom")
46
+
47
+ # Audio is a mono float list at 16 kHz. This model decodes in the one language it was trained for and
48
+ # takes no `language=` argument -- passing one warns and is ignored, because nothing in its decode
49
+ # could act on it.
50
+ result = model.speech2text.infer(audio, timestamps=True)
51
+ print(result.text)
52
+
53
+ # It emits no timestamp tokens, so `segments` is one span covering the whole clip and
54
+ # `result.timestamped` is False. Check that before treating a start/end as a boundary the model chose.
55
+ for segment in result.segments:
56
+ print(segment.start, segment.end, segment.text)
57
+ ```
58
+
59
+ ### The layer underneath
60
+
61
+ The call above is the high-level door: one per task, named for the modality pair it maps between, with
62
+ the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)`
63
+ passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
64
+ door does not name.
65
+
66
+ `model.driver_source` prints that driver, including a header comment documenting every argument it
67
+ accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and
68
+ [loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two.
69
+
70
+ ## Known limitations
71
+
72
+ **At most 81.9 s per call.** The decoder's position table has 4096 encoder rows (50 per second); a longer clip is refused with an error rather than truncated. Split long audio at its pauses -- a VAD such as `silero-vad-loom` finds them -- and transcribe each part.
73
+
74
+ Decoding is the model card's own: greedy, and capped at 6.5 tokens per second of audio "to avoid hallucination loops". A clip shorter than about 0.3 s therefore returns no text. Like other encoder-decoder recognizers it can still repeat or invent words on noisy or very short audio.
75
+
76
+ The whole clip is one pass, every encoder layer attending through the sliding window the model was trained with -- what the model card's usage computes. This export does not run the encoder incrementally (live streaming); English only, mono 16 kHz.
77
+
78
+ ## Files
79
+
80
+ - `moonshine-streaming-tiny.gguf` -- the model, exported with loom-exporter.
moonshine-streaming-tiny.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:79a10caf9e0b0a76596163a32c71ad6355be607cb9a03e4bf50e5733cee22035
3
+ size 136094304