fdemelo commited on
Commit
7ae1e36
·
verified ·
1 Parent(s): eb43a6f

Re-export flan-t5-small (loom-exporter f85a071-dirty)

Browse files
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +89 -0
  3. flan-t5-small.gguf +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ flan-t5-small.gguf filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ - fr
6
+ - ro
7
+ - de
8
+ - multilingual
9
+ base_model:
10
+ - google/flan-t5-small
11
+ pipeline_tag: text-generation
12
+ tags:
13
+ - loom
14
+ - text2text-generation
15
+ library_name: loom-py-rt
16
+ ---
17
+ # FLAN-T5 Small
18
+
19
+ Google's instruction-tuned T5-small, exported for loom.cpp. Family 6: text in, text out, through an encoder read once and a KV-cached decoder that cross-attends to it -- the first encoder-decoder TEXT model and the first SentencePiece Unigram vocabulary in this collection.
20
+
21
+ This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF
22
+ that carries its own graph topologies, tokenizer (if any) and driver script, produced by
23
+ [loom-exporter](https://github.com/loom-ai-org/loom-exporter).
24
+
25
+ ## Original model
26
+
27
+ Exported from [`google/flan-t5-small`](https://huggingface.co/google/flan-t5-small). Weights are unmodified; this repo packages the same parameters into
28
+ loom.cpp's GGUF format.
29
+
30
+ ## License
31
+
32
+ `apache-2.0`, inherited from the base model above.
33
+
34
+ ## Language(s)
35
+
36
+ `en`, `fr`, `ro`, `de`, `multilingual`
37
+
38
+ ## Usage
39
+
40
+ Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI:
41
+
42
+ ```sh
43
+ pip install -U "loom-py-rt[hub]"
44
+ ```
45
+
46
+ ```python
47
+ import loom
48
+
49
+ model = loom.Model.from_pretrained("loom-ai-org/flan-t5-small-loom")
50
+
51
+ # The instruction is part of the text -- this is an instruction-TUNED encoder-decoder, not a chat
52
+ # model, and it has no template to apply.
53
+ print(model.text2text.infer(
54
+ "translate English to German: The weather today is cold and rainy in the north.",
55
+ max_new_tokens=32))
56
+
57
+ # The same door, on the other tasks this checkpoint was tuned for:
58
+ print(model.text2text.infer("Answer the following question. What is the capital of France?"))
59
+ print(model.text2text.infer(
60
+ "summarize: The committee met on Tuesday to discuss the annual budget. After three hours of "
61
+ "debate, the members agreed to postpone the vote until the following month, citing incomplete "
62
+ "figures from the finance office.",
63
+ max_new_tokens=64))
64
+ ```
65
+
66
+ ### The layer underneath
67
+
68
+ The call above is the high-level door: one per task, named for the modality pair it maps between, with
69
+ the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)`
70
+ passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
71
+ door does not name.
72
+
73
+ `model.driver_source` prints that driver, including a header comment documenting every argument it
74
+ accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and
75
+ [loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two.
76
+
77
+ ## Known limitations
78
+
79
+ **It is 80M parameters, and it answers like it.** flan-t5-small is the smallest member of the flan collection: it follows an instruction's SHAPE reliably and is often wrong about the content. Asked for the capital of France it says `london`, and `transformers` says the same thing on the same checkpoint -- that is the model, not the export. Use it to try the interface, and step up to `flan-t5-base` or larger for answers you intend to keep.
80
+
81
+ **The instruction is part of the text.** There is no chat template and no system turn; prompts look like `translate English to German: ...` or `Answer the following question. ...`, which is how the checkpoint was tuned. A bare sentence with no instruction is out of distribution and usually comes back as a fragment of itself.
82
+
83
+ **Greedy, not beam search.** The upstream `config.json` names 4 beams for its translation and summarization presets. The export decodes one token at a time from the argmax (or from the sampler, if you pass `temperature=`), so translations here are the greedy path through the same model rather than what the published presets would give.
84
+
85
+ Sequences are capped at 512 tokens on each side. T5 has no learned position table, so nothing in the file stops a longer source -- but its relative-position buckets saturate at 128 tokens of distance and the KV cache is built at 512, which is the length this export is honest about.
86
+
87
+ ## Files
88
+
89
+ - `flan-t5-small.gguf` -- the model, exported with loom-exporter.
flan-t5-small.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:25f9c5cf10fecf771cb45c6fc521bc3ff03d7a93655c5f96e9737af448e5aafd
3
+ size 309134336