Add link to paper

#1
by nielsr HF Staff - opened
Files changed (1) hide show
  1. README.md +13 -11
README.md CHANGED
@@ -1,18 +1,18 @@
1
  ---
2
- license: apache-2.0
 
3
  library_name: transformers
 
4
  pipeline_tag: text-generation
5
- language:
6
- - en
7
  tags:
8
- - daedalus
9
- - cpu-inference
10
- - lfm2
11
- - hybrid
12
  widget:
13
- - text: "What is the capital of France?"
14
- - text: "Explain photosynthesis in one sentence."
15
- - text: "What is the difference between a CPU and a GPU?"
16
  inference:
17
  parameters:
18
  max_new_tokens: 96
@@ -23,6 +23,8 @@ inference:
23
 
24
  # Daedalus-150M — Instruct
25
 
 
 
26
  A 150M-parameter language model built for **CPU inference**. Full attention is
27
  kept in only 6 of its 18 layers; the other 12 use short convolutions whose
28
  memory is two timesteps wide however long the conversation gets. Decoding
@@ -85,4 +87,4 @@ lines. English only, 2048-token context, single seed.
85
  The 4-bit build costs about 6% perplexity — quantisation-aware training was
86
  built but did not run. Roughly 48% of the convolution channels are inert and
87
  cannot be pruned, and the 49,152-entry vocabulary is larger than this model size
88
- warrants. All three are documented in the paper.
 
1
  ---
2
+ language:
3
+ - en
4
  library_name: transformers
5
+ license: apache-2.0
6
  pipeline_tag: text-generation
 
 
7
  tags:
8
+ - daedalus
9
+ - cpu-inference
10
+ - lfm2
11
+ - hybrid
12
  widget:
13
+ - text: What is the capital of France?
14
+ - text: Explain photosynthesis in one sentence.
15
+ - text: What is the difference between a CPU and a GPU?
16
  inference:
17
  parameters:
18
  max_new_tokens: 96
 
23
 
24
  # Daedalus-150M — Instruct
25
 
26
+ This model was presented in the paper [Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference](https://huggingface.co/papers/2608.20210).
27
+
28
  A 150M-parameter language model built for **CPU inference**. Full attention is
29
  kept in only 6 of its 18 layers; the other 12 use short convolutions whose
30
  memory is two timesteps wide however long the conversation gets. Decoding
 
87
  The 4-bit build costs about 6% perplexity — quantisation-aware training was
88
  built but did not run. Roughly 48% of the convolution channels are inert and
89
  cannot be pruned, and the 49,152-entry vocabulary is larger than this model size
90
+ warrants. All three are documented in the paper.