PsiACE commited on
Commit
959828a
·
verified ·
1 Parent(s): 0e69360

Add benchmark overview and training dataset associations

Browse files
Files changed (4) hide show
  1. .gitattributes +1 -0
  2. README.md +38 -0
  3. SHA256SUMS +2 -1
  4. assets/overview.png +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ assets/overview.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -10,6 +10,20 @@ tags:
10
  - rlcd
11
  - lora
12
  - research
 
 
 
 
 
 
 
 
 
 
 
 
 
 
13
  ---
14
 
15
  # Dohnuts-0.1.0-0.8B
@@ -20,6 +34,8 @@ ordered scores. Independent questions about one input share computation in a
20
  single forward pass. The interface returns decisions without generating reasoning
21
  or free-form answers.
22
 
 
 
23
  ## Model details
24
 
25
  | Property | Value |
@@ -130,6 +146,28 @@ samples perturb decision logits; they are not generated trajectories. LoRA is
130
  merged before temperature fitting and final evaluation. The
131
  [RLCD specification](https://github.com/PsiACE/dohnuts/blob/main/docs/rlcd.md) gives the objective and fixed schedule.
132
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
133
  ## Limitations
134
 
135
  - Quality depends on the task. Jev leads on the public JevBench tasks; Laya
 
10
  - rlcd
11
  - lora
12
  - research
13
+ datasets:
14
+ - AmazonScience/massive
15
+ - google/boolq
16
+ - fancyzhx/ag_news
17
+ - dair-ai/emotion
18
+ - PolyAI/banking77
19
+ - facebook/xnli
20
+ - HuggingFaceM4/A-OKVQA
21
+ - derek-thomas/ScienceQA
22
+ - lmms-lab-encoder/VQAv2
23
+ - LocalLLaMA/typed-decisions
24
+ - microsoft/wiki_qa
25
+ - UCLNLP/sharc
26
+ - ucirvine/sms_spam
27
  ---
28
 
29
  # Dohnuts-0.1.0-0.8B
 
34
  single forward pass. The interface returns decisions without generating reasoning
35
  or free-form answers.
36
 
37
+ ![Dohnuts 0.1.0 model overview and benchmarks](https://huggingface.co/PsiACE/Dohnuts-0.1.0-0.8B/resolve/main/assets/overview.png)
38
+
39
  ## Model details
40
 
41
  | Property | Value |
 
146
  merged before temperature fitting and final evaluation. The
147
  [RLCD specification](https://github.com/PsiACE/dohnuts/blob/main/docs/rlcd.md) gives the objective and fixed schedule.
148
 
149
+ ### Training datasets
150
+
151
+ The 26 training groups are derived from the following source datasets. Hub links
152
+ identify the datasets; the [download manifests](https://github.com/PsiACE/dohnuts/tree/main/data/manifests/) pin the files,
153
+ revisions, and checksums actually used, including official archives downloaded
154
+ outside the Hub.
155
+
156
+ | Task family | Sources |
157
+ | --- | --- |
158
+ | Intent and topic classification | [MASSIVE 1.1](https://huggingface.co/datasets/AmazonScience/massive) (en-US, zh-CN), [AG News](https://huggingface.co/datasets/fancyzhx/ag_news), [BANKING77](https://huggingface.co/datasets/PolyAI/banking77) |
159
+ | Entailment, emotion, and Boolean QA | [XNLI](https://huggingface.co/datasets/facebook/xnli) (en, zh), [emotion](https://huggingface.co/datasets/dair-ai/emotion), [BoolQ](https://huggingface.co/datasets/google/boolq) (SuperGLUE distribution) |
160
+ | Visual decisions | [CLEVR 1.0](https://cs.stanford.edu/people/jcjohns/clevr/), [A-OKVQA](https://huggingface.co/datasets/HuggingFaceM4/A-OKVQA), [ScienceQA](https://huggingface.co/datasets/derek-thomas/ScienceQA) (image subset), [VQAv2](https://huggingface.co/datasets/lmms-lab-encoder/VQAv2) (yes/no) |
161
+ | Screen region decisions | [ScreenQA](https://github.com/google-research-datasets/screen_qa), with [Rico](https://www.interactionmining.org/archive/rico) screenshots and view hierarchies |
162
+ | Typed decisions | [LocalLLaMA/typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions), using public soft teacher distributions |
163
+ | Retrieval and relevance | [Amazon ESCI](https://github.com/amazon-science/esci-data) (en, es, ja), [WikiQA](https://huggingface.co/datasets/microsoft/wiki_qa) |
164
+ | Policy and contract decisions | [ShARC](https://huggingface.co/datasets/UCLNLP/sharc), [ContractNLI](https://stanfordnlp.github.io/contract-nli/) |
165
+ | Spam and phishing | [SpamAssassin](https://spamassassin.apache.org/old/publiccorpus/), [Nazario phishing corpus](https://monkey.org/~jose/phishing/), [UCI SMS Spam Collection](https://huggingface.co/datasets/ucirvine/sms_spam) |
166
+
167
+ JevBench tasks and the frozen Laya benchmark inputs are evaluation-only.
168
+ The [data protocol](https://github.com/PsiACE/dohnuts/blob/main/docs/data-and-evaluation.md) describes source-specific
169
+ conversions, grouped partitions, exclusions, and terms.
170
+
171
  ## Limitations
172
 
173
  - Quality depends on the task. Jev leads on the public JevBench tasks; Laya
SHA256SUMS CHANGED
@@ -1,5 +1,6 @@
1
  600ca4e25fe11762b75a97e714707fab48bb778374e92d24c6ca068791661c11 LICENSE
2
  f3b4840eb0568e827f4ba3096850aecbf5347a0fa3307542cf2b8e40ba9367d5 NOTICE
3
- 6dec0a5d8401c41b0b0c8ba6146e8a306c2327d6fb2e67a8cc32486b5a03f6c9 README.md
4
  196be33a0282537bcd821e2115643b352d0ad2a0bbaf7b242a1b7fe5bd96cfdf adapter.safetensors
5
  379b05ae34a1f5758a5ccec793b5b059273ab4f4496434f6edefeedd22e2baa4 dohnuts.json
 
 
1
  600ca4e25fe11762b75a97e714707fab48bb778374e92d24c6ca068791661c11 LICENSE
2
  f3b4840eb0568e827f4ba3096850aecbf5347a0fa3307542cf2b8e40ba9367d5 NOTICE
3
+ b2a8fb161d262d35bef46a3a2788ae50da77df8b19c15a2a4d5814866c34221c README.md
4
  196be33a0282537bcd821e2115643b352d0ad2a0bbaf7b242a1b7fe5bd96cfdf adapter.safetensors
5
  379b05ae34a1f5758a5ccec793b5b059273ab4f4496434f6edefeedd22e2baa4 dohnuts.json
6
+ 8b9cdf495fa2b5b70f05c3b015f4f1013e8cd5891be0b9cf8181590608558312 assets/overview.png
assets/overview.png ADDED

Git LFS Details

  • SHA256: 8b9cdf495fa2b5b70f05c3b015f4f1013e8cd5891be0b9cf8181590608558312
  • Pointer size: 131 Bytes
  • Size of remote file: 939 kB