Add benchmark overview and training dataset associations
Browse files- .gitattributes +1 -0
- README.md +38 -0
- SHA256SUMS +2 -1
- assets/overview.png +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
assets/overview.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -10,6 +10,20 @@ tags:
|
|
| 10 |
- rlcd
|
| 11 |
- lora
|
| 12 |
- research
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 13 |
---
|
| 14 |
|
| 15 |
# Dohnuts-0.1.0-0.8B
|
|
@@ -20,6 +34,8 @@ ordered scores. Independent questions about one input share computation in a
|
|
| 20 |
single forward pass. The interface returns decisions without generating reasoning
|
| 21 |
or free-form answers.
|
| 22 |
|
|
|
|
|
|
|
| 23 |
## Model details
|
| 24 |
|
| 25 |
| Property | Value |
|
|
@@ -130,6 +146,28 @@ samples perturb decision logits; they are not generated trajectories. LoRA is
|
|
| 130 |
merged before temperature fitting and final evaluation. The
|
| 131 |
[RLCD specification](https://github.com/PsiACE/dohnuts/blob/main/docs/rlcd.md) gives the objective and fixed schedule.
|
| 132 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 133 |
## Limitations
|
| 134 |
|
| 135 |
- Quality depends on the task. Jev leads on the public JevBench tasks; Laya
|
|
|
|
| 10 |
- rlcd
|
| 11 |
- lora
|
| 12 |
- research
|
| 13 |
+
datasets:
|
| 14 |
+
- AmazonScience/massive
|
| 15 |
+
- google/boolq
|
| 16 |
+
- fancyzhx/ag_news
|
| 17 |
+
- dair-ai/emotion
|
| 18 |
+
- PolyAI/banking77
|
| 19 |
+
- facebook/xnli
|
| 20 |
+
- HuggingFaceM4/A-OKVQA
|
| 21 |
+
- derek-thomas/ScienceQA
|
| 22 |
+
- lmms-lab-encoder/VQAv2
|
| 23 |
+
- LocalLLaMA/typed-decisions
|
| 24 |
+
- microsoft/wiki_qa
|
| 25 |
+
- UCLNLP/sharc
|
| 26 |
+
- ucirvine/sms_spam
|
| 27 |
---
|
| 28 |
|
| 29 |
# Dohnuts-0.1.0-0.8B
|
|
|
|
| 34 |
single forward pass. The interface returns decisions without generating reasoning
|
| 35 |
or free-form answers.
|
| 36 |
|
| 37 |
+

|
| 38 |
+
|
| 39 |
## Model details
|
| 40 |
|
| 41 |
| Property | Value |
|
|
|
|
| 146 |
merged before temperature fitting and final evaluation. The
|
| 147 |
[RLCD specification](https://github.com/PsiACE/dohnuts/blob/main/docs/rlcd.md) gives the objective and fixed schedule.
|
| 148 |
|
| 149 |
+
### Training datasets
|
| 150 |
+
|
| 151 |
+
The 26 training groups are derived from the following source datasets. Hub links
|
| 152 |
+
identify the datasets; the [download manifests](https://github.com/PsiACE/dohnuts/tree/main/data/manifests/) pin the files,
|
| 153 |
+
revisions, and checksums actually used, including official archives downloaded
|
| 154 |
+
outside the Hub.
|
| 155 |
+
|
| 156 |
+
| Task family | Sources |
|
| 157 |
+
| --- | --- |
|
| 158 |
+
| Intent and topic classification | [MASSIVE 1.1](https://huggingface.co/datasets/AmazonScience/massive) (en-US, zh-CN), [AG News](https://huggingface.co/datasets/fancyzhx/ag_news), [BANKING77](https://huggingface.co/datasets/PolyAI/banking77) |
|
| 159 |
+
| Entailment, emotion, and Boolean QA | [XNLI](https://huggingface.co/datasets/facebook/xnli) (en, zh), [emotion](https://huggingface.co/datasets/dair-ai/emotion), [BoolQ](https://huggingface.co/datasets/google/boolq) (SuperGLUE distribution) |
|
| 160 |
+
| Visual decisions | [CLEVR 1.0](https://cs.stanford.edu/people/jcjohns/clevr/), [A-OKVQA](https://huggingface.co/datasets/HuggingFaceM4/A-OKVQA), [ScienceQA](https://huggingface.co/datasets/derek-thomas/ScienceQA) (image subset), [VQAv2](https://huggingface.co/datasets/lmms-lab-encoder/VQAv2) (yes/no) |
|
| 161 |
+
| Screen region decisions | [ScreenQA](https://github.com/google-research-datasets/screen_qa), with [Rico](https://www.interactionmining.org/archive/rico) screenshots and view hierarchies |
|
| 162 |
+
| Typed decisions | [LocalLLaMA/typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions), using public soft teacher distributions |
|
| 163 |
+
| Retrieval and relevance | [Amazon ESCI](https://github.com/amazon-science/esci-data) (en, es, ja), [WikiQA](https://huggingface.co/datasets/microsoft/wiki_qa) |
|
| 164 |
+
| Policy and contract decisions | [ShARC](https://huggingface.co/datasets/UCLNLP/sharc), [ContractNLI](https://stanfordnlp.github.io/contract-nli/) |
|
| 165 |
+
| Spam and phishing | [SpamAssassin](https://spamassassin.apache.org/old/publiccorpus/), [Nazario phishing corpus](https://monkey.org/~jose/phishing/), [UCI SMS Spam Collection](https://huggingface.co/datasets/ucirvine/sms_spam) |
|
| 166 |
+
|
| 167 |
+
JevBench tasks and the frozen Laya benchmark inputs are evaluation-only.
|
| 168 |
+
The [data protocol](https://github.com/PsiACE/dohnuts/blob/main/docs/data-and-evaluation.md) describes source-specific
|
| 169 |
+
conversions, grouped partitions, exclusions, and terms.
|
| 170 |
+
|
| 171 |
## Limitations
|
| 172 |
|
| 173 |
- Quality depends on the task. Jev leads on the public JevBench tasks; Laya
|
SHA256SUMS
CHANGED
|
@@ -1,5 +1,6 @@
|
|
| 1 |
600ca4e25fe11762b75a97e714707fab48bb778374e92d24c6ca068791661c11 LICENSE
|
| 2 |
f3b4840eb0568e827f4ba3096850aecbf5347a0fa3307542cf2b8e40ba9367d5 NOTICE
|
| 3 |
-
|
| 4 |
196be33a0282537bcd821e2115643b352d0ad2a0bbaf7b242a1b7fe5bd96cfdf adapter.safetensors
|
| 5 |
379b05ae34a1f5758a5ccec793b5b059273ab4f4496434f6edefeedd22e2baa4 dohnuts.json
|
|
|
|
|
|
| 1 |
600ca4e25fe11762b75a97e714707fab48bb778374e92d24c6ca068791661c11 LICENSE
|
| 2 |
f3b4840eb0568e827f4ba3096850aecbf5347a0fa3307542cf2b8e40ba9367d5 NOTICE
|
| 3 |
+
b2a8fb161d262d35bef46a3a2788ae50da77df8b19c15a2a4d5814866c34221c README.md
|
| 4 |
196be33a0282537bcd821e2115643b352d0ad2a0bbaf7b242a1b7fe5bd96cfdf adapter.safetensors
|
| 5 |
379b05ae34a1f5758a5ccec793b5b059273ab4f4496434f6edefeedd22e2baa4 dohnuts.json
|
| 6 |
+
8b9cdf495fa2b5b70f05c3b015f4f1013e8cd5891be0b9cf8181590608558312 assets/overview.png
|
assets/overview.png
ADDED
|
Git LFS Details
|