Submission 76053

#21
by Quazim0t0 - opened

Base model link: https://huggingface.co/Quazim0t0/Escarda-86M-Base

Escarda-86M-Base — benchmark scores

Multiple-choice (zero-shot, full test/validation splits)

Task acc acc_stderr acc_norm acc_norm_stderr
arc_easy 0.3801 0.0100 0.3615 0.0099
arc_challenge 0.1886 0.0114 0.2235 0.0122
hellaswag 0.2759 0.0045 0.2832 0.0045
winogrande 0.5162 0.0140 — —
piqa 0.5843 0.0115 0.5631 0.0116
openbookqa 0.1300 0.0150 0.2500 0.0194
boolq 0.5138 0.0087 — —

ArithMark-2.0 (AxiomicLabs, n=2500, chance=0.25)

Metric Value
acc 0.2536 ± 0.0087
acc_norm 0.2348 ± 0.0085

Language modeling

Metric Value
WikiText-2 byte_ppl ↓ 2.2228
BLiMP acc ↑ 0.7144 (12 paradigms × 150 pairs)

GlintResearch org

Please open this as a PR so we don't have to guess at the color or brand name you want

GlintResearch org

Edit: ill just do it

I apologize for not doing it correctly, I just saw this. Thank you.

CompactAI changed discussion status to closed

Sign up or log in to comment