Add Pollock 1.5 Mini LM 128M (Fabryka AI)
Add Pollock 1.5 Mini LM 128M (Fabryka AI)
Adds Pollock 1.5 Mini LM 128M, a 127.43M-parameter English base language
model trained from scratch by Fabryka AI.
- Model: https://huggingface.co/SlayerLab/pollock-mini-lm-125m
- Evaluated revision:
f2a1ec072098d257c36b129b0e554e2e6de9c46e - Architecture: GPT-2-compatible decoder, 16 layers, 768 hidden width, 12 attention heads, 3,200-wide MLP, 2,048-token context
- Training tokens: 21,503,919,992
- License: mixed upstream dataset terms, documented in the model repository
Results
All results are zero-shot, complete-split evaluations withlm-evaluation-harness 0.4.12 and Transformers 5.15.1.
| Benchmark | Metric | Score |
|---|---|---|
| BLiMP | acc |
78.4970 |
| ARC-Easy | acc |
47.7273 |
| WikiText-2 | byte_perplexity |
1.9435 |
BLiMP and ARC-Easy were evaluated in BF16 with batch size 8 and maximum context
1,024 on Apple M1 Max MPS. The full raw result includes both ARC-Easy acc
(0.4772727273) and acc_norm (0.4183501684):
WikiText-2 was evaluated separately on an NVIDIA RTX 4090 using BF16, batch
size 8, the model's full 2,048-token context, all 62 test documents, and no
sample limit. The evaluation took 22.92 seconds.
lm_eval \
--model hf \
--model_args "pretrained=SlayerLab/pollock-mini-lm-125m,revision=f2a1ec072098d257c36b129b0e554e2e6de9c46e,dtype=bfloat16,max_length=2048" \
--tasks wikitext \
--num_fewshot 0 \
--batch_size 8 \
--device cuda:0 \
--seed "0,1234,1234,1234"
The WikiText-2 run also produced bits_per_byte = 0.9586784660 andword_perplexity = 34.9322042476.
Parameter count
The leaderboard entry uses the native model's 127,427,328 unique trainable
parameters. The Transformers compatibility artifact contains 127,565,312
serialized parameters because GPT2LMHeadModel requires an additional 137,984
zero-valued bias parameters that were absent from the trained native model.
Changes
- Adds one model object to
models. - Uses
Fabryka AIdirectly as the organization name. - Adds the organization color
#963200.