Add Pollock 1.5 Mini LM 128M (Fabryka AI)
Browse files# Add Pollock 1.5 Mini LM 128M (Fabryka AI)
Adds **Pollock 1.5 Mini LM 128M**, a 127.43M-parameter English base language
model trained from scratch by Fabryka AI.
- Model: https://huggingface.co/SlayerLab/pollock-mini-lm-125m
- Evaluated revision: [`f2a1ec072098d257c36b129b0e554e2e6de9c46e`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/tree/f2a1ec072098d257c36b129b0e554e2e6de9c46e)
- Architecture: GPT-2-compatible decoder, 16 layers, 768 hidden width, 12 attention heads, 3,200-wide MLP, 2,048-token context
- Training tokens: 21,503,919,992
- License: mixed upstream dataset terms, documented in the model repository
## Results
All results are zero-shot, complete-split evaluations with
`lm-evaluation-harness 0.4.12` and Transformers `5.15.1`.
| Benchmark | Metric | Score |
| --- | --- | ---: |
| BLiMP | `acc` | 78.4970 |
| ARC-Easy | `acc` | 47.7273 |
| WikiText-2 | `byte_perplexity` | 1.9435 |
BLiMP and ARC-Easy were evaluated in BF16 with batch size 8 and maximum context
1,024 on Apple M1 Max MPS. The full raw result includes both ARC-Easy `acc`
(`0.4772727273`) and `acc_norm` (`0.4183501684`):
https://huggingface.co/SlayerLab/pollock-mini-lm-125m/blob/f2a1ec072098d257c36b129b0e554e2e6de9c46e/benchmarks/english.json
WikiText-2 was evaluated separately on an NVIDIA RTX 4090 using BF16, batch
size 8, the model's full 2,048-token context, all 62 test documents, and no
sample limit. The evaluation took 22.92 seconds.
```bash
lm_eval \
--model hf \
--model_args "pretrained=SlayerLab/pollock-mini-lm-125m,revision=f2a1ec072098d257c36b129b0e554e2e6de9c46e,dtype=bfloat16,max_length=2048" \
--tasks wikitext \
--num_fewshot 0 \
--batch_size 8 \
--device cuda:0 \
--seed "0,1234,1234,1234"
```
The WikiText-2 run also produced `bits_per_byte = 0.9586784660` and
`word_perplexity = 34.9322042476`.
## Parameter count
The leaderboard entry uses the native model's 127,427,328 unique trainable
parameters. The Transformers compatibility artifact contains 127,565,312
serialized parameters because `GPT2LMHeadModel` requires an additional 137,984
zero-valued bias parameters that were absent from the trained native model.
## Changes
- Adds one model object to `models`.
- Uses `Fabryka AI` directly as the organization name.
- Adds the organization color `#963200`.
- index.html +14 -0
|
@@ -588,6 +588,19 @@ footer a:hover { text-decoration: underline; }
|
|
| 588 |
</div>
|
| 589 |
<script>
|
| 590 |
const models = [
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 591 |
{
|
| 592 |
name: "JugnuLM-53M",
|
| 593 |
org: "altslate",
|
|
@@ -1643,6 +1656,7 @@ const orgNameMap = {
|
|
| 1643 |
};
|
| 1644 |
|
| 1645 |
const colorMap = {
|
|
|
|
| 1646 |
glintresearch: '#3fb950',
|
| 1647 |
supralabs: '#58a6ff',
|
| 1648 |
axiomiclabs: '#c2b6ff',
|
|
|
|
| 588 |
</div>
|
| 589 |
<script>
|
| 590 |
const models = [
|
| 591 |
+
{
|
| 592 |
+
name: "Pollock 1.5 Mini LM 128M",
|
| 593 |
+
org: "Fabryka AI",
|
| 594 |
+
params: "127.43M",
|
| 595 |
+
blimp: 78.497,
|
| 596 |
+
arc: 47.7273,
|
| 597 |
+
wiki: 1.9435,
|
| 598 |
+
tokens: "21.50B",
|
| 599 |
+
releaseDate: "2026-09-20",
|
| 600 |
+
links: {
|
| 601 |
+
card: "https://huggingface.co/SlayerLab/pollock-mini-lm-125m"
|
| 602 |
+
}
|
| 603 |
+
},
|
| 604 |
{
|
| 605 |
name: "JugnuLM-53M",
|
| 606 |
org: "altslate",
|
|
|
|
| 1656 |
};
|
| 1657 |
|
| 1658 |
const colorMap = {
|
| 1659 |
+
'Fabryka AI': '#963200',
|
| 1660 |
glintresearch: '#3fb950',
|
| 1661 |
supralabs: '#58a6ff',
|
| 1662 |
axiomiclabs: '#c2b6ff',
|