Add peacebell-v1-148M (locally benchmarked, Base Bench 1.1 Elo 902)
Browse filesAdds **peacebell-v1-148M**, benchmarked locally with the official runners: https://huggingface.co/wayneworkman2012/peacebell-v1-148M
148,553,302 parameters (tied head counted once), pretrained from scratch by me, open weights, Apache-2.0, 16K context.
**Please note: this is a single-domain model.** It was trained only on World War II text (301 Wikipedia articles plus synthetic data derived from them), with no general web text. Chance-level scores on the general suites are the real result, not a loading problem. I am adding it as a data point for what a from-scratch, one-domain model looks like on this index.
**Results**
| Benchmark | Score |
|---|---|
| BananaMind Base Bench 1.1, Overall Elo | **902** (131 / 350 correct) |
| ARC Easy (acc_norm) | 27.27 |
| HellaSwag (acc_norm) | 26.16 |
| PIQA (acc_norm) | 47.61 |
| Arithmark 3 (acc_norm) | 27.10 |
Base Bench category Elo: language 1035, commonsense 826, world knowledge 929, context 954, quantitative 774, logic 891, code 924.
**Method**, same as your `LOCAL_RUN_METHOD`:
- Base Bench 1.1: official `benchmark.py`, unmodified, `--device cuda --dtype bfloat16`, data pulled from the Hub (SHA-256 `2f563bb4...abea21`, matches the official value). For transparency: in float32 the same run gives Elo 893 (127 / 350); four near-tie items flip between the two dtypes. I report the bfloat16 number because that is the documented command, and the provenance line states both.
- ARC Easy, HellaSwag, PIQA: lm-eval 0.4.13, `hf` backend, zero-shot, acc_norm, float32.
- Arithmark 3: official script, acc_norm, bfloat16 default.
To verify: the model needs `trust_remote_code=True` (the default in your runner) and `pip install sentencepiece`.
```
python benchmark.py --model wayneworkman2012/peacebell-v1-148M --device cuda --dtype bfloat16 --out-dir runs/peacebell-v1-148M
```
**What this PR changes**
Three added lines in `matched-models.js`, nothing removed: one `MATCHED_ROWS` row, one `ORGANIZATION_ROWS` entry (`wayneworkman`), and one `LOCAL_RUNS` entry. The `LOCAL_RUNS` entry is written as its own literal instead of calling `localRun()`, because that helper stamps 2026-09-13 and lm-eval 0.4.10, and I did not want to touch a helper other entries share. I ran your `data.js`, `matched-models.js`, `organizations.js` and `scoring.js` under node before and after: no row throws, every org resolves, and every other model's index, frontier status, profile bars and provenance text are identical.
I accepted the Base Bench terms; the dataset is used for this evaluation only and is kept away from my training data.
馃 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01Kpv2a5EA4htxnrqzE1VdGG
- matched-models.js +3 -0
|
@@ -63,6 +63,7 @@ const MATCHED_ROWS = [
|
|
| 63 |
['NanoDex-1M','dedeprogames',1062272,'DedeProGames/NanoDex-1M',29.17,27.01,53.48,26.50,827,892,850,828,774,795,939,729],
|
| 64 |
['BananaMind-Sundae','bananamindresearch',20156544,'bananamind-research-community/BananaMind-Sundae',33.92,26.08,54.46,34.20,891,1167,789,874,928,785,920,831],
|
| 65 |
['BackKiyo-10M','dedebckp',9976832,'DedeBckp/BackKiyo-10M',35.40,28.08,55.66,35.10,924,1070,860,920,815,940,973,924],
|
|
|
|
| 66 |
];
|
| 67 |
|
| 68 |
const ORGANIZATION_ROWS = [
|
|
@@ -77,6 +78,7 @@ const ORGANIZATION_ROWS = [
|
|
| 77 |
['ucr','Universal Computing Research','UniversalComputingResearch','#87bde1'],['allura','Allura','allura-org','#d8aaf3'],
|
| 78 |
['codesoft','CodeSoft','CodeSoft','#c09d80'],['dalab','DALab Community','DALabCommunity','#91c49c'],
|
| 79 |
['dedeprogames','DedeProGames','DedeProGames','#fa9696'],['sz14','sz14','sz14','#83b7a6'],
|
|
|
|
| 80 |
];
|
| 81 |
for (const [id,name,repo,color] of ORGANIZATION_ROWS) {
|
| 82 |
ORGS[id] = {name, logo:null, chartColor:color, chartBorder:color, url:`https://huggingface.co/${repo}`};
|
|
@@ -100,6 +102,7 @@ const LOCAL_RUNS = {
|
|
| 100 |
'LowOnMind-5M':localRun(113,512), 'LowOnMind-1M':localRun(100,512),
|
| 101 |
'LowOnMind-300k':localRun(93,512), 'NanoDex-Test-500K-200M':localRun(88,512),
|
| 102 |
'NanoDex-1M':localRun(98,512), 'BananaMind-Sundae':localRun(134,1024), 'BackKiyo-10M':localRun(144,4096),
|
|
|
|
| 103 |
};
|
| 104 |
const RECOMMENDED_MODELS = new Set(['BananaMind-2-Pro','BananaMind-2-Medium','BananaMind-2-Mini',
|
| 105 |
'SmolLM2-135M','SmolLM-135M','GPT-X2.5-135M','GPT-X2-125M','GPT-2','Kiyo-135M','Kiyo-65M',
|
|
|
|
| 63 |
['NanoDex-1M','dedeprogames',1062272,'DedeProGames/NanoDex-1M',29.17,27.01,53.48,26.50,827,892,850,828,774,795,939,729],
|
| 64 |
['BananaMind-Sundae','bananamindresearch',20156544,'bananamind-research-community/BananaMind-Sundae',33.92,26.08,54.46,34.20,891,1167,789,874,928,785,920,831],
|
| 65 |
['BackKiyo-10M','dedebckp',9976832,'DedeBckp/BackKiyo-10M',35.40,28.08,55.66,35.10,924,1070,860,920,815,940,973,924],
|
| 66 |
+
['peacebell-v1-148M','wayneworkman',148553302,'wayneworkman2012/peacebell-v1-148M',27.27,26.16,47.61,27.10,902,1035,826,929,954,774,891,924],
|
| 67 |
];
|
| 68 |
|
| 69 |
const ORGANIZATION_ROWS = [
|
|
|
|
| 78 |
['ucr','Universal Computing Research','UniversalComputingResearch','#87bde1'],['allura','Allura','allura-org','#d8aaf3'],
|
| 79 |
['codesoft','CodeSoft','CodeSoft','#c09d80'],['dalab','DALab Community','DALabCommunity','#91c49c'],
|
| 80 |
['dedeprogames','DedeProGames','DedeProGames','#fa9696'],['sz14','sz14','sz14','#83b7a6'],
|
| 81 |
+
['wayneworkman','Wayne Workman','wayneworkman2012','#14b8a6'],
|
| 82 |
];
|
| 83 |
for (const [id,name,repo,color] of ORGANIZATION_ROWS) {
|
| 84 |
ORGS[id] = {name, logo:null, chartColor:color, chartBorder:color, url:`https://huggingface.co/${repo}`};
|
|
|
|
| 102 |
'LowOnMind-5M':localRun(113,512), 'LowOnMind-1M':localRun(100,512),
|
| 103 |
'LowOnMind-300k':localRun(93,512), 'NanoDex-Test-500K-200M':localRun(88,512),
|
| 104 |
'NanoDex-1M':localRun(98,512), 'BananaMind-Sundae':localRun(134,1024), 'BackKiyo-10M':localRun(144,4096),
|
| 105 |
+
'peacebell-v1-148M':{contextWindow:16384, provenance:'Local official complete run 路 2026-09-19 路 131 / 350 correct (bfloat16; float32 gives 127 / 350, Elo 893). ARC Easy, HellaSwag and PIQA: lm-eval 0.4.13 zero-shot acc_norm (float32). Arithmark 3: official script acc_norm (bfloat16). Domain-specific model: trained from scratch on World War II text only.'},
|
| 106 |
};
|
| 107 |
const RECOMMENDED_MODELS = new Set(['BananaMind-2-Pro','BananaMind-2-Medium','BananaMind-2-Mini',
|
| 108 |
'SmolLM2-135M','SmolLM-135M','GPT-X2.5-135M','GPT-X2-125M','GPT-2','Kiyo-135M','Kiyo-65M',
|