Add GPT-U-20M (DedeProGames) and show it by default
#10
by DedeProGames - opened
DedeProGames/GPT-U-20M is a 20,453,760-parameter Llama-architecture base model trained from scratch on 2.6B tokens (DCLM 45% / FineWeb-Edu 35% / The Stack v3 20%).
Changes (matched-models.js only)
- New
MATCHED_ROWSrow with the measurements below. LOCAL_RUNSentry (context window 1024 + provenance).GPT-U-20Madded toRECOMMENDED_MODELS, so it is selected by default when the page opens.
| Benchmark | Score | Method |
|---|---|---|
| BananaMind Base Bench 1.1 | 945 Elo (152/350, 43.43%) | official benchmark.py, complete 350-item split, bfloat16, no BOS |
| ARC Easy | 37.25 | lm-eval 0.4.13, zero-shot acc_norm, float32 |
| HellaSwag | 27.85 | lm-eval 0.4.13, zero-shot acc_norm, float32 |
| PIQA | 58.38 | lm-eval 0.4.13, zero-shot acc_norm, float32 |
| Arithmark 3.0 | 35.10 | official bencharithmark-3.py, acc_norm, bfloat16 |
Base Bench category Elo:
| Category | Elo |
|---|---|
| Language completion | 1174 |
| Commonsense | 879 |
| World knowledge | 883 |
| Context tracking | 761 |
| Quantitative | 861 |
| Logical reasoning | 1073 |
| Code completion | 1012 |
Checked locally by evaluating data.js, matched-models.js, organizations.js and scoring.js in Node: the model is recommended (21 default models), index 41.09, size frontier within ±2M (BananaMind-Sundae: 37.2), no duplicate ids.
Reviewed automatically and merged. Adds a single new leaderboard entry for DedeProGames/GPT-U-20M (20.5M params) with a matching local-run provenance record and marks it recommended.
Automated review by BananaMindBot using z-ai/glm-5.3-flash. A maintainer can override this.
BananaMindBot changed pull request status to merged