Adds TinyMoE-100m-2x8 to the leaderboard
#36
by FlameF0X - opened
Model: FlameF0X/TinyMoE-100m-2x8
Params: 99.8M (Mixtral-style MoE, 8 local experts / 2 active per token, 10 layers, hidden=384)
Training data: ~70-80% TinyStories, ~20-30% wikitext-103 (~0.6B tokens, 1 epoch)
Hardware for eval: Tesla T4
Eval harness: lm-evaluation-harness v0.4.12
Results:
- BLiMP: 61.13%
- ARC-Easy (acc): 26.68%
- WikiText-2 bits-per-byte: 1.955
Model card: https://huggingface.co/FlameF0X/TinyMoE-100m-2x8
CompactAI changed pull request status to merged
hmm trained on wikitext?
its mostly trained on tinystories. it cant recall Wikipedia facts.
if this makes you feel better, i have a the same model trained on wikitext only and it performs worse
combined_dataset = interleave_datasets(
[tinystories, wikitext],
probabilities=[0.8, 0.2],
seed=42,
)
im working rn at a version thats trained on cosmopedia-v2 and fineweb-edu-dedup from HuggingFaceTB/smollm-corpus