Adds TinyMoE-100m-2x8 to the leaderboard

#36

Model: FlameF0X/TinyMoE-100m-2x8
Params: 99.8M (Mixtral-style MoE, 8 local experts / 2 active per token, 10 layers, hidden=384)
Training data: ~70-80% TinyStories, ~20-30% wikitext-103 (~0.6B tokens, 1 epoch)
Hardware for eval: Tesla T4
Eval harness: lm-evaluation-harness v0.4.12

Results:

  • BLiMP: 61.13%
  • ARC-Easy (acc): 26.68%
  • WikiText-2 bits-per-byte: 1.955

Model card: https://huggingface.co/FlameF0X/TinyMoE-100m-2x8

CompactAI changed pull request status to merged

hmm trained on wikitext?

its mostly trained on tinystories. it cant recall Wikipedia facts.

if this makes you feel better, i have a the same model trained on wikitext only and it performs worse

 combined_dataset = interleave_datasets(
    [tinystories, wikitext],
    probabilities=[0.8, 0.2],
     seed=42,
 )

im working rn at a version thats trained on cosmopedia-v2 and fineweb-edu-dedup from HuggingFaceTB/smollm-corpus

Sign up or log in to comment