Vela-Lumen-31M v1.2

A 35.5M parameter language model trained from scratch on diverse data.

Performance (Full Benchmarks, All Samples)

Benchmark V1.2 V1.1 Notes
GSM8K (8,792) 50.5% 60.2% Math word problems
ARC-Challenge (2,590) 26.0% 22.5% Science reasoning (MC)
HellaSwag (5,000) 24.9% 24.9% Commonsense (MC)
TruthfulQA (817) 25.5% 29.8% Truthfulness
WinoGrande (5,000) 51.2% 51.2% Coreference (MC)

Key Improvements over V1.1

  • Faster training β€” 3.4h vs 6.5h (2x faster)
  • More diverse data β€” Trained on chat, math, and code data
  • Better ARC-C β€” 26.0% vs 22.5% (16% improvement)
  • Comparable GSM8K β€” 50.5% vs 60.2% (slightly lower but still strong)

Honest Assessment

V1.2 trades GSM8K performance for faster training and better ARC-C. The model was trained on 1.5B pretrain tokens + 89K SFT samples in 3.4h, compared to V1.1's 7.1B tokens in 6.5h. This shows the importance of training data quantity for reasoning tasks.

Architecture

Component Value
Parameters 35,473,665 (35.5M)
Layers 8
Hidden dim 512
Q heads 8
KV heads 4 (GQA)
FFN dim 2,048
Norm RMSNorm
Position RoPE (theta=10,000)
Vocab 24,189 (BPE)
Max seq 128 tokens
Weight tying Yes

Training

Phase Data Tokens Steps Time
Pretrain FineMath + Web 1.5B 160K ~2.5h
SFT Chat + Math + Code 89K samples 40K ~0.9h
Total 1.5B tokens 200K ~3.4h
  • Hardware: Single RTX 5060 Laptop GPU (8GB VRAM)
  • Optimizer: AdamW (lr=3e-4, cosine schedule, weight_decay=0.05)
  • Mixed precision: FP16 with gradient scaling
  • Tokenizer: Real BPE (24,189 vocab)

Downloads

Quick Start (Ollama)

wget https://huggingface.co/ParallaxOpen/Vela-Lumen-31M-v1.2/resolve/main/vela-lumen-31m-v1.2-f16.gguf

cat > Modelfile << 'EOF'
FROM vela-lumen-31m-v1.2-f16.gguf
TEMPLATE "{{ .System }}{{ .Prompt }}"
SYSTEM "You are a helpful assistant."
PARAMETER temperature 0.7
EOF

ollama create vela-lumen-31m-v1.2 -f Modelfile
ollama run vela-lumen-31m-v1.2

Why V1.2?

V1.2 represents an important step in our research:

  • Faster training β€” 2x faster than V1.1
  • More diverse data β€” Trained on chat, math, and code
  • Better reasoning β€” Improved ARC-C performance
  • Honest comparison β€” Shows the trade-off between training time and performance

License

CC BY-NC 4.0 (non-commercial use with attribution)

Citation

@article{parallaxopen2026vela,
  title={Vela-Lumen-31M v1.2: Training a 35M Parameter Language Model from Scratch},
  author={ParallaxOpen Team},
  year={2026},
  note={Trained on single RTX 5060 in 3.4 hours with 1.5B tokens}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support