Vela-Lumen-31M v1.2
A 35.5M parameter language model trained from scratch on diverse data.
Performance (Full Benchmarks, All Samples)
| Benchmark | V1.2 | V1.1 | Notes |
|---|---|---|---|
| GSM8K (8,792) | 50.5% | 60.2% | Math word problems |
| ARC-Challenge (2,590) | 26.0% | 22.5% | Science reasoning (MC) |
| HellaSwag (5,000) | 24.9% | 24.9% | Commonsense (MC) |
| TruthfulQA (817) | 25.5% | 29.8% | Truthfulness |
| WinoGrande (5,000) | 51.2% | 51.2% | Coreference (MC) |
Key Improvements over V1.1
- Faster training β 3.4h vs 6.5h (2x faster)
- More diverse data β Trained on chat, math, and code data
- Better ARC-C β 26.0% vs 22.5% (16% improvement)
- Comparable GSM8K β 50.5% vs 60.2% (slightly lower but still strong)
Honest Assessment
V1.2 trades GSM8K performance for faster training and better ARC-C. The model was trained on 1.5B pretrain tokens + 89K SFT samples in 3.4h, compared to V1.1's 7.1B tokens in 6.5h. This shows the importance of training data quantity for reasoning tasks.
Architecture
| Component | Value |
|---|---|
| Parameters | 35,473,665 (35.5M) |
| Layers | 8 |
| Hidden dim | 512 |
| Q heads | 8 |
| KV heads | 4 (GQA) |
| FFN dim | 2,048 |
| Norm | RMSNorm |
| Position | RoPE (theta=10,000) |
| Vocab | 24,189 (BPE) |
| Max seq | 128 tokens |
| Weight tying | Yes |
Training
| Phase | Data | Tokens | Steps | Time |
|---|---|---|---|---|
| Pretrain | FineMath + Web | 1.5B | 160K | ~2.5h |
| SFT | Chat + Math + Code | 89K samples | 40K | ~0.9h |
| Total | 1.5B tokens | 200K | ~3.4h |
- Hardware: Single RTX 5060 Laptop GPU (8GB VRAM)
- Optimizer: AdamW (lr=3e-4, cosine schedule, weight_decay=0.05)
- Mixed precision: FP16 with gradient scaling
- Tokenizer: Real BPE (24,189 vocab)
Downloads
- HuggingFace
- GGUF, SafeTensors, tokenizer included
Quick Start (Ollama)
wget https://huggingface.co/ParallaxOpen/Vela-Lumen-31M-v1.2/resolve/main/vela-lumen-31m-v1.2-f16.gguf
cat > Modelfile << 'EOF'
FROM vela-lumen-31m-v1.2-f16.gguf
TEMPLATE "{{ .System }}{{ .Prompt }}"
SYSTEM "You are a helpful assistant."
PARAMETER temperature 0.7
EOF
ollama create vela-lumen-31m-v1.2 -f Modelfile
ollama run vela-lumen-31m-v1.2
Why V1.2?
V1.2 represents an important step in our research:
- Faster training β 2x faster than V1.1
- More diverse data β Trained on chat, math, and code
- Better reasoning β Improved ARC-C performance
- Honest comparison β Shows the trade-off between training time and performance
License
CC BY-NC 4.0 (non-commercial use with attribution)
Citation
@article{parallaxopen2026vela,
title={Vela-Lumen-31M v1.2: Training a 35M Parameter Language Model from Scratch},
author={ParallaxOpen Team},
year={2026},
note={Trained on single RTX 5060 in 3.4 hours with 1.5B tokens}
}
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support