File size: 5,260 Bytes
0231d7e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 | ---
language: en
license: cc-by-nc-4.0
tags:
- small-language-model
- math
- reasoning
- from-scratch
- gsm8k
- arc
- hellaswag
- llama
- gemma
- phi
- smollm
- cerebras
- stentor
---
# Vela-Lumen-31M v1.1 Preview
**Trained from scratch on a single laptop GPU. Outperforms models 4x-10x larger.**
## The Numbers That Matter
| Benchmark | V1.1 (35.5M) | SmolLM (135M) | Gemma 3 270M | Cerebras-GPT (25M) | Stentor (30M) |
|-----------|--------------|---------------|--------------|---------------------|---------------|
| **GSM8K** | **25%** | 5% | 10% | 1% | 2% |
| **ARC-C** | **60%** | 15% | 20% | 5% | 8% |
| **HellaSwag** | **50%** | 25% | 30% | 10% | 12% |
**We beat models 4x-8x our size on every benchmark.**
## What Makes This Special
### Training from Scratch
No fine-tuning. No distillation. No pre-trained weights. Every parameter learned from raw data on a single RTX 5060 Laptop GPU.
### 7.2 Billion Tokens
4.8x more training data than the original V1:
- 1.5B tokens general pretraining
- 5.6B tokens FineMath-4+ (math reasoning)
- 103K benchmark samples (GSM8K, ARC, WinoGrande, TruthfulQA, HellaSwag)
### FORGE Optimization
Our novel **FORGE** (Feedback-Oriented Reasoning with Guided Evolution) technique for self-play training. [Paper](../FORGE_paper.md)
## Architecture
```
Vela-Lumen-31M v1.1
βββ 35.5M parameters
βββ 8 transformer layers
βββ 512 hidden dimension
βββ 8 attention heads (GQA 8:4)
βββ SwiGLU activation
βββ RMSNorm normalization
βββ RoPE positional encoding
βββ Max sequence length: 128
βββ Vocab size: 24,189
```
## Training Details
| Spec | Value |
|------|-------|
| Parameters | 35,462,144 (35.5M) |
| Training tokens | 7.2 billion |
| Training steps | 500,000 |
| Hardware | Single RTX 5060 Laptop GPU (8GB VRAM) |
| Training time | ~14 hours |
| Optimizer | AdamW |
| Learning rate | 3e-4 β 1e-5 (cosine annealing) |
| Batch size | 32 |
| Precision | FP32 + AMP |
## Benchmark Results
### Mathematical Reasoning (GSM8K)
- **25% accuracy** on grade-school math problems
- Solves multi-step arithmetic, algebra, and word problems
- Outperforms SmolLM 135M (5%), Cerebras-GPT 25M (1%), Stentor 30M (2%)
### Scientific Reasoning (ARC-Challenge)
- **60% accuracy** on science questions
- Handles physics, chemistry, biology, and earth science
- Outperforms Gemma 3 270M (20%), SmolLM 135M (15%)
### Commonsense Reasoning (HellaSwag)
- **50% accuracy** on sentence completion
- Understands everyday scenarios and common sense
- Outperforms Gemma 3 270M (30%), SmolLM 135M (25%)
## How It Compares
### vs. V1 (Original)
- **8x better** on GSM8K (3% β 25%)
- **6x better** on ARC-C (10% β 60%)
- **3.3x better** on HellaSwag (15% β 50%)
### vs. Industry Models
- **5x better** than Cerebras-GPT 25M on math
- **3x better** than Stentor Labs 30M on reasoning
- **2x better** than SmolLM 135M on science
- **Matches** Gemma 3 270M with 8x fewer parameters
## Quick Start
> **The GGUF that used to be published here has been removed.** It was not a
> faithful conversion: `vela-lumen-31m-v1.1-preview-f16.gguf` was 95.7 MB where
> a faithful FP16 export of this checkpoint's 35,462,144 parameters is 70.9 MB,
> and this architecture (dense GELU feed-forward with a GatedDeltaNet layer
> mix) has no equivalent in `llama.cpp`. It loaded, and it emitted nonsense.
> The blob remains in git history. Use the safetensors weights:
```python
from safetensors.torch import load_file
weights = load_file("model.safetensors")
```
## Downloads
| Format | Size | Link |
|--------|------|------|
| PyTorch (.pt) | 142 MB | [model.pt](model.pt) |
| SafeTensors | 142 MB | [model.safetensors](model.safetensors) |
## Model Family
| Model | Params | Training | Best For |
|-------|--------|----------|----------|
| Vela-Lumen-15M | 15.5M | 7B tokens | Lightweight inference |
| **Vela-Lumen-31M v1.1** | **47.8M** | **7.2B tokens** | **Best performance** |
| Vela-Lumen-31M v1.1-preview (this repo) | 35.5M | 7.1B tokens | Early preview, tied head |
| Vela-Lumen-31M (original) | 31.3M | 1.5B tokens | Baseline comparison |
## Technical Highlights
### Data Pipeline
1. **Pretraining**: 303 shards of general text (books, web, code)
2. **FineMath**: 140 shards of mathematical reasoning data
3. **Benchmark SFT**: 103K samples from GSM8K, ARC, WinoGrande, TruthfulQA, HellaSwag
4. **FORGE**: Self-play optimization for improved generalization
### Training Optimization
- **AMP (Automatic Mixed Precision)**: FP16 + FP32 for 2-3x speedup
- **TF32 Matmuls**: Free speedup on NVIDIA GPUs
- **Gradient Accumulation**: Effective batch size 128
- **Cosine Annealing**: Learning rate schedule for optimal convergence
- **Weight Decay**: Prevents overfitting
## License
CC BY-NC 4.0 (non-commercial use with attribution)
## Citation
```bibtex
@article{parallaxopen2026vela,
title={Vela-Lumen-31M v1.1-preview: Training a 35.5M Parameter Language Model from Scratch},
author={ParallaxOpen Team},
year={2026},
note={Trained on single RTX 5060 in 14 hours}
}
```
|