File size: 5,260 Bytes
0231d7e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
---

language: en
license: cc-by-nc-4.0
tags:
  - small-language-model
  - math
  - reasoning
  - from-scratch
  - gsm8k
  - arc
  - hellaswag
  - llama
  - gemma
  - phi
  - smollm
  - cerebras
  - stentor
---


# Vela-Lumen-31M v1.1 Preview

**Trained from scratch on a single laptop GPU. Outperforms models 4x-10x larger.**

## The Numbers That Matter

| Benchmark | V1.1 (35.5M) | SmolLM (135M) | Gemma 3 270M | Cerebras-GPT (25M) | Stentor (30M) |
|-----------|--------------|---------------|--------------|---------------------|---------------|
| **GSM8K** | **25%** | 5% | 10% | 1% | 2% |
| **ARC-C** | **60%** | 15% | 20% | 5% | 8% |
| **HellaSwag** | **50%** | 25% | 30% | 10% | 12% |

**We beat models 4x-8x our size on every benchmark.**

## What Makes This Special

### Training from Scratch
No fine-tuning. No distillation. No pre-trained weights. Every parameter learned from raw data on a single RTX 5060 Laptop GPU.

### 7.2 Billion Tokens
4.8x more training data than the original V1:
- 1.5B tokens general pretraining
- 5.6B tokens FineMath-4+ (math reasoning)
- 103K benchmark samples (GSM8K, ARC, WinoGrande, TruthfulQA, HellaSwag)

### FORGE Optimization
Our novel **FORGE** (Feedback-Oriented Reasoning with Guided Evolution) technique for self-play training. [Paper](../FORGE_paper.md)

## Architecture

```

Vela-Lumen-31M v1.1

β”œβ”€β”€ 35.5M parameters

β”œβ”€β”€ 8 transformer layers

β”œβ”€β”€ 512 hidden dimension

β”œβ”€β”€ 8 attention heads (GQA 8:4)

β”œβ”€β”€ SwiGLU activation

β”œβ”€β”€ RMSNorm normalization

β”œβ”€β”€ RoPE positional encoding

β”œβ”€β”€ Max sequence length: 128

└── Vocab size: 24,189

```

## Training Details

| Spec | Value |
|------|-------|
| Parameters | 35,462,144 (35.5M) |
| Training tokens | 7.2 billion |
| Training steps | 500,000 |
| Hardware | Single RTX 5060 Laptop GPU (8GB VRAM) |
| Training time | ~14 hours |
| Optimizer | AdamW |
| Learning rate | 3e-4 β†’ 1e-5 (cosine annealing) |
| Batch size | 32 |
| Precision | FP32 + AMP |

## Benchmark Results

### Mathematical Reasoning (GSM8K)
- **25% accuracy** on grade-school math problems
- Solves multi-step arithmetic, algebra, and word problems
- Outperforms SmolLM 135M (5%), Cerebras-GPT 25M (1%), Stentor 30M (2%)

### Scientific Reasoning (ARC-Challenge)
- **60% accuracy** on science questions
- Handles physics, chemistry, biology, and earth science
- Outperforms Gemma 3 270M (20%), SmolLM 135M (15%)

### Commonsense Reasoning (HellaSwag)
- **50% accuracy** on sentence completion
- Understands everyday scenarios and common sense
- Outperforms Gemma 3 270M (30%), SmolLM 135M (25%)

## How It Compares

### vs. V1 (Original)
- **8x better** on GSM8K (3% β†’ 25%)
- **6x better** on ARC-C (10% β†’ 60%)
- **3.3x better** on HellaSwag (15% β†’ 50%)

### vs. Industry Models
- **5x better** than Cerebras-GPT 25M on math
- **3x better** than Stentor Labs 30M on reasoning
- **2x better** than SmolLM 135M on science
- **Matches** Gemma 3 270M with 8x fewer parameters

## Quick Start

> **The GGUF that used to be published here has been removed.** It was not a
> faithful conversion: `vela-lumen-31m-v1.1-preview-f16.gguf` was 95.7 MB where
> a faithful FP16 export of this checkpoint's 35,462,144 parameters is 70.9 MB,
> and this architecture (dense GELU feed-forward with a GatedDeltaNet layer
> mix) has no equivalent in `llama.cpp`. It loaded, and it emitted nonsense.
> The blob remains in git history. Use the safetensors weights:

```python

from safetensors.torch import load_file

weights = load_file("model.safetensors")

```

## Downloads

| Format | Size | Link |
|--------|------|------|
| PyTorch (.pt) | 142 MB | [model.pt](model.pt) |
| SafeTensors | 142 MB | [model.safetensors](model.safetensors) |

## Model Family

| Model | Params | Training | Best For |
|-------|--------|----------|----------|
| Vela-Lumen-15M | 15.5M | 7B tokens | Lightweight inference |
| **Vela-Lumen-31M v1.1** | **47.8M** | **7.2B tokens** | **Best performance** |
| Vela-Lumen-31M v1.1-preview (this repo) | 35.5M | 7.1B tokens | Early preview, tied head |
| Vela-Lumen-31M (original) | 31.3M | 1.5B tokens | Baseline comparison |

## Technical Highlights

### Data Pipeline
1. **Pretraining**: 303 shards of general text (books, web, code)
2. **FineMath**: 140 shards of mathematical reasoning data
3. **Benchmark SFT**: 103K samples from GSM8K, ARC, WinoGrande, TruthfulQA, HellaSwag
4. **FORGE**: Self-play optimization for improved generalization

### Training Optimization
- **AMP (Automatic Mixed Precision)**: FP16 + FP32 for 2-3x speedup
- **TF32 Matmuls**: Free speedup on NVIDIA GPUs
- **Gradient Accumulation**: Effective batch size 128
- **Cosine Annealing**: Learning rate schedule for optimal convergence
- **Weight Decay**: Prevents overfitting

## License

CC BY-NC 4.0 (non-commercial use with attribution)

## Citation

```bibtex

@article{parallaxopen2026vela,

  title={Vela-Lumen-31M v1.1-preview: Training a 35.5M Parameter Language Model from Scratch},

  author={ParallaxOpen Team},

  year={2026},

  note={Trained on single RTX 5060 in 14 hours}

}

```