Upload docs/LOSS_VS_GEN.md with huggingface_hub
Browse files- docs/LOSS_VS_GEN.md +34 -0
docs/LOSS_VS_GEN.md
ADDED
|
@@ -0,0 +1,34 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Why Train Loss Was Not Accurate vs Generation
|
| 2 |
+
|
| 3 |
+
**Date:** 2026-08-16 23:31 UTC
|
| 4 |
+
|
| 5 |
+
## Diagnosis
|
| 6 |
+
|
| 7 |
+
Teacher-forced CE (train metric) and free-run generation measure different things.
|
| 8 |
+
|
| 9 |
+
| Metric | GPU1-ish value | Meaning |
|
| 10 |
+
|--------|----------------|---------|
|
| 11 |
+
| Train TF CE (warm continuous) | ~1.3-1.5 | Next-token prediction with ground-truth history + carried thought state |
|
| 12 |
+
| Cold TF CE (reset state) | ~7-40 | Same objective but without warm dynamical state |
|
| 13 |
+
| Free-run AR CE vs ground truth | huge / match~0 | Model leaves data manifold when feeding its own tokens |
|
| 14 |
+
|
| 15 |
+
So a loss of 1.5 was **accurate for teacher forcing**, not for generation.
|
| 16 |
+
|
| 17 |
+
## Fix in training
|
| 18 |
+
|
| 19 |
+
Scheduled sampling (sequential, VRAM-safe):
|
| 20 |
+
|
| 21 |
+
1. Pass A: teacher-forced CE + LB (always)
|
| 22 |
+
2. Pass B (30% of steps): mix 20% model-sampled tokens into inputs, CE again against true targets
|
| 23 |
+
|
| 24 |
+
Logged now:
|
| 25 |
+
|
| 26 |
+
- `tf=` / `ema_tf=` teacher-forced
|
| 27 |
+
- `ss=` / `ema_ss=` scheduled-sampling pass (gen-closer)
|
| 28 |
+
|
| 29 |
+
Script: `fast4gpu_stage2_ss.py`
|
| 30 |
+
|
| 31 |
+
## Live (after switch)
|
| 32 |
+
|
| 33 |
+
GPU1 example: tf~1.45 while ss~2.9 — gap is visible and optimized.
|
| 34 |
+
|