thefinalboss commited on
Commit
ba5cf64
·
verified ·
1 Parent(s): 26a6b48

Upload docs/LOSS_VS_GEN.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. docs/LOSS_VS_GEN.md +34 -0
docs/LOSS_VS_GEN.md ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Why Train Loss Was Not Accurate vs Generation
2
+
3
+ **Date:** 2026-08-16 23:31 UTC
4
+
5
+ ## Diagnosis
6
+
7
+ Teacher-forced CE (train metric) and free-run generation measure different things.
8
+
9
+ | Metric | GPU1-ish value | Meaning |
10
+ |--------|----------------|---------|
11
+ | Train TF CE (warm continuous) | ~1.3-1.5 | Next-token prediction with ground-truth history + carried thought state |
12
+ | Cold TF CE (reset state) | ~7-40 | Same objective but without warm dynamical state |
13
+ | Free-run AR CE vs ground truth | huge / match~0 | Model leaves data manifold when feeding its own tokens |
14
+
15
+ So a loss of 1.5 was **accurate for teacher forcing**, not for generation.
16
+
17
+ ## Fix in training
18
+
19
+ Scheduled sampling (sequential, VRAM-safe):
20
+
21
+ 1. Pass A: teacher-forced CE + LB (always)
22
+ 2. Pass B (30% of steps): mix 20% model-sampled tokens into inputs, CE again against true targets
23
+
24
+ Logged now:
25
+
26
+ - `tf=` / `ema_tf=` teacher-forced
27
+ - `ss=` / `ema_ss=` scheduled-sampling pass (gen-closer)
28
+
29
+ Script: `fast4gpu_stage2_ss.py`
30
+
31
+ ## Live (after switch)
32
+
33
+ GPU1 example: tf~1.45 while ss~2.9 — gap is visible and optimized.
34
+