Probe READMEs: decay is 1-sqrt (as in the trainer's WSD), not linear
Browse files- probe0_s12A_420k/README.md +1 -1
- probe_s12A_420k/README.md +1 -1
- probe_s12B_420k/README.md +1 -1
probe0_s12A_420k/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
# probe0_s12A_420k — decay-to-zero probe of `s12A_arcmix_pool/` at step 420,000
|
| 2 |
|
| 3 |
-
**Why it exists:** a short side run that measures what the arm would score if training stopped here. The arm (ARC-MIX pool, building pair 1 arm A) trains at a constant learning rate, and constant-LR checkpoints are not comparable with decayed models. This probe copies the arm's step-420,000 checkpoint and runs only the decay phase: the learning rate
|
| 4 |
|
| 5 |
- **Use:** identical to `probe_s12A_420k/` (same checkpoint, data, random stream and 4,000 steps) except that the learning rate decays to 0 instead of 6e-5. Comparing the two probes on the held-out selection sets tests whether decaying fully to zero helps.
|
| 6 |
- **Status:** research checkpoint, **not a leaderboard submission**.
|
|
|
|
| 1 |
# probe0_s12A_420k — decay-to-zero probe of `s12A_arcmix_pool/` at step 420,000
|
| 2 |
|
| 3 |
+
**Why it exists:** a short side run that measures what the arm would score if training stopped here. The arm (ARC-MIX pool, building pair 1 arm A) trains at a constant learning rate, and constant-LR checkpoints are not comparable with decayed models. This probe copies the arm's step-420,000 checkpoint and runs only the decay phase: the learning rate decays from 3e-4 to 0 over 4,000 steps with a 1−sqrt schedule (as in the trainer's WSD decay) (420,000 → 424,000), on the same data and with the same optimizer state. The arm itself is not affected and keeps training.
|
| 4 |
|
| 5 |
- **Use:** identical to `probe_s12A_420k/` (same checkpoint, data, random stream and 4,000 steps) except that the learning rate decays to 0 instead of 6e-5. Comparing the two probes on the held-out selection sets tests whether decaying fully to zero helps.
|
| 6 |
- **Status:** research checkpoint, **not a leaderboard submission**.
|
probe_s12A_420k/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
# probe_s12A_420k — decay probe of `s12A_arcmix_pool/` at step 420,000
|
| 2 |
|
| 3 |
-
**Why it exists:** a short side run that measures what the arm would score if training stopped here. The arm (ARC-MIX pool, building pair 1 arm A) trains at a constant learning rate, and constant-LR checkpoints are not comparable with decayed models. This probe copies the arm's step-420,000 checkpoint and runs only the decay phase: the learning rate
|
| 4 |
|
| 5 |
- **Use:** compared with the flagship on the held-out selection sets, it checks that the higher learning rate is not damaging the model (a rule fixed before the result: WikiText-2 validation worse than the flagship by more than 0.010 byte-ppl, or half-BLiMP worse by more than 0.5 pp, means the pair switches to a 2e-4 peak).
|
| 6 |
- **Status:** research checkpoint, **not a leaderboard submission**.
|
|
|
|
| 1 |
# probe_s12A_420k — decay probe of `s12A_arcmix_pool/` at step 420,000
|
| 2 |
|
| 3 |
+
**Why it exists:** a short side run that measures what the arm would score if training stopped here. The arm (ARC-MIX pool, building pair 1 arm A) trains at a constant learning rate, and constant-LR checkpoints are not comparable with decayed models. This probe copies the arm's step-420,000 checkpoint and runs only the decay phase: the learning rate decays from 3e-4 to 6e-5 over 4,000 steps with a 1−sqrt schedule (as in the trainer's WSD decay) (420,000 → 424,000), on the same data and with the same optimizer state. The arm itself is not affected and keeps training.
|
| 4 |
|
| 5 |
- **Use:** compared with the flagship on the held-out selection sets, it checks that the higher learning rate is not damaging the model (a rule fixed before the result: WikiText-2 validation worse than the flagship by more than 0.010 byte-ppl, or half-BLiMP worse by more than 0.5 pp, means the pair switches to a 2e-4 peak).
|
| 6 |
- **Status:** research checkpoint, **not a leaderboard submission**.
|
probe_s12B_420k/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
# probe_s12B_420k — decay probe of `s12B_arcmix_edu/` at step 420,000
|
| 2 |
|
| 3 |
-
**Why it exists:** a short side run that measures what the arm would score if training stopped here. The arm (ARC-MIX + FineWeb-Edu pool, building pair 1 arm B) trains at a constant learning rate, and constant-LR checkpoints are not comparable with decayed models. This probe copies the arm's step-420,000 checkpoint and runs only the decay phase: the learning rate
|
| 4 |
|
| 5 |
- **Use:** compared with the flagship and with the decay probe of arm A (`probe_s12A_420k/`) on the held-out selection sets, it shows how the two arms would compare after decay at this point.
|
| 6 |
- **Status:** research checkpoint, **not a leaderboard submission**.
|
|
|
|
| 1 |
# probe_s12B_420k — decay probe of `s12B_arcmix_edu/` at step 420,000
|
| 2 |
|
| 3 |
+
**Why it exists:** a short side run that measures what the arm would score if training stopped here. The arm (ARC-MIX + FineWeb-Edu pool, building pair 1 arm B) trains at a constant learning rate, and constant-LR checkpoints are not comparable with decayed models. This probe copies the arm's step-420,000 checkpoint and runs only the decay phase: the learning rate decays from 3e-4 to 6e-5 over 4,000 steps with a 1−sqrt schedule (as in the trainer's WSD decay) (420,000 → 424,000), on the same data and with the same optimizer state. The arm itself is not affected and keeps training.
|
| 4 |
|
| 5 |
- **Use:** compared with the flagship and with the decay probe of arm A (`probe_s12A_420k/`) on the held-out selection sets, it shows how the two arms would compare after decay at this point.
|
| 6 |
- **Status:** research checkpoint, **not a leaderboard submission**.
|