jihwan1205 commited on
Commit
df7cdca
·
verified ·
1 Parent(s): c94e146

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +2 -4
README.md CHANGED
@@ -20,15 +20,13 @@ A [COCONUT](https://huggingface.co/ModalityDance/latent-tts-coconut) GPT-2 (124M
20
  latent-reasoning model post-trained with **SVP-V-GRPO** — reinforcement learning whose
21
  rollout exploration comes from *weight-space* perturbation of the attention value
22
  projections. For each layer and rollout, SVP independently samples Gaussian noise and
23
- rescales coefficients in the fixed SVD basis of `W_V` as `sᵢ → sᵢ(1+αgᵢ)`. The resulting
24
- coefficients can be negative and are not necessarily singular values of the perturbed
25
- matrix.
26
 
27
  Only the value-projection weights (the V columns of each `c_attn`) differ from the base
28
  checkpoint. The perturbation is a training-time exploration mechanism and is **not** used
29
  at deployment: inference is ordinary greedy decoding.
30
 
31
- This repository contains the paper's GPT-2 SVP-V-GRPO checkpoint from epoch 9 `B=32`, `lr=6e-5`, seed-0 run.
32
 
33
  ## Results
34
 
 
20
  latent-reasoning model post-trained with **SVP-V-GRPO** — reinforcement learning whose
21
  rollout exploration comes from *weight-space* perturbation of the attention value
22
  projections. For each layer and rollout, SVP independently samples Gaussian noise and
23
+ rescales coefficients in the fixed SVD basis of `W_V` as `sᵢ → sᵢ(1+αgᵢ)`.
 
 
24
 
25
  Only the value-projection weights (the V columns of each `c_attn`) differ from the base
26
  checkpoint. The perturbation is a training-time exploration mechanism and is **not** used
27
  at deployment: inference is ordinary greedy decoding.
28
 
29
+ This repository contains the paper's GPT-2 SVP-V-GRPO checkpoint from epoch 9 `B=32`, `lr=6e-5` run.
30
 
31
  ## Results
32