jihwan1205 commited on
Commit
a996045
·
verified ·
1 Parent(s): 0efa1f0

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -11,7 +11,7 @@ tags:
11
  - grpo
12
  - reinforcement-learning
13
  datasets:
14
- - gsm8k
15
  ---
16
 
17
  # SVP-V-GRPO · COCONUT GPT-2
@@ -74,7 +74,7 @@ inference code.
74
  | | |
75
  |---|---|
76
  | Base | `ModalityDance/latent-tts-coconut` (COCONUT GPT-2, 124M) |
77
- | Data | GSM8K-Aug training stream (385k), 9 epochs |
78
  | Algorithm | GRPO — G=32 rollouts/prompt, B=8 prompts/update, μ=2 inner epochs, DAPO mixed-outcome filter, Dr.GRPO advantage, k3 KL (β=0.02) to the frozen base |
79
  | Exploration | SVP: per-rollout multiplicative Gaussian jitter on the singular values of `W_V` (σᵢ → σᵢ(1+αgᵢ), α=0.6), re-anchored every 50 iterations |
80
  | Trained parameters | V columns of every `c_attn` (≈ 7M of 124M) |
 
11
  - grpo
12
  - reinforcement-learning
13
  datasets:
14
+ - zen-E/GSM8k-Aug
15
  ---
16
 
17
  # SVP-V-GRPO · COCONUT GPT-2
 
74
  | | |
75
  |---|---|
76
  | Base | `ModalityDance/latent-tts-coconut` (COCONUT GPT-2, 124M) |
77
+ | Data | [GSM8K-Aug](https://huggingface.co/datasets/zen-E/GSM8k-Aug) training split (385,620 problems), 9 epochs |
78
  | Algorithm | GRPO — G=32 rollouts/prompt, B=8 prompts/update, μ=2 inner epochs, DAPO mixed-outcome filter, Dr.GRPO advantage, k3 KL (β=0.02) to the frozen base |
79
  | Exploration | SVP: per-rollout multiplicative Gaussian jitter on the singular values of `W_V` (σᵢ → σᵢ(1+αgᵢ), α=0.6), re-anchored every 50 iterations |
80
  | Trained parameters | V columns of every `c_attn` (≈ 7M of 124M) |