jihwan1205 commited on
Commit
0efa1f0
·
verified ·
1 Parent(s): c6c2170

Upload folder using huggingface_hub

Browse files
README.md ADDED
@@ -0,0 +1,95 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model:
3
+ - ModalityDance/latent-tts-coconut
4
+ license: mit
5
+ pipeline_tag: text-generation
6
+ library_name: transformers
7
+ tags:
8
+ - latent-reasoning
9
+ - continuous-thought
10
+ - coconut
11
+ - grpo
12
+ - reinforcement-learning
13
+ datasets:
14
+ - gsm8k
15
+ ---
16
+
17
+ # SVP-V-GRPO · COCONUT GPT-2
18
+
19
+ A [COCONUT](https://huggingface.co/ModalityDance/latent-tts-coconut) GPT-2 (124M)
20
+ latent-reasoning model post-trained with **SVP-V-GRPO** — reinforcement learning
21
+ whose exploration comes from *weight-space* perturbation of the attention value
22
+ projections rather than from token sampling.
23
+
24
+ Only the value-projection weights (`W_V`, the V columns of each `c_attn`) differ
25
+ from the base checkpoint; everything else is untouched. Inference is ordinary
26
+ greedy decoding — the perturbation is a training-time exploration mechanism and
27
+ is **not** used at deployment.
28
+
29
+ Part of the **SVP-V** family: `svp-v-coconut-gpt2` (this model) and
30
+ `svp-v-coconut-llama1b` (LLaMA-3.2-1B, in training).
31
+
32
+ ## Results
33
+
34
+ Clean greedy decoding, `max_new_tokens=64` (six latent steps + marker + answer),
35
+ canonical answer extraction. All rows measured in one harness.
36
+
37
+ | Model (GPT-2 124M) | GSM8K | GSM-Hard | SVAMP | ASDiv-A | MultiArith | GSM-Plus |
38
+ |---|---|---|---|---|---|---|
39
+ | COCONUT (base) | 34.1 | 7.7 | 35.6 | 60.2 | 80.9 | 17.4 |
40
+ | SLPO-COCONUT | 34.9 | 7.6 | 34.3 | 58.7 | 82.8 | 18.2 |
41
+ | SIM-CoT (COCONUT) | 44.7 | 9.3 | 40.6 | 67.2 | 90.5 | 21.5 |
42
+ | SIM-CoT (CODI) | 39.3 | 8.7 | 40.1 | 64.3 | 91.0 | 22.6 |
43
+ | CODI | 42.5 | 9.3 | 40.0 | 65.4 | 91.9 | 23.1 |
44
+ | SLPO-CODI | 42.9 | 9.5 | 43.0 | 67.5 | 90.5 | 24.1 |
45
+ | **This model** | **50.5** | **11.2** | **45.2** | **72.5** | **94.1** | **28.3** |
46
+
47
+ GSM8K coverage under the 16-sample deployment ensemble rises alongside single-shot
48
+ accuracy (63.2 → 70.7 over training), i.e. RL here does not collapse the sampling
49
+ diversity it was trained under.
50
+
51
+ ## Usage
52
+
53
+ This is a fixed-length continuous-thought model: the prompt ends with
54
+ `<|start-latent|>`, six latent steps feed each step's last hidden state back as
55
+ the next input embedding, `<|end-latent|>` closes the latent phase, and the answer
56
+ is then decoded as ordinary tokens. A plain `generate()` call will **not**
57
+ reproduce the numbers above — the two-phase loop is required.
58
+
59
+ ```python
60
+ from transformers import AutoTokenizer, GPT2LMHeadModel
61
+
62
+ REPO = "jihwan1205/svp-v-coconut-gpt2"
63
+ tok = AutoTokenizer.from_pretrained(REPO)
64
+ model = GPT2LMHeadModel.from_pretrained(REPO).eval().cuda()
65
+ # then run the two-phase latent loop (6 latent steps, then decode)
66
+ ```
67
+
68
+ Reference implementation of the loop: `generate_batch` in
69
+ `src/generation/engine.py` of the SVP repository, or the original COCONUT
70
+ inference code.
71
+
72
+ ## Training
73
+
74
+ | | |
75
+ |---|---|
76
+ | Base | `ModalityDance/latent-tts-coconut` (COCONUT GPT-2, 124M) |
77
+ | Data | GSM8K-Aug training stream (385k), 9 epochs |
78
+ | Algorithm | GRPO — G=32 rollouts/prompt, B=8 prompts/update, μ=2 inner epochs, DAPO mixed-outcome filter, Dr.GRPO advantage, k3 KL (β=0.02) to the frozen base |
79
+ | Exploration | SVP: per-rollout multiplicative Gaussian jitter on the singular values of `W_V` (σᵢ → σᵢ(1+αgᵢ), α=0.6), re-anchored every 50 iterations |
80
+ | Trained parameters | V columns of every `c_attn` (≈ 7M of 124M) |
81
+ | Optimizer | AdamW, lr 3e-5 constant, grad-clip 1.0 |
82
+ | Reward | Final-answer correctness only |
83
+
84
+ ## Limitations
85
+
86
+ - English grade-school arithmetic word problems only; the four transfer benchmarks
87
+ above (SVAMP, ASDiv-A, MultiArith, GSM-Plus) are the extent of tested generalization.
88
+ - GSM-Hard remains low (11.2) — large-magnitude arithmetic is a limit of the
89
+ 124M backbone, not something RL fixed.
90
+ - Requires the COCONUT two-phase inference loop; it is not a drop-in chat model.
91
+
92
+ ## License
93
+
94
+ MIT, inherited from the base checkpoint (which derives from
95
+ `openai-community/gpt2`).
added_tokens.json ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {
2
+ "<|end-latent|>": 50258,
3
+ "<|latent|>": 50259,
4
+ "<|start-latent|>": 50257
5
+ }
config.json ADDED
@@ -0,0 +1,45 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "activation_function": "gelu_new",
3
+ "add_cross_attention": false,
4
+ "architectures": [
5
+ "GPT2LMHeadModel"
6
+ ],
7
+ "attn_pdrop": 0.1,
8
+ "bos_token_id": 50256,
9
+ "dtype": "float32",
10
+ "embd_pdrop": 0.1,
11
+ "eos_token_id": 50256,
12
+ "initializer_range": 0.02,
13
+ "latent_end_id": -100,
14
+ "latent_id": -100,
15
+ "latent_start_id": -100,
16
+ "layer_norm_epsilon": 1e-05,
17
+ "model_type": "gpt2",
18
+ "n_ctx": 1024,
19
+ "n_embd": 768,
20
+ "n_head": 12,
21
+ "n_inner": null,
22
+ "n_layer": 12,
23
+ "n_positions": 1024,
24
+ "pad_token_id": 50256,
25
+ "reorder_and_upcast_attn": false,
26
+ "resid_pdrop": 0.1,
27
+ "scale_attn_by_inverse_layer_idx": false,
28
+ "scale_attn_weights": true,
29
+ "summary_activation": null,
30
+ "summary_first_dropout": 0.1,
31
+ "summary_proj_to_labels": true,
32
+ "summary_type": "cls_index",
33
+ "summary_use_proj": true,
34
+ "target_id": -100,
35
+ "task_specific_params": {
36
+ "text-generation": {
37
+ "do_sample": true,
38
+ "max_length": 50
39
+ }
40
+ },
41
+ "tie_word_embeddings": true,
42
+ "transformers_version": "5.13.0",
43
+ "use_cache": true,
44
+ "vocab_size": 50260
45
+ }
generation_config.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 50256,
4
+ "eos_token_id": 50256,
5
+ "pad_token_id": 50256,
6
+ "transformers_version": "5.13.0"
7
+ }
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e8e113a71f6ba0d83d9c407964d159e1c72518c4d56935d779fc5d741a69304b
3
+ size 497783424
special_tokens_map.json ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<|endoftext|>",
4
+ "lstrip": false,
5
+ "normalized": true,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "eos_token": {
10
+ "content": "<|endoftext|>",
11
+ "lstrip": false,
12
+ "normalized": true,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "pad_token": {
17
+ "content": "<|endoftext|>",
18
+ "lstrip": false,
19
+ "normalized": true,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ },
23
+ "unk_token": {
24
+ "content": "<|endoftext|>",
25
+ "lstrip": false,
26
+ "normalized": true,
27
+ "rstrip": false,
28
+ "single_word": false
29
+ }
30
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": "<|endoftext|>",
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "is_local": true,
9
+ "local_files_only": false,
10
+ "model_max_length": 1024,
11
+ "pad_token": "<|endoftext|>",
12
+ "tokenizer_class": "GPT2Tokenizer",
13
+ "unk_token": "<|endoftext|>"
14
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff