LakshyAAAgrawal commited on
Commit
bd51128
Β·
verified Β·
1 Parent(s): 4305b15

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +12 -10
README.md CHANGED
@@ -4,11 +4,11 @@ language:
4
  license: apache-2.0
5
  base_model: Qwen/Qwen3-1.7B
6
  tags:
 
7
  - continuous-thought
8
  - latent-reasoning
9
  - distillation
10
  - gsm8k
11
- - codi
12
  datasets:
13
  - openai/gsm8k
14
  metrics:
@@ -40,13 +40,15 @@ model-index:
40
 
41
  ## Overview
42
 
43
- This model implements **Continuous Thought Distillation (CODI)** β€” an autoregressive latent
44
- reasoning loop that processes K=6 continuous thought steps before generating a text answer.
45
- Teacher hidden states are extracted from multiple chain-of-thought rollouts generated by the
46
- base model and distilled into the latent representations.
 
47
 
48
  **Key idea**: Instead of distilling from a single reasoning trace, we aggregate hidden states
49
- from multiple rollouts (16 per problem). The teacher signal is the uniform average of ALL rollouts (correct + incorrect).
 
50
 
51
  ## Architecture
52
 
@@ -178,9 +180,9 @@ chain-of-thought text after the latent steps, requiring more tokens than standar
178
  | **This model** | **codi_uniform** | **True** | **2.0** | **256** | **83.2%** |
179
  | Qwen3-1.7B (base) | β€” | β€” | β€” | β€” | 77.3% |
180
  | Discrete SFT | sft | β€” | β€” | β€” | 80.7% |
181
- | CODI RW final-step | rw | no | 1.0 | 128 | 81.0% |
182
- | CODI Uniform per-step ans256 | uniform | yes | 2.0 | 256 | 83.2% |
183
- | CODI RW per-step ans256 | rw | yes | 1.0 | 256 | 82.7% |
184
 
185
  ## Citation
186
 
@@ -188,7 +190,7 @@ If you use this model, please cite:
188
 
189
  ```bibtex
190
  @misc{continuous-thought-2025,
191
- title={Continuous Thought Distillation: Reward-Weighted Multi-Trace Reasoning},
192
  author={Lakshya Agrawal},
193
  year={2025},
194
  url={https://huggingface.co/LakshyAAAgrawal/continuous-thought-r11_uniform_perstep_g2_ans256}
 
4
  license: apache-2.0
5
  base_model: Qwen/Qwen3-1.7B
6
  tags:
7
+ - qthink
8
  - continuous-thought
9
  - latent-reasoning
10
  - distillation
11
  - gsm8k
 
12
  datasets:
13
  - openai/gsm8k
14
  metrics:
 
40
 
41
  ## Overview
42
 
43
+ This model implements **QThink** (Parallel Latent Reasoning via Per-Step Distillation of
44
+ Multiple Rollouts) β€” an autoregressive latent reasoning loop that processes K=6 continuous
45
+ thought steps before generating a text answer. Teacher hidden states are extracted from
46
+ multiple chain-of-thought rollouts generated by the base model and distilled into the latent
47
+ representations at every step.
48
 
49
  **Key idea**: Instead of distilling from a single reasoning trace, we aggregate hidden states
50
+ from multiple rollouts (16 per problem) and supervise every latent step. The teacher signal
51
+ is the uniform average of ALL rollouts (correct + incorrect).
52
 
53
  ## Architecture
54
 
 
180
  | **This model** | **codi_uniform** | **True** | **2.0** | **256** | **83.2%** |
181
  | Qwen3-1.7B (base) | β€” | β€” | β€” | β€” | 77.3% |
182
  | Discrete SFT | sft | β€” | β€” | β€” | 80.7% |
183
+ | QThink RW final-step | rw | no | 1.0 | 128 | 81.0% |
184
+ | QThink Uniform per-step ans256 | uniform | yes | 2.0 | 256 | 83.2% |
185
+ | QThink RW per-step ans256 | rw | yes | 1.0 | 256 | 82.7% |
186
 
187
  ## Citation
188
 
 
190
 
191
  ```bibtex
192
  @misc{continuous-thought-2025,
193
+ title={QThink: Parallel Latent Reasoning via Per-Step Distillation of Multiple Rollouts},
194
  author={Lakshya Agrawal},
195
  year={2025},
196
  url={https://huggingface.co/LakshyAAAgrawal/continuous-thought-r11_uniform_perstep_g2_ans256}