LakshyAAAgrawal commited on
Commit
07e5d5e
Β·
verified Β·
1 Parent(s): ac0ff86

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +12 -10
README.md CHANGED
@@ -4,11 +4,11 @@ language:
4
  license: apache-2.0
5
  base_model: Qwen/Qwen3-1.7B
6
  tags:
 
7
  - continuous-thought
8
  - latent-reasoning
9
  - distillation
10
  - gsm8k
11
- - codi
12
  datasets:
13
  - openai/gsm8k
14
  metrics:
@@ -39,13 +39,15 @@ model-index:
39
 
40
  ## Overview
41
 
42
- This model implements **Continuous Thought Distillation (CODI)** β€” an autoregressive latent
43
- reasoning loop that processes K=6 continuous thought steps before generating a text answer.
44
- Teacher hidden states are extracted from multiple chain-of-thought rollouts generated by the
45
- base model and distilled into the latent representations.
 
46
 
47
  **Key idea**: Instead of distilling from a single reasoning trace, we aggregate hidden states
48
- from multiple rollouts (16 per problem). The teacher signal is the hidden state from the single best rollout.
 
49
 
50
  ## Architecture
51
 
@@ -177,9 +179,9 @@ chain-of-thought text after the latent steps, requiring more tokens than standar
177
  | **This model** | **codi_single** | **True** | **1.0** | **128** | **80.4%** |
178
  | Qwen3-1.7B (base) | β€” | β€” | β€” | β€” | 77.3% |
179
  | Discrete SFT | sft | β€” | β€” | β€” | 80.7% |
180
- | CODI RW final-step | rw | no | 1.0 | 128 | 81.0% |
181
- | CODI Uniform per-step ans256 | uniform | yes | 2.0 | 256 | 83.2% |
182
- | CODI RW per-step ans256 | rw | yes | 1.0 | 256 | 82.7% |
183
 
184
  ## Citation
185
 
@@ -187,7 +189,7 @@ If you use this model, please cite:
187
 
188
  ```bibtex
189
  @misc{continuous-thought-2025,
190
- title={Continuous Thought Distillation: Reward-Weighted Multi-Trace Reasoning},
191
  author={Lakshya Agrawal},
192
  year={2025},
193
  url={https://huggingface.co/LakshyAAAgrawal/continuous-thought-r11_single_perstep_g1}
 
4
  license: apache-2.0
5
  base_model: Qwen/Qwen3-1.7B
6
  tags:
7
+ - qthink
8
  - continuous-thought
9
  - latent-reasoning
10
  - distillation
11
  - gsm8k
 
12
  datasets:
13
  - openai/gsm8k
14
  metrics:
 
39
 
40
  ## Overview
41
 
42
+ This model implements **QThink** (Parallel Latent Reasoning via Per-Step Distillation of
43
+ Multiple Rollouts) β€” an autoregressive latent reasoning loop that processes K=6 continuous
44
+ thought steps before generating a text answer. Teacher hidden states are extracted from
45
+ multiple chain-of-thought rollouts generated by the base model and distilled into the latent
46
+ representations at every step.
47
 
48
  **Key idea**: Instead of distilling from a single reasoning trace, we aggregate hidden states
49
+ from multiple rollouts (16 per problem) and supervise every latent step. The teacher signal
50
+ is the hidden state from the single best rollout.
51
 
52
  ## Architecture
53
 
 
179
  | **This model** | **codi_single** | **True** | **1.0** | **128** | **80.4%** |
180
  | Qwen3-1.7B (base) | β€” | β€” | β€” | β€” | 77.3% |
181
  | Discrete SFT | sft | β€” | β€” | β€” | 80.7% |
182
+ | QThink RW final-step | rw | no | 1.0 | 128 | 81.0% |
183
+ | QThink Uniform per-step ans256 | uniform | yes | 2.0 | 256 | 83.2% |
184
+ | QThink RW per-step ans256 | rw | yes | 1.0 | 256 | 82.7% |
185
 
186
  ## Citation
187
 
 
189
 
190
  ```bibtex
191
  @misc{continuous-thought-2025,
192
+ title={QThink: Parallel Latent Reasoning via Per-Step Distillation of Multiple Rollouts},
193
  author={Lakshya Agrawal},
194
  year={2025},
195
  url={https://huggingface.co/LakshyAAAgrawal/continuous-thought-r11_single_perstep_g1}