Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -4,11 +4,11 @@ language:
|
|
| 4 |
license: apache-2.0
|
| 5 |
base_model: Qwen/Qwen3-1.7B
|
| 6 |
tags:
|
|
|
|
| 7 |
- continuous-thought
|
| 8 |
- latent-reasoning
|
| 9 |
- distillation
|
| 10 |
- gsm8k
|
| 11 |
-
- codi
|
| 12 |
datasets:
|
| 13 |
- openai/gsm8k
|
| 14 |
metrics:
|
|
@@ -40,13 +40,15 @@ model-index:
|
|
| 40 |
|
| 41 |
## Overview
|
| 42 |
|
| 43 |
-
This model implements **
|
| 44 |
-
reasoning loop that processes K=6 continuous
|
| 45 |
-
|
| 46 |
-
base model and distilled into the latent
|
|
|
|
| 47 |
|
| 48 |
**Key idea**: Instead of distilling from a single reasoning trace, we aggregate hidden states
|
| 49 |
-
from multiple rollouts (16 per problem)
|
|
|
|
| 50 |
|
| 51 |
## Architecture
|
| 52 |
|
|
@@ -178,9 +180,9 @@ chain-of-thought text after the latent steps, requiring more tokens than standar
|
|
| 178 |
| **This model** | **codi_uniform** | **True** | **2.0** | **256** | **83.2%** |
|
| 179 |
| Qwen3-1.7B (base) | β | β | β | β | 77.3% |
|
| 180 |
| Discrete SFT | sft | β | β | β | 80.7% |
|
| 181 |
-
|
|
| 182 |
-
|
|
| 183 |
-
|
|
| 184 |
|
| 185 |
## Citation
|
| 186 |
|
|
@@ -188,7 +190,7 @@ If you use this model, please cite:
|
|
| 188 |
|
| 189 |
```bibtex
|
| 190 |
@misc{continuous-thought-2025,
|
| 191 |
-
title={
|
| 192 |
author={Lakshya Agrawal},
|
| 193 |
year={2025},
|
| 194 |
url={https://huggingface.co/LakshyAAAgrawal/continuous-thought-r11_uniform_perstep_g2_ans256}
|
|
|
|
| 4 |
license: apache-2.0
|
| 5 |
base_model: Qwen/Qwen3-1.7B
|
| 6 |
tags:
|
| 7 |
+
- qthink
|
| 8 |
- continuous-thought
|
| 9 |
- latent-reasoning
|
| 10 |
- distillation
|
| 11 |
- gsm8k
|
|
|
|
| 12 |
datasets:
|
| 13 |
- openai/gsm8k
|
| 14 |
metrics:
|
|
|
|
| 40 |
|
| 41 |
## Overview
|
| 42 |
|
| 43 |
+
This model implements **QThink** (Parallel Latent Reasoning via Per-Step Distillation of
|
| 44 |
+
Multiple Rollouts) β an autoregressive latent reasoning loop that processes K=6 continuous
|
| 45 |
+
thought steps before generating a text answer. Teacher hidden states are extracted from
|
| 46 |
+
multiple chain-of-thought rollouts generated by the base model and distilled into the latent
|
| 47 |
+
representations at every step.
|
| 48 |
|
| 49 |
**Key idea**: Instead of distilling from a single reasoning trace, we aggregate hidden states
|
| 50 |
+
from multiple rollouts (16 per problem) and supervise every latent step. The teacher signal
|
| 51 |
+
is the uniform average of ALL rollouts (correct + incorrect).
|
| 52 |
|
| 53 |
## Architecture
|
| 54 |
|
|
|
|
| 180 |
| **This model** | **codi_uniform** | **True** | **2.0** | **256** | **83.2%** |
|
| 181 |
| Qwen3-1.7B (base) | β | β | β | β | 77.3% |
|
| 182 |
| Discrete SFT | sft | β | β | β | 80.7% |
|
| 183 |
+
| QThink RW final-step | rw | no | 1.0 | 128 | 81.0% |
|
| 184 |
+
| QThink Uniform per-step ans256 | uniform | yes | 2.0 | 256 | 83.2% |
|
| 185 |
+
| QThink RW per-step ans256 | rw | yes | 1.0 | 256 | 82.7% |
|
| 186 |
|
| 187 |
## Citation
|
| 188 |
|
|
|
|
| 190 |
|
| 191 |
```bibtex
|
| 192 |
@misc{continuous-thought-2025,
|
| 193 |
+
title={QThink: Parallel Latent Reasoning via Per-Step Distillation of Multiple Rollouts},
|
| 194 |
author={Lakshya Agrawal},
|
| 195 |
year={2025},
|
| 196 |
url={https://huggingface.co/LakshyAAAgrawal/continuous-thought-r11_uniform_perstep_g2_ans256}
|