Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -4,11 +4,11 @@ language:
|
|
| 4 |
license: apache-2.0
|
| 5 |
base_model: Qwen/Qwen3-1.7B
|
| 6 |
tags:
|
|
|
|
| 7 |
- continuous-thought
|
| 8 |
- latent-reasoning
|
| 9 |
- distillation
|
| 10 |
- gsm8k
|
| 11 |
-
- codi
|
| 12 |
datasets:
|
| 13 |
- openai/gsm8k
|
| 14 |
metrics:
|
|
@@ -39,13 +39,15 @@ model-index:
|
|
| 39 |
|
| 40 |
## Overview
|
| 41 |
|
| 42 |
-
This model implements **
|
| 43 |
-
reasoning loop that processes K=6 continuous
|
| 44 |
-
|
| 45 |
-
base model and distilled into the latent
|
|
|
|
| 46 |
|
| 47 |
**Key idea**: Instead of distilling from a single reasoning trace, we aggregate hidden states
|
| 48 |
-
from multiple rollouts (16 per problem)
|
|
|
|
| 49 |
|
| 50 |
## Architecture
|
| 51 |
|
|
@@ -177,9 +179,9 @@ chain-of-thought text after the latent steps, requiring more tokens than standar
|
|
| 177 |
| **This model** | **codi_single** | **True** | **1.0** | **128** | **80.4%** |
|
| 178 |
| Qwen3-1.7B (base) | β | β | β | β | 77.3% |
|
| 179 |
| Discrete SFT | sft | β | β | β | 80.7% |
|
| 180 |
-
|
|
| 181 |
-
|
|
| 182 |
-
|
|
| 183 |
|
| 184 |
## Citation
|
| 185 |
|
|
@@ -187,7 +189,7 @@ If you use this model, please cite:
|
|
| 187 |
|
| 188 |
```bibtex
|
| 189 |
@misc{continuous-thought-2025,
|
| 190 |
-
title={
|
| 191 |
author={Lakshya Agrawal},
|
| 192 |
year={2025},
|
| 193 |
url={https://huggingface.co/LakshyAAAgrawal/continuous-thought-r11_single_perstep_g1}
|
|
|
|
| 4 |
license: apache-2.0
|
| 5 |
base_model: Qwen/Qwen3-1.7B
|
| 6 |
tags:
|
| 7 |
+
- qthink
|
| 8 |
- continuous-thought
|
| 9 |
- latent-reasoning
|
| 10 |
- distillation
|
| 11 |
- gsm8k
|
|
|
|
| 12 |
datasets:
|
| 13 |
- openai/gsm8k
|
| 14 |
metrics:
|
|
|
|
| 39 |
|
| 40 |
## Overview
|
| 41 |
|
| 42 |
+
This model implements **QThink** (Parallel Latent Reasoning via Per-Step Distillation of
|
| 43 |
+
Multiple Rollouts) β an autoregressive latent reasoning loop that processes K=6 continuous
|
| 44 |
+
thought steps before generating a text answer. Teacher hidden states are extracted from
|
| 45 |
+
multiple chain-of-thought rollouts generated by the base model and distilled into the latent
|
| 46 |
+
representations at every step.
|
| 47 |
|
| 48 |
**Key idea**: Instead of distilling from a single reasoning trace, we aggregate hidden states
|
| 49 |
+
from multiple rollouts (16 per problem) and supervise every latent step. The teacher signal
|
| 50 |
+
is the hidden state from the single best rollout.
|
| 51 |
|
| 52 |
## Architecture
|
| 53 |
|
|
|
|
| 179 |
| **This model** | **codi_single** | **True** | **1.0** | **128** | **80.4%** |
|
| 180 |
| Qwen3-1.7B (base) | β | β | β | β | 77.3% |
|
| 181 |
| Discrete SFT | sft | β | β | β | 80.7% |
|
| 182 |
+
| QThink RW final-step | rw | no | 1.0 | 128 | 81.0% |
|
| 183 |
+
| QThink Uniform per-step ans256 | uniform | yes | 2.0 | 256 | 83.2% |
|
| 184 |
+
| QThink RW per-step ans256 | rw | yes | 1.0 | 256 | 82.7% |
|
| 185 |
|
| 186 |
## Citation
|
| 187 |
|
|
|
|
| 189 |
|
| 190 |
```bibtex
|
| 191 |
@misc{continuous-thought-2025,
|
| 192 |
+
title={QThink: Parallel Latent Reasoning via Per-Step Distillation of Multiple Rollouts},
|
| 193 |
author={Lakshya Agrawal},
|
| 194 |
year={2025},
|
| 195 |
url={https://huggingface.co/LakshyAAAgrawal/continuous-thought-r11_single_perstep_g1}
|