File size: 9,591 Bytes
7a578f3
679b37d
7a578f3
679b37d
7a578f3
679b37d
7a578f3
679b37d
 
 
 
 
3ccf008
 
 
e97c886
 
679b37d
 
e97c886
 
7a578f3
 
679b37d
7a578f3
679b37d
7a578f3
679b37d
7a578f3
679b37d
7a578f3
679b37d
7a578f3
679b37d
7a578f3
679b37d
7a578f3
679b37d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7a578f3
679b37d
7a578f3
679b37d
 
 
 
7a578f3
679b37d
 
 
 
 
 
7a578f3
679b37d
7a578f3
679b37d
7a578f3
679b37d
7a578f3
679b37d
 
7a578f3
679b37d
 
 
 
 
 
 
 
7a578f3
679b37d
 
 
7a578f3
679b37d
 
7a578f3
679b37d
 
 
 
 
 
 
 
7a578f3
679b37d
 
7a578f3
679b37d
7a578f3
679b37d
 
 
 
7a578f3
679b37d
7a578f3
679b37d
 
 
 
 
6460d08
679b37d
6460d08
679b37d
 
 
 
 
 
 
6460d08
679b37d
6460d08
679b37d
6460d08
91265b1
 
 
 
 
 
 
679b37d
 
6460d08
679b37d
6460d08
679b37d
6460d08
679b37d
6460d08
679b37d
6460d08
679b37d
6460d08
be3ed35
 
 
390bc5f
be3ed35
 
 
 
 
 
 
 
 
 
 
 
 
679b37d
b8400ec
679b37d
390bc5f
679b37d
390bc5f
679b37d
 
 
 
 
b8400ec
 
 
91265b1
41160dd
 
 
 
 
 
 
 
 
 
91265b1
 
 
 
 
 
 
 
41160dd
 
 
 
 
 
 
 
 
 
2c3524d
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
---
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:
- qwen3
- sft
- trl
- knowledge-distillation
- thinking
- longwriter
- convergent-intelligence
- convergentintel
- edge
- distillation
base_model:
- reaperdoesntknow/Disctil-Qwen3-1.7B
datasets:
- longwriter-6k
- 0xZee/dataset-CoT-Differential-Equations-636
- 0xZee/dataset-CoT-Linear-Algebra-667
---

# Qwen3-1.7B-Thinking-Distil

**Extended Reasoning Distillation from Qwen3-30B-A3B-Thinking β†’ 1.7B**

*Convergent Intelligence LLC: Research Division*

---

## What This Is

The most downloaded model in the Convergent Intelligence portfolio. Qwen3-1.7B-Thinking-Distil captures extended deliberation patterns from the Qwen3-30B-A3B **Thinking** teacher β€” the variant that generates long-form reasoning chains before committing to an answer β€” and compresses them into a 1.7B student via supervised fine-tuning on the [longwriter-6k](https://huggingface.co/datasets/longwriter-6k) dataset.

The Thinking teacher produces the **richest signal** of the three teacher variants in the DistilQwen family (Instruct, Thinking, Coder). Where Instruct distillation captures clean instruction-following and Coder captures hierarchical decomposition, Thinking distillation captures the extended internal monologue β€” the model reasoning through uncertainty, backtracking, and re-evaluating before arriving at a conclusion. That deliberative depth is what makes this variant the highest-download model in the collection.

## Architecture

| Parameter | Value |
|-----------|-------|
| Architecture | Qwen3ForCausalLM |
| Parameters | ~2.03B (1.7B effective) |
| Hidden Size | 2048 |
| Layers | 28 |
| Attention Heads | 16 (Q) / 8 (KV) β€” GQA |
| Intermediate | 6144 |
| Head Dimension | 128 |
| Context Length | 40,960 tokens (max position) |
| Vocabulary | 151,936 |
| Precision | BF16 |
| Activation | SiLU |

## Training

**Teacher:** Qwen3-30B-A3B-Thinking
**Student:** Qwen3-1.7B
**Dataset:** longwriter-6k β€” long-form generation samples that preserve extended reasoning chains
**Method:** Supervised Fine-Tuning (SFT) via TRL

| Parameter | Value |
|-----------|-------|
| Max Sequence Length | 4,096 |
| Precision | BF16 |
| Framework | TRL (SFTTrainer) |
| Hardware | NVIDIA H100 |

The training captures the teacher's extended thinking traces through direct SFT rather than logit-level KD. This is a deliberate design choice β€” the longwriter-6k dataset provides naturally long reasoning samples where the signal is in the structure of the generation (how the teacher approaches, reconsiders, and resolves), not just the final token probabilities.

For the full topology-aware distillation pipeline (BV decomposition, jump detection, curriculum ordering), see [TopologicalQwen](https://huggingface.co/reaperdoesntknow/TopologicalQwen). This model is the SFT-direct variant β€” simpler, faster to train, and empirically the most downloaded for a reason: the Thinking teacher's extended chains transfer well through pure SFT.

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "reaperdoesntknow/Qwen3-1.7B-Thinking-Distil",
    torch_dtype="auto",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
    "reaperdoesntknow/Qwen3-1.7B-Thinking-Distil"
)

messages = [
    {"role": "user", "content": "Explain why gradient descent can get stuck in saddle points but not local minima in high dimensions."}
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

output = model.generate(
    **inputs,
    max_new_tokens=2048,
    do_sample=True,
    top_p=0.9,
    temperature=0.7,
    repetition_penalty=1.15
)

print(tokenizer.decode(output[0], skip_special_tokens=True))
```

### Generation Tips

- **Temperature 0.6–0.8** works best for reasoning tasks β€” low enough for coherence, high enough to activate the extended deliberation patterns from the Thinking teacher.
- **Repetition penalty 1.1–1.2** prevents the model from getting caught in reasoning loops during long generations.
- **Max tokens 1024–2048** β€” the model was trained on 4096 max seq, so it can generate long. Give it room.
- The model inherits the Thinking teacher's tendency to reason before answering. Let it.

## Distillation Position

```
Qwen3-30B-A3B-Thinking (teacher)
  ↓ SFT on longwriter-6k (4096 max seq)
Qwen3-1.7B-Thinking-Distil ← you are here
```

This model is the **direct SFT** path. The DistilQwen collection also includes models that go through additional refinement stages:

```
Qwen3-1.7B (base)
  β†’ Qwen3-1.7B-Distilled-30B-A3B (Instruct teacher KD)
    β†’ DiStil (uncensored SFT)
      β†’ Disctil (DISC refinement)
        β†’ TopologicalQwen (full TKD pipeline)
```

Different paths, different capabilities. This model prioritizes extended reasoning. TopologicalQwen prioritizes structural precision. The Coder variant prioritizes hierarchical decomposition. They're complementary.

## DistilQwen Collection

| Model | What It Does |
|-------|-------------|
| **[Qwen3-1.7B-Thinking-Distil](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Thinking-Distil)** | **← this model. Thinking teacher SFT.** |
| [TopologicalQwen](https://huggingface.co/reaperdoesntknow/TopologicalQwen) | Full TKD pipeline. BV decomposition + DualMind format. |
| [DiStil-Qwen3-1.7B-uncensored](https://huggingface.co/reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored) | DISC-informed uncensored distillation. |
| [Qwen3-1.7B-Coder-Distilled-SFT](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT) | Coder teacher. Hierarchical problem solving. |
| [DistilQwen3-1.7B-uncensored](https://huggingface.co/reaperdoesntknow/DistilQwen3-1.7B-uncensored) | Base uncensored variant. |

Full collection: [DistilQwen on HuggingFace](https://huggingface.co/collections/reaperdoesntknow/distilqwen-69bf40ec669117e3f069ef1c)

## Methodology

Full methodology paper: **[Structure Over Scale: Proof-Weighted Knowledge Distillation](https://doi.org/10.57967/hf/8165)** (DOI: 10.57967/hf/8165)

Companion paper: **[Three Teachers to Dual Cognition](https://doi.org/10.57967/hf/8184)** (DOI: 10.57967/hf/8184) β€” covers the DualMind extension and ghost imprinting phenomenon.

## License

Apache 2.0 β€” same as the base Qwen3 model.


## Mathematical Foundations: Discrepancy Calculus (DISC)

This model's training pipeline is grounded in Discrepancy Calculus β€” a measure-theoretic framework that treats singularities as primary structure rather than pathology. Full theory: *"On the Formal Analysis of Discrepancy Calculus"* (CIx, 2026; Convergent Intelligence LLC: Research Division).

**The Core Operator:**

$$Df(x) = \lim_{\varepsilon \downarrow 0} \frac{1}{\varepsilon} \int_x^{x+\varepsilon} \frac{|f(t) - f(x)|}{|t - x|}\, dt$$

For smooth $f$: $Df(x) = |f'(x)|$. For rough $f$: $D$ localizes irregularity to null sets while preserving integral structure.

**The Mesh Fundamental Identity** β€” every BV function decomposes as:

$$f(b) - f(a) = \underbrace{\int_a^b f'(x)\,dx}_{\text{smooth (AC)}} + \underbrace{\sum_{x \in J_f} \Delta f(x)}_{\text{jumps}} + \underbrace{D^c f(I)}_{\text{Cantor drift}}$$

Standard knowledge distillation captures only term 1. Topological Knowledge Distillation (TKD) preserves all three by treating the teacher's output distribution as a BV function and computing discrepancy energy, jump sets, and gap energy density before training begins.

## Citation

```bibtex
@misc{cix2026distilqwen,
  title={Structure Over Scale: Proof-Weighted Knowledge Distillation from Qwen3-30B to 1.7B},
  author={Convergent Intelligence},
  year={2026},
  doi={10.57967/hf/8165},
  publisher={Convergent Intelligence LLC: Research Division}
}
```

---

*Convergent Intelligence LLC: Research Division.*
*[Full portfolio](https://huggingface.co/reaperdoesntknow) | [DistilQwen Collection](https://huggingface.co/collections/reaperdoesntknow/distilqwen-69bf40ec669117e3f069ef1c) | [DualMind Collection](https://huggingface.co/collections/reaperdoesntknow/dualmind-69c93f888c6e79ecc69cf41e)*

---

## Convergent Intelligence Portfolio

*Part of the [DistilQwen Series](https://huggingface.co/collections/reaperdoesntknow/distilqwen-69bf40ec669117e3f069ef1c) by [Convergent Intelligence LLC: Research Division](https://huggingface.co/reaperdoesntknow)*

### Related Models

| Model | Format |
|-------|--------|
| [TopologicalQwen](https://huggingface.co/reaperdoesntknow/TopologicalQwen) | BF16 |
| [Qwen3-1.7B-Thinking-Distil](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Thinking-Distil) | BF16 |
| [Qwen3-1.7B-Coder-Distilled-SFT](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT) | BF16 |
| [DiStil-Qwen3-1.7B-uncensored](https://huggingface.co/reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored) | BF16 |
| [DistilQwen3-1.7B-uncensored](https://huggingface.co/reaperdoesntknow/DistilQwen3-1.7B-uncensored) | BF16 |
| [Qwen3-1.7B-Distilled-30B-A3B](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B) | BF16 |

### Papers

| Paper | DOI |
|-------|-----|
| [Structure Over Scale](https://huggingface.co/reaperdoesntknow/Structure-Over-Scale) | 10.57967/hf/8165 |
| [Three Teachers to Dual Cognition](https://huggingface.co/reaperdoesntknow/DualMind_Methodolgy) | 10.57967/hf/8184 |
| [Discrepancy Calculus](https://huggingface.co/reaperdoesntknow/Discrepancy_Calculus) | 10.57967/hf/8194 |

---
<!-- cix-keeper-ts:2026-09-29T13:16:36Z -->