Text Generation
Transformers
Safetensors
qwen3
sft
trl
knowledge-distillation
thinking
longwriter
convergent-intelligence
convergentintel
edge
distillation
conversational
text-generation-inference
Instructions to use reaperdoesntknow/Qwen3-1.7B-Thinking-Distil with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use reaperdoesntknow/Qwen3-1.7B-Thinking-Distil with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="reaperdoesntknow/Qwen3-1.7B-Thinking-Distil") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("reaperdoesntknow/Qwen3-1.7B-Thinking-Distil") model = AutoModelForCausalLM.from_pretrained("reaperdoesntknow/Qwen3-1.7B-Thinking-Distil", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use reaperdoesntknow/Qwen3-1.7B-Thinking-Distil with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "reaperdoesntknow/Qwen3-1.7B-Thinking-Distil" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "reaperdoesntknow/Qwen3-1.7B-Thinking-Distil", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/reaperdoesntknow/Qwen3-1.7B-Thinking-Distil
- SGLang
How to use reaperdoesntknow/Qwen3-1.7B-Thinking-Distil with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "reaperdoesntknow/Qwen3-1.7B-Thinking-Distil" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "reaperdoesntknow/Qwen3-1.7B-Thinking-Distil", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "reaperdoesntknow/Qwen3-1.7B-Thinking-Distil" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "reaperdoesntknow/Qwen3-1.7B-Thinking-Distil", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use reaperdoesntknow/Qwen3-1.7B-Thinking-Distil with Docker Model Runner:
docker model run hf.co/reaperdoesntknow/Qwen3-1.7B-Thinking-Distil
File size: 9,591 Bytes
7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 3ccf008 e97c886 679b37d e97c886 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 7a578f3 679b37d 6460d08 679b37d 6460d08 679b37d 6460d08 679b37d 6460d08 679b37d 6460d08 91265b1 679b37d 6460d08 679b37d 6460d08 679b37d 6460d08 679b37d 6460d08 679b37d 6460d08 679b37d 6460d08 be3ed35 390bc5f be3ed35 679b37d b8400ec 679b37d 390bc5f 679b37d 390bc5f 679b37d b8400ec 91265b1 41160dd 91265b1 41160dd 2c3524d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 | ---
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:
- qwen3
- sft
- trl
- knowledge-distillation
- thinking
- longwriter
- convergent-intelligence
- convergentintel
- edge
- distillation
base_model:
- reaperdoesntknow/Disctil-Qwen3-1.7B
datasets:
- longwriter-6k
- 0xZee/dataset-CoT-Differential-Equations-636
- 0xZee/dataset-CoT-Linear-Algebra-667
---
# Qwen3-1.7B-Thinking-Distil
**Extended Reasoning Distillation from Qwen3-30B-A3B-Thinking β 1.7B**
*Convergent Intelligence LLC: Research Division*
---
## What This Is
The most downloaded model in the Convergent Intelligence portfolio. Qwen3-1.7B-Thinking-Distil captures extended deliberation patterns from the Qwen3-30B-A3B **Thinking** teacher β the variant that generates long-form reasoning chains before committing to an answer β and compresses them into a 1.7B student via supervised fine-tuning on the [longwriter-6k](https://huggingface.co/datasets/longwriter-6k) dataset.
The Thinking teacher produces the **richest signal** of the three teacher variants in the DistilQwen family (Instruct, Thinking, Coder). Where Instruct distillation captures clean instruction-following and Coder captures hierarchical decomposition, Thinking distillation captures the extended internal monologue β the model reasoning through uncertainty, backtracking, and re-evaluating before arriving at a conclusion. That deliberative depth is what makes this variant the highest-download model in the collection.
## Architecture
| Parameter | Value |
|-----------|-------|
| Architecture | Qwen3ForCausalLM |
| Parameters | ~2.03B (1.7B effective) |
| Hidden Size | 2048 |
| Layers | 28 |
| Attention Heads | 16 (Q) / 8 (KV) β GQA |
| Intermediate | 6144 |
| Head Dimension | 128 |
| Context Length | 40,960 tokens (max position) |
| Vocabulary | 151,936 |
| Precision | BF16 |
| Activation | SiLU |
## Training
**Teacher:** Qwen3-30B-A3B-Thinking
**Student:** Qwen3-1.7B
**Dataset:** longwriter-6k β long-form generation samples that preserve extended reasoning chains
**Method:** Supervised Fine-Tuning (SFT) via TRL
| Parameter | Value |
|-----------|-------|
| Max Sequence Length | 4,096 |
| Precision | BF16 |
| Framework | TRL (SFTTrainer) |
| Hardware | NVIDIA H100 |
The training captures the teacher's extended thinking traces through direct SFT rather than logit-level KD. This is a deliberate design choice β the longwriter-6k dataset provides naturally long reasoning samples where the signal is in the structure of the generation (how the teacher approaches, reconsiders, and resolves), not just the final token probabilities.
For the full topology-aware distillation pipeline (BV decomposition, jump detection, curriculum ordering), see [TopologicalQwen](https://huggingface.co/reaperdoesntknow/TopologicalQwen). This model is the SFT-direct variant β simpler, faster to train, and empirically the most downloaded for a reason: the Thinking teacher's extended chains transfer well through pure SFT.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"reaperdoesntknow/Qwen3-1.7B-Thinking-Distil",
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
"reaperdoesntknow/Qwen3-1.7B-Thinking-Distil"
)
messages = [
{"role": "user", "content": "Explain why gradient descent can get stuck in saddle points but not local minima in high dimensions."}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(
**inputs,
max_new_tokens=2048,
do_sample=True,
top_p=0.9,
temperature=0.7,
repetition_penalty=1.15
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
```
### Generation Tips
- **Temperature 0.6β0.8** works best for reasoning tasks β low enough for coherence, high enough to activate the extended deliberation patterns from the Thinking teacher.
- **Repetition penalty 1.1β1.2** prevents the model from getting caught in reasoning loops during long generations.
- **Max tokens 1024β2048** β the model was trained on 4096 max seq, so it can generate long. Give it room.
- The model inherits the Thinking teacher's tendency to reason before answering. Let it.
## Distillation Position
```
Qwen3-30B-A3B-Thinking (teacher)
β SFT on longwriter-6k (4096 max seq)
Qwen3-1.7B-Thinking-Distil β you are here
```
This model is the **direct SFT** path. The DistilQwen collection also includes models that go through additional refinement stages:
```
Qwen3-1.7B (base)
β Qwen3-1.7B-Distilled-30B-A3B (Instruct teacher KD)
β DiStil (uncensored SFT)
β Disctil (DISC refinement)
β TopologicalQwen (full TKD pipeline)
```
Different paths, different capabilities. This model prioritizes extended reasoning. TopologicalQwen prioritizes structural precision. The Coder variant prioritizes hierarchical decomposition. They're complementary.
## DistilQwen Collection
| Model | What It Does |
|-------|-------------|
| **[Qwen3-1.7B-Thinking-Distil](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Thinking-Distil)** | **β this model. Thinking teacher SFT.** |
| [TopologicalQwen](https://huggingface.co/reaperdoesntknow/TopologicalQwen) | Full TKD pipeline. BV decomposition + DualMind format. |
| [DiStil-Qwen3-1.7B-uncensored](https://huggingface.co/reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored) | DISC-informed uncensored distillation. |
| [Qwen3-1.7B-Coder-Distilled-SFT](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT) | Coder teacher. Hierarchical problem solving. |
| [DistilQwen3-1.7B-uncensored](https://huggingface.co/reaperdoesntknow/DistilQwen3-1.7B-uncensored) | Base uncensored variant. |
Full collection: [DistilQwen on HuggingFace](https://huggingface.co/collections/reaperdoesntknow/distilqwen-69bf40ec669117e3f069ef1c)
## Methodology
Full methodology paper: **[Structure Over Scale: Proof-Weighted Knowledge Distillation](https://doi.org/10.57967/hf/8165)** (DOI: 10.57967/hf/8165)
Companion paper: **[Three Teachers to Dual Cognition](https://doi.org/10.57967/hf/8184)** (DOI: 10.57967/hf/8184) β covers the DualMind extension and ghost imprinting phenomenon.
## License
Apache 2.0 β same as the base Qwen3 model.
## Mathematical Foundations: Discrepancy Calculus (DISC)
This model's training pipeline is grounded in Discrepancy Calculus β a measure-theoretic framework that treats singularities as primary structure rather than pathology. Full theory: *"On the Formal Analysis of Discrepancy Calculus"* (CIx, 2026; Convergent Intelligence LLC: Research Division).
**The Core Operator:**
$$Df(x) = \lim_{\varepsilon \downarrow 0} \frac{1}{\varepsilon} \int_x^{x+\varepsilon} \frac{|f(t) - f(x)|}{|t - x|}\, dt$$
For smooth $f$: $Df(x) = |f'(x)|$. For rough $f$: $D$ localizes irregularity to null sets while preserving integral structure.
**The Mesh Fundamental Identity** β every BV function decomposes as:
$$f(b) - f(a) = \underbrace{\int_a^b f'(x)\,dx}_{\text{smooth (AC)}} + \underbrace{\sum_{x \in J_f} \Delta f(x)}_{\text{jumps}} + \underbrace{D^c f(I)}_{\text{Cantor drift}}$$
Standard knowledge distillation captures only term 1. Topological Knowledge Distillation (TKD) preserves all three by treating the teacher's output distribution as a BV function and computing discrepancy energy, jump sets, and gap energy density before training begins.
## Citation
```bibtex
@misc{cix2026distilqwen,
title={Structure Over Scale: Proof-Weighted Knowledge Distillation from Qwen3-30B to 1.7B},
author={Convergent Intelligence},
year={2026},
doi={10.57967/hf/8165},
publisher={Convergent Intelligence LLC: Research Division}
}
```
---
*Convergent Intelligence LLC: Research Division.*
*[Full portfolio](https://huggingface.co/reaperdoesntknow) | [DistilQwen Collection](https://huggingface.co/collections/reaperdoesntknow/distilqwen-69bf40ec669117e3f069ef1c) | [DualMind Collection](https://huggingface.co/collections/reaperdoesntknow/dualmind-69c93f888c6e79ecc69cf41e)*
---
## Convergent Intelligence Portfolio
*Part of the [DistilQwen Series](https://huggingface.co/collections/reaperdoesntknow/distilqwen-69bf40ec669117e3f069ef1c) by [Convergent Intelligence LLC: Research Division](https://huggingface.co/reaperdoesntknow)*
### Related Models
| Model | Format |
|-------|--------|
| [TopologicalQwen](https://huggingface.co/reaperdoesntknow/TopologicalQwen) | BF16 |
| [Qwen3-1.7B-Thinking-Distil](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Thinking-Distil) | BF16 |
| [Qwen3-1.7B-Coder-Distilled-SFT](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT) | BF16 |
| [DiStil-Qwen3-1.7B-uncensored](https://huggingface.co/reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored) | BF16 |
| [DistilQwen3-1.7B-uncensored](https://huggingface.co/reaperdoesntknow/DistilQwen3-1.7B-uncensored) | BF16 |
| [Qwen3-1.7B-Distilled-30B-A3B](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B) | BF16 |
### Papers
| Paper | DOI |
|-------|-----|
| [Structure Over Scale](https://huggingface.co/reaperdoesntknow/Structure-Over-Scale) | 10.57967/hf/8165 |
| [Three Teachers to Dual Cognition](https://huggingface.co/reaperdoesntknow/DualMind_Methodolgy) | 10.57967/hf/8184 |
| [Discrepancy Calculus](https://huggingface.co/reaperdoesntknow/Discrepancy_Calculus) | 10.57967/hf/8194 |
---
<!-- cix-keeper-ts:2026-09-29T13:16:36Z -->
|