summerMC commited on
Commit
81604c2
·
verified ·
1 Parent(s): 68b2f3e

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +193 -87
README.md CHANGED
@@ -1,107 +1,176 @@
1
- [---
2
- library_name: transformers
3
  pipeline_tag: text-generation
 
4
  tags:
5
- - qwen
6
- - qwen3.5
7
- - text-generation
8
- - recurrent
9
- - linear-attention
10
- - cuda
11
- - custom-code
12
- base_model: Qwen/Qwen3.5-9B
 
 
 
13
  ---
14
 
15
- # summerMC/Qwen3.5-9B-SpeedX9-GDN32
 
 
 
 
 
 
16
 
17
- This repository contains a UNI MAX conversion of
18
- `Qwen/Qwen3.5-9B`.
19
 
20
- The source hybrid attention topology is converted to a recurrent
21
- linear-attention runtime. Full-attention layers are replaced through
22
- layer-wise distillation from neighboring native GDN/linear-attention
23
- layers.
24
 
25
- ## Architecture
26
 
27
- - Base model: `Qwen/Qwen3.5-9B`
28
- - Hidden layers: `32`
29
- - Runtime topology: linear attention
30
- - Converted source full-attention layers: `3, 7, 11, 15, 19, 23, 27, 31`
31
- - Custom runtime: `qwen35_unimax_engine`
32
- - Kernel acceleration: enabled for the benchmark below
33
 
34
- The checkpoint should be used with the matching UNI MAX runtime code.
35
- It is not claimed to be numerically identical to the original
36
- full-attention model.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
 
38
- ## Distillation
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
39
 
40
- The converted full-attention layers were initialized from neighboring
41
- native GDN layers and then optimized with sequential layer-wise
42
- distillation.
43
 
44
- The current build used 20 optimization steps per converted layer.
45
 
46
- ## Benchmark environment
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
47
 
48
  | Item | Value |
49
- |---|---|
50
  | GPU | NVIDIA RTX PRO 6000 Blackwell Server Edition |
51
  | GPU VRAM | 94.97 GiB |
52
  | Peak allocated VRAM during benchmark | 70.14 GiB |
53
  | PyTorch | 2.11.0+cu130 |
54
- | CUDA | 13.0 |
55
- | CUDA capability | 12.0 |
56
  | NVIDIA driver | 580.82.07 |
57
  | Decode steps | 64 |
58
- | Warmup runs | 2 |
59
- | Measured runs | 5 |
60
  | Context lengths | 256, 1024, 4096, 16384 |
61
- | Benchmark wall time | 90.62 s |
62
 
63
- ## Benchmark results
64
 
65
  | Metric | Value |
66
- |---|---:|
67
- | verified | None |
68
- | selected_mode | exact |
69
- | selected_block | 8 |
 
 
70
 
71
- The benchmark above was generated directly from
72
- `benchmark_uni_max()` on the hardware shown above.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
73
 
74
- Because throughput, latency, kernel selection, quantization support,
75
- and memory consumption depend strongly on the GPU, CUDA/PyTorch
76
- versions, context length, and runtime configuration, these numbers
77
- should not be treated as hardware-independent performance claims.
78
 
79
- ## Raw benchmark output
80
 
 
81
 
82
- ![image](https://cdn-uploads.huggingface.co/production/uploads/69e94913e947e9e6a0ae9d77/VYA_rtMMqGs3Auhdt2lyz.png)
83
 
84
- The machine-readable benchmark record is also included as
85
- `benchmark_results.json`.
86
 
87
- ## Runtime configuration
88
 
89
- The benchmark evaluated:
 
 
 
 
 
 
 
 
90
 
91
- - CUDA kernels enabled
92
- - graph candidates: `(2, 4, 8)`
93
- - quantization candidates: `('fp8', 'int8')`
94
- - quality verification steps: `64`
95
- - quality context: `64`
96
- - minimum fast-path agreement: `1.0`
97
 
98
- The runtime may select a different execution path depending on
99
- hardware support and quality verification.
100
 
101
- ## Usage
102
 
103
- Install the matching `qwen35_unimax_engine` runtime before loading the
104
- checkpoint.
105
 
106
  ```python
107
  from qwen35_unimax_engine import UniMaxPipeline
@@ -122,34 +191,71 @@ output = pipe(
122
  "Explain recurrent linear attention:",
123
  max_new_tokens=128,
124
  )
125
-
126
  print(output)
127
  ```
128
 
129
- ## Validation
130
 
131
- Before publication, the local checkpoint was checked with the UNI MAX
132
- checkpoint validator and benchmark path.
133
 
134
- Users should independently evaluate task quality before deploying the
135
- converted model. Layer-wise agreement tests are narrower than a full
136
- language-model evaluation suite.
 
 
 
 
 
137
 
138
- ## Limitations
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
139
 
140
- This is a converted and distilled model, not the unmodified
141
- `Qwen/Qwen3.5-9B` checkpoint.
142
 
143
- Replacing full attention with recurrent linear attention changes the
144
- model architecture and can affect accuracy, long-context behavior,
145
- reasoning, generation quality, and numerical output.
146
 
147
- Benchmark results are specific to the hardware and software
148
- environment documented above.
 
 
 
149
 
150
- ## Base model
151
 
152
- Derived from `Qwen/Qwen3.5-9B`. Refer to the upstream model repository
153
- for the original model documentation, license, intended use, and
154
- limitations.
155
- ](https://huggingface.co/summerMC/Qwen3.5-9B-SpeedX9-GDN32)
 
1
+ ---
2
+ base_model: Qwen/Qwen3.5-9B
3
  pipeline_tag: text-generation
4
+ library_name: transformers
5
  tags:
6
+ - qwen
7
+ - qwen3.5
8
+ - recurrent
9
+ - linear-attention
10
+ - gdn
11
+ - cuda
12
+ - custom_code
13
+ language:
14
+ - en
15
+ - ja
16
+ # license: TODO — match the upstream Qwen/Qwen3.5-9B license
17
  ---
18
 
19
+ **Language / 言語:** [English](#english) | [日本語](#日本語)
20
+
21
+ ---
22
+
23
+ <a id="english"></a>
24
+
25
+ # Qwen3.5-9B-SpeedX9-GDN32
26
 
27
+ A recurrent linear-attention conversion of [`Qwen/Qwen3.5-9B`](https://huggingface.co/Qwen/Qwen3.5-9B), produced with the UNI MAX toolchain and designed for fast decoding.
 
28
 
29
+ The source model's hybrid attention layout is converted into a fully recurrent linear-attention (GDN) runtime. Each full-attention layer is replaced by a layer initialized from a neighboring native GDN layer, then tuned with sequential layer-wise distillation.
 
 
 
30
 
31
+ > **Important:** This is a converted and distilled model, **not** the original `Qwen/Qwen3.5-9B` checkpoint. It is not numerically identical to the original, and quality has not been fully benchmarked (see [Limitations](#limitations)).
32
 
33
+ ## Overview
 
 
 
 
 
34
 
35
+ | Item | Value |
36
+ | --- | --- |
37
+ | Base model | `Qwen/Qwen3.5-9B` |
38
+ | Parameters / dtype | ~9B / BF16 |
39
+ | Hidden layers | 32 |
40
+ | Runtime topology | Linear attention (all layers) |
41
+ | Converted full-attention layers | 3, 7, 11, 15, 19, 23, 27, 31 (8 layers) |
42
+ | Custom runtime | `qwen35_unimax_engine` |
43
+ | Kernel acceleration | CUDA kernels (enabled in the benchmark) |
44
+
45
+ ## How it was made
46
+
47
+ 1. **Layer replacement.** The 8 full-attention layers listed above were replaced with GDN/linear-attention layers, each initialized from a neighboring native GDN layer.
48
+ 2. **Distillation.** The replaced layers were optimized one after another with layer-wise distillation, using **20 optimization steps per converted layer** in this build.
49
+
50
+ ## Quick start
51
+
52
+ The checkpoint must be loaded with the matching `qwen35_unimax_engine` runtime. Install it first.
53
 
54
+ ```python
55
+ from qwen35_unimax_engine import UniMaxPipeline
56
+
57
+ pipe = UniMaxPipeline.from_pretrained(
58
+ "summerMC/Qwen3.5-9B-SpeedX9-GDN32",
59
+ mode="auto",
60
+ graph_candidates=(2, 4, 8),
61
+ quant_modes=("fp8", "int8"),
62
+ quality_steps=64,
63
+ min_agreement=1.0,
64
+ use_kernels=True,
65
+ )
66
+
67
+ print(pipe.runtime_summary())
68
+
69
+ output = pipe(
70
+ "Explain recurrent linear attention:",
71
+ max_new_tokens=128,
72
+ )
73
+ print(output)
74
+ ```
75
 
76
+ The repository is tagged `custom_code`; loading through plain `transformers` requires `trust_remote_code=True` and has not been verified here. Use the UNI MAX runtime above as the supported path.
 
 
77
 
78
+ ### Runtime configuration
79
 
80
+ | Option | Value used in the benchmark |
81
+ | --- | --- |
82
+ | CUDA kernels | enabled |
83
+ | Graph candidates | `(2, 4, 8)` |
84
+ | Quantization candidates | `("fp8", "int8")` |
85
+ | Quality verification steps | 64 |
86
+ | Quality context | 64 |
87
+ | Minimum fast-path agreement | 1.0 |
88
+
89
+ With `mode="auto"`, the runtime may choose a different execution path depending on your hardware and the quality-verification result.
90
+
91
+ ## Benchmark
92
+
93
+ Generated directly with `benchmark_uni_max()`.
94
+
95
+ **Environment**
96
 
97
  | Item | Value |
98
+ | --- | --- |
99
  | GPU | NVIDIA RTX PRO 6000 Blackwell Server Edition |
100
  | GPU VRAM | 94.97 GiB |
101
  | Peak allocated VRAM during benchmark | 70.14 GiB |
102
  | PyTorch | 2.11.0+cu130 |
103
+ | CUDA | 13.0 (compute capability 12.0) |
 
104
  | NVIDIA driver | 580.82.07 |
105
  | Decode steps | 64 |
106
+ | Warmup / measured runs | 2 / 5 |
 
107
  | Context lengths | 256, 1024, 4096, 16384 |
108
+ | Total wall time | 90.62 s |
109
 
110
+ **Selected configuration**
111
 
112
  | Metric | Value |
113
+ | --- | --- |
114
+ | Selected mode | `exact` |
115
+ | Selected block | 8 |
116
+ | Verified | _(not recorded in this build)_ |
117
+
118
+ <!-- TODO: Replace the screenshot below with a table of tokens/s and latency per context length, taken from benchmark_results.json. The current card exposes the numbers only as an image. -->
119
 
120
+ ![Raw benchmark output](https://cdn-uploads.huggingface.co/production/uploads/69e94913e947e9e6a0ae9d77/VYA_rtMMqGs3Auhdt2lyz.png)
121
+
122
+ The machine-readable record is included as `benchmark_results.json`.
123
+
124
+ Throughput, latency, kernel selection, quantization support, and memory use depend heavily on the GPU, CUDA/PyTorch versions, context length, and runtime settings. **Do not treat these numbers as hardware-independent performance claims.**
125
+
126
+ ## Validation
127
+
128
+ Before publication, the checkpoint was checked with the UNI MAX checkpoint validator and the benchmark path. These layer-wise agreement checks are narrower than a full language-model evaluation suite, so please evaluate task quality yourself before deploying.
129
+
130
+ ## Limitations
131
+
132
+ - Replacing full attention with recurrent linear attention changes the architecture. It can affect accuracy, long-context behavior, reasoning, generation quality, and numerical outputs relative to the original model.
133
+ - Only 20 distillation steps per converted layer were used in this build.
134
+ - No downstream task benchmarks (e.g., MMLU, long-context retrieval) are reported.
135
+ - Benchmark results apply only to the hardware and software environment documented above.
136
+ - Requires the matching custom runtime and an NVIDIA GPU with CUDA.
137
+
138
+ ## Base model and license
139
+
140
+ Derived from `Qwen/Qwen3.5-9B`. Refer to the [upstream repository](https://huggingface.co/Qwen/Qwen3.5-9B) for the original documentation, license, intended use, and limitations.
141
+
142
+ ---
143
 
144
+ <a id="日本語"></a>
 
 
 
145
 
146
+ # Qwen3.5-9B-SpeedX9-GDN32
147
 
148
+ [`Qwen/Qwen3.5-9B`](https://huggingface.co/Qwen/Qwen3.5-9B) を、UNI MAX ツールチェーンで再帰型リニアアテンション(GDN)に変換した、高速デコード向けモデルです。
149
 
150
+ 元モデルのハイブリッドアテンション構成を、完全な再帰型リニアアテンションのランタイムに変換しています。フルアテンション層は、隣接するネイティブ GDN 層から初期化した層に置き換え、層ごとの逐次蒸留(layer-wise distillation)で調整しました。
151
 
152
+ > **重要:** 本モデルは変換・蒸留された派生モデルであり、元の `Qwen/Qwen3.5-9B` そのものでは**ありません**。数値的に元モデルと一致するものではなく、品質も十分には評価されていません([制限事項](#制限事項)を参照)。
 
153
 
154
+ ## 概要
155
 
156
+ | 項目 | 値 |
157
+ | --- | --- |
158
+ | ベースモデル | `Qwen/Qwen3.5-9B` |
159
+ | パラメータ数 / dtype | 約9B / BF16 |
160
+ | 隠れ層数 | 32 |
161
+ | ランタイム構成 | リニアアテンション(全層) |
162
+ | 変換したフルアテンション層 | 3, 7, 11, 15, 19, 23, 27, 31(計8層) |
163
+ | 専用ランタイム | `qwen35_unimax_engine` |
164
+ | カーネル高速化 | CUDA カーネル(ベンチマーク時は有効) |
165
 
166
+ ## 作成方法
 
 
 
 
 
167
 
168
+ 1. **層の置き換え:** 上記8つのフルアテンション層を GDN/リニアアテンション層に置き換え、各層は隣接するネイティブ GDN 層から初期化しました。
169
+ 2. **蒸留:** 置き換えた層を順番に層ごとの蒸留で最適化しました。今回のビルドでは**変換した各層につき 20 ステップ**です。
170
 
171
+ ## クイックスタート
172
 
173
+ チェックポイントの読み込みには、対応する `qwen35_unimax_engine` ランタイムが必要です。先にインストールしてください。
 
174
 
175
  ```python
176
  from qwen35_unimax_engine import UniMaxPipeline
 
191
  "Explain recurrent linear attention:",
192
  max_new_tokens=128,
193
  )
 
194
  print(output)
195
  ```
196
 
197
+ このリポジトリには `custom_code` タグが付いており、素の `transformers` で読み込む場合は `trust_remote_code=True` が必要ですが、動作は未検証です。上記の UNI MAX ランタイムを推奨パスとしてください。
198
 
199
+ ### ランタイム設定
 
200
 
201
+ | オプション | ベンチマークでの値 |
202
+ | --- | --- |
203
+ | CUDA カーネル | 有効 |
204
+ | グラフ候補 | `(2, 4, 8)` |
205
+ | 量子化候補 | `("fp8", "int8")` |
206
+ | 品質検証ステップ数 | 64 |
207
+ | 品質検証コンテキスト長 | 64 |
208
+ | 高速パスの最小一致率 | 1.0 |
209
 
210
+ `mode="auto"` では、ハードウェアや品質検証の結果に応じて、別の実行パスが選ばれることがあります。
211
+
212
+ ## ベンチマーク
213
+
214
+ `benchmark_uni_max()` から直接生成した結果です。
215
+
216
+ **実行環境**
217
+
218
+ | 項目 | 値 |
219
+ | --- | --- |
220
+ | GPU | NVIDIA RTX PRO 6000 Blackwell Server Edition |
221
+ | GPU VRAM | 94.97 GiB |
222
+ | ベンチマーク中の最大確保 VRAM | 70.14 GiB |
223
+ | PyTorch | 2.11.0+cu130 |
224
+ | CUDA | 13.0(compute capability 12.0) |
225
+ | NVIDIA ドライバ | 580.82.07 |
226
+ | デコードステップ数 | 64 |
227
+ | ウォームアップ / 計測回数 | 2 / 5 |
228
+ | コンテキスト長 | 256, 1024, 4096, 16384 |
229
+ | 総実行時間 | 90.62 秒 |
230
+
231
+ **選択された構成**
232
+
233
+ | 指標 | 値 |
234
+ | --- | --- |
235
+ | 選択モード | `exact` |
236
+ | 選択ブロック | 8 |
237
+ | verified | (このビルドでは記録なし) |
238
+
239
+ <!-- TODO: 下の画像を、benchmark_results.json に基づくコンテキスト長ごとの tokens/s・レイテンシの表に置き換える。 -->
240
+
241
+ ![ベンチマーク生出力](https://cdn-uploads.huggingface.co/production/uploads/69e94913e947e9e6a0ae9d77/VYA_rtMMqGs3Auhdt2lyz.png)
242
+
243
+ 機械可読な結果は `benchmark_results.json` に含まれています。
244
+
245
+ スループット、レイテンシ、カーネル選択、量子化対応、メモリ使用量は、GPU、CUDA/PyTorch のバージョン、コンテキスト長、ランタイム設定に大きく依存します。**これらの数値をハードウェア非依存の性能保証として扱わないでください。**
246
+
247
+ ## 検証
248
 
249
+ 公開前に、UNI MAX のチェックポイント検証ツールとベンチマーク経路で確認しています。ただし、層単位の一致確認は言語モデル全体の評価スイートよりも範囲が狭いため、運用前にご自身のタスクで品質を評価してください。
 
250
 
251
+ ## 制限事項
 
 
252
 
253
+ - フルアテンションを再帰型リニアアテンションに置き換えているため、元モデルと比べて、精度、長文脈での挙動、推論能力、生成品質、数値出力が変化する可能性があります。
254
+ - 今回のビルドでは、変換した各層の蒸留は 20 ステップのみです。
255
+ - 下流タスクのベンチマーク(MMLU、長文脈検索など)は報告していません。
256
+ - ベンチマーク結果は、上記に記載したハードウェア・ソフトウェア環境に限られます。
257
+ - 専用ランタイムと、CUDA 対応の NVIDIA GPU が必要です。
258
 
259
+ ## ベースモデルとライセンス
260
 
261
+ `Qwen/Qwen3.5-9B` から派生しています。元のドキュメント、ライセンス、想定用途、制限事項は[上流リポジトリ](https://huggingface.co/Qwen/Qwen3.5-9B)をご確認ください。