Language / 言語: English | 日本語


Qwen3.5-9B-SpeedX9-GDN32

A recurrent linear-attention conversion of Qwen/Qwen3.5-9B, produced with the UNI MAX toolchain and designed for fast decoding.

The source model's hybrid attention layout is converted into a fully recurrent linear-attention (GDN) runtime. Each full-attention layer is replaced by a layer initialized from a neighboring native GDN layer, then tuned with sequential layer-wise distillation.

Important: This is a converted and distilled model, not the original Qwen/Qwen3.5-9B checkpoint. It is not numerically identical to the original, and quality has not been fully benchmarked (see Limitations).

Overview

Item Value
Base model Qwen/Qwen3.5-9B
Parameters / dtype ~9B / BF16
Hidden layers 32
Runtime topology Linear attention (all layers)
Converted full-attention layers 3, 7, 11, 15, 19, 23, 27, 31 (8 layers)
Custom runtime qwen35_unimax_engine
Kernel acceleration CUDA kernels (enabled in the benchmark)

How it was made

  1. Layer replacement. The 8 full-attention layers listed above were replaced with GDN/linear-attention layers, each initialized from a neighboring native GDN layer.
  2. Distillation. The replaced layers were optimized one after another with layer-wise distillation, using 20 optimization steps per converted layer in this build.

Quick start

The checkpoint must be loaded with the matching qwen35_unimax_engine runtime. Install it first.

from qwen35_unimax_engine import UniMaxPipeline

pipe = UniMaxPipeline.from_pretrained(
    "summerMC/Qwen3.5-9B-SpeedX9-GDN32",
    mode="auto",
    graph_candidates=(2, 4, 8),
    quant_modes=("fp8", "int8"),
    quality_steps=64,
    min_agreement=1.0,
    use_kernels=True,
)

print(pipe.runtime_summary())

output = pipe(
    "Explain recurrent linear attention:",
    max_new_tokens=128,
)
print(output)

The repository is tagged custom_code; loading through plain transformers requires trust_remote_code=True and has not been verified here. Use the UNI MAX runtime above as the supported path.

Runtime configuration

Option Value used in the benchmark
CUDA kernels enabled
Graph candidates (2, 4, 8)
Quantization candidates ("fp8", "int8")
Quality verification steps 64
Quality context 64
Minimum fast-path agreement 1.0

With mode="auto", the runtime may choose a different execution path depending on your hardware and the quality-verification result.

Benchmark

Generated directly with benchmark_uni_max().

Environment

Item Value
GPU NVIDIA RTX PRO 6000 Blackwell Server Edition
GPU VRAM 94.97 GiB
Peak allocated VRAM during benchmark 70.14 GiB
PyTorch 2.11.0+cu130
CUDA 13.0 (compute capability 12.0)
NVIDIA driver 580.82.07
Decode steps 64
Warmup / measured runs 2 / 5
Context lengths 256, 1024, 4096, 16384
Total wall time 90.62 s

Selected configuration

Metric Value
Selected mode exact
Selected block 8
Verified (not recorded in this build)

Raw benchmark output

The machine-readable record is included as benchmark_results.json.

Throughput, latency, kernel selection, quantization support, and memory use depend heavily on the GPU, CUDA/PyTorch versions, context length, and runtime settings. Do not treat these numbers as hardware-independent performance claims.

Validation

Before publication, the checkpoint was checked with the UNI MAX checkpoint validator and the benchmark path. These layer-wise agreement checks are narrower than a full language-model evaluation suite, so please evaluate task quality yourself before deploying.

Limitations

  • Replacing full attention with recurrent linear attention changes the architecture. It can affect accuracy, long-context behavior, reasoning, generation quality, and numerical outputs relative to the original model.
  • Only 20 distillation steps per converted layer were used in this build.
  • No downstream task benchmarks (e.g., MMLU, long-context retrieval) are reported.
  • Benchmark results apply only to the hardware and software environment documented above.
  • Requires the matching custom runtime and an NVIDIA GPU with CUDA.

Base model and license

Derived from Qwen/Qwen3.5-9B. Refer to the upstream repository for the original documentation, license, intended use, and limitations.


Qwen3.5-9B-SpeedX9-GDN32

Qwen/Qwen3.5-9B を、UNI MAX ツールチェーンで再帰型リニアアテンション(GDN)に変換した、高速デコード向けモデルです。

元モデルのハイブリッドアテンション構成を、完全な再帰型リニアアテンションのランタイムに変換しています。フルアテンション層は、隣接するネイティブ GDN 層から初期化した層に置き換え、層ごとの逐次蒸留(layer-wise distillation)で調整しました。

重要: 本モデルは変換・蒸留された派生モデルであり、元の Qwen/Qwen3.5-9B そのものではありません。数値的に元モデルと一致するものではなく、品質も十分には評価されていません(制限事項を参照)。

概要

項目 値
ベースモデル Qwen/Qwen3.5-9B
パラメータ数 / dtype 約9B / BF16
隠れ層数 32
ランタイム構成 リニアアテンション(全層)
変換したフルアテンション層 3, 7, 11, 15, 19, 23, 27, 31(計8層)
専用ランタイム qwen35_unimax_engine
カーネル高速化 CUDA カーネル(ベンチマーク時は有効)

作成方法

  1. 層の置き換え: 上記8つのフルアテンション層を GDN/リニアアテンション層に置き換え、各層は隣接するネイティブ GDN 層から初期化しました。
  2. 蒸留: 置き換えた層を順番に層ごとの蒸留で最適化しました。今回のビルドでは変換した各層につき 20 ステップです。

クイックスタート

チェックポイントの読み込みには、対応する qwen35_unimax_engine ランタイムが必要です。先にインストールしてください。

from qwen35_unimax_engine import UniMaxPipeline

pipe = UniMaxPipeline.from_pretrained(
    "summerMC/Qwen3.5-9B-SpeedX9-GDN32",
    mode="auto",
    graph_candidates=(2, 4, 8),
    quant_modes=("fp8", "int8"),
    quality_steps=64,
    min_agreement=1.0,
    use_kernels=True,
)

print(pipe.runtime_summary())

output = pipe(
    "Explain recurrent linear attention:",
    max_new_tokens=128,
)
print(output)

このリポジトリには custom_code タグが付いており、素の transformers で読み込む場合は trust_remote_code=True が必要ですが、動作は未検証です。上記の UNI MAX ランタイムを推奨パスとしてください。

ランタイム設定

オプション ベンチマークでの値
CUDA カーネル 有効
グラフ候補 (2, 4, 8)
量子化候補 ("fp8", "int8")
品質検証ステップ数 64
品質検証コンテキスト長 64
高速パスの最小一致率 1.0

mode="auto" では、ハードウェアや品質検証の結果に応じて、別の実行パスが選ばれることがあります。

ベンチマーク

benchmark_uni_max() から直接生成した結果です。

実行環境

項目 値
GPU NVIDIA RTX PRO 6000 Blackwell Server Edition
GPU VRAM 94.97 GiB
ベンチマーク中の最大確保 VRAM 70.14 GiB
PyTorch 2.11.0+cu130
CUDA 13.0(compute capability 12.0)
NVIDIA ドライバ 580.82.07
デコードステップ数 64
ウォームアップ / 計測回数 2 / 5
コンテキスト長 256, 1024, 4096, 16384
総実行時間 90.62 秒

選択された構成

指標 値
選択モード exact
選択ブロック 8
verified (このビルドでは記録なし)

ベンチマーク生出力

機械可読な結果は benchmark_results.json に含まれています。

スループット、レイテンシ、カーネル選択、量子化対応、メモリ使用量は、GPU、CUDA/PyTorch のバージョン、コンテキスト長、ランタイム設定に大きく依存します。これらの数値をハードウェア非依存の性能保証として扱わないでください。

検証

公開前に、UNI MAX のチェックポイント検証ツールとベンチマーク経路で確認しています。ただし、層単位の一致確認は言語モデル全体の評価スイートよりも範囲が狭いため、運用前にご自身のタスクで品質を評価してください。

制限事項

  • フルアテンションを再帰型リニアアテンションに置き換えているため、元モデルと比べて、精度、長文脈での挙動、推論能力、生成品質、数値出力が変化する可能性があります。
  • 今回のビルドでは、変換した各層の蒸留は 20 ステップのみです。
  • 下流タスクのベンチマーク(MMLU、長文脈検索など)は報告していません。
  • ベンチマーク結果は、上記に記載したハードウェア・ソフトウェア環境に限られます。
  • 専用ランタイムと、CUDA 対応の NVIDIA GPU が必要です。

ベースモデルとライセンス

Qwen/Qwen3.5-9B から派生しています。元のドキュメント、ライセンス、想定用途、制限事項は上流リポジトリをご確認ください。

Downloads last month
435
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for summerMC/Qwen3.5-9B-SpeedX9-GDN32

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(1027)
this model

Collection including summerMC/Qwen3.5-9B-SpeedX9-GDN32