Instructions to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="summerMC/Qwen3.5-9B-SpeedX9-GDN32", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("summerMC/Qwen3.5-9B-SpeedX9-GDN32", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "summerMC/Qwen3.5-9B-SpeedX9-GDN32" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summerMC/Qwen3.5-9B-SpeedX9-GDN32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/summerMC/Qwen3.5-9B-SpeedX9-GDN32
- SGLang
How to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "summerMC/Qwen3.5-9B-SpeedX9-GDN32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summerMC/Qwen3.5-9B-SpeedX9-GDN32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "summerMC/Qwen3.5-9B-SpeedX9-GDN32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summerMC/Qwen3.5-9B-SpeedX9-GDN32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with Docker Model Runner:
docker model run hf.co/summerMC/Qwen3.5-9B-SpeedX9-GDN32
Qwen3.5-9B-SpeedX9-GDN32
A recurrent linear-attention conversion of Qwen/Qwen3.5-9B, produced with the UNI MAX toolchain and designed for fast decoding.
The source model's hybrid attention layout is converted into a fully recurrent linear-attention (GDN) runtime. Each full-attention layer is replaced by a layer initialized from a neighboring native GDN layer, then tuned with sequential layer-wise distillation.
Important: This is a converted and distilled model, not the original
Qwen/Qwen3.5-9Bcheckpoint. It is not numerically identical to the original, and quality has not been fully benchmarked (see Limitations).
Overview
| Item | Value |
|---|---|
| Base model | Qwen/Qwen3.5-9B |
| Parameters / dtype | ~9B / BF16 |
| Hidden layers | 32 |
| Runtime topology | Linear attention (all layers) |
| Converted full-attention layers | 3, 7, 11, 15, 19, 23, 27, 31 (8 layers) |
| Custom runtime | qwen35_unimax_engine |
| Kernel acceleration | CUDA kernels (enabled in the benchmark) |
How it was made
- Layer replacement. The 8 full-attention layers listed above were replaced with GDN/linear-attention layers, each initialized from a neighboring native GDN layer.
- Distillation. The replaced layers were optimized one after another with layer-wise distillation, using 20 optimization steps per converted layer in this build.
Quick start
The checkpoint must be loaded with the matching qwen35_unimax_engine runtime. Install it first.
from qwen35_unimax_engine import UniMaxPipeline
pipe = UniMaxPipeline.from_pretrained(
"summerMC/Qwen3.5-9B-SpeedX9-GDN32",
mode="auto",
graph_candidates=(2, 4, 8),
quant_modes=("fp8", "int8"),
quality_steps=64,
min_agreement=1.0,
use_kernels=True,
)
print(pipe.runtime_summary())
output = pipe(
"Explain recurrent linear attention:",
max_new_tokens=128,
)
print(output)
The repository is tagged custom_code; loading through plain transformers requires trust_remote_code=True and has not been verified here. Use the UNI MAX runtime above as the supported path.
Runtime configuration
| Option | Value used in the benchmark |
|---|---|
| CUDA kernels | enabled |
| Graph candidates | (2, 4, 8) |
| Quantization candidates | ("fp8", "int8") |
| Quality verification steps | 64 |
| Quality context | 64 |
| Minimum fast-path agreement | 1.0 |
With mode="auto", the runtime may choose a different execution path depending on your hardware and the quality-verification result.
Benchmark
Generated directly with benchmark_uni_max().
Environment
| Item | Value |
|---|---|
| GPU | NVIDIA RTX PRO 6000 Blackwell Server Edition |
| GPU VRAM | 94.97 GiB |
| Peak allocated VRAM during benchmark | 70.14 GiB |
| PyTorch | 2.11.0+cu130 |
| CUDA | 13.0 (compute capability 12.0) |
| NVIDIA driver | 580.82.07 |
| Decode steps | 64 |
| Warmup / measured runs | 2 / 5 |
| Context lengths | 256, 1024, 4096, 16384 |
| Total wall time | 90.62 s |
Selected configuration
| Metric | Value |
|---|---|
| Selected mode | exact |
| Selected block | 8 |
| Verified | (not recorded in this build) |
The machine-readable record is included as benchmark_results.json.
Throughput, latency, kernel selection, quantization support, and memory use depend heavily on the GPU, CUDA/PyTorch versions, context length, and runtime settings. Do not treat these numbers as hardware-independent performance claims.
Validation
Before publication, the checkpoint was checked with the UNI MAX checkpoint validator and the benchmark path. These layer-wise agreement checks are narrower than a full language-model evaluation suite, so please evaluate task quality yourself before deploying.
Limitations
- Replacing full attention with recurrent linear attention changes the architecture. It can affect accuracy, long-context behavior, reasoning, generation quality, and numerical outputs relative to the original model.
- Only 20 distillation steps per converted layer were used in this build.
- No downstream task benchmarks (e.g., MMLU, long-context retrieval) are reported.
- Benchmark results apply only to the hardware and software environment documented above.
- Requires the matching custom runtime and an NVIDIA GPU with CUDA.
Base model and license
Derived from Qwen/Qwen3.5-9B. Refer to the upstream repository for the original documentation, license, intended use, and limitations.
Qwen3.5-9B-SpeedX9-GDN32
Qwen/Qwen3.5-9B を、UNI MAX ツールチェーンで再帰型リニアアテンション(GDN)に変換した、高速デコード向けモデルです。
元モデルのハイブリッドアテンション構成を、完全な再帰型リニアアテンションのランタイムに変換しています。フルアテンション層は、隣接するネイティブ GDN 層から初期化した層に置き換え、層ごとの逐次蒸留(layer-wise distillation)で調整しました。
重要: 本モデルは変換・蒸留された派生モデルであり、元の
Qwen/Qwen3.5-9Bそのものではありません。数値的に元モデルと一致するものではなく、品質も十分には評価されていません(制限事項を参照)。
概要
| 項目 | 値 |
|---|---|
| ベースモデル | Qwen/Qwen3.5-9B |
| パラメータ数 / dtype | 約9B / BF16 |
| 隠れ層数 | 32 |
| ランタイム構成 | リニアアテンション(全層) |
| 変換したフルアテンション層 | 3, 7, 11, 15, 19, 23, 27, 31(計8層) |
| 専用ランタイム | qwen35_unimax_engine |
| カーネル高速化 | CUDA カーネル(ベンチマーク時は有効) |
作成方法
- 層の置き換え: 上記8つのフルアテンション層を GDN/リニアアテンション層に置き換え、各層は隣接するネイティブ GDN 層から初期化しました。
- 蒸留: 置き換えた層を順番に層ごとの蒸留で最適化しました。今回のビルドでは変換した各層につき 20 ステップです。
クイックスタート
チェックポイントの読み込みには、対応する qwen35_unimax_engine ランタイムが必要です。先にインストールしてください。
from qwen35_unimax_engine import UniMaxPipeline
pipe = UniMaxPipeline.from_pretrained(
"summerMC/Qwen3.5-9B-SpeedX9-GDN32",
mode="auto",
graph_candidates=(2, 4, 8),
quant_modes=("fp8", "int8"),
quality_steps=64,
min_agreement=1.0,
use_kernels=True,
)
print(pipe.runtime_summary())
output = pipe(
"Explain recurrent linear attention:",
max_new_tokens=128,
)
print(output)
このリポジトリには custom_code タグが付いており、素の transformers で読み込む場合は trust_remote_code=True が必要ですが、動作は未検証です。上記の UNI MAX ランタイムを推奨パスとしてください。
ランタイム設定
| オプション | ベンチマークでの値 |
|---|---|
| CUDA カーネル | 有効 |
| グラフ候補 | (2, 4, 8) |
| 量子化候補 | ("fp8", "int8") |
| 品質検証ステップ数 | 64 |
| 品質検証コンテキスト長 | 64 |
| 高速パスの最小一致率 | 1.0 |
mode="auto" では、ハードウェアや品質検証の結果に応じて、別の実行パスが選ばれることがあります。
ベンチマーク
benchmark_uni_max() から直接生成した結果です。
実行環境
| 項目 | 値 |
|---|---|
| GPU | NVIDIA RTX PRO 6000 Blackwell Server Edition |
| GPU VRAM | 94.97 GiB |
| ベンチマーク中の最大確保 VRAM | 70.14 GiB |
| PyTorch | 2.11.0+cu130 |
| CUDA | 13.0(compute capability 12.0) |
| NVIDIA ドライバ | 580.82.07 |
| デコードステップ数 | 64 |
| ウォームアップ / 計測回数 | 2 / 5 |
| コンテキスト長 | 256, 1024, 4096, 16384 |
| 総実行時間 | 90.62 秒 |
選択された構成
| 指標 | 値 |
|---|---|
| 選択モード | exact |
| 選択ブロック | 8 |
| verified | (このビルドでは記録なし) |
機械可読な結果は benchmark_results.json に含まれています。
スループット、レイテンシ、カーネル選択、量子化対応、メモリ使用量は、GPU、CUDA/PyTorch のバージョン、コンテキスト長、ランタイム設定に大きく依存します。これらの数値をハードウェア非依存の性能保証として扱わないでください。
検証
公開前に、UNI MAX のチェックポイント検証ツールとベンチマーク経路で確認しています。ただし、層単位の一致確認は言語モデル全体の評価スイートよりも範囲が狭いため、運用前にご自身のタスクで品質を評価してください。
制限事項
- フルアテンションを再帰型リニアアテンションに置き換えているため、元モデルと比べて、精度、長文脈での挙動、推論能力、生成品質、数値出力が変化する可能性があります。
- 今回のビルドでは、変換した各層の蒸留は 20 ステップのみです。
- 下流タスクのベンチマーク(MMLU、長文脈検索など)は報告していません。
- ベンチマーク結果は、上記に記載したハードウェア・ソフトウェア環境に限られます。
- 専用ランタイムと、CUDA 対応の NVIDIA GPU が必要です。
ベースモデルとライセンス
Qwen/Qwen3.5-9B から派生しています。元のドキュメント、ライセンス、想定用途、制限事項は上流リポジトリをご確認ください。
- Downloads last month
- 435
