Text Generation
Transformers
Safetensors
English
Japanese
qwen3_5_gdn24
qwen
qwen3.5
recurrent
linear-attention
gdn
cuda
custom_code
conversational
Instructions to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="summerMC/Qwen3.5-9B-SpeedX9-GDN32", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("summerMC/Qwen3.5-9B-SpeedX9-GDN32", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "summerMC/Qwen3.5-9B-SpeedX9-GDN32" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summerMC/Qwen3.5-9B-SpeedX9-GDN32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/summerMC/Qwen3.5-9B-SpeedX9-GDN32
- SGLang
How to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "summerMC/Qwen3.5-9B-SpeedX9-GDN32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summerMC/Qwen3.5-9B-SpeedX9-GDN32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "summerMC/Qwen3.5-9B-SpeedX9-GDN32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "summerMC/Qwen3.5-9B-SpeedX9-GDN32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use summerMC/Qwen3.5-9B-SpeedX9-GDN32 with Docker Model Runner:
docker model run hf.co/summerMC/Qwen3.5-9B-SpeedX9-GDN32
Update README.md
Browse files
README.md
CHANGED
|
@@ -1,107 +1,176 @@
|
|
| 1 |
-
|
| 2 |
-
|
| 3 |
pipeline_tag: text-generation
|
|
|
|
| 4 |
tags:
|
| 5 |
-
- qwen
|
| 6 |
-
- qwen3.5
|
| 7 |
-
-
|
| 8 |
-
-
|
| 9 |
-
-
|
| 10 |
-
- cuda
|
| 11 |
-
-
|
| 12 |
-
|
|
|
|
|
|
|
|
|
|
| 13 |
---
|
| 14 |
|
| 15 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
|
| 17 |
-
|
| 18 |
-
`Qwen/Qwen3.5-9B`.
|
| 19 |
|
| 20 |
-
The source hybrid attention
|
| 21 |
-
linear-attention runtime. Full-attention layers are replaced through
|
| 22 |
-
layer-wise distillation from neighboring native GDN/linear-attention
|
| 23 |
-
layers.
|
| 24 |
|
| 25 |
-
|
| 26 |
|
| 27 |
-
|
| 28 |
-
- Hidden layers: `32`
|
| 29 |
-
- Runtime topology: linear attention
|
| 30 |
-
- Converted source full-attention layers: `3, 7, 11, 15, 19, 23, 27, 31`
|
| 31 |
-
- Custom runtime: `qwen35_unimax_engine`
|
| 32 |
-
- Kernel acceleration: enabled for the benchmark below
|
| 33 |
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
|
| 40 |
-
The
|
| 41 |
-
native GDN layers and then optimized with sequential layer-wise
|
| 42 |
-
distillation.
|
| 43 |
|
| 44 |
-
|
| 45 |
|
| 46 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 47 |
|
| 48 |
| Item | Value |
|
| 49 |
-
|---|---|
|
| 50 |
| GPU | NVIDIA RTX PRO 6000 Blackwell Server Edition |
|
| 51 |
| GPU VRAM | 94.97 GiB |
|
| 52 |
| Peak allocated VRAM during benchmark | 70.14 GiB |
|
| 53 |
| PyTorch | 2.11.0+cu130 |
|
| 54 |
-
| CUDA | 13.0 |
|
| 55 |
-
| CUDA capability | 12.0 |
|
| 56 |
| NVIDIA driver | 580.82.07 |
|
| 57 |
| Decode steps | 64 |
|
| 58 |
-
| Warmup runs | 2 |
|
| 59 |
-
| Measured runs | 5 |
|
| 60 |
| Context lengths | 256, 1024, 4096, 16384 |
|
| 61 |
-
|
|
| 62 |
|
| 63 |
-
|
| 64 |
|
| 65 |
| Metric | Value |
|
| 66 |
-
|---|---
|
| 67 |
-
|
|
| 68 |
-
|
|
| 69 |
-
|
|
|
|
|
|
|
|
| 70 |
|
| 71 |
-
|
| 72 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 73 |
|
| 74 |
-
|
| 75 |
-
and memory consumption depend strongly on the GPU, CUDA/PyTorch
|
| 76 |
-
versions, context length, and runtime configuration, these numbers
|
| 77 |
-
should not be treated as hardware-independent performance claims.
|
| 78 |
|
| 79 |
-
#
|
| 80 |
|
|
|
|
| 81 |
|
| 82 |
-
|
| 83 |
|
| 84 |
-
|
| 85 |
-
`benchmark_results.json`.
|
| 86 |
|
| 87 |
-
##
|
| 88 |
|
| 89 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 90 |
|
| 91 |
-
|
| 92 |
-
- graph candidates: `(2, 4, 8)`
|
| 93 |
-
- quantization candidates: `('fp8', 'int8')`
|
| 94 |
-
- quality verification steps: `64`
|
| 95 |
-
- quality context: `64`
|
| 96 |
-
- minimum fast-path agreement: `1.0`
|
| 97 |
|
| 98 |
-
|
| 99 |
-
|
| 100 |
|
| 101 |
-
##
|
| 102 |
|
| 103 |
-
|
| 104 |
-
checkpoint.
|
| 105 |
|
| 106 |
```python
|
| 107 |
from qwen35_unimax_engine import UniMaxPipeline
|
|
@@ -122,34 +191,71 @@ output = pipe(
|
|
| 122 |
"Explain recurrent linear attention:",
|
| 123 |
max_new_tokens=128,
|
| 124 |
)
|
| 125 |
-
|
| 126 |
print(output)
|
| 127 |
```
|
| 128 |
|
| 129 |
-
|
| 130 |
|
| 131 |
-
|
| 132 |
-
checkpoint validator and benchmark path.
|
| 133 |
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 137 |
|
| 138 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 139 |
|
| 140 |
-
|
| 141 |
-
`Qwen/Qwen3.5-9B` checkpoint.
|
| 142 |
|
| 143 |
-
|
| 144 |
-
model architecture and can affect accuracy, long-context behavior,
|
| 145 |
-
reasoning, generation quality, and numerical output.
|
| 146 |
|
| 147 |
-
|
| 148 |
-
|
|
|
|
|
|
|
|
|
|
| 149 |
|
| 150 |
-
##
|
| 151 |
|
| 152 |
-
|
| 153 |
-
for the original model documentation, license, intended use, and
|
| 154 |
-
limitations.
|
| 155 |
-
](https://huggingface.co/summerMC/Qwen3.5-9B-SpeedX9-GDN32)
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3.5-9B
|
| 3 |
pipeline_tag: text-generation
|
| 4 |
+
library_name: transformers
|
| 5 |
tags:
|
| 6 |
+
- qwen
|
| 7 |
+
- qwen3.5
|
| 8 |
+
- recurrent
|
| 9 |
+
- linear-attention
|
| 10 |
+
- gdn
|
| 11 |
+
- cuda
|
| 12 |
+
- custom_code
|
| 13 |
+
language:
|
| 14 |
+
- en
|
| 15 |
+
- ja
|
| 16 |
+
# license: TODO — match the upstream Qwen/Qwen3.5-9B license
|
| 17 |
---
|
| 18 |
|
| 19 |
+
**Language / 言語:** [English](#english) | [日本語](#日本語)
|
| 20 |
+
|
| 21 |
+
---
|
| 22 |
+
|
| 23 |
+
<a id="english"></a>
|
| 24 |
+
|
| 25 |
+
# Qwen3.5-9B-SpeedX9-GDN32
|
| 26 |
|
| 27 |
+
A recurrent linear-attention conversion of [`Qwen/Qwen3.5-9B`](https://huggingface.co/Qwen/Qwen3.5-9B), produced with the UNI MAX toolchain and designed for fast decoding.
|
|
|
|
| 28 |
|
| 29 |
+
The source model's hybrid attention layout is converted into a fully recurrent linear-attention (GDN) runtime. Each full-attention layer is replaced by a layer initialized from a neighboring native GDN layer, then tuned with sequential layer-wise distillation.
|
|
|
|
|
|
|
|
|
|
| 30 |
|
| 31 |
+
> **Important:** This is a converted and distilled model, **not** the original `Qwen/Qwen3.5-9B` checkpoint. It is not numerically identical to the original, and quality has not been fully benchmarked (see [Limitations](#limitations)).
|
| 32 |
|
| 33 |
+
## Overview
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
|
| 35 |
+
| Item | Value |
|
| 36 |
+
| --- | --- |
|
| 37 |
+
| Base model | `Qwen/Qwen3.5-9B` |
|
| 38 |
+
| Parameters / dtype | ~9B / BF16 |
|
| 39 |
+
| Hidden layers | 32 |
|
| 40 |
+
| Runtime topology | Linear attention (all layers) |
|
| 41 |
+
| Converted full-attention layers | 3, 7, 11, 15, 19, 23, 27, 31 (8 layers) |
|
| 42 |
+
| Custom runtime | `qwen35_unimax_engine` |
|
| 43 |
+
| Kernel acceleration | CUDA kernels (enabled in the benchmark) |
|
| 44 |
+
|
| 45 |
+
## How it was made
|
| 46 |
+
|
| 47 |
+
1. **Layer replacement.** The 8 full-attention layers listed above were replaced with GDN/linear-attention layers, each initialized from a neighboring native GDN layer.
|
| 48 |
+
2. **Distillation.** The replaced layers were optimized one after another with layer-wise distillation, using **20 optimization steps per converted layer** in this build.
|
| 49 |
+
|
| 50 |
+
## Quick start
|
| 51 |
+
|
| 52 |
+
The checkpoint must be loaded with the matching `qwen35_unimax_engine` runtime. Install it first.
|
| 53 |
|
| 54 |
+
```python
|
| 55 |
+
from qwen35_unimax_engine import UniMaxPipeline
|
| 56 |
+
|
| 57 |
+
pipe = UniMaxPipeline.from_pretrained(
|
| 58 |
+
"summerMC/Qwen3.5-9B-SpeedX9-GDN32",
|
| 59 |
+
mode="auto",
|
| 60 |
+
graph_candidates=(2, 4, 8),
|
| 61 |
+
quant_modes=("fp8", "int8"),
|
| 62 |
+
quality_steps=64,
|
| 63 |
+
min_agreement=1.0,
|
| 64 |
+
use_kernels=True,
|
| 65 |
+
)
|
| 66 |
+
|
| 67 |
+
print(pipe.runtime_summary())
|
| 68 |
+
|
| 69 |
+
output = pipe(
|
| 70 |
+
"Explain recurrent linear attention:",
|
| 71 |
+
max_new_tokens=128,
|
| 72 |
+
)
|
| 73 |
+
print(output)
|
| 74 |
+
```
|
| 75 |
|
| 76 |
+
The repository is tagged `custom_code`; loading through plain `transformers` requires `trust_remote_code=True` and has not been verified here. Use the UNI MAX runtime above as the supported path.
|
|
|
|
|
|
|
| 77 |
|
| 78 |
+
### Runtime configuration
|
| 79 |
|
| 80 |
+
| Option | Value used in the benchmark |
|
| 81 |
+
| --- | --- |
|
| 82 |
+
| CUDA kernels | enabled |
|
| 83 |
+
| Graph candidates | `(2, 4, 8)` |
|
| 84 |
+
| Quantization candidates | `("fp8", "int8")` |
|
| 85 |
+
| Quality verification steps | 64 |
|
| 86 |
+
| Quality context | 64 |
|
| 87 |
+
| Minimum fast-path agreement | 1.0 |
|
| 88 |
+
|
| 89 |
+
With `mode="auto"`, the runtime may choose a different execution path depending on your hardware and the quality-verification result.
|
| 90 |
+
|
| 91 |
+
## Benchmark
|
| 92 |
+
|
| 93 |
+
Generated directly with `benchmark_uni_max()`.
|
| 94 |
+
|
| 95 |
+
**Environment**
|
| 96 |
|
| 97 |
| Item | Value |
|
| 98 |
+
| --- | --- |
|
| 99 |
| GPU | NVIDIA RTX PRO 6000 Blackwell Server Edition |
|
| 100 |
| GPU VRAM | 94.97 GiB |
|
| 101 |
| Peak allocated VRAM during benchmark | 70.14 GiB |
|
| 102 |
| PyTorch | 2.11.0+cu130 |
|
| 103 |
+
| CUDA | 13.0 (compute capability 12.0) |
|
|
|
|
| 104 |
| NVIDIA driver | 580.82.07 |
|
| 105 |
| Decode steps | 64 |
|
| 106 |
+
| Warmup / measured runs | 2 / 5 |
|
|
|
|
| 107 |
| Context lengths | 256, 1024, 4096, 16384 |
|
| 108 |
+
| Total wall time | 90.62 s |
|
| 109 |
|
| 110 |
+
**Selected configuration**
|
| 111 |
|
| 112 |
| Metric | Value |
|
| 113 |
+
| --- | --- |
|
| 114 |
+
| Selected mode | `exact` |
|
| 115 |
+
| Selected block | 8 |
|
| 116 |
+
| Verified | _(not recorded in this build)_ |
|
| 117 |
+
|
| 118 |
+
<!-- TODO: Replace the screenshot below with a table of tokens/s and latency per context length, taken from benchmark_results.json. The current card exposes the numbers only as an image. -->
|
| 119 |
|
| 120 |
+

|
| 121 |
+
|
| 122 |
+
The machine-readable record is included as `benchmark_results.json`.
|
| 123 |
+
|
| 124 |
+
Throughput, latency, kernel selection, quantization support, and memory use depend heavily on the GPU, CUDA/PyTorch versions, context length, and runtime settings. **Do not treat these numbers as hardware-independent performance claims.**
|
| 125 |
+
|
| 126 |
+
## Validation
|
| 127 |
+
|
| 128 |
+
Before publication, the checkpoint was checked with the UNI MAX checkpoint validator and the benchmark path. These layer-wise agreement checks are narrower than a full language-model evaluation suite, so please evaluate task quality yourself before deploying.
|
| 129 |
+
|
| 130 |
+
## Limitations
|
| 131 |
+
|
| 132 |
+
- Replacing full attention with recurrent linear attention changes the architecture. It can affect accuracy, long-context behavior, reasoning, generation quality, and numerical outputs relative to the original model.
|
| 133 |
+
- Only 20 distillation steps per converted layer were used in this build.
|
| 134 |
+
- No downstream task benchmarks (e.g., MMLU, long-context retrieval) are reported.
|
| 135 |
+
- Benchmark results apply only to the hardware and software environment documented above.
|
| 136 |
+
- Requires the matching custom runtime and an NVIDIA GPU with CUDA.
|
| 137 |
+
|
| 138 |
+
## Base model and license
|
| 139 |
+
|
| 140 |
+
Derived from `Qwen/Qwen3.5-9B`. Refer to the [upstream repository](https://huggingface.co/Qwen/Qwen3.5-9B) for the original documentation, license, intended use, and limitations.
|
| 141 |
+
|
| 142 |
+
---
|
| 143 |
|
| 144 |
+
<a id="日本語"></a>
|
|
|
|
|
|
|
|
|
|
| 145 |
|
| 146 |
+
# Qwen3.5-9B-SpeedX9-GDN32
|
| 147 |
|
| 148 |
+
[`Qwen/Qwen3.5-9B`](https://huggingface.co/Qwen/Qwen3.5-9B) を、UNI MAX ツールチェーンで再帰型リニアアテンション(GDN)に変換した、高速デコード向けモデルです。
|
| 149 |
|
| 150 |
+
元モデルのハイブリッドアテンション構成を、完全な再帰型リニアアテンションのランタイムに変換しています。フルアテンション層は、隣接するネイティブ GDN 層から初期化した層に置き換え、層ごとの逐次蒸留(layer-wise distillation)で調整しました。
|
| 151 |
|
| 152 |
+
> **重要:** 本モデルは変換・蒸留された派生モデルであり、元の `Qwen/Qwen3.5-9B` そのものでは**ありません**。数値的に元モデルと一致するものではなく、品質も十分には評価されていません([制限事項](#制限事項)を参照)。
|
|
|
|
| 153 |
|
| 154 |
+
## 概要
|
| 155 |
|
| 156 |
+
| 項目 | 値 |
|
| 157 |
+
| --- | --- |
|
| 158 |
+
| ベースモデル | `Qwen/Qwen3.5-9B` |
|
| 159 |
+
| パラメータ数 / dtype | 約9B / BF16 |
|
| 160 |
+
| 隠れ層数 | 32 |
|
| 161 |
+
| ランタイム構成 | リニアアテンション(全層) |
|
| 162 |
+
| 変換したフルアテンション層 | 3, 7, 11, 15, 19, 23, 27, 31(計8層) |
|
| 163 |
+
| 専用ランタイム | `qwen35_unimax_engine` |
|
| 164 |
+
| カーネル高速化 | CUDA カーネル(ベンチマーク時は有効) |
|
| 165 |
|
| 166 |
+
## 作成方法
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 167 |
|
| 168 |
+
1. **層の置き換え:** 上記8つのフルアテンション層を GDN/リニアアテンション層に置き換え、各層は隣接するネイティブ GDN 層から初期化しました。
|
| 169 |
+
2. **蒸留:** 置き換えた層を順番に層ごとの蒸留で最適化しました。今回のビルドでは**変換した各層につき 20 ステップ**です。
|
| 170 |
|
| 171 |
+
## クイックスタート
|
| 172 |
|
| 173 |
+
チェックポイントの読み込みには、対応する `qwen35_unimax_engine` ランタイムが必要です。先にインストールしてください。
|
|
|
|
| 174 |
|
| 175 |
```python
|
| 176 |
from qwen35_unimax_engine import UniMaxPipeline
|
|
|
|
| 191 |
"Explain recurrent linear attention:",
|
| 192 |
max_new_tokens=128,
|
| 193 |
)
|
|
|
|
| 194 |
print(output)
|
| 195 |
```
|
| 196 |
|
| 197 |
+
このリポジトリには `custom_code` タグが付いており、素の `transformers` で読み込む場合は `trust_remote_code=True` が必要ですが、動作は未検証です。上記の UNI MAX ランタイムを推奨パスとしてください。
|
| 198 |
|
| 199 |
+
### ランタイム設定
|
|
|
|
| 200 |
|
| 201 |
+
| オプション | ベンチマークでの値 |
|
| 202 |
+
| --- | --- |
|
| 203 |
+
| CUDA カーネル | 有効 |
|
| 204 |
+
| グラフ候補 | `(2, 4, 8)` |
|
| 205 |
+
| 量子化候補 | `("fp8", "int8")` |
|
| 206 |
+
| 品質検証ステップ数 | 64 |
|
| 207 |
+
| 品質検証コンテキスト長 | 64 |
|
| 208 |
+
| 高速パスの最小一致率 | 1.0 |
|
| 209 |
|
| 210 |
+
`mode="auto"` では、ハードウェアや品質検証の結果に応じて、別の実行パスが選ばれることがあります。
|
| 211 |
+
|
| 212 |
+
## ベンチマーク
|
| 213 |
+
|
| 214 |
+
`benchmark_uni_max()` から直接生成した結果です。
|
| 215 |
+
|
| 216 |
+
**実行環境**
|
| 217 |
+
|
| 218 |
+
| 項目 | 値 |
|
| 219 |
+
| --- | --- |
|
| 220 |
+
| GPU | NVIDIA RTX PRO 6000 Blackwell Server Edition |
|
| 221 |
+
| GPU VRAM | 94.97 GiB |
|
| 222 |
+
| ベンチマーク中の最大確保 VRAM | 70.14 GiB |
|
| 223 |
+
| PyTorch | 2.11.0+cu130 |
|
| 224 |
+
| CUDA | 13.0(compute capability 12.0) |
|
| 225 |
+
| NVIDIA ドライバ | 580.82.07 |
|
| 226 |
+
| デコードステップ数 | 64 |
|
| 227 |
+
| ウォームアップ / 計測回数 | 2 / 5 |
|
| 228 |
+
| コンテキスト長 | 256, 1024, 4096, 16384 |
|
| 229 |
+
| 総実行時間 | 90.62 秒 |
|
| 230 |
+
|
| 231 |
+
**選択された構成**
|
| 232 |
+
|
| 233 |
+
| 指標 | 値 |
|
| 234 |
+
| --- | --- |
|
| 235 |
+
| 選択モード | `exact` |
|
| 236 |
+
| 選択ブロック | 8 |
|
| 237 |
+
| verified | (このビルドでは記録なし) |
|
| 238 |
+
|
| 239 |
+
<!-- TODO: 下の画像を、benchmark_results.json に基づくコンテキスト長ごとの tokens/s・レイテンシの表に置き換える。 -->
|
| 240 |
+
|
| 241 |
+

|
| 242 |
+
|
| 243 |
+
機械可読な結果は `benchmark_results.json` に含まれています。
|
| 244 |
+
|
| 245 |
+
スループット、レイテンシ、カーネル選択、量子化対応、メモリ使用量は、GPU、CUDA/PyTorch のバージョン、コンテキスト長、ランタイム設定に大きく依存します。**これらの数値をハードウェア非依存の性能保証として扱わないでください。**
|
| 246 |
+
|
| 247 |
+
## 検証
|
| 248 |
|
| 249 |
+
公開前に、UNI MAX のチェックポイント検証ツールとベンチマーク経路で確認しています。ただし、層単位の一致確認は言語モデル全体の評価スイートよりも範囲が狭いため、運用前にご自身のタスクで品質を評価してください。
|
|
|
|
| 250 |
|
| 251 |
+
## 制限事項
|
|
|
|
|
|
|
| 252 |
|
| 253 |
+
- フルアテンションを再帰型リニアアテンションに置き換えているため、元モデルと比べて、精度、長文脈での挙動、推論能力、生成品質、数値出力が変化する可能性があります。
|
| 254 |
+
- 今回のビルドでは、変換した各層の蒸留は 20 ステップのみです。
|
| 255 |
+
- 下流タスクのベンチマーク(MMLU、長文脈検索など)は報告していません。
|
| 256 |
+
- ベンチマーク結果は、上記に記載したハードウェア・ソフトウェア環境に限られます。
|
| 257 |
+
- 専用ランタイムと、CUDA 対応の NVIDIA GPU が必要です。
|
| 258 |
|
| 259 |
+
## ベースモデルとライセンス
|
| 260 |
|
| 261 |
+
`Qwen/Qwen3.5-9B` から派生しています。元のドキュメント、ライセンス、想定用途、制限事項は[上流リポジトリ](https://huggingface.co/Qwen/Qwen3.5-9B)をご確認ください。
|
|
|
|
|
|
|
|
|