Rename Swift 1.5 fast variant to INT4 heads
Browse files- QUANTIZATION_MANIFEST.json +3 -2
- README.md +6 -6
- RELEASE_MANIFEST.json +6 -6
- RUNTIME.md +4 -4
QUANTIZATION_MANIFEST.json
CHANGED
|
@@ -5,7 +5,7 @@
|
|
| 5 |
"revision": "9dba8a05877150d587215a519ce6befec3c9978a",
|
| 6 |
"downloaded_at": 1790420447.6224113
|
| 7 |
},
|
| 8 |
-
"variant": "
|
| 9 |
"preflight_only": false,
|
| 10 |
"unchanged_body_tensors_verified": 2382,
|
| 11 |
"body": "unchanged asymmetric AWQ INT4 group128",
|
|
@@ -138,5 +138,6 @@
|
|
| 138 |
"drafter/capture.py": "6bb7cff9e235999923fa83b994034be6c9246b45fa7d89cca6fabf067f030e9b",
|
| 139 |
"drafter/train_mtp.py": "900951feef2e0ff7d729f9baac02e9125b5ec16de5680009e368e20139b79faf",
|
| 140 |
"prepare/build_draft_vocab.py": "426f371a8db9f465c48b56ae05cbc782f0613980005405deac2ad95504d088bc"
|
| 141 |
-
}
|
|
|
|
| 142 |
}
|
|
|
|
| 5 |
"revision": "9dba8a05877150d587215a519ce6befec3c9978a",
|
| 6 |
"downloaded_at": 1790420447.6224113
|
| 7 |
},
|
| 8 |
+
"variant": "INT4 heads",
|
| 9 |
"preflight_only": false,
|
| 10 |
"unchanged_body_tensors_verified": 2382,
|
| 11 |
"body": "unchanged asymmetric AWQ INT4 group128",
|
|
|
|
| 138 |
"drafter/capture.py": "6bb7cff9e235999923fa83b994034be6c9246b45fa7d89cca6fabf067f030e9b",
|
| 139 |
"drafter/train_mtp.py": "900951feef2e0ff7d729f9baac02e9125b5ec16de5680009e368e20139b79faf",
|
| 140 |
"prepare/build_draft_vocab.py": "426f371a8db9f465c48b56ae05cbc782f0613980005405deac2ad95504d088bc"
|
| 141 |
+
},
|
| 142 |
+
"original_variant": "fast"
|
| 143 |
}
|
README.md
CHANGED
|
@@ -16,7 +16,7 @@ tags:
|
|
| 16 |
|
| 17 |
Model card written by GPT-6 Astra, refined with human feedback:
|
| 18 |
|
| 19 |
-
# Swift 1.5 HyperQwen
|
| 20 |
|
| 21 |
Swift 1.5's efficient reasoning, combined with HyperQwen's optimized serving, **delivered 37% lower average task completion time** than HyperQwen serving the Qwen fast checkpoint in our RTX 3090 evaluation.
|
| 22 |
|
|
@@ -36,7 +36,7 @@ Upstream checkpoint: [ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ](https://huggingfac
|
|
| 36 |
|
| 37 |
## Performance
|
| 38 |
|
| 39 |
-
| Measurement | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5
|
| 40 |
|---|---:|---:|---:|---:|
|
| 41 |
| **Average request time ↓** | **108.1 s** | **66.2 s** | **72.2 s** | **68.2 s** |
|
| 42 |
| **Average output tokens/task ↓** | **8,985** | **5,245** | **5,751** | **5,669** |
|
|
@@ -45,7 +45,7 @@ Upstream checkpoint: [ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ](https://huggingfac
|
|
| 45 |
|
| 46 |
## Quality
|
| 47 |
|
| 48 |
-
| Test | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5
|
| 49 |
|---|---:|---:|---:|---:|
|
| 50 |
| **GSM8K — 200-question subset** | 97.5% | 98.0% | 98.0% | 97.5% |
|
| 51 |
| **IFBench — 300 prompts, strict** | 74.0% | 73.3% | 73.7% | 72.3% |
|
|
@@ -60,7 +60,7 @@ Upstream checkpoint: [ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ](https://huggingfac
|
|
| 60 |
- **Tool-call/JSON checks:** 20 custom weather-tool tasks checking the function name, arguments, Celsius-to-Fahrenheit conversion and final JSON; plus 10 JSON inventory-filtering tasks. These are integration checks, not an external agent benchmark.
|
| 61 |
- **Perplexity:** 18,729 scored tokens from English Wikipedia and Python source; lower is better. Recomputed from saved token log-probabilities, excluding Danish; see [subset results](evaluation/perplexity-english-python.json).
|
| 62 |
|
| 63 |
-
Compared with the Swift 1.5 INT8-head variant, this
|
| 64 |
|
| 65 |
**Evaluation setup:** RTX 3090 24 GB; FP8 KV cache; 150,000-token configured context; 128,000 output tokens per call. These are runtime settings, not fixed model properties, and the TPS test uses short prompts. Task evaluation used two concurrent requests; TPS was measured with one request at a time. All models used the same serving settings and task budgets. Thinking tests used xhigh effort, temperature 1.0, top_p 0.95, top_k 20 and seed 15027; GSM8K/tool checks were greedy.
|
| 66 |
|
|
@@ -69,8 +69,8 @@ Compared with the Swift 1.5 INT8-head variant, this fast conversion measured **3
|
|
| 69 |
Requires the **patched HyperQwen runtime**, not stock vLLM or GGUF tools. From an installed HyperQwen checkout:
|
| 70 |
|
| 71 |
```bash
|
| 72 |
-
hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-
|
| 73 |
-
MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-
|
| 74 |
bash single-user/start_qwen.sh
|
| 75 |
```
|
| 76 |
|
|
|
|
| 16 |
|
| 17 |
Model card written by GPT-6 Astra, refined with human feedback:
|
| 18 |
|
| 19 |
+
# Swift 1.5 HyperQwen — INT4 heads
|
| 20 |
|
| 21 |
Swift 1.5's efficient reasoning, combined with HyperQwen's optimized serving, **delivered 37% lower average task completion time** than HyperQwen serving the Qwen fast checkpoint in our RTX 3090 evaluation.
|
| 22 |
|
|
|
|
| 36 |
|
| 37 |
## Performance
|
| 38 |
|
| 39 |
+
| Measurement | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 INT4 heads |
|
| 40 |
|---|---:|---:|---:|---:|
|
| 41 |
| **Average request time ↓** | **108.1 s** | **66.2 s** | **72.2 s** | **68.2 s** |
|
| 42 |
| **Average output tokens/task ↓** | **8,985** | **5,245** | **5,751** | **5,669** |
|
|
|
|
| 45 |
|
| 46 |
## Quality
|
| 47 |
|
| 48 |
+
| Test | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 INT4 heads |
|
| 49 |
|---|---:|---:|---:|---:|
|
| 50 |
| **GSM8K — 200-question subset** | 97.5% | 98.0% | 98.0% | 97.5% |
|
| 51 |
| **IFBench — 300 prompts, strict** | 74.0% | 73.3% | 73.7% | 72.3% |
|
|
|
|
| 60 |
- **Tool-call/JSON checks:** 20 custom weather-tool tasks checking the function name, arguments, Celsius-to-Fahrenheit conversion and final JSON; plus 10 JSON inventory-filtering tasks. These are integration checks, not an external agent benchmark.
|
| 61 |
- **Perplexity:** 18,729 scored tokens from English Wikipedia and Python source; lower is better. Recomputed from saved token log-probabilities, excluding Danish; see [subset results](evaluation/perplexity-english-python.json).
|
| 62 |
|
| 63 |
+
Compared with the Swift 1.5 INT8-head variant, this INT4-head conversion measured **3.2% higher decode TPS** and **5.5% lower average request time**; benchmark score changes were mixed.
|
| 64 |
|
| 65 |
**Evaluation setup:** RTX 3090 24 GB; FP8 KV cache; 150,000-token configured context; 128,000 output tokens per call. These are runtime settings, not fixed model properties, and the TPS test uses short prompts. Task evaluation used two concurrent requests; TPS was measured with one request at a time. All models used the same serving settings and task budgets. Thinking tests used xhigh effort, temperature 1.0, top_p 0.95, top_k 20 and seed 15027; GSM8K/tool checks were greedy.
|
| 66 |
|
|
|
|
| 69 |
Requires the **patched HyperQwen runtime**, not stock vLLM or GGUF tools. From an installed HyperQwen checkout:
|
| 70 |
|
| 71 |
```bash
|
| 72 |
+
hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4 --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4
|
| 73 |
+
MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4" CTX=long MAX_LEN=150000 SPEC=mtp \
|
| 74 |
bash single-user/start_qwen.sh
|
| 75 |
```
|
| 76 |
|
RELEASE_MANIFEST.json
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
{
|
| 2 |
-
"repo": "daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-
|
| 3 |
"files": {
|
| 4 |
"LICENSE": {
|
| 5 |
"size": 13306,
|
|
@@ -14,16 +14,16 @@
|
|
| 14 |
"sha256": "e0e788911888bdcc087d34fc12dc5bf3818034762334de3575ea34d9559b186f"
|
| 15 |
},
|
| 16 |
"QUANTIZATION_MANIFEST.json": {
|
| 17 |
-
"size":
|
| 18 |
-
"sha256": "
|
| 19 |
},
|
| 20 |
"README.md": {
|
| 21 |
-
"size":
|
| 22 |
-
"sha256": "
|
| 23 |
},
|
| 24 |
"RUNTIME.md": {
|
| 25 |
"size": 2608,
|
| 26 |
-
"sha256": "
|
| 27 |
},
|
| 28 |
"chat_template.jinja": {
|
| 29 |
"size": 8952,
|
|
|
|
| 1 |
{
|
| 2 |
+
"repo": "daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4",
|
| 3 |
"files": {
|
| 4 |
"LICENSE": {
|
| 5 |
"size": 13306,
|
|
|
|
| 14 |
"sha256": "e0e788911888bdcc087d34fc12dc5bf3818034762334de3575ea34d9559b186f"
|
| 15 |
},
|
| 16 |
"QUANTIZATION_MANIFEST.json": {
|
| 17 |
+
"size": 8434,
|
| 18 |
+
"sha256": "40ebbba877dd55e1c512e68785f3127673d57ac799c5ea7139761b563ed7de4b"
|
| 19 |
},
|
| 20 |
"README.md": {
|
| 21 |
+
"size": 5620,
|
| 22 |
+
"sha256": "de13ff8b19a75a4d6dcda0c9a36e276d3a5b19b74a44ef012bbfe39562f24a85"
|
| 23 |
},
|
| 24 |
"RUNTIME.md": {
|
| 25 |
"size": 2608,
|
| 26 |
+
"sha256": "4a6a153fdf00ece6ddd976ee745951db61c5a36a69ba06eaf102a1ec1f5c3c74"
|
| 27 |
},
|
| 28 |
"chat_template.jinja": {
|
| 29 |
"size": 8952,
|
RUNTIME.md
CHANGED
|
@@ -15,13 +15,13 @@ An independent clean-machine installation has not been tested.
|
|
| 15 |
## Download and launch from an installed HyperQwen checkout
|
| 16 |
|
| 17 |
```bash
|
| 18 |
-
hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-
|
| 19 |
-
--local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-
|
| 20 |
|
| 21 |
# In the HyperQwen checkout; use the launcher snapshot accompanying the model.
|
| 22 |
-
cp models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-
|
| 23 |
|
| 24 |
-
MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-
|
| 25 |
CTX=long MAX_LEN=150000 SPEC=mtp DRAFT_TOKENS=3 \
|
| 26 |
PREFIX_CACHE=1 TOOLS=1 VISION=1 VISION_OFFLOAD=1 \
|
| 27 |
GPU_UTIL=0.93 MAX_SEQS=8 API_SERVERS=1 HOST=127.0.0.1 PORT=18020 \
|
|
|
|
| 15 |
## Download and launch from an installed HyperQwen checkout
|
| 16 |
|
| 17 |
```bash
|
| 18 |
+
hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4 \
|
| 19 |
+
--local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4
|
| 20 |
|
| 21 |
# In the HyperQwen checkout; use the launcher snapshot accompanying the model.
|
| 22 |
+
cp models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4/runtime/single-user/start_qwen.sh single-user/start_qwen.sh
|
| 23 |
|
| 24 |
+
MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4" \
|
| 25 |
CTX=long MAX_LEN=150000 SPEC=mtp DRAFT_TOKENS=3 \
|
| 26 |
PREFIX_CACHE=1 TOOLS=1 VISION=1 VISION_OFFLOAD=1 \
|
| 27 |
GPU_UTIL=0.93 MAX_SEQS=8 API_SERVERS=1 HOST=127.0.0.1 PORT=18020 \
|