daavidhauser commited on
Commit
64515d3
·
verified ·
1 Parent(s): 478e487

Rename Swift 1.5 fast variant to INT4 heads

Browse files
QUANTIZATION_MANIFEST.json CHANGED
@@ -5,7 +5,7 @@
5
  "revision": "9dba8a05877150d587215a519ce6befec3c9978a",
6
  "downloaded_at": 1790420447.6224113
7
  },
8
- "variant": "fast",
9
  "preflight_only": false,
10
  "unchanged_body_tensors_verified": 2382,
11
  "body": "unchanged asymmetric AWQ INT4 group128",
@@ -138,5 +138,6 @@
138
  "drafter/capture.py": "6bb7cff9e235999923fa83b994034be6c9246b45fa7d89cca6fabf067f030e9b",
139
  "drafter/train_mtp.py": "900951feef2e0ff7d729f9baac02e9125b5ec16de5680009e368e20139b79faf",
140
  "prepare/build_draft_vocab.py": "426f371a8db9f465c48b56ae05cbc782f0613980005405deac2ad95504d088bc"
141
- }
 
142
  }
 
5
  "revision": "9dba8a05877150d587215a519ce6befec3c9978a",
6
  "downloaded_at": 1790420447.6224113
7
  },
8
+ "variant": "INT4 heads",
9
  "preflight_only": false,
10
  "unchanged_body_tensors_verified": 2382,
11
  "body": "unchanged asymmetric AWQ INT4 group128",
 
138
  "drafter/capture.py": "6bb7cff9e235999923fa83b994034be6c9246b45fa7d89cca6fabf067f030e9b",
139
  "drafter/train_mtp.py": "900951feef2e0ff7d729f9baac02e9125b5ec16de5680009e368e20139b79faf",
140
  "prepare/build_draft_vocab.py": "426f371a8db9f465c48b56ae05cbc782f0613980005405deac2ad95504d088bc"
141
+ },
142
+ "original_variant": "fast"
143
  }
README.md CHANGED
@@ -16,7 +16,7 @@ tags:
16
 
17
  Model card written by GPT-6 Astra, refined with human feedback:
18
 
19
- # Swift 1.5 HyperQwen-fast
20
 
21
  Swift 1.5's efficient reasoning, combined with HyperQwen's optimized serving, **delivered 37% lower average task completion time** than HyperQwen serving the Qwen fast checkpoint in our RTX 3090 evaluation.
22
 
@@ -36,7 +36,7 @@ Upstream checkpoint: [ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ](https://huggingfac
36
 
37
  ## Performance
38
 
39
- | Measurement | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 fast |
40
  |---|---:|---:|---:|---:|
41
  | **Average request time ↓** | **108.1 s** | **66.2 s** | **72.2 s** | **68.2 s** |
42
  | **Average output tokens/task ↓** | **8,985** | **5,245** | **5,751** | **5,669** |
@@ -45,7 +45,7 @@ Upstream checkpoint: [ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ](https://huggingfac
45
 
46
  ## Quality
47
 
48
- | Test | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 fast |
49
  |---|---:|---:|---:|---:|
50
  | **GSM8K — 200-question subset** | 97.5% | 98.0% | 98.0% | 97.5% |
51
  | **IFBench — 300 prompts, strict** | 74.0% | 73.3% | 73.7% | 72.3% |
@@ -60,7 +60,7 @@ Upstream checkpoint: [ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ](https://huggingfac
60
  - **Tool-call/JSON checks:** 20 custom weather-tool tasks checking the function name, arguments, Celsius-to-Fahrenheit conversion and final JSON; plus 10 JSON inventory-filtering tasks. These are integration checks, not an external agent benchmark.
61
  - **Perplexity:** 18,729 scored tokens from English Wikipedia and Python source; lower is better. Recomputed from saved token log-probabilities, excluding Danish; see [subset results](evaluation/perplexity-english-python.json).
62
 
63
- Compared with the Swift 1.5 INT8-head variant, this fast conversion measured **3.2% higher decode TPS** and **5.5% lower average request time**; benchmark score changes were mixed.
64
 
65
  **Evaluation setup:** RTX 3090 24 GB; FP8 KV cache; 150,000-token configured context; 128,000 output tokens per call. These are runtime settings, not fixed model properties, and the TPS test uses short prompts. Task evaluation used two concurrent requests; TPS was measured with one request at a time. All models used the same serving settings and task budgets. Thinking tests used xhigh effort, temperature 1.0, top_p 0.95, top_k 20 and seed 15027; GSM8K/tool checks were greedy.
66
 
@@ -69,8 +69,8 @@ Compared with the Swift 1.5 INT8-head variant, this fast conversion measured **3
69
  Requires the **patched HyperQwen runtime**, not stock vLLM or GGUF tools. From an installed HyperQwen checkout:
70
 
71
  ```bash
72
- hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-fast --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-fast
73
- MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-fast" CTX=long MAX_LEN=150000 SPEC=mtp \
74
  bash single-user/start_qwen.sh
75
  ```
76
 
 
16
 
17
  Model card written by GPT-6 Astra, refined with human feedback:
18
 
19
+ # Swift 1.5 HyperQwen — INT4 heads
20
 
21
  Swift 1.5's efficient reasoning, combined with HyperQwen's optimized serving, **delivered 37% lower average task completion time** than HyperQwen serving the Qwen fast checkpoint in our RTX 3090 evaluation.
22
 
 
36
 
37
  ## Performance
38
 
39
+ | Measurement | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 INT4 heads |
40
  |---|---:|---:|---:|---:|
41
  | **Average request time ↓** | **108.1 s** | **66.2 s** | **72.2 s** | **68.2 s** |
42
  | **Average output tokens/task ↓** | **8,985** | **5,245** | **5,751** | **5,669** |
 
45
 
46
  ## Quality
47
 
48
+ | Test | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 INT4 heads |
49
  |---|---:|---:|---:|---:|
50
  | **GSM8K — 200-question subset** | 97.5% | 98.0% | 98.0% | 97.5% |
51
  | **IFBench — 300 prompts, strict** | 74.0% | 73.3% | 73.7% | 72.3% |
 
60
  - **Tool-call/JSON checks:** 20 custom weather-tool tasks checking the function name, arguments, Celsius-to-Fahrenheit conversion and final JSON; plus 10 JSON inventory-filtering tasks. These are integration checks, not an external agent benchmark.
61
  - **Perplexity:** 18,729 scored tokens from English Wikipedia and Python source; lower is better. Recomputed from saved token log-probabilities, excluding Danish; see [subset results](evaluation/perplexity-english-python.json).
62
 
63
+ Compared with the Swift 1.5 INT8-head variant, this INT4-head conversion measured **3.2% higher decode TPS** and **5.5% lower average request time**; benchmark score changes were mixed.
64
 
65
  **Evaluation setup:** RTX 3090 24 GB; FP8 KV cache; 150,000-token configured context; 128,000 output tokens per call. These are runtime settings, not fixed model properties, and the TPS test uses short prompts. Task evaluation used two concurrent requests; TPS was measured with one request at a time. All models used the same serving settings and task budgets. Thinking tests used xhigh effort, temperature 1.0, top_p 0.95, top_k 20 and seed 15027; GSM8K/tool checks were greedy.
66
 
 
69
  Requires the **patched HyperQwen runtime**, not stock vLLM or GGUF tools. From an installed HyperQwen checkout:
70
 
71
  ```bash
72
+ hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4 --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4
73
+ MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4" CTX=long MAX_LEN=150000 SPEC=mtp \
74
  bash single-user/start_qwen.sh
75
  ```
76
 
RELEASE_MANIFEST.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "repo": "daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-fast",
3
  "files": {
4
  "LICENSE": {
5
  "size": 13306,
@@ -14,16 +14,16 @@
14
  "sha256": "e0e788911888bdcc087d34fc12dc5bf3818034762334de3575ea34d9559b186f"
15
  },
16
  "QUANTIZATION_MANIFEST.json": {
17
- "size": 8398,
18
- "sha256": "895d4e0dd27f896ed2a3a0e2160c59967cd10e8506f690b6a85c5cf90c529cec"
19
  },
20
  "README.md": {
21
- "size": 5593,
22
- "sha256": "100dec87d8127bb5d22cbb467c8442a24ef070c852b9089ce51fffb09435c7df"
23
  },
24
  "RUNTIME.md": {
25
  "size": 2608,
26
- "sha256": "07b3d0091fee7fb8edf752fa2f3bb9118afe514eaba86291b88796aac0f906f5"
27
  },
28
  "chat_template.jinja": {
29
  "size": 8952,
 
1
  {
2
+ "repo": "daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4",
3
  "files": {
4
  "LICENSE": {
5
  "size": 13306,
 
14
  "sha256": "e0e788911888bdcc087d34fc12dc5bf3818034762334de3575ea34d9559b186f"
15
  },
16
  "QUANTIZATION_MANIFEST.json": {
17
+ "size": 8434,
18
+ "sha256": "40ebbba877dd55e1c512e68785f3127673d57ac799c5ea7139761b563ed7de4b"
19
  },
20
  "README.md": {
21
+ "size": 5620,
22
+ "sha256": "de13ff8b19a75a4d6dcda0c9a36e276d3a5b19b74a44ef012bbfe39562f24a85"
23
  },
24
  "RUNTIME.md": {
25
  "size": 2608,
26
+ "sha256": "4a6a153fdf00ece6ddd976ee745951db61c5a36a69ba06eaf102a1ec1f5c3c74"
27
  },
28
  "chat_template.jinja": {
29
  "size": 8952,
RUNTIME.md CHANGED
@@ -15,13 +15,13 @@ An independent clean-machine installation has not been tested.
15
  ## Download and launch from an installed HyperQwen checkout
16
 
17
  ```bash
18
- hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-fast \
19
- --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-fast
20
 
21
  # In the HyperQwen checkout; use the launcher snapshot accompanying the model.
22
- cp models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-fast/runtime/single-user/start_qwen.sh single-user/start_qwen.sh
23
 
24
- MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-fast" \
25
  CTX=long MAX_LEN=150000 SPEC=mtp DRAFT_TOKENS=3 \
26
  PREFIX_CACHE=1 TOOLS=1 VISION=1 VISION_OFFLOAD=1 \
27
  GPU_UTIL=0.93 MAX_SEQS=8 API_SERVERS=1 HOST=127.0.0.1 PORT=18020 \
 
15
  ## Download and launch from an installed HyperQwen checkout
16
 
17
  ```bash
18
+ hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4 \
19
+ --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4
20
 
21
  # In the HyperQwen checkout; use the launcher snapshot accompanying the model.
22
+ cp models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4/runtime/single-user/start_qwen.sh single-user/start_qwen.sh
23
 
24
+ MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4" \
25
  CTX=long MAX_LEN=150000 SPEC=mtp DRAFT_TOKENS=3 \
26
  PREFIX_CACHE=1 TOOLS=1 VISION=1 VISION_OFFLOAD=1 \
27
  GPU_UTIL=0.93 MAX_SEQS=8 API_SERVERS=1 HOST=127.0.0.1 PORT=18020 \