rovangju commited on
Commit
bf37ab8
·
verified ·
1 Parent(s): 43ddb52

Update README.md

Browse files

Bit of cleanup of llm fluff.

Files changed (1) hide show
  1. README.md +3 -22
README.md CHANGED
@@ -6,7 +6,7 @@ library_name: transformers
6
  base_model: ukisai/Swift-Qwen3.8-27b
7
  quantized_with:
8
  name: llm-compressor
9
- version: "0.14"
10
  tags:
11
  - compressed-tensors
12
  - w8a8
@@ -75,11 +75,6 @@ Two things were handled deliberately during quantization:
75
  original index, per-channel scale shapes checked, and a hard failure if any
76
  activation scale tensors were written) passed for all shards.
77
 
78
- ## Requirements
79
-
80
- - vLLM with compressed-tensors W8A8 INT8 support (compute capability **≥ 7.5**)
81
- - Fits on a single 64 GB GPU (loaded at ~28.5 GiB)
82
-
83
  ## Usage — vLLM
84
 
85
  ```bash
@@ -120,16 +115,6 @@ are wide; treat them as a sanity gate, not a final benchmark. The strict-match
120
  GSM8K score is a known artifact of the harness's strict filter on this
121
  template, not a model failure.
122
 
123
- ## Reproducing the quantization
124
-
125
- ```bash
126
- CUDA_VISIBLE_DEVICES=1 .venv/bin/python scripts/quantize_swift_w8a8.py
127
- ```
128
-
129
- - Base model resident in RAM (~95 GB free needed); quant ops stream each
130
- module onto a free GPU (needs ~10 GB free on `cuda:0`)
131
- - ~45 GB free disk; ~5–10 min wall time
132
-
133
  ## Appendix — benchmark serve script
134
 
135
  The vLLM run used for the benchmark numbers above (identical for both
@@ -137,12 +122,8 @@ models; the only difference between the two runs was the model path):
137
 
138
  ```bash
139
  #!/usr/bin/env bash
140
- # 03-w8a8-champ: kv-cache-dtype removed, attn backend -> FLASH_ATTN, K -> 5
141
- # generated from study.yaml by bench/study.py — do not hand-edit;
142
- # edit the run entry in study.yaml and rerun `python bench/study.py materialize`.
143
- set -euo pipefail
144
 
145
- export CUDA_VISIBLE_DEVICES=0
146
 
147
  vllm serve <MODEL> \
148
  --dtype bfloat16 \
@@ -169,4 +150,4 @@ Swift distribution.
169
  `swift-open-license-1.0`, inherited from the base model
170
  [ukisai/Swift-Qwen3.8-27b](https://huggingface.co/ukisai/Swift-Qwen3.8-27b) —
171
  see the [base model's LICENSE](https://huggingface.co/ukisai/Swift-Qwen3.8-27b/blob/main/LICENSE)
172
- for the exact terms (this quantization adds no restrictions of its own).
 
6
  base_model: ukisai/Swift-Qwen3.8-27b
7
  quantized_with:
8
  name: llm-compressor
9
+ version: '0.14'
10
  tags:
11
  - compressed-tensors
12
  - w8a8
 
75
  original index, per-channel scale shapes checked, and a hard failure if any
76
  activation scale tensors were written) passed for all shards.
77
 
 
 
 
 
 
78
  ## Usage — vLLM
79
 
80
  ```bash
 
115
  GSM8K score is a known artifact of the harness's strict filter on this
116
  template, not a model failure.
117
 
 
 
 
 
 
 
 
 
 
 
118
  ## Appendix — benchmark serve script
119
 
120
  The vLLM run used for the benchmark numbers above (identical for both
 
122
 
123
  ```bash
124
  #!/usr/bin/env bash
 
 
 
 
125
 
126
+ set -euo pipefail
127
 
128
  vllm serve <MODEL> \
129
  --dtype bfloat16 \
 
150
  `swift-open-license-1.0`, inherited from the base model
151
  [ukisai/Swift-Qwen3.8-27b](https://huggingface.co/ukisai/Swift-Qwen3.8-27b) —
152
  see the [base model's LICENSE](https://huggingface.co/ukisai/Swift-Qwen3.8-27b/blob/main/LICENSE)
153
+ for the exact terms (this quantization adds no restrictions of its own).