File size: 2,608 Bytes
1613e65 64515d3 1613e65 64515d3 1613e65 64515d3 1613e65 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 | # HyperQwen runtime
This checkpoint requires the patched HyperQwen vLLM runtime. It is not a GGUF.
Tested HyperQwen base revision: `253c76aea0a240bf7cb2bd3ed92b672c0f258d0d`. Exact local launcher snapshots and
patch files are in `runtime/`; calibration/evaluation scripts are included separately.
Tested packages: vLLM 0.27.1, PyTorch 2.13.0, Transformers 5.15.0,
compressed-tensors 0.17.0, safetensors 0.8.0. Follow the
[pinned HyperQwen setup instructions](https://github.com/syv-ai/HyperQwen/blob/253c76aea0a240bf7cb2bd3ed92b672c0f258d0d/README.md#setup)
for the compiler, attention libraries, and patch installation. Use the supplied
`runtime/patches/` set when applying vLLM patches, and the supplied launcher for
the final serve command. Do not requantize this already prepared model.
An independent clean-machine installation has not been tested.
## Download and launch from an installed HyperQwen checkout
```bash
hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4 \
--local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4
# In the HyperQwen checkout; use the launcher snapshot accompanying the model.
cp models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4/runtime/single-user/start_qwen.sh single-user/start_qwen.sh
MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4" \
CTX=long MAX_LEN=150000 SPEC=mtp DRAFT_TOKENS=3 \
PREFIX_CACHE=1 TOOLS=1 VISION=1 VISION_OFFLOAD=1 \
GPU_UTIL=0.93 MAX_SEQS=8 API_SERVERS=1 HOST=127.0.0.1 PORT=18020 \
VLLM_MAMBA_ALIGN_KEEP_CHECKPOINTS=1 FLASHINFER_DISABLE_VERSION_CHECK=1 \
EXTRA_ARGS='--limit-mm-per-prompt {"image":{"count":10}}' \
bash single-user/start_qwen.sh
```
The tested run did not enable INT8 activation quantization. Remove conflicting
INT8_ACT, PREFILL_ATTN, or other performance overrides from your local environment
and `.env` when reproducing it. The launcher uses FP8 KV and FP16 recurrent state
in this profile. Test memory on your own stack before assuming 150k capacity.
Use model name `qwen3.8-27b` in API requests. Authentication follows HyperQwen's
normal `api_key.txt` / `VLLM_API_KEY` configuration.
## Evaluation requests
Thinking: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0,
repetition_penalty 1.0, seed 15027, enable_thinking true, reasoning_effort xhigh.
Output limit: 128000 tokens. Nonthinking GSM8K/tool tests were greedy.
The quality suite used two concurrent requests; the speed test used one.
The final published campaign validates this single-user configuration only.
Separate batch launchers/INT8 activation settings are not benchmarked by these results.
|