daavidhauser's picture
Publish Swift HyperQwen collection with performance and quality comparisons
2bc6021 verified
|
Raw History Blame Contribute Delete
2.59 kB

HyperQwen runtime

This checkpoint requires the patched HyperQwen vLLM runtime. It is not a GGUF. Tested HyperQwen base revision: 253c76aea0a240bf7cb2bd3ed92b672c0f258d0d. Exact local launcher snapshots and patch files are in runtime/; calibration/evaluation scripts are included separately.

Tested packages: vLLM 0.27.1, PyTorch 2.13.0, Transformers 5.15.0, compressed-tensors 0.17.0, safetensors 0.8.0. Follow the pinned HyperQwen setup instructions for the compiler, attention libraries, and patch installation. Use the supplied runtime/patches/ set when applying vLLM patches, and the supplied launcher for the final serve command. Do not requantize this already prepared model. An independent clean-machine installation has not been tested.

Download and launch from an installed HyperQwen checkout

hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen \
  --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen

# In the HyperQwen checkout; use the launcher snapshot accompanying the model.
cp models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/runtime/single-user/start_qwen.sh single-user/start_qwen.sh

MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen" \
CTX=long MAX_LEN=150000 SPEC=mtp DRAFT_TOKENS=3 \
PREFIX_CACHE=1 TOOLS=1 VISION=1 VISION_OFFLOAD=1 \
GPU_UTIL=0.93 MAX_SEQS=8 API_SERVERS=1 HOST=127.0.0.1 PORT=18020 \
VLLM_MAMBA_ALIGN_KEEP_CHECKPOINTS=1 FLASHINFER_DISABLE_VERSION_CHECK=1 \
EXTRA_ARGS='--limit-mm-per-prompt {"image":{"count":10}}' \
  bash single-user/start_qwen.sh

The tested run did not enable INT8 activation quantization. Remove conflicting INT8_ACT, PREFILL_ATTN, or other performance overrides from your local environment and .env when reproducing it. The launcher uses FP8 KV and FP16 recurrent state in this profile. Test memory on your own stack before assuming 150k capacity. Use model name qwen3.8-27b in API requests. Authentication follows HyperQwen's normal api_key.txt / VLLM_API_KEY configuration.

Evaluation requests

Thinking: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0, repetition_penalty 1.0, seed 15027, enable_thinking true, reasoning_effort xhigh. Output limit: 128000 tokens. Nonthinking GSM8K/tool tests were greedy. The quality suite used two concurrent requests; the speed test used one.

The final published campaign validates this single-user configuration only. Separate batch launchers/INT8 activation settings are not benchmarked by these results.