Download RUNTIME.md from daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen: direct link, hf CLI and curl.
- Browser
- Download file 2.59 kB
-
https://huggingface.co/daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/resolve/main/RUNTIME.md
- Command line
-
hf download hf://daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/RUNTIME.md
-
curl -L -o RUNTIME.md https://huggingface.co/daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/resolve/main/RUNTIME.md
HyperQwen runtime
This checkpoint requires the patched HyperQwen vLLM runtime. It is not a GGUF.
Tested HyperQwen base revision: 253c76aea0a240bf7cb2bd3ed92b672c0f258d0d. Exact local launcher snapshots and
patch files are in runtime/; calibration/evaluation scripts are included separately.
Tested packages: vLLM 0.27.1, PyTorch 2.13.0, Transformers 5.15.0,
compressed-tensors 0.17.0, safetensors 0.8.0. Follow the
pinned HyperQwen setup instructions
for the compiler, attention libraries, and patch installation. Use the supplied
runtime/patches/ set when applying vLLM patches, and the supplied launcher for
the final serve command. Do not requantize this already prepared model.
An independent clean-machine installation has not been tested.
Download and launch from an installed HyperQwen checkout
hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen \
--local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen
# In the HyperQwen checkout; use the launcher snapshot accompanying the model.
cp models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/runtime/single-user/start_qwen.sh single-user/start_qwen.sh
MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen" \
CTX=long MAX_LEN=150000 SPEC=mtp DRAFT_TOKENS=3 \
PREFIX_CACHE=1 TOOLS=1 VISION=1 VISION_OFFLOAD=1 \
GPU_UTIL=0.93 MAX_SEQS=8 API_SERVERS=1 HOST=127.0.0.1 PORT=18020 \
VLLM_MAMBA_ALIGN_KEEP_CHECKPOINTS=1 FLASHINFER_DISABLE_VERSION_CHECK=1 \
EXTRA_ARGS='--limit-mm-per-prompt {"image":{"count":10}}' \
bash single-user/start_qwen.sh
The tested run did not enable INT8 activation quantization. Remove conflicting
INT8_ACT, PREFILL_ATTN, or other performance overrides from your local environment
and .env when reproducing it. The launcher uses FP8 KV and FP16 recurrent state
in this profile. Test memory on your own stack before assuming 150k capacity.
Use model name qwen3.8-27b in API requests. Authentication follows HyperQwen's
normal api_key.txt / VLLM_API_KEY configuration.
Evaluation requests
Thinking: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0, repetition_penalty 1.0, seed 15027, enable_thinking true, reasoning_effort xhigh. Output limit: 128000 tokens. Nonthinking GSM8K/tool tests were greedy. The quality suite used two concurrent requests; the speed test used one.
The final published campaign validates this single-user configuration only. Separate batch launchers/INT8 activation settings are not benchmarked by these results.