# HyperQwen runtime This checkpoint requires the patched HyperQwen vLLM runtime. It is not a GGUF. Tested HyperQwen base revision: `253c76aea0a240bf7cb2bd3ed92b672c0f258d0d`. Exact local launcher snapshots and patch files are in `runtime/`; calibration/evaluation scripts are included separately. Tested packages: vLLM 0.27.1, PyTorch 2.13.0, Transformers 5.15.0, compressed-tensors 0.17.0, safetensors 0.8.0. Follow the [pinned HyperQwen setup instructions](https://github.com/syv-ai/HyperQwen/blob/253c76aea0a240bf7cb2bd3ed92b672c0f258d0d/README.md#setup) for the compiler, attention libraries, and patch installation. Use the supplied `runtime/patches/` set when applying vLLM patches, and the supplied launcher for the final serve command. Do not requantize this already prepared model. An independent clean-machine installation has not been tested. ## Download and launch from an installed HyperQwen checkout ```bash hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen \ --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen # In the HyperQwen checkout; use the launcher snapshot accompanying the model. cp models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/runtime/single-user/start_qwen.sh single-user/start_qwen.sh MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen" \ CTX=long MAX_LEN=150000 SPEC=mtp DRAFT_TOKENS=3 \ PREFIX_CACHE=1 TOOLS=1 VISION=1 VISION_OFFLOAD=1 \ GPU_UTIL=0.93 MAX_SEQS=8 API_SERVERS=1 HOST=127.0.0.1 PORT=18020 \ VLLM_MAMBA_ALIGN_KEEP_CHECKPOINTS=1 FLASHINFER_DISABLE_VERSION_CHECK=1 \ EXTRA_ARGS='--limit-mm-per-prompt {"image":{"count":10}}' \ bash single-user/start_qwen.sh ``` The tested run did not enable INT8 activation quantization. Remove conflicting INT8_ACT, PREFILL_ATTN, or other performance overrides from your local environment and `.env` when reproducing it. The launcher uses FP8 KV and FP16 recurrent state in this profile. Test memory on your own stack before assuming 150k capacity. Use model name `qwen3.8-27b` in API requests. Authentication follows HyperQwen's normal `api_key.txt` / `VLLM_API_KEY` configuration. ## Evaluation requests Thinking: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0, repetition_penalty 1.0, seed 15027, enable_thinking true, reasoning_effort xhigh. Output limit: 128000 tokens. Nonthinking GSM8K/tool tests were greedy. The quality suite used two concurrent requests; the speed test used one. The final published campaign validates this single-user configuration only. Separate batch launchers/INT8 activation settings are not benchmarked by these results.