|
Download RUNTIME.md from daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen: direct link, hf CLI and curl.
- Browser
- Download file 2.59 kB
-
https://huggingface.co/daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/resolve/main/RUNTIME.md
- Command line
-
hf download hf://daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/RUNTIME.md
-
curl -L -o RUNTIME.md https://huggingface.co/daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/resolve/main/RUNTIME.md
2.59 kB
| # HyperQwen runtime | |
| This checkpoint requires the patched HyperQwen vLLM runtime. It is not a GGUF. | |
| Tested HyperQwen base revision: `253c76aea0a240bf7cb2bd3ed92b672c0f258d0d`. Exact local launcher snapshots and | |
| patch files are in `runtime/`; calibration/evaluation scripts are included separately. | |
| Tested packages: vLLM 0.27.1, PyTorch 2.13.0, Transformers 5.15.0, | |
| compressed-tensors 0.17.0, safetensors 0.8.0. Follow the | |
| [pinned HyperQwen setup instructions](https://github.com/syv-ai/HyperQwen/blob/253c76aea0a240bf7cb2bd3ed92b672c0f258d0d/README.md#setup) | |
| for the compiler, attention libraries, and patch installation. Use the supplied | |
| `runtime/patches/` set when applying vLLM patches, and the supplied launcher for | |
| the final serve command. Do not requantize this already prepared model. | |
| An independent clean-machine installation has not been tested. | |
| ## Download and launch from an installed HyperQwen checkout | |
| ```bash | |
| hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen \ | |
| --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen | |
| # In the HyperQwen checkout; use the launcher snapshot accompanying the model. | |
| cp models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen/runtime/single-user/start_qwen.sh single-user/start_qwen.sh | |
| MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen" \ | |
| CTX=long MAX_LEN=150000 SPEC=mtp DRAFT_TOKENS=3 \ | |
| PREFIX_CACHE=1 TOOLS=1 VISION=1 VISION_OFFLOAD=1 \ | |
| GPU_UTIL=0.93 MAX_SEQS=8 API_SERVERS=1 HOST=127.0.0.1 PORT=18020 \ | |
| VLLM_MAMBA_ALIGN_KEEP_CHECKPOINTS=1 FLASHINFER_DISABLE_VERSION_CHECK=1 \ | |
| EXTRA_ARGS='--limit-mm-per-prompt {"image":{"count":10}}' \ | |
| bash single-user/start_qwen.sh | |
| ``` | |
| The tested run did not enable INT8 activation quantization. Remove conflicting | |
| INT8_ACT, PREFILL_ATTN, or other performance overrides from your local environment | |
| and `.env` when reproducing it. The launcher uses FP8 KV and FP16 recurrent state | |
| in this profile. Test memory on your own stack before assuming 150k capacity. | |
| Use model name `qwen3.8-27b` in API requests. Authentication follows HyperQwen's | |
| normal `api_key.txt` / `VLLM_API_KEY` configuration. | |
| ## Evaluation requests | |
| Thinking: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0, | |
| repetition_penalty 1.0, seed 15027, enable_thinking true, reasoning_effort xhigh. | |
| Output limit: 128000 tokens. Nonthinking GSM8K/tool tests were greedy. | |
| The quality suite used two concurrent requests; the speed test used one. | |
| The final published campaign validates this single-user configuration only. | |
| Separate batch launchers/INT8 activation settings are not benchmarked by these results. | |