--- license: apache-2.0 base_model: Umranz/Hydra-Turbo tags: - gguf - llama.cpp - qwen35 - quantized - text-generation pipeline_tag: text-generation --- # Hydra-Turbo-GGUF This repository contains official GGUF quantized weights for [Umranz/Hydra-Turbo](https://huggingface.co/Umranz/Hydra-Turbo), a 9B parameter model based on the Qwen 3.5 architecture. The GGUF files in this repository were converted using the latest build of `llama.cpp` with the `--no-nextn` conversion flag to ensure compatibility across local inference runtimes. --- ## Model Overview - **Base Model**: Umranz/Hydra-Turbo - **Architecture**: Qwen 3.5 (Hybrid Linear Attention + Full Attention) - **Parameters**: 8.95B (9.0B label) - **Context Length**: 262,144 tokens - **Vocabulary Size**: 248,320 - **License**: Apache 2.0 --- ## Available Files and Quantization Formats | File Name | Quantization Type | File Size | Description | | :--- | :--- | :--- | :--- | | `hydra-turbo-q4_0.gguf` | Q4_0 | 4.95 GB | Legacy 4-bit quantization. Fast execution, low VRAM footprint. | | `hydra-turbo-q4_k_m.gguf` | Q4_K_M | 5.24 GB | Recommended medium 4-bit quant. Balanced accuracy and memory usage. | | `hydra-turbo-q5_k_m.gguf` | Q5_K_M | 6.02 GB | High-precision 5-bit quant. Reduced perplexity loss with minimal speed overhead. | | `hydra-turbo-q8_0.gguf` | Q8_0 | 8.87 GB | Near-lossless 8-bit quantization. Maximum output quality. | --- ## Technical Specifications - **Layers**: 32 Main Transformer Blocks (SSM / Gated Delta Net + Full Attention) - **Embedding Length**: 4096 - **Feed Forward Dimension**: 12288 - **Attention Heads**: 16 (Key-Value Heads: 4) - **Rotary Position Embedding (RoPE)**: Frequency Base 10,000,000 - **Layer Norm Epsilon**: 1e-6 --- ## Prompt Template Hydra-Turbo uses the standard Qwen Chat Template: ```text <|im_start|>system You are a helpful assistant.<|im_end|> <|im_start|>user Your query here<|im_end|> <|im_start|>assistant