File size: 7,336 Bytes
dce8120
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
---
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:
- vLLM
- AWQ
- autoawq
- INT4
- W4A16
- qwen3.5
- reasoning
- long-context
- 1M-context
- function-calling
- tool-use
- cybersecurity
- biomedical
- agentic
base_model:
- empero-ai/Qwythos-9B-Claude-Mythos-5-1M
base_model_relation: quantized
---

# Qwythos-9B-Claude-Mythos-5-1M-AWQ

Base model: [empero-ai/Qwythos-9B-Claude-Mythos-5-1M](https://huggingface.co/empero-ai/Qwythos-9B-Claude-Mythos-5-1M)

This repository contains an **AWQ INT4 / W4A16** quantized export of Qwythos-9B-Claude-Mythos-5-1M.

The model was quantized on an **NVIDIA GeForce RTX 4090** with a self-implemented AWQ exporter that writes a vLLM / AutoAWQ-style weight layout:

- `.qweight`
- `.qzeros`
- `.scales`
- sharded `model-xxxxx-of-xxxxx.safetensors`
- `model.safetensors.index.json`

The quantization path is **data-free**: no external calibration dataset was used. The exporter estimates per-layer input importance from the original weights, then searches weighted group clipping / minmax scales to reduce reconstruction error.

This release is primarily prepared for **vLLM** with AWQ GEMM weight-only inference. It does not claim compatibility with every runtime.

## Quantization Policy

Qwythos is based on Qwen3.5, whose hybrid attention stack contains linear-attention and attention modules that should remain in native precision for vLLM loading. For this reason, this export intentionally quantizes only the text MLP projection weights.

| Scope | Format |
|-------|--------|
| Text MLP `gate_proj` | W4A16 AWQ, group size 128, zero point |
| Text MLP `up_proj` | W4A16 AWQ, group size 128, zero point |
| Text MLP `down_proj` | W4A16 AWQ, group size 128, zero point |
| First text layer | Native BF16/FP16 |
| `linear_attn.*` | Native BF16/FP16 |
| `self_attn.*` | Native BF16/FP16 |
| Embeddings, norms, rotary embeddings | Native BF16/FP16 |
| `lm_head` | Native BF16/FP16 |
| Vision / visual modules | Native BF16/FP16 |
| MTP / special heads, if present | Native BF16/FP16 |

Quantization config used by the exporter:

```yaml
w_bit: 4
q_group_size: 128
zero_point: true
awq_grid_size: 8
awq_samples: 8
calibration_mode: data_free
max_shard_size: 3GB
```

The exporter keeps the Qwen3.5 linear-attention tensors in their official split layout rather than fusing them, which is important for vLLM compatibility.

## Dependencies

Recommended vLLM serving stack:

```bash
pip install --upgrade vllm transformers accelerate safetensors
```

Qwen3.5 support in serving frameworks is still relatively new. If your local vLLM build cannot load `qwen3_5`, upgrade vLLM and Transformers first.

## vLLM Startup Command

Start with a conservative context length on a single RTX 4090. The source model advertises a 1M context window, but KV cache memory, not AWQ weight size, becomes the limiting factor at long context.

```bash
vllm serve JunHowie/Qwythos-9B-Claude-Mythos-5-1M-AWQ \
  --host 0.0.0.0 \
  --port 8000 \
  --served-model-name Qwythos-9B-AWQ \
  --trust-remote-code \
  --quantization awq \
  --dtype float16 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.85
```

If memory headroom is sufficient, increase `--max-model-len` gradually.

## Example Request

```bash
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwythos-9B-AWQ",
    "messages": [
      {
        "role": "user",
        "content": "Explain how AWQ quantization protects salient channels, and why Qwen3.5 linear-attention tensors are kept native."
      }
    ],
    "temperature": 0.6,
    "top_p": 0.95,
    "max_tokens": 4096
  }'
```

## Model Download

```python
from huggingface_hub import snapshot_download

snapshot_download(
    "JunHowie/Qwythos-9B-Claude-Mythos-5-1M-AWQ",
    cache_dir="your_local_path",
)
```

## Notes

- This is a quantized derivative of Qwythos-9B-Claude-Mythos-5-1M, not a new fine-tune.
- The original model is a reasoning model and may emit `<think>...</think>` blocks before the final answer.
- Recommended sampling follows the base model card: `temperature=0.6`, `top_p=0.95`, `top_k=20`, `repetition_penalty=1.05`.
- The source model advertises a 1,048,576-token context window with YaRN rope scaling. For this AWQ release, start smaller on consumer GPUs and raise context length after memory testing.
- No benchmark numbers are changed or claimed by this AWQ export. Use the original model card for upstream capability references.

## Quantization Log

```text
2026-06-30
1. Quantized empero-ai/Qwythos-9B-Claude-Mythos-5-1M on RTX 4090.
2. Exported AWQ GEMM W4A16 weights with group size 128 and zero point.
3. Used data-free AWQ search; no calibration dataset was involved.
4. Kept Qwen3.5 attention, linear-attention, embeddings, norms, lm_head, visual modules, and first text layer in native precision for vLLM compatibility.
```

## Original Model Overview

Qwythos-9B is developed by [Empero](https://empero.org). It is a full-parameter reasoning model built on top of a deeply uncensored Qwen3.5-9B base and post-trained on over 500 million tokens of Claude Mythos and Claude Fable traces, with chain-of-thought generated by Empero AI's internal tool `rethink`.

Key upstream features:

- **1,048,576-token context** via YaRN rope scaling enabled by default.
- **Native function calling** following Qwen3.5's tool-use format.
- **Reasoning-oriented behavior** with `<think>` blocks.
- **Strong gains over the base Qwen3.5-9B** in the original evaluation setup, especially on MMLU and GSM8K.
- **Domain emphasis** on cybersecurity, biomedical, quantitative reasoning, coding, and agentic workflows.

Original evaluation headline from the base model card:

| Task | Metric | Base Qwen3.5-9B | Qwythos-9B | Delta |
|---|---|---:|---:|---:|
| gsm8k | exact_match flexible | 0.670 | 0.860 | +0.190 |
| gsm8k | exact_match strict | 0.510 | 0.810 | +0.300 |
| mmlu | acc | 0.232 | 0.575 | +0.343 |
| arc_challenge | acc | 0.470 | 0.490 | +0.020 |
| arc_challenge | acc_norm | 0.400 | 0.410 | +0.010 |
| gpqa_diamond | exact_match flexible | 0.630 | 0.580 | -0.050 |

These numbers belong to the original FP/BF16 model card and are included only as upstream reference.

## Base Model Details

- **Upstream model:** [empero-ai/Qwythos-9B-Claude-Mythos-5-1M](https://huggingface.co/empero-ai/Qwythos-9B-Claude-Mythos-5-1M)
- **Underlying base:** [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)
- **Fine-tune type:** full-parameter supervised fine-tune
- **Context length:** 1,048,576 tokens with YaRN rope scaling
- **License:** Apache-2.0

## Limitations

- This AWQ export has not been presented as a full formal benchmark release.
- Very long context still requires substantial KV-cache memory.
- For exact identifiers, safety-critical medical details, live facts, or security-sensitive workflows, pair the model with retrieval/tools and verify outputs.
- Qwythos is intentionally uncensored. Add application-level policy and review layers for public-facing deployments.

## Acknowledgements

- Original model: [Empero Qwythos-9B-Claude-Mythos-5-1M](https://huggingface.co/empero-ai/Qwythos-9B-Claude-Mythos-5-1M)
- Underlying base: [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)
- Quantized release: `JunHowie/Qwythos-9B-Claude-Mythos-5-1M-AWQ`