Card: add Vision section (--mmproj + --image-min-tokens 1024) and bundled chat_template.jinja note
Browse files
README.md
CHANGED
|
@@ -153,6 +153,33 @@ before relying on it.
|
|
| 153 |
- [Qwen3.6-27B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF) · [Qwen3.6-35B-A3B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-35B-A3B-MTP-ROCmFP4-GGUF) · [Qwopus3.6-27B-Coder-MTP](https://huggingface.co/plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF)
|
| 154 |
- [Qwen3.6-40B-Deckard-MTP](https://huggingface.co/plunderstruck/Qwen3.6-40B-Deckard-MTP-ROCmFP4-GGUF) · [Qwen3-Coder-Next](https://huggingface.co/plunderstruck/Qwen3-Coder-Next-ROCmFP4-GGUF) · [Nex-N2-mini](https://huggingface.co/plunderstruck/Nex-N2-mini-ROCmFP4-GGUF)
|
| 155 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 156 |
## Credits & license
|
| 157 |
|
| 158 |
- **Base model:** [`OBLITERATUS/Qwen3.6-27B-OBLITERATED`](https://huggingface.co/OBLITERATUS/Qwen3.6-27B-OBLITERATED) (Apache-2.0), abliterated from [`Qwen/Qwen3.6-27B`](https://huggingface.co/Qwen/Qwen3.6-27B) (Qwen team). A derivative quantization — verify the base terms before redistribution/use.
|
|
|
|
| 153 |
- [Qwen3.6-27B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF) · [Qwen3.6-35B-A3B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-35B-A3B-MTP-ROCmFP4-GGUF) · [Qwopus3.6-27B-Coder-MTP](https://huggingface.co/plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF)
|
| 154 |
- [Qwen3.6-40B-Deckard-MTP](https://huggingface.co/plunderstruck/Qwen3.6-40B-Deckard-MTP-ROCmFP4-GGUF) · [Qwen3-Coder-Next](https://huggingface.co/plunderstruck/Qwen3-Coder-Next-ROCmFP4-GGUF) · [Nex-N2-mini](https://huggingface.co/plunderstruck/Nex-N2-mini-ROCmFP4-GGUF)
|
| 155 |
|
| 156 |
+
## Vision (multimodal)
|
| 157 |
+
|
| 158 |
+
This is a Qwen3-VL-lineage model — it does **vision** the standard llama.cpp way: supply a **Qwen3-VL
|
| 159 |
+
`mmproj` projector** at launch with `--mmproj` (**no different LLM GGUF needed**). Any Qwen3-VL projector
|
| 160 |
+
whose `projection_dim` is **5120** matches this model's hidden size (e.g. the Qwen3.6-27B projector); f16 or
|
| 161 |
+
f32 both work.
|
| 162 |
+
|
| 163 |
+
> **The flag that matters: `--image-min-tokens 1024`.** Without it the server feeds too few image tokens and
|
| 164 |
+
> the model **describes images incorrectly** (right gist, wrong detail). Qwen-VL needs **≥1024 image tokens**
|
| 165 |
+
> for correct grounding — the server even logs a warning at load. Verified on this model: a code label
|
| 166 |
+
> misread at default tokens read correctly once the flag was set.
|
| 167 |
+
|
| 168 |
+
```bash
|
| 169 |
+
# add to your llama-server launch:
|
| 170 |
+
--mmproj mmproj-Qwen3.6-27B-f16.gguf \
|
| 171 |
+
--image-min-tokens 1024
|
| 172 |
+
```
|
| 173 |
+
|
| 174 |
+
Two caveats: (1) it's a **thinking model** — for one-shot image Q&A, allow enough tokens to finish `<think>`
|
| 175 |
+
or disable thinking (e.g. this repo's `chat_template.jinja` inline `<|think_off|>`), or the visible answer
|
| 176 |
+
can come back empty; (2) loading `--mmproj` **disables cross-turn prompt-cache reuse** on this runtime — use
|
| 177 |
+
vision for one-shot image tasks, not long cache-dependent sessions.
|
| 178 |
+
|
| 179 |
+
> **Bundled fixed template:** this repo includes [`chat_template.jinja`](./chat_template.jinja) (froggeric's
|
| 180 |
+
> unified Qwen3.6 template) — pass `--chat-template-file chat_template.jinja` for reliable tool calls, inline
|
| 181 |
+
> `<|think_off|>`/`<|think_on|>`, and vision in one template.
|
| 182 |
+
|
| 183 |
## Credits & license
|
| 184 |
|
| 185 |
- **Base model:** [`OBLITERATUS/Qwen3.6-27B-OBLITERATED`](https://huggingface.co/OBLITERATUS/Qwen3.6-27B-OBLITERATED) (Apache-2.0), abliterated from [`Qwen/Qwen3.6-27B`](https://huggingface.co/Qwen/Qwen3.6-27B) (Qwen team). A derivative quantization — verify the base terms before redistribution/use.
|