plunderstruck commited on
Commit
e49818d
·
verified ·
1 Parent(s): 701cc87

Card: add Vision section (--mmproj + --image-min-tokens 1024) and bundled chat_template.jinja note

Browse files
Files changed (1) hide show
  1. README.md +27 -0
README.md CHANGED
@@ -153,6 +153,33 @@ before relying on it.
153
  - [Qwen3.6-27B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF) · [Qwen3.6-35B-A3B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-35B-A3B-MTP-ROCmFP4-GGUF) · [Qwopus3.6-27B-Coder-MTP](https://huggingface.co/plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF)
154
  - [Qwen3.6-40B-Deckard-MTP](https://huggingface.co/plunderstruck/Qwen3.6-40B-Deckard-MTP-ROCmFP4-GGUF) · [Qwen3-Coder-Next](https://huggingface.co/plunderstruck/Qwen3-Coder-Next-ROCmFP4-GGUF) · [Nex-N2-mini](https://huggingface.co/plunderstruck/Nex-N2-mini-ROCmFP4-GGUF)
155
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
156
  ## Credits & license
157
 
158
  - **Base model:** [`OBLITERATUS/Qwen3.6-27B-OBLITERATED`](https://huggingface.co/OBLITERATUS/Qwen3.6-27B-OBLITERATED) (Apache-2.0), abliterated from [`Qwen/Qwen3.6-27B`](https://huggingface.co/Qwen/Qwen3.6-27B) (Qwen team). A derivative quantization — verify the base terms before redistribution/use.
 
153
  - [Qwen3.6-27B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF) · [Qwen3.6-35B-A3B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-35B-A3B-MTP-ROCmFP4-GGUF) · [Qwopus3.6-27B-Coder-MTP](https://huggingface.co/plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF)
154
  - [Qwen3.6-40B-Deckard-MTP](https://huggingface.co/plunderstruck/Qwen3.6-40B-Deckard-MTP-ROCmFP4-GGUF) · [Qwen3-Coder-Next](https://huggingface.co/plunderstruck/Qwen3-Coder-Next-ROCmFP4-GGUF) · [Nex-N2-mini](https://huggingface.co/plunderstruck/Nex-N2-mini-ROCmFP4-GGUF)
155
 
156
+ ## Vision (multimodal)
157
+
158
+ This is a Qwen3-VL-lineage model — it does **vision** the standard llama.cpp way: supply a **Qwen3-VL
159
+ `mmproj` projector** at launch with `--mmproj` (**no different LLM GGUF needed**). Any Qwen3-VL projector
160
+ whose `projection_dim` is **5120** matches this model's hidden size (e.g. the Qwen3.6-27B projector); f16 or
161
+ f32 both work.
162
+
163
+ > **The flag that matters: `--image-min-tokens 1024`.** Without it the server feeds too few image tokens and
164
+ > the model **describes images incorrectly** (right gist, wrong detail). Qwen-VL needs **≥1024 image tokens**
165
+ > for correct grounding — the server even logs a warning at load. Verified on this model: a code label
166
+ > misread at default tokens read correctly once the flag was set.
167
+
168
+ ```bash
169
+ # add to your llama-server launch:
170
+ --mmproj mmproj-Qwen3.6-27B-f16.gguf \
171
+ --image-min-tokens 1024
172
+ ```
173
+
174
+ Two caveats: (1) it's a **thinking model** — for one-shot image Q&A, allow enough tokens to finish `<think>`
175
+ or disable thinking (e.g. this repo's `chat_template.jinja` inline `<|think_off|>`), or the visible answer
176
+ can come back empty; (2) loading `--mmproj` **disables cross-turn prompt-cache reuse** on this runtime — use
177
+ vision for one-shot image tasks, not long cache-dependent sessions.
178
+
179
+ > **Bundled fixed template:** this repo includes [`chat_template.jinja`](./chat_template.jinja) (froggeric's
180
+ > unified Qwen3.6 template) — pass `--chat-template-file chat_template.jinja` for reliable tool calls, inline
181
+ > `<|think_off|>`/`<|think_on|>`, and vision in one template.
182
+
183
  ## Credits & license
184
 
185
  - **Base model:** [`OBLITERATUS/Qwen3.6-27B-OBLITERATED`](https://huggingface.co/OBLITERATUS/Qwen3.6-27B-OBLITERATED) (Apache-2.0), abliterated from [`Qwen/Qwen3.6-27B`](https://huggingface.co/Qwen/Qwen3.6-27B) (Qwen team). A derivative quantization — verify the base terms before redistribution/use.