Instructions to use google/gemma-4-12B-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use google/gemma-4-12B-it with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("google/gemma-4-12B-it") model = AutoModelForMultimodalLM.from_pretrained("google/gemma-4-12B-it", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- AMD Developer Cloud
Emit an empty thought after a tool response when enable_thinking=false
Browse filesSame change as https://huggingface.co/google/gemma-4-26B-A4B-it/discussions/67: this repository ships a byte-identical `chat_template.jinja`.
With `enable_thinking=false` the generation prompt after a user turn ends with an empty thought
(`<|channel>thought\n<channel|>`), but after a tool response it ends right after `<tool_response|>`. The model then
opens the thought channel on its own; on long tool-calling conversations it sometimes never closes it, and a
reasoning parser returns an empty `content`. This change emits the same empty thought after a tool response. The
renders after a user turn, and with `enable_thinking=true`, are unchanged.
Reproduction, measurements and a workaround are in the linked PR. They were made with gemma-4-26B-A4B-it; I did not
run this checkpoint.
- chat_template.jinja +2 -0
|
@@ -386,5 +386,7 @@
|
|
| 386 |
{%- endif -%}
|
| 387 |
{%- elif ns.prev_message_type == 'tool_response' and enable_thinking -%}
|
| 388 |
{{- '<|channel>thought\n' -}}
|
|
|
|
|
|
|
| 389 |
{%- endif -%}
|
| 390 |
{%- endif -%}
|
|
|
|
| 386 |
{%- endif -%}
|
| 387 |
{%- elif ns.prev_message_type == 'tool_response' and enable_thinking -%}
|
| 388 |
{{- '<|channel>thought\n' -}}
|
| 389 |
+
{%- elif ns.prev_message_type == 'tool_response' -%}
|
| 390 |
+
{{- '<|channel>thought\n<channel|>' -}}
|
| 391 |
{%- endif -%}
|
| 392 |
{%- endif -%}
|