File size: 1,120 Bytes
8e017e8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
Gemma 4 31B Instruct
Copyright (c) Google DeepMind

Original model weights: https://huggingface.co/google/gemma-4-31B-it
Distributed by Google DeepMind under the Apache License 2.0
(https://ai.google.dev/gemma/apache_2).

This repository contains a derivative work: an INT8 W8A8 post-training quantized
version of the above model, produced with AMD Quark
(https://github.com/amd/quark). The original BF16 weights have been transformed
into INT8 per-channel weights with per-token dynamic INT8 activations; the
embedding, lm_head and the entire vision tower remain in BF16.

Modifications made:
  - Linear weights of the language tower converted from BF16 to INT8 with
    per-output-channel symmetric scales.
  - quantization_config block appended to config.json (custom_mode='quark',
    pack_method='order', weight_format='real_quantized').
  - All other tokenizer / processor / chat_template files are unchanged from
    the upstream google/gemma-4-31B-it release.

The license, attribution and disclaimer of warranty terms of the Apache License
2.0 (see LICENSE) apply to both the original work and this derivative.