Gemma 4 31B Instruct Copyright (c) Google DeepMind Original model weights: https://huggingface.co/google/gemma-4-31B-it Distributed by Google DeepMind under the Apache License 2.0 (https://ai.google.dev/gemma/apache_2). This repository contains a derivative work: an INT8 W8A8 post-training quantized version of the above model, produced with AMD Quark (https://github.com/amd/quark). The original BF16 weights have been transformed into INT8 per-channel weights with per-token dynamic INT8 activations; the embedding, lm_head and the entire vision tower remain in BF16. Modifications made: - Linear weights of the language tower converted from BF16 to INT8 with per-output-channel symmetric scales. - quantization_config block appended to config.json (custom_mode='quark', pack_method='order', weight_format='real_quantized'). - All other tokenizer / processor / chat_template files are unchanged from the upstream google/gemma-4-31B-it release. The license, attribution and disclaimer of warranty terms of the Apache License 2.0 (see LICENSE) apply to both the original work and this derivative.