nameistoken's picture
Add files using upload-large-folder tool
8e017e8 verified
Raw History Blame Contribute Delete
1.12 kB
Gemma 4 31B Instruct
Copyright (c) Google DeepMind
Original model weights: https://huggingface.co/google/gemma-4-31B-it
Distributed by Google DeepMind under the Apache License 2.0
(https://ai.google.dev/gemma/apache_2).
This repository contains a derivative work: an INT8 W8A8 post-training quantized
version of the above model, produced with AMD Quark
(https://github.com/amd/quark). The original BF16 weights have been transformed
into INT8 per-channel weights with per-token dynamic INT8 activations; the
embedding, lm_head and the entire vision tower remain in BF16.
Modifications made:
- Linear weights of the language tower converted from BF16 to INT8 with
per-output-channel symmetric scales.
- quantization_config block appended to config.json (custom_mode='quark',
pack_method='order', weight_format='real_quantized').
- All other tokenizer / processor / chat_template files are unchanged from
the upstream google/gemma-4-31B-it release.
The license, attribution and disclaimer of warranty terms of the Apache License
2.0 (see LICENSE) apply to both the original work and this derivative.