Hy-MT2 30B A3B MLX Translation Quantizations
Collection
Public MLX translation model quantizations for Hy-MT2 30B A3B. • 7 items • Updated
How to use dawncr0w/Hy-MT2-30B-A3B-oQ8-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Hy-MT2-30B-A3B-oQ8-MLX dawncr0w/Hy-MT2-30B-A3B-oQ8-MLX
This is an oQ8 MLX quantized derivative of tencent/Hy-MT2-30B-A3B.
The model was quantized locally with oMLX oQ8. The resulting config targets approximately 8.50 bpw and includes a packaged hy_v3.py model file so MLX-LM can load the HYV3 architecture from the model directory.
Local validation completed with the oMLX app bundle on macOS:
mlx_lm load: passed
generation smoke test: passed
prompt: Hello
max tokens: 4
peak memory: 29.815 GB
Use an MLX-LM build that supports the APIs used by the packaged hy_v3.py file.
python -m mlx_lm generate \
--model /path/to/Hy-MT2-30B-A3B-oQ8-MLX \
--ignore-chat-template \
--prompt "Hello" \
--max-tokens 32 \
--temp 0
The base model is distributed under the Tencent HY Community License Agreement. This distribution includes LICENSE.txt and NOTICE from/for the Tencent HY license requirements.
This repository is not affiliated with, sponsored by, or endorsed by Tencent.
8-bit
Base model
tencent/Hy-MT2-30B-A3B