Hy-MT2-30B-A3B-oQ8-MLX

This is an oQ8 MLX quantized derivative of tencent/Hy-MT2-30B-A3B.

The model was quantized locally with oMLX oQ8. The resulting config targets approximately 8.50 bpw and includes a packaged hy_v3.py model file so MLX-LM can load the HYV3 architecture from the model directory.

Validation

Local validation completed with the oMLX app bundle on macOS:

mlx_lm load: passed
generation smoke test: passed
prompt: Hello
max tokens: 4
peak memory: 29.815 GB

Usage

Use an MLX-LM build that supports the APIs used by the packaged hy_v3.py file.

python -m mlx_lm generate \
  --model /path/to/Hy-MT2-30B-A3B-oQ8-MLX \
  --ignore-chat-template \
  --prompt "Hello" \
  --max-tokens 32 \
  --temp 0

License And Notice

The base model is distributed under the Tencent HY Community License Agreement. This distribution includes LICENSE.txt and NOTICE from/for the Tencent HY license requirements.

This repository is not affiliated with, sponsored by, or endorsed by Tencent.

Downloads last month
99
Safetensors
Model size
30B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dawncr0w/Hy-MT2-30B-A3B-oQ8-MLX

Quantized
(20)
this model

Collection including dawncr0w/Hy-MT2-30B-A3B-oQ8-MLX