Instructions to use wezzel98765/gemma-4-12B-it-oQ8-fp16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use wezzel98765/gemma-4-12B-it-oQ8-fp16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir gemma-4-12B-it-oQ8-fp16 wezzel98765/gemma-4-12B-it-oQ8-fp16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
oLMX FAILS TO LOAD
M1 MAX 32GB
ERROR:
gemma-4-12B-it-oQ8-fp16
Error: {"error":{"message":"Internal server error","type":"server_error","param":null,"code":null}})
Thanks for your message - let me look into it shortly. I’ll re check all quants .
Which OMLX are you running ?
Thanks for the reply - I'm running v0.4.2.dev1
Gotcha! Dev2 has since come out. Ill re quant both of the new Gemma4 models again and have them uploaded approx 11pm GMT
Re-done. Please try again :)
Downloaded but still same problem - did the model work for you?
FYI: I tried another model oQ8-fp16 in oMLX which works fine without problems
Thanks for your message. I too am having issues now actually on this.
Having checked here on HF, I see one other person with an OQ8 version, it is the BF16 one versus FP16 ... that might be the issue. I will look into it further