How to use from
Hermes Agent
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "salohcin714/granite-4.1-3b-8bit-mlx"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default salohcin714/granite-4.1-3b-8bit-mlx
Run Hermes
hermes
Quick Links

granite-4.1-3b-8bit-mlx

Provenance

Converted from ibm-granite/granite-4.1-3b using mlx-lm 0.31.3.

Usage

from mlx_lm import load, generate

model, tokenizer = load("salohcin714/granite-4.1-3b-8bit-mlx")
messages = [{"role": "user", "content": "Hello"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
text = generate(model, tokenizer, prompt=prompt, verbose=True)

Modifications

Weights converted to MLX safetensors layout and quantized (8-bit affine quantization, group size 64, via round-to-nearest, no calibration). Redundant tied lm_head.weight dropped where the model ties input/output embeddings. No fine-tuning; no added training data.

License and attribution

Licensed under Apache 2.0. Original weights by the Granite Team, IBM. See the upstream model card and the included LICENSE file for the full text.

Disclaimer

This repository is not affiliated with or endorsed by IBM. "Granite" is an IBM trademark, used here descriptively to identify the origin of the base model. IBM's published benchmarks describe the original weights, not this quantized/converted artifact, and must not be read as claims about this repo.

Downloads last month
21
Safetensors
Model size
3B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for salohcin714/granite-4.1-3b-8bit-mlx

Quantized
(64)
this model