Holo4
Collection
Holo4 model family releases, including GGUF and MLX conversions. • 6 items • Updated
How to use abenzerps/Holotron4-30B-A3B-MLX with MLX:
# Make sure mlx-vlm is installed
# pip install --upgrade mlx-vlm
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
# Load the model
model, processor = load("abenzerps/Holotron4-30B-A3B-MLX")
config = load_config("abenzerps/Holotron4-30B-A3B-MLX")
# Prepare input
image = ["http://images.cocodataset.org/val2017/000000039769.jpg"]
prompt = "Describe this image."
# Apply chat template
formatted_prompt = apply_chat_template(
processor, config, prompt, num_images=1
)
# Generate output
output = generate(model, processor, formatted_prompt, image)
print(output)How to use abenzerps/Holotron4-30B-A3B-MLX with Pi:
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "abenzerps/Holotron4-30B-A3B-MLX"
# Install Pi:
npm install -g @earendil-works/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
"providers": {
"mlx-lm": {
"baseUrl": "http://localhost:8080/v1",
"api": "openai-completions",
"apiKey": "none",
"models": [
{
"id": "abenzerps/Holotron4-30B-A3B-MLX"
}
]
}
}
}# Start Pi in your project directory: pi
How to use abenzerps/Holotron4-30B-A3B-MLX with Hermes Agent:
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "abenzerps/Holotron4-30B-A3B-MLX"
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default abenzerps/Holotron4-30B-A3B-MLX
hermes
How to use abenzerps/Holotron4-30B-A3B-MLX with OpenClaw:
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "abenzerps/Holotron4-30B-A3B-MLX"
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "abenzerps/Holotron4-30B-A3B-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
openclaw agent --local --agent main --message "Hello from Hugging Face"
MLX quantizations of Hcompany/Holotron4-30B-A3B, a 30B hybrid NemotronH MoE vision-language model (VLM) for Computer Use, tool-driven work, and agentic workflows.
Holotron4 improves over its base model, Nemotron 3 Nano Omni, on GUI workflows and in environments with MCP tools, APIs, or code sandboxes. Gains are absolute percentage points.
| Benchmark | Interface | Nemotron 3 Nano Omni | Holotron4-30B-A3B | Gain |
|---|---|---|---|---|
| OSWorld | GUI | 21.0 | 76.3 | +55.3 |
| OSWorld 2.0 | GUI and code | 0.2 | 7.9 | +7.7 |
| AutomationBench | MCP | 19.4 | 35.6 | +16.2 |
| PinchBench | Terminal | 84.7 | 88.6 | +3.9 |
| ALE (Linux, code) | Terminal | 0.6 | 8.5 | +7.9 |
| Quantization | File | Size |
|---|---|---|
| 4-bit | Holotron4-30B-A3B-MLX-4bit | 18.30 GB |
| 6-bit | Holotron4-30B-A3B-MLX-6bit | 25.82 GB |
| 8-bit | Holotron4-30B-A3B-MLX-8bit | 33.33 GB |
pip install -U mlx-vlm
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
# Load 4-bit (or subfolder="Holotron4-30B-A3B-MLX-6bit", subfolder="Holotron4-30B-A3B-MLX-8bit")
model_path = "abenzerps/Holotron4-30B-A3B-MLX"
subfolder = "Holotron4-30B-A3B-MLX-4bit"
model, processor = load(model_path, subfolder=subfolder)
config = load_config(model_path, subfolder=subfolder)
prompt = "Describe the user interface elements shown in this screenshot."
image = ["screenshot.png"]
formatted_prompt = apply_chat_template(
processor, config, prompt, num_images=len(image)
)
output = generate(model, processor, formatted_prompt, image, verbose=True)
print(output)
# 4-bit
python -m mlx_vlm.generate \
--model abenzerps/Holotron4-30B-A3B-MLX --subfolder Holotron4-30B-A3B-MLX-4bit \
--image screenshot.png \
--prompt "What action should be taken next to achieve the user goal?"
# 6-bit
python -m mlx_vlm.generate \
--model abenzerps/Holotron4-30B-A3B-MLX --subfolder Holotron4-30B-A3B-MLX-6bit \
--image screenshot.png \
--prompt "What action should be taken next to achieve the user goal?"
# 8-bit
python -m mlx_vlm.generate \
--model abenzerps/Holotron4-30B-A3B-MLX --subfolder Holotron4-30B-A3B-MLX-8bit \
--image screenshot.png \
--prompt "What action should be taken next to achieve the user goal?"
Quantized
Base model
Hcompany/Holotron4-30B-A3B