How to use from
Pi
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "abenzerps/Holotron4-30B-A3B-MLX"
Configure the model in Pi
# Install Pi:
npm install -g @earendil-works/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "mlx-lm": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "abenzerps/Holotron4-30B-A3B-MLX"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

Holotron4-30B-A3B MLX

MLX quantizations of Hcompany/Holotron4-30B-A3B, a 30B hybrid NemotronH MoE vision-language model (VLM) for Computer Use, tool-driven work, and agentic workflows.

Upstream benchmarks

Holotron4 improves over its base model, Nemotron 3 Nano Omni, on GUI workflows and in environments with MCP tools, APIs, or code sandboxes. Gains are absolute percentage points.

Benchmark Interface Nemotron 3 Nano Omni Holotron4-30B-A3B Gain
OSWorld GUI 21.0 76.3 +55.3
OSWorld 2.0 GUI and code 0.2 7.9 +7.7
AutomationBench MCP 19.4 35.6 +16.2
PinchBench Terminal 84.7 88.6 +3.9
ALE (Linux, code) Terminal 0.6 8.5 +7.9

MLX Files

Quantization File Size
4-bit Holotron4-30B-A3B-MLX-4bit 18.30 GB
6-bit Holotron4-30B-A3B-MLX-6bit 25.82 GB
8-bit Holotron4-30B-A3B-MLX-8bit 33.33 GB

Usage with MLX-VLM

Installation

pip install -U mlx-vlm

Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

# Load 4-bit (or subfolder="Holotron4-30B-A3B-MLX-6bit", subfolder="Holotron4-30B-A3B-MLX-8bit")
model_path = "abenzerps/Holotron4-30B-A3B-MLX"
subfolder = "Holotron4-30B-A3B-MLX-4bit"

model, processor = load(model_path, subfolder=subfolder)
config = load_config(model_path, subfolder=subfolder)

prompt = "Describe the user interface elements shown in this screenshot."
image = ["screenshot.png"]

formatted_prompt = apply_chat_template(
    processor, config, prompt, num_images=len(image)
)

output = generate(model, processor, formatted_prompt, image, verbose=True)
print(output)

Command Line Interface

# 4-bit
python -m mlx_vlm.generate \
    --model abenzerps/Holotron4-30B-A3B-MLX --subfolder Holotron4-30B-A3B-MLX-4bit \
    --image screenshot.png \
    --prompt "What action should be taken next to achieve the user goal?"

# 6-bit
python -m mlx_vlm.generate \
    --model abenzerps/Holotron4-30B-A3B-MLX --subfolder Holotron4-30B-A3B-MLX-6bit \
    --image screenshot.png \
    --prompt "What action should be taken next to achieve the user goal?"

# 8-bit
python -m mlx_vlm.generate \
    --model abenzerps/Holotron4-30B-A3B-MLX --subfolder Holotron4-30B-A3B-MLX-8bit \
    --image screenshot.png \
    --prompt "What action should be taken next to achieve the user goal?"

Source

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for abenzerps/Holotron4-30B-A3B-MLX

Quantized
(5)
this model

Collection including abenzerps/Holotron4-30B-A3B-MLX