Instructions to use divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16 with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16") config = load_config("divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Nemotron-3-Nano-Omni-30B Abliterated (full, MLX)
Abliterated build of NVIDIA's tri-modal Nemotron-3-Nano-Omni-30B for Apple Silicon. Text, vision, and audio all work — the refusal direction was removed from the language stack only; the vision and sound towers are copied through untouched.
As far as I can tell this is the first abliterated Omni in MLX — the other uncensored builds are NVFP4 / GGUF and don't run on a Mac.
- Precision: full (~62GB)
- Refusal removed: verified 6/6 on a held-out harmful set
- Vision: describes images correctly
- Audio: transcribes speech
- Tensors: 1,481 — 401 language (abliterated) + 390 vision + 684 sound + 6 projection, name-for-name identical to the base
Run it
pip install -U mlx-vlm
python -m mlx_vlm generate --model divinetribe/Nemotron-3-Nano-Omni-30B-Abliterated-MM-bf16 \
--image your_image.jpg --prompt "Describe this image."
Needs a recent mlx-vlm (0.6.12+) with Omni support. Runtime reference:
https://github.com/nicedreamzapp/nemotron-omni-mlx
GGUF / llama.cpp
Not MLX-only. @wacomctl672 converted this to GGUF and ran it
in llama.cpp at Q4_K_M with the combined audio+vision mmproj from the ggml-org repository —
text, vision and audio all working. Ready-made weights:
wacomctl672/Nemotron-3-Nano-Omni-30B-Abliterated-MM-GGUF
If you convert it yourself, it needs one patch to llama.cpp/conversion/nemotron.py. Add this at
the top of modify_tensors (around line 431):
if "switch_mlp.fc1.weight" in name:
yield f"blk.{bid}.ffn_up_exps.weight", data_torch
return
elif "switch_mlp.fc2.weight" in name:
yield f"blk.{bid}.ffn_down_exps.weight", data_torch
return
Credit and details in discussion #1.
Method
Directional ablation (Arditi et al.). The refusal direction on this model is spread
across two blocks (16 and 31), not one — a single-layer ablation leaves it refusing.
Both directions are Gram-Schmidt'd and orthogonalized out of every residual-writing
projection: mamba out_proj, attention o_proj, MoE routed-expert fc2, shared-expert
down_proj, plus the token embeddings.
Use responsibly
Safety alignment has been removed. You are responsible for what you generate.
Part of Claude Code Local
This model is one of the fighters in Claude Code Local (3.2k★), which runs Claude Code 100% on-device on Apple Silicon through an MLX-native Anthropic-API server. Not sure which local model to run as an agent? Check the Agent-12 local agent leaderboard: real agent tasks, judged by the filesystem, same hardware for every row.
Built by Matt Macosko in Arcata, CA. Open to work on local-AI and Apple Silicon inference: matt@ineedhemp.com.
- Downloads last month
- 239
Quantized