Instructions to use Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF with Ollama:
ollama run hf.co/Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF with Docker Model Runner:
docker model run hf.co/Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
- Lemonade
How to use Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Abiray/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Nemotron-3-Nano-Omni-30B-A3B-Reasoning - GGUF
This repository contains high-fidelity GGUF quantizations of NVIDIA's Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16.
These quantizations were explicitly compiled to maximize logic, reasoning, and narrative consistency for local deployments, making them highly suitable for text-based RPG engines, structured JSON output.
🧠 High-Fidelity Quantization Strategy: FP16 Output Head
Unlike standard GGUF conversions, these models were quantized using the --leave-output-tensor flag.
What does this mean?
The final projection layer (lm_head / output.weight), which maps the model's internal states to the vocabulary, has been preserved in pristine FP16 precision (~2.1 GB). While this slightly increases the overall file size and initial HDD load time, it completely eliminates the "numerical noise" introduced when crushing the output head to 4-bit or 5-bit.
The Result: Smaller quantizations (like Q3_K_M or Q4_K_M) retain the sharp logic, precise tool-calling, and chain-of-thought (<think>) capabilities of the massive uncompressed model, all while fitting comfortably into local RAM constraints.
⚙️ Model Architecture & Hardware Requirements
- Architecture: Mamba2-Transformer Hybrid Mixture of Experts (MoE)
- Parameters: 30 Billion total
- Active Parameters: ~3 Billion per token (A3B)
- Context Length: Up to 256k tokens
Because this is a sparse MoE model, it requires significantly less RAM bandwidth and compute power than a dense 30B model. A Q4_K_M variant will easily run on machines with 8GB to 16GB of system RAM. The primary bottleneck will be the initial model loading time.
📂 Available Quantizations
| Quantization | Bits / Weight | Use Case / Notes |
|---|---|---|
| Q8_0 | 8.5 | Extreme fidelity. Best if you have high RAM but limited VRAM. |
| Q6_K | 6.5 | Excellent balance for 16GB+ systems. Near-perfect F16 parity. |
| Q5_K_M | 5.5 | High quality, slightly faster inference than Q6. |
| Q4_K_M | 4.8 | [RECOMMENDED] The sweet spot for performance vs. intelligence. FP16 head ensures reasoning stays intact. |
| Q4_K_S | 4.5 | Slightly smaller than K_M, minimal quality loss. |
| Q3_K_M | 3.5 | Maximum compression. Great for severely resource-constrained setups (8GB RAM). |
💬 Prompt Format (Reasoning Mode)
This model is trained to utilize a chain-of-thought reasoning budget. It natively supports <think> tags before generating its final response.
Chat Template Example:
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
What is 2+2?<|im_end|>
<|im_start|>assistant
<think>
1. The user is asking for a simple arithmetic operation.
2. The operation is addition: 2 + 2.
3. The result of 2 + 2 is 4.
</think>
The answer is 4.<|im_end|>
- Downloads last month
- 537
3-bit
4-bit
5-bit
6-bit
8-bit