Instructions to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
- SGLang
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with Ollama:
ollama run hf.co/WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with Docker Model Runner:
docker model run hf.co/WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
- Lemonade
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Nemotron-3.5-Lightning-30B-A3B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "WhiskyAKM/Nemotron-3.5-Lightning-30B-A3B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload folder using huggingface_hub
Browse files- .gitattributes +12 -0
- README.md +194 -0
- chat_template.jinja +190 -0
- nemotron-3.5-lightning-30b-a3b-Q4_0.gguf +3 -0
- nemotron-3.5-lightning-30b-a3b-Q4_K_M.gguf +3 -0
- nemotron-3.5-lightning-30b-a3b-Q4_K_S.gguf +3 -0
- nemotron-3.5-lightning-30b-a3b-Q5_K_M.gguf +3 -0
- nemotron-3.5-lightning-30b-a3b-Q5_K_S.gguf +3 -0
- nemotron-3.5-lightning-30b-a3b-Q6_K.gguf +3 -0
- nemotron-3.5-lightning-30b-a3b-Q8_0.gguf +3 -0
- nemotron-3.5-lightning-30b-a3b-bf16.gguf +3 -0
- nemotron-3.5-lightning-30b-a3b-dflash.gguf +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,15 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
accuracy_plot.png filter=lfs diff=lfs merge=lfs -text
|
| 38 |
+
agentic_coding_benchmarks.png filter=lfs diff=lfs merge=lfs -text
|
| 39 |
+
nemotron-3.5-lightning-30b-a3b-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
|
| 40 |
+
nemotron-3.5-lightning-30b-a3b-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
| 41 |
+
nemotron-3.5-lightning-30b-a3b-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
| 42 |
+
nemotron-3.5-lightning-30b-a3b-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
| 43 |
+
nemotron-3.5-lightning-30b-a3b-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
| 44 |
+
nemotron-3.5-lightning-30b-a3b-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
|
| 45 |
+
nemotron-3.5-lightning-30b-a3b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
|
| 46 |
+
nemotron-3.5-lightning-30b-a3b-bf16.gguf filter=lfs diff=lfs merge=lfs -text
|
| 47 |
+
nemotron-3.5-lightning-30b-a3b-dflash.gguf filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,194 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model:
|
| 3 |
+
- nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
|
| 4 |
+
language:
|
| 5 |
+
- en
|
| 6 |
+
- es
|
| 7 |
+
- fr
|
| 8 |
+
- de
|
| 9 |
+
- it
|
| 10 |
+
- ja
|
| 11 |
+
library_name: transformers
|
| 12 |
+
license: openmdw-1.1
|
| 13 |
+
license_link: https://openmdw.ai/license/1-1/
|
| 14 |
+
pipeline_tag: text-generation
|
| 15 |
+
tags:
|
| 16 |
+
- nvidia
|
| 17 |
+
- nemotron-3.5
|
| 18 |
+
- gguf
|
| 19 |
+
- llama.cpp
|
| 20 |
+
- text-generation
|
| 21 |
+
- moe
|
| 22 |
+
---
|
| 23 |
+
|
| 24 |
+
# NVIDIA-Nemotron-3.5-Lightning-30B-A3B - GGUF
|
| 25 |
+
|
| 26 |
+
This repository contains GGUF format model files for [NVIDIA's NVIDIA-Nemotron-3.5-Lightning-30B-A3B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16).
|
| 27 |
+
|
| 28 |
+
These files were converted and quantized using [llama.cpp](https://github.com/ggerganov/llama.cpp).
|
| 29 |
+
|
| 30 |
+
## Available Files
|
| 31 |
+
|
| 32 |
+
| Filename | Quant Method | Description |
|
| 33 |
+
| --- | --- | --- |
|
| 34 |
+
| `nemotron-3.5-lightning-30b-a3b-bf16.gguf` | BF16 | Full-precision reference weights (unquantized) |
|
| 35 |
+
| `nemotron-3.5-lightning-30b-a3b-Q8_0.gguf` | Q8_0 | Extremely high quality, fast, high resource usage |
|
| 36 |
+
| `nemotron-3.5-lightning-30b-a3b-Q6_K.gguf` | Q6_K | Very high quality, near-lossless quantization |
|
| 37 |
+
| `nemotron-3.5-lightning-30b-a3b-Q5_K_M.gguf` | Q5_K_M | High quality, balanced performance and memory |
|
| 38 |
+
| `nemotron-3.5-lightning-30b-a3b-Q5_K_S.gguf` | Q5_K_S | High quality, slightly smaller footprint than Q5_K_M |
|
| 39 |
+
| `nemotron-3.5-lightning-30b-a3b-Q4_K_M.gguf` | Q4_K_M | Recommended balance of size, speed, and quality |
|
| 40 |
+
| `nemotron-3.5-lightning-30b-a3b-Q4_K_S.gguf` | Q4_K_S | 4-bit quantization with small memory footprint |
|
| 41 |
+
| `nemotron-3.5-lightning-30b-a3b-Q4_0.gguf` | Q4_0 | Standard 4-bit quantization |
|
| 42 |
+
| `nemotron-3.5-lightning-30b-a3b-dflash.gguf` | — | DFlash speculative decoding draft model |
|
| 43 |
+
|
| 44 |
+
## Model Summary
|
| 45 |
+
|
| 46 |
+
| Total Parameters | 30B (3B active) |
|
| 47 |
+
| --- | --- |
|
| 48 |
+
| Architecture | MoE — Mamba-2 + MoE + Attention hybrid |
|
| 49 |
+
| Context Length | Up to 1M tokens (256K native default) |
|
| 50 |
+
| Supported Languages | English (and coding languages), Spanish, French, German, Italian, Japanese |
|
| 51 |
+
| Speculative Decoding | DSpark, DFlash, MTP (Multi-Token Prediction) |
|
| 52 |
+
| Reasoning Mode | Configurable on/off via chat template (`enable_thinking=True/False`) |
|
| 53 |
+
| Recommended Sampling | Temperature 1.0, Top_P 0.95 |
|
| 54 |
+
| License | OpenMDW License Agreement, version 1.1 |
|
| 55 |
+
| Original Model | [nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16) |
|
| 56 |
+
|
| 57 |
+
## Model Overview
|
| 58 |
+
|
| 59 |
+
Model Developer: NVIDIA Corporation
|
| 60 |
+
|
| 61 |
+
Model Dates: December 2025 - May 2026
|
| 62 |
+
|
| 63 |
+
Data Freshness:
|
| 64 |
+
- The pre-training data has a cutoff date of September 2025.
|
| 65 |
+
- The post-training data has a cutoff date of May 2026.
|
| 66 |
+
|
| 67 |
+
### What is Nemotron?
|
| 68 |
+
|
| 69 |
+
NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.
|
| 70 |
+
|
| 71 |
+
## Description
|
| 72 |
+
|
| 73 |
+
NVIDIA-Nemotron-3.5-Lightning-30B-A3B is a large language model (LLM) trained by NVIDIA.
|
| 74 |
+
|
| 75 |
+
The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside speculative decoding methods (DSpark, DFlash, MTP) for faster text generation. The model has 3B active parameters and 30B parameters in total.
|
| 76 |
+
|
| 77 |
+
This model is ready for commercial use under the OpenMDW-1.1 license.
|
| 78 |
+
|
| 79 |
+
## Usage with llama.cpp
|
| 80 |
+
|
| 81 |
+
### CLI / llama-cli
|
| 82 |
+
|
| 83 |
+
Reasoning ON (default):
|
| 84 |
+
|
| 85 |
+
```bash
|
| 86 |
+
llama-cli \
|
| 87 |
+
-m nemotron-3.5-lightning-30b-a3b-Q4_K_M.gguf \
|
| 88 |
+
--jinja \
|
| 89 |
+
--chat-template-file chat_template.jinja \
|
| 90 |
+
-p "Write a Python function to compute Fibonacci numbers." \
|
| 91 |
+
--temp 1.0 --top-p 0.95 \
|
| 92 |
+
-ngl 99
|
| 93 |
+
```
|
| 94 |
+
|
| 95 |
+
### llama-server
|
| 96 |
+
|
| 97 |
+
Start the OpenAI-compatible server:
|
| 98 |
+
|
| 99 |
+
```bash
|
| 100 |
+
llama-server \
|
| 101 |
+
-m nemotron-3.5-lightning-30b-a3b-Q4_K_M.gguf \
|
| 102 |
+
--temp 1.0 --top-p 0.95 \
|
| 103 |
+
-np 1 \
|
| 104 |
+
-c 40960 \
|
| 105 |
+
--port 8000 \
|
| 106 |
+
-ngl 99 \
|
| 107 |
+
-fa on \
|
| 108 |
+
--jinja \
|
| 109 |
+
--chat-template-file chat_template.jinja \
|
| 110 |
+
--no-webui \
|
| 111 |
+
--fit off
|
| 112 |
+
```
|
| 113 |
+
|
| 114 |
+
#### With DFlash Speculative Decoding
|
| 115 |
+
|
| 116 |
+
Accelerate token generation using the DFlash draft model:
|
| 117 |
+
|
| 118 |
+
```bash
|
| 119 |
+
llama-server \
|
| 120 |
+
-m nemotron-3.5-lightning-30b-a3b-Q4_K_M.gguf \
|
| 121 |
+
-md nemotron-3.5-lightning-30b-a3b-dflash.gguf \
|
| 122 |
+
--draft-max 6 \
|
| 123 |
+
--temp 1.0 --top-p 0.95 \
|
| 124 |
+
-np 1 \
|
| 125 |
+
-c 40960 \
|
| 126 |
+
--port 8000 \
|
| 127 |
+
-ngl 99 \
|
| 128 |
+
-ngld 99 \
|
| 129 |
+
-fa on \
|
| 130 |
+
--jinja \
|
| 131 |
+
--chat-template-file chat_template.jinja \
|
| 132 |
+
--no-webui \
|
| 133 |
+
--fit off
|
| 134 |
+
```
|
| 135 |
+
|
| 136 |
+
### API Client Example (OpenAI SDK)
|
| 137 |
+
|
| 138 |
+
```python
|
| 139 |
+
from openai import OpenAI
|
| 140 |
+
|
| 141 |
+
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
|
| 142 |
+
|
| 143 |
+
# Reasoning ON (default)
|
| 144 |
+
response = client.chat.completions.create(
|
| 145 |
+
model="nemotron-3.5-lightning-30b-a3b",
|
| 146 |
+
messages=[{"role": "user", "content": "Write a haiku about GPUs"}],
|
| 147 |
+
max_tokens=4096,
|
| 148 |
+
temperature=1.0,
|
| 149 |
+
top_p=0.95,
|
| 150 |
+
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
|
| 151 |
+
)
|
| 152 |
+
print(response.choices[0].message.content)
|
| 153 |
+
|
| 154 |
+
# Reasoning OFF (direct answer)
|
| 155 |
+
response = client.chat.completions.create(
|
| 156 |
+
model="nemotron-3.5-lightning-30b-a3b",
|
| 157 |
+
messages=[{"role": "user", "content": "What is the capital of Japan?"}],
|
| 158 |
+
max_tokens=128,
|
| 159 |
+
temperature=1.0,
|
| 160 |
+
top_p=0.95,
|
| 161 |
+
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
|
| 162 |
+
)
|
| 163 |
+
print(response.choices[0].message.content)
|
| 164 |
+
```
|
| 165 |
+
|
| 166 |
+
## Benchmarks
|
| 167 |
+
|
| 168 |
+
### Reasoning Benchmark Evaluations
|
| 169 |
+
|
| 170 |
+
| Task | Nemotron-3.5-Lightning-30B-A3B-BF16 | Nemotron-3.5-Lightning-30B-A3B-NVFP4 |
|
| 171 |
+
| --- | --- | --- |
|
| 172 |
+
| **General Knowledge** | | |
|
| 173 |
+
| MMLU Pro | 81.94 | 81.62 |
|
| 174 |
+
| AA-Omniscience | 17.50 | 16.63 |
|
| 175 |
+
| **Reasoning** | | |
|
| 176 |
+
| GPQA Diamond (no tools) | 75.44 | 75.57 |
|
| 177 |
+
| HLE (text-only, no tools) | 11.72 | 10.47 |
|
| 178 |
+
| SciCode | 32.60 | 31.38 |
|
| 179 |
+
| **Coding & Agentic** | | |
|
| 180 |
+
| SWE-bench Verified | 51.56 | 52.80 |
|
| 181 |
+
| SWE-bench Multilingual | 39.33 | 36.47 |
|
| 182 |
+
| Terminal-Bench 2.1 | 24.58 | 23.46 |
|
| 183 |
+
| PinchBench | 85.37 | 83.43 |
|
| 184 |
+
| BrowseComp | 36.97 | 36.81 |
|
| 185 |
+
| τ³-bench (Banking) | 9.28 | 9.48 |
|
| 186 |
+
| GDPval-AA-V2 | 832 | 865 |
|
| 187 |
+
| **Instruction Following** | | |
|
| 188 |
+
| IFBench (loose) | 71.88 | 72.88 |
|
| 189 |
+
| **Long Context** | | |
|
| 190 |
+
| AA-LCR | 52.00 | 49.19 |
|
| 191 |
+
|
| 192 |
+
## License and Terms of Use
|
| 193 |
+
|
| 194 |
+
Governing Download Terms: Use of this model is governed by the OpenMDW-1.1 model license.
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,190 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{% macro render_extra_keys(json_dict, handled_keys) %}
|
| 2 |
+
{%- if json_dict is mapping %}
|
| 3 |
+
{%- for json_key in json_dict if json_key not in handled_keys %}
|
| 4 |
+
{%- if json_dict[json_key] is mapping or (json_dict[json_key] is sequence and json_dict[json_key] is not string) %}
|
| 5 |
+
{{- '\n<' ~ json_key ~ '>' ~ (json_dict[json_key] | tojson | safe) ~ '</' ~ json_key ~ '>' }}
|
| 6 |
+
{%- else %}
|
| 7 |
+
{{-'\n<' ~ json_key ~ '>' ~ (json_dict[json_key] | string) ~ '</' ~ json_key ~ '>' }}
|
| 8 |
+
{%- endif %}
|
| 9 |
+
{%- endfor %}
|
| 10 |
+
{%- endif %}
|
| 11 |
+
{% endmacro %}
|
| 12 |
+
{%- set enable_thinking = enable_thinking if enable_thinking is defined else True %}
|
| 13 |
+
{%- set truncate_history_thinking = truncate_history_thinking if truncate_history_thinking is defined else True %}
|
| 14 |
+
{%- set ns = namespace(last_user_idx = -1) %}
|
| 15 |
+
{%- set loop_messages = messages %}
|
| 16 |
+
{%- for m in loop_messages %}
|
| 17 |
+
{%- if m["role"] == "user" %}
|
| 18 |
+
{%- set ns.last_user_idx = loop.index0 %}
|
| 19 |
+
{%- endif %}
|
| 20 |
+
{%- endfor %}
|
| 21 |
+
{%- if messages[0]["role"] == "system" %}
|
| 22 |
+
{%- set system_message = messages[0]["content"] %}
|
| 23 |
+
{%- set loop_messages = messages[1:] %}
|
| 24 |
+
{%- else %}
|
| 25 |
+
{%- set system_message = "" %}
|
| 26 |
+
{%- set loop_messages = messages %}
|
| 27 |
+
{%- endif %}
|
| 28 |
+
{%- if not tools is defined %}
|
| 29 |
+
{%- set tools = [] %}
|
| 30 |
+
{%- endif %}
|
| 31 |
+
{%- set ns = namespace(last_user_idx = -1) %}
|
| 32 |
+
{%- for m in loop_messages %}
|
| 33 |
+
{%- if m["role"] == "user" %}
|
| 34 |
+
{%- set ns.last_user_idx = loop.index0 %}
|
| 35 |
+
{%- endif %}
|
| 36 |
+
{%- endfor %}
|
| 37 |
+
{%- if system_message is defined %}
|
| 38 |
+
{{- "<|im_start|>system\n" + system_message }}
|
| 39 |
+
{%- else %}
|
| 40 |
+
{%- if tools is iterable and tools | length > 0 %}
|
| 41 |
+
{{- "<|im_start|>system\n" }}
|
| 42 |
+
{%- endif %}
|
| 43 |
+
{%- endif %}
|
| 44 |
+
{%- if tools is iterable and tools | length > 0 %}
|
| 45 |
+
{%- if system_message is defined and system_message | length > 0 %}
|
| 46 |
+
{{- "\n\n" }}
|
| 47 |
+
{%- endif %}
|
| 48 |
+
{{- "# Tools\n\nYou have access to the following functions:\n\n" }}
|
| 49 |
+
{{- "<tools>" }}
|
| 50 |
+
{%- for tool in tools %}
|
| 51 |
+
{%- if tool.function is defined %}
|
| 52 |
+
{%- set tool = tool.function %}
|
| 53 |
+
{%- endif %}
|
| 54 |
+
{{- "\n<function>\n<name>" ~ tool.name ~ "</name>" }}
|
| 55 |
+
{%- if tool.description is defined %}
|
| 56 |
+
{{- '\n<description>' ~ (tool.description | trim) ~ '</description>' }}
|
| 57 |
+
{%- endif %}
|
| 58 |
+
{{- '\n<parameters>' }}
|
| 59 |
+
{%- if tool.parameters is defined and tool.parameters is mapping and tool.parameters.properties is defined and tool.parameters.properties is mapping %}
|
| 60 |
+
{%- for param_name, param_fields in tool.parameters.properties|items %}
|
| 61 |
+
{{- '\n<parameter>' }}
|
| 62 |
+
{{- '\n<name>' ~ param_name ~ '</name>' }}
|
| 63 |
+
{%- if param_fields.type is defined %}
|
| 64 |
+
{{- '\n<type>' ~ (param_fields.type | string) ~ '</type>' }}
|
| 65 |
+
{%- endif %}
|
| 66 |
+
{%- if param_fields.description is defined %}
|
| 67 |
+
{{- '\n<description>' ~ (param_fields.description | trim) ~ '</description>' }}
|
| 68 |
+
{%- endif %}
|
| 69 |
+
{%- if param_fields.enum is defined %}
|
| 70 |
+
{{- '\n<enum>' ~ (param_fields.enum | tojson | safe) ~ '</enum>' }}
|
| 71 |
+
{%- endif %}
|
| 72 |
+
{%- set handled_keys = ['name', 'type', 'description', 'enum'] %}
|
| 73 |
+
{{- render_extra_keys(param_fields, handled_keys) }}
|
| 74 |
+
{{- '\n</parameter>' }}
|
| 75 |
+
{%- endfor %}
|
| 76 |
+
{%- endif %}
|
| 77 |
+
{% set handled_keys = ['type', 'properties', 'required'] %}
|
| 78 |
+
{{- render_extra_keys(tool.parameters, handled_keys) }}
|
| 79 |
+
{%- if tool.parameters is defined and tool.parameters.required is defined %}
|
| 80 |
+
{{- '\n<required>' ~ (tool.parameters.required | tojson | safe) ~ '</required>' }}
|
| 81 |
+
{%- endif %}
|
| 82 |
+
{{- '\n</parameters>' }}
|
| 83 |
+
{%- set handled_keys = ['type', 'name', 'description', 'parameters'] %}
|
| 84 |
+
{{- render_extra_keys(tool, handled_keys) }}
|
| 85 |
+
{{- '\n</function>' }}
|
| 86 |
+
{%- endfor %}
|
| 87 |
+
{{- "\n</tools>" }}
|
| 88 |
+
{{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
|
| 89 |
+
{%- endif %}
|
| 90 |
+
{%- if system_message is defined %}
|
| 91 |
+
{{- '<|im_end|>\n' }}
|
| 92 |
+
{%- else %}
|
| 93 |
+
{%- if tools is iterable and tools | length > 0 %}
|
| 94 |
+
{{- '<|im_end|>\n' }}
|
| 95 |
+
{%- endif %}
|
| 96 |
+
{%- endif %}
|
| 97 |
+
{%- for message in loop_messages %}
|
| 98 |
+
{%- if message.role == "assistant" %}
|
| 99 |
+
{%- if message.reasoning_content is defined and message.reasoning_content is string and message.reasoning_content | trim | length > 0 %}
|
| 100 |
+
{%- set content = "<think>\n" ~ message.reasoning_content ~ "</think>" ~ (message.content | default('', true)) %}
|
| 101 |
+
{%- else %}
|
| 102 |
+
{%- set content = message.content | default('', true) %}
|
| 103 |
+
{%- if content is string -%}
|
| 104 |
+
{%- if '<think>' not in content and '</think>' not in content -%}
|
| 105 |
+
{%- set content = "<think></think>" ~ content -%}
|
| 106 |
+
{%- endif -%}
|
| 107 |
+
{%- else -%}
|
| 108 |
+
{%- set content = content -%}
|
| 109 |
+
{%- endif -%}
|
| 110 |
+
{%- endif %}
|
| 111 |
+
{%- if message.tool_calls is defined and message.tool_calls is iterable and message.tool_calls | length > 0 %}
|
| 112 |
+
{{- '<|im_start|>assistant\n' }}
|
| 113 |
+
{%- set include_content = not (truncate_history_thinking and loop.index0 < ns.last_user_idx) %}
|
| 114 |
+
{%- if content is string and content | trim | length > 0 %}
|
| 115 |
+
{%- if include_content %}
|
| 116 |
+
{{- (content | trim) ~ '\n' -}}
|
| 117 |
+
{%- else %}
|
| 118 |
+
{%- set c = (content | string) %}
|
| 119 |
+
{%- if '</think>' in c %}
|
| 120 |
+
{%- set c = c.split('</think>')[-1] %}
|
| 121 |
+
{%- elif '<think>' in c %}
|
| 122 |
+
{%- set c = c.split('<think>')[0] %}
|
| 123 |
+
{%- endif %}
|
| 124 |
+
{%- set c = "<think></think>" ~ c %}
|
| 125 |
+
{%- if c | length > 0 %}
|
| 126 |
+
{{- c ~ '\n' -}}
|
| 127 |
+
{%- endif %}
|
| 128 |
+
{%- endif %}
|
| 129 |
+
{%- else %}
|
| 130 |
+
{{- "<think></think>" -}}
|
| 131 |
+
{%- endif %}
|
| 132 |
+
{%- for tool_call in message.tool_calls %}
|
| 133 |
+
{%- if tool_call.function is defined %}
|
| 134 |
+
{%- set tool_call = tool_call.function %}
|
| 135 |
+
{%- endif %}
|
| 136 |
+
{{- '<tool_call>\n<function=' ~ tool_call.name ~ '>\n' -}}
|
| 137 |
+
{%- if tool_call.arguments is defined %}
|
| 138 |
+
{%- for args_name, args_value in tool_call.arguments|items %}
|
| 139 |
+
{{- '<parameter=' ~ args_name ~ '>\n' -}}
|
| 140 |
+
{%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
|
| 141 |
+
{{- args_value ~ '\n</parameter>\n' -}}
|
| 142 |
+
{%- endfor %}
|
| 143 |
+
{%- endif %}
|
| 144 |
+
{{- '</function>\n</tool_call>\n' -}}
|
| 145 |
+
{%- endfor %}
|
| 146 |
+
{{- '<|im_end|>\n' }}
|
| 147 |
+
{%- else %}
|
| 148 |
+
{%- if not (truncate_history_thinking and loop.index0 < ns.last_user_idx) %}
|
| 149 |
+
{{- '<|im_start|>assistant\n' ~ (content | default('', true) | string | trim) ~ '<|im_end|>\n' }}
|
| 150 |
+
{%- else %}
|
| 151 |
+
{%- set c = (content | default('', true) | string) %}
|
| 152 |
+
{%- if '<think>' in c and '</think>' in c %}
|
| 153 |
+
{%- set c = "<think></think>" ~ c.split('</think>')[-1] %}
|
| 154 |
+
{%- endif %}
|
| 155 |
+
{%- set c = c | trim %}
|
| 156 |
+
{%- if c | length > 0 %}
|
| 157 |
+
{{- '<|im_start|>assistant\n' ~ c ~ '<|im_end|>\n' }}
|
| 158 |
+
{%- else %}
|
| 159 |
+
{{- '<|im_start|>assistant\n<|im_end|>\n' }}
|
| 160 |
+
{%- endif %}
|
| 161 |
+
{%- endif %}
|
| 162 |
+
{%- endif %}
|
| 163 |
+
{%- elif message.role == "user" or message.role == "system" %}
|
| 164 |
+
{{- '<|im_start|>' + message.role + '\n' }}
|
| 165 |
+
{%- set content = message.content | string %}
|
| 166 |
+
{{- content }}
|
| 167 |
+
{{- '<|im_end|>\n' }}
|
| 168 |
+
{%- elif message.role == "tool" %}
|
| 169 |
+
{%- if loop.previtem and loop.previtem.role != "tool" %}
|
| 170 |
+
{{- '<|im_start|>user\n' }}
|
| 171 |
+
{%- endif %}
|
| 172 |
+
{{- '<tool_response>\n' }}
|
| 173 |
+
{{- message.content }}
|
| 174 |
+
{{- '\n</tool_response>\n' }}
|
| 175 |
+
{%- if not loop.last and loop.nextitem.role != "tool" %}
|
| 176 |
+
{{- '<|im_end|>\n' }}
|
| 177 |
+
{%- elif loop.last %}
|
| 178 |
+
{{- '<|im_end|>\n' }}
|
| 179 |
+
{%- endif %}
|
| 180 |
+
{%- else %}
|
| 181 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>\n' }}
|
| 182 |
+
{%- endif %}
|
| 183 |
+
{%- endfor %}
|
| 184 |
+
{%- if add_generation_prompt %}
|
| 185 |
+
{%- if enable_thinking %}
|
| 186 |
+
{{- '<|im_start|>assistant\n<think>\n' }}
|
| 187 |
+
{%- else %}
|
| 188 |
+
{{- '<|im_start|>assistant\n<think></think>' }}
|
| 189 |
+
{%- endif %}
|
| 190 |
+
{%- endif %}
|
nemotron-3.5-lightning-30b-a3b-Q4_0.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ef78f377618e020d3bb50fc831fd0edc113b468aea87a9996aa99e683c3a3a0c
|
| 3 |
+
size 18728781472
|
nemotron-3.5-lightning-30b-a3b-Q4_K_M.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:50d2ce39590c44b2db38531faf03eb9956658d58b8e778add38eb8f74d67959b
|
| 3 |
+
size 25430739616
|
nemotron-3.5-lightning-30b-a3b-Q4_K_S.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:12294213e4a9001140c75f01be7ead4ebcfba443bfd99707c6ddcb7d5414c181
|
| 3 |
+
size 22835894944
|
nemotron-3.5-lightning-30b-a3b-Q5_K_M.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:920fa53845bca15e0461244c39b86d09c8d5eb639e7e58abfbeba1c6bf739caa
|
| 3 |
+
size 27040754848
|
nemotron-3.5-lightning-30b-a3b-Q5_K_S.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0cfd27ccf850ee7ad8534b7eec46ec2b13b9f11c628a9965ab1b699232512506
|
| 3 |
+
size 24810682528
|
nemotron-3.5-lightning-30b-a3b-Q6_K.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:56bcf2d5062f69593a42da794714da9a29a7c7dd3f7b78d8ed1cca1f4a192738
|
| 3 |
+
size 34921148320
|
nemotron-3.5-lightning-30b-a3b-Q8_0.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:de89886ff4c2c9c8dcc9049bff00732f2b393f2af4161b69ec4a0fefbedeafbf
|
| 3 |
+
size 35004642976
|
nemotron-3.5-lightning-30b-a3b-bf16.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7e0330212140b1c6f9b20bcb208b26b4cccdcfb51dc2c3bf1b5afb7663ed928e
|
| 3 |
+
size 65852184736
|
nemotron-3.5-lightning-30b-a3b-dflash.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:1d8344244f363ec678ce17eb93443761cbb9226779a76bf162243fc3f73f9f8a
|
| 3 |
+
size 1185032704
|