File size: 1,708 Bytes
84b7252
 
 
 
 
 
 
 
 
 
 
 
 
e11536e
e2918ac
84b7252
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
069037f
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
---
library_name: mlx
base_model: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
license: other
pipeline_tag: text-generation
tags:
- mlx
- mlx-lm
- quantized
- 6-bit
- base_model:nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
---

[![Open in MLXHub](https://mlxhub.app/assets/badge-open-in-mlxhub.svg)](https://mlxhub.app/open-model?repo=DreamFoundries/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-6bit)

# NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 MLX 6-bit

This repository contains an MLX-LM conversion of [nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16).

## Conversion Details

- Original model: `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16`
- Model family: Nemotron 3
- Source model type: `nemotron_h`
- Model size: 31,577,937,344 parameters
- Quantization: MLX-LM affine quantization
- Bits: 6-bit
- Group size: 64
- Local MLX folder size at upload time: 23.92 GiB
- Local safetensors weight size at upload time: 23.90 GiB

## Usage

```bash
mlx_lm.generate --model DreamFoundries/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-6bit --prompt "Hello" --max-tokens 64
```

## Benchmarks

No comparative benchmarks have been run yet. The repository does not currently provide quality, speed, memory, or benchmark comparisons against the original weights or other quantizations.

## License

This is a converted/quantized derivative of the original model. Please refer to the original model repository for the upstream license and usage terms: https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

---

[![Download MLXHub](https://mlxhub.app/assets/badge-download-mlxhub.svg)](https://apps.apple.com/app/apple-store/id6766485144?pt=121945436&ct=HuggingFace&mt=8)