jamesdumay commited on
Commit
c27196a
·
verified ·
1 Parent(s): 961f341

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +126 -0
README.md ADDED
@@ -0,0 +1,126 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: mesh-llm
3
+ license: "other"
4
+ base_model:
5
+ - "unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF"
6
+ pipeline_tag: "text-generation"
7
+ tags:
8
+ - gguf
9
+ - mesh-llm
10
+ - layer-package
11
+ - skippy
12
+ - distributed-inference
13
+ - local-inference
14
+ - openai-compatible
15
+ ---
16
+
17
+ <div align="center">
18
+ <a href="https://www.meshllm.cloud">
19
+ <img src="https://meshllm.cloud/assets/images/jelly-logo-wordmark.png" alt="Mesh LLM" width="220">
20
+ </a>
21
+
22
+ <h1>NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_XL</h1>
23
+
24
+ <p>
25
+ <strong>Distributed GGUF inference package for Mesh LLM</strong>
26
+ </p>
27
+
28
+ <p>
29
+ <a href="https://www.meshllm.cloud"><img alt="Website" src="https://img.shields.io/badge/Website-meshllm.cloud-111111?style=for-the-badge"></a>
30
+ <a href="https://github.com/Mesh-LLM/mesh-llm"><img alt="GitHub" src="https://img.shields.io/badge/GitHub-Mesh--LLM-24292f?style=for-the-badge&logo=github"></a>
31
+ <a href="https://discord.gg/rs6fmc63eN"><img alt="Discord" src="https://img.shields.io/badge/Discord-Join-5865F2?style=for-the-badge&logo=discord&logoColor=white"></a>
32
+ </p>
33
+ </div>
34
+
35
+ GGUF layer package for running **NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_XL** across a local Mesh LLM cluster.
36
+
37
+ This package is derived from [unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF](https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF) and keeps the original GGUF distribution split into per-layer artifacts for distributed inference.
38
+
39
+ ## Highlights
40
+
41
+ | Run locally | Pool multiple machines | OpenAI-compatible | Package variant |
42
+ |---|---|---|---|
43
+ | Private inference on your hardware | Split layers across peers | Serve `/v1/chat/completions` locally | `UD-Q4_K_XL` layer package |
44
+
45
+ ## Model Overview
46
+
47
+ | Property | Value |
48
+ |---|---|
49
+ | **Source model** | [unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF](https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF) |
50
+ | **Model id** | `unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF:UD-Q4_K_XL` |
51
+ | **Family** | NVIDIA |
52
+ | **Parameter scale** | 120B-A12B |
53
+ | **Quantization** | `UD-Q4_K_XL` |
54
+ | **Layer count** | 88 |
55
+ | **Activation width** | not recorded |
56
+ | **Package size** | 0 B |
57
+ | **Source file** | `UD-Q4_K_XL/NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_XL-00001-of-00003.gguf` |
58
+ | **Package repo** | [meshllm/NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_XL-layers](https://huggingface.co/meshllm/NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_XL-layers) |
59
+ | **License** | `other` from [unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF](https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF) |
60
+
61
+ ## Recommended Use
62
+
63
+ - Local and private inference with Mesh LLM.
64
+ - Multi-machine serving when the full GGUF is too large for one host.
65
+ - OpenAI-compatible chat/completions workflows through Mesh LLM's local API.
66
+
67
+ For upstream architecture details, chat template guidance, sampling recommendations, license terms, and benchmark notes, see the source model card: [unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF](https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF).
68
+
69
+ ## Quickstart
70
+
71
+ ```bash
72
+ # Run this on each machine that should contribute memory/compute.
73
+ mesh-llm serve --model "meshllm/NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_XL-layers" --split
74
+ ```
75
+
76
+ ```bash
77
+ # Check the mesh and discover the OpenAI-compatible model name.
78
+ curl -s http://localhost:3131/api/status
79
+ curl -s http://localhost:3131/v1/models
80
+ ```
81
+
82
+ ```bash
83
+ # Send an OpenAI-compatible chat request.
84
+ curl -s http://localhost:3131/v1/chat/completions \
85
+ -H "Content-Type: application/json" \
86
+ -d '{
87
+ "model": "unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF:UD-Q4_K_XL",
88
+ "messages": [{"role": "user", "content": "Write a tiny hello-world function in Rust."}],
89
+ "max_tokens": 128
90
+ }'
91
+ ```
92
+
93
+ ## Package Variant
94
+
95
+ | Property | Value |
96
+ |---|---|
97
+ | **Format** | `gguf` |
98
+ | **Canonical source ref** | `unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF@036038fb30334a2d56a146c6f0d4871ab5edccbb/UD-Q4_K_XL/NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_XL-00001-of-00003.gguf` |
99
+ | **Source revision** | `036038fb30334a2d56a146c6f0d4871ab5edccbb` |
100
+ | **Source SHA-256** | `f22083eb6b15acb52905308ab083e8b0cc38897005cc45e8881abd164580aac2` |
101
+ | **Skippy ABI** | `not recorded` |
102
+ | **Package manifest SHA-256** | `61e345feb4d3b4b435ae040bb94c6da9f8a05d94f54cbbdaf5bc7d7b2e781b60` |
103
+
104
+ ## What Is Included
105
+
106
+ | Artifact | Path | Contents | SHA-256 |
107
+ |---|---|---|---|
108
+ | Manifest | `model-package.json` | Package schema, source identity, checksums | `61e345feb4d3b4b435ae040bb94c6da9f8a05d94f54cbbdaf5bc7d7b2e781b60` |
109
+
110
+ ## Validation
111
+
112
+ Generated by the Mesh LLM HF Jobs splitter from `mesh-llm` ref `f932c4d1dc12b3e3a670d5f470cedd5cdcc5db39`.
113
+ Each artifact is checksummed as it is written, uploaded to this repository, and removed from the job workspace before the next artifact is produced.
114
+
115
+ ```bash
116
+ skippy-model-package write-package "/hf-cache/UD-Q4_K_XL/NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_XL-00001-of-00003.gguf" --out-dir "/tmp/meshllm-layer-job-meshllm_NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_XL-layers-1/package"
117
+ ```
118
+
119
+ ## Links
120
+
121
+ - Source model: [unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF](https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF)
122
+ - Mesh LLM website: [meshllm.cloud](https://www.meshllm.cloud)
123
+ - Mesh LLM: [github.com/Mesh-LLM/mesh-llm](https://github.com/Mesh-LLM/mesh-llm)
124
+ - Discord: [discord.gg/rs6fmc63eN](https://discord.gg/rs6fmc63eN)
125
+ - Package catalog: [meshllm/catalog](https://huggingface.co/datasets/meshllm/catalog)
126
+ - Package format: [layer-package-repos.md](https://github.com/Mesh-LLM/mesh-llm/blob/main/docs/specs/layer-package-repos.md)