NANI-Nithin commited on
Commit
99ca745
·
verified ·
1 Parent(s): 3fb3aca

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +159 -0
README.md ADDED
@@ -0,0 +1,159 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model:
4
+ - ai9stars/G9v3-3B
5
+ library_name: llama.cpp
6
+ tags:
7
+ - gguf
8
+ - llama-cpp
9
+ - text-generation
10
+ - conversational
11
+ - quantized
12
+ - 3b
13
+ - ai9stars
14
+ language:
15
+ - en
16
+ pipeline_tag: text-generation
17
+ quantized_by: NANI-Nithin
18
+ ---
19
+
20
+ # G9v3-3B-GGUF
21
+
22
+ GGUF quantized releases of **[ai9stars/G9v3-3B](https://huggingface.co/ai9stars/G9v3-3B)** for llama.cpp and compatible runtimes.
23
+
24
+ ## Model Information
25
+
26
+ - **Base Model:** ai9stars/G9v3-3B
27
+ - **Parameter Size:** 3B
28
+ - **Format:** GGUF
29
+ - **Quantized By:** NANI-Nithin
30
+ - **Quantization Tool:** llama.cpp
31
+
32
+ ## Available Files
33
+
34
+ ### 2-bit
35
+
36
+ - Q2_K
37
+ - IQ2_M
38
+ - Q2_K_L
39
+
40
+ ### 3-bit
41
+
42
+ - IQ3_XXS
43
+ - IQ3_XS
44
+ - Q3_K_S
45
+ - IQ3_M
46
+ - Q3_K_M
47
+ - Q3_K_L
48
+ - Q3_K_XL
49
+
50
+ ### 4-bit
51
+
52
+ - IQ4_XS
53
+ - IQ4_NL
54
+ - Q4_0
55
+ - Q4_1
56
+ - Q4_K_S
57
+ - Q4_K_M
58
+
59
+ ### 5-bit
60
+
61
+ - Q5_K_S
62
+ - Q5_K_M
63
+
64
+ ### 6-bit
65
+
66
+ - Q6_K
67
+ - Q6_K_L
68
+
69
+ ### 8-bit
70
+
71
+ - Q8_0
72
+
73
+ ### Full Precision
74
+
75
+ - F16/BF16 GGUF
76
+
77
+ ## Recommended Quantizations
78
+
79
+ ### Best Overall
80
+
81
+ **Q4_K_M**
82
+
83
+ Recommended for most users. Excellent balance of quality, memory usage, and speed.
84
+
85
+ ### Higher Quality
86
+
87
+ **Q5_K_M** or **Q6_K**
88
+
89
+ For users seeking maximum quality while still benefiting from quantization.
90
+
91
+ ### Best IQ Quant
92
+
93
+ **IQ4_NL**
94
+
95
+ Excellent quality-per-GB and one of the strongest modern importance-aware quantizations.
96
+
97
+ ### Low Memory Systems
98
+
99
+ **IQ3_M** or **Q3_K_M**
100
+
101
+ Good balance of usability and reduced memory requirements.
102
+
103
+ ## IQ Quantizations
104
+
105
+ The following importance-aware quantizations are included:
106
+
107
+ - IQ2_M
108
+ - IQ3_XXS
109
+ - IQ3_XS
110
+ - IQ3_M
111
+ - IQ4_XS
112
+ - IQ4_NL
113
+
114
+ These quantizations were generated using an importance matrix (imatrix) calibration pass and typically provide improved quality retention compared to traditional quantization methods at similar file sizes.
115
+
116
+ ## Usage
117
+
118
+ ### llama.cpp
119
+
120
+ ```bash
121
+ ./llama-cli \
122
+ -m G9v3-3B-Q4_K_M.gguf \
123
+ -p "Hello!"
124
+ ```
125
+
126
+ ### Ollama
127
+
128
+ Create a Modelfile:
129
+
130
+ ```text
131
+ FROM ./G9v3-3B-Q4_K_M.gguf
132
+ ```
133
+
134
+ Then:
135
+
136
+ ```bash
137
+ ollama create g9v3-3b -f Modelfile
138
+ ollama run g9v3-3b
139
+ ```
140
+
141
+ ### LM Studio
142
+
143
+ Download the desired GGUF file and import it directly into LM Studio.
144
+
145
+ ## Credits
146
+
147
+ - Original Model: **ai9stars/G9v3-3B**
148
+ - GGUF Conversion & Quantization: **NANI-Nithin**
149
+ - Quantization Framework: **llama.cpp**
150
+
151
+ ## Disclaimer
152
+
153
+ This repository contains converted GGUF files only.
154
+
155
+ Please refer to the original model repository for licensing terms, training methodology, benchmark results, intended use, limitations, and safety information.
156
+
157
+ Original model:
158
+
159
+ https://huggingface.co/ai9stars/G9v3-3B