--- license: apache-2.0 base_model: - ai9stars/G9v3-3B library_name: llama.cpp tags: - gguf - llama-cpp - text-generation - conversational - quantized - 3b - ai9stars language: - en pipeline_tag: text-generation quantized_by: NANI-Nithin --- # G9v3-3B-GGUF GGUF quantized releases of **[ai9stars/G9v3-3B](https://huggingface.co/ai9stars/G9v3-3B)** for llama.cpp and compatible runtimes. ## Model Information - **Base Model:** ai9stars/G9v3-3B - **Parameter Size:** 3B - **Format:** GGUF - **Quantized By:** NANI-Nithin - **Quantization Tool:** llama.cpp ## Available Files ### 2-bit - Q2_K - IQ2_M - Q2_K_L ### 3-bit - IQ3_XXS - IQ3_XS - Q3_K_S - IQ3_M - Q3_K_M - Q3_K_L - Q3_K_XL ### 4-bit - IQ4_XS - IQ4_NL - Q4_0 - Q4_1 - Q4_K_S - Q4_K_M ### 5-bit - Q5_K_S - Q5_K_M ### 6-bit - Q6_K - Q6_K_L ### 8-bit - Q8_0 ### Full Precision - F16/BF16 GGUF ## Recommended Quantizations ### Best Overall **Q4_K_M** Recommended for most users. Excellent balance of quality, memory usage, and speed. ### Higher Quality **Q5_K_M** or **Q6_K** For users seeking maximum quality while still benefiting from quantization. ### Best IQ Quant **IQ4_NL** Excellent quality-per-GB and one of the strongest modern importance-aware quantizations. ### Low Memory Systems **IQ3_M** or **Q3_K_M** Good balance of usability and reduced memory requirements. ## IQ Quantizations The following importance-aware quantizations are included: - IQ2_M - IQ3_XXS - IQ3_XS - IQ3_M - IQ4_XS - IQ4_NL These quantizations were generated using an importance matrix (imatrix) calibration pass and typically provide improved quality retention compared to traditional quantization methods at similar file sizes. ## Usage ### llama.cpp ```bash ./llama-cli \ -m G9v3-3B-Q4_K_M.gguf \ -p "Hello!" ``` ### Ollama Create a Modelfile: ```text FROM ./G9v3-3B-Q4_K_M.gguf ``` Then: ```bash ollama create g9v3-3b -f Modelfile ollama run g9v3-3b ``` ### LM Studio Download the desired GGUF file and import it directly into LM Studio. ## Credits - Original Model: **ai9stars/G9v3-3B** - GGUF Conversion & Quantization: **NANI-Nithin** - Quantization Framework: **llama.cpp** ## Disclaimer This repository contains converted GGUF files only. Please refer to the original model repository for licensing terms, training methodology, benchmark results, intended use, limitations, and safety information. Original model: https://huggingface.co/ai9stars/G9v3-3B