How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf MISHANM/google-gemma-2-2b-it.gguf
# Run inference directly in the terminal:
llama cli -hf MISHANM/google-gemma-2-2b-it.gguf
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf MISHANM/google-gemma-2-2b-it.gguf
# Run inference directly in the terminal:
llama cli -hf MISHANM/google-gemma-2-2b-it.gguf
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf MISHANM/google-gemma-2-2b-it.gguf
# Run inference directly in the terminal:
./llama-cli -hf MISHANM/google-gemma-2-2b-it.gguf
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf MISHANM/google-gemma-2-2b-it.gguf
# Run inference directly in the terminal:
./build/bin/llama-cli -hf MISHANM/google-gemma-2-2b-it.gguf
Use Docker
docker model run hf.co/MISHANM/google-gemma-2-2b-it.gguf
Quick Links

MISHANM/google-gemma-2-2b-it.gguf

This model is a GGUF version of the Google gemma-2-2b-it model, optimized for use with the llama.cpp framework. It is designed to run efficiently on CPUs and can be used for various natural language processing tasks.

Model Details

  1. Language: English
  2. Tasks: Text generation
  3. Base Model: google/gemma-2-2b-it

Building and Running the Model

To build and run the model using llama.cpp, follow these steps:

Build llama.cpp Locally

git clone https://github.com/ggerganov/llama.cpp  
cd llama.cpp  
cmake -B build  
cmake --build build --config Release  

Run the Model

Navigate to the build directory and run the model with a prompt:

cd llama.cpp/build/bin   

Inference with llama.cpp

./llama-cli -m /path/to/model/ -p "Your prompt here" -n 128  

Citation Information

@misc{MISHANM/google-gemma-2-2b-it.gguf,
  author = {Mishan Maurya},
  title = {Introducing Google gemma-2-2b-it GGUF Model},
  year = {2025},
  publisher = {Hugging Face},
  journal = {Hugging Face repository},
  
}
Downloads last month
44
GGUF
Model size
3B params
Architecture
gemma2
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for MISHANM/google-gemma-2-2b-it.gguf

Quantized
(201)
this model