How to use from
Pi
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf Zoont/InternVL3-2B-4-Bit-GGUF-with-mmproj:
Configure the model in Pi
# Install Pi:
npm install -g @earendil-works/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "llama-cpp": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "Zoont/InternVL3-2B-4-Bit-GGUF-with-mmproj:"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

GGUF Conversion & Quantization of OpenGVLab/InternVL3-2B (4-Bit Quantization)

This model is converted & quantized from OpenGVLab/InternVL3-2B using llama.cpp version 6217 (7a6e91ad)

All quants made using imatrix option with Bartowski's dataset

Model Details

For more details about the model, see its original model card

Downloads last month
163
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Zoont/InternVL3-2B-4-Bit-GGUF-with-mmproj

Quantized
(4)
this model