Instructions to use Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M # Run inference directly in the terminal: llama cli -hf Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M # Run inference directly in the terminal: llama cli -hf Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M
Use Docker
docker model run hf.co/Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix with Ollama:
ollama run hf.co/Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix with Docker Model Runner:
docker model run hf.co/Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M
- Lemonade
How to use Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Lewdiculous/Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix:Q4_K_M
Run and chat with the model
lemonade run user.Eris_PrimeV4-Vision-7B-GGUF-IQ-Imatrix-Q4_K_M
List all available models
lemonade list
- Atomic Chat
These are quants for an experimental model.
"Q4_K_M", "Q4_K_S", "IQ4_XS", "Q5_K_M", "Q5_K_S",
"Q6_K", "Q8_0", "IQ3_M", "IQ3_S", "IQ3_XXS"
Original model weights:
https://huggingface.co/Nitral-AI/Eris_PrimeV4-Vision-7B
Vision/multimodal capabilities:
If you want to use vision functionality:
- Make sure you are using the latest version of KoboldCpp.
To use the multimodal capabilities of this model, such as vision, you also need to load the specified mmproj file, you can get it here, it's also hosted in this repository inside the mmproj folder.
- You can load the mmproj by using the corresponding section in the interface:
- For CLI users, you can load the mmproj file by adding the respective flag to your usual command:
--mmproj your-mmproj-file.gguf
Quantization information:
Steps performed:
Base⇢ GGUF(F16)⇢ Imatrix-Data(F16)⇢ GGUF(Imatrix-Quants)
Using the latest llama.cpp at the time.
- Downloads last month
- 477
3-bit
4-bit
5-bit
6-bit
8-bit
16-bit



