GGUF
conversational

Llama-3.1-Swallow-8B-Instruct-v0.5-Q4_K_M-GGUF

This is a Q4_K_M GGUF quantization of tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.5, converted using llama.cpp.

Usage with Ollama

ollama run hf.co/FalconSuzuki/Llama-3.1-Swallow-8B-Instruct-v0.5-Q4_K_M-GGUF:Q4_K_M

Usage with llama.cpp

llama-cli -m llama-3.1-swallow-8b-instruct-v0.5-Q4_K_M.gguf -cnv -c 4096

Note: the model's native context length is 131072, but a smaller context size (e.g. 4096) is recommended for typical local use to avoid excessive memory usage.

License

Please refer to the original model's license terms (Llama 3.3 license and Gemma Terms of Use).

Downloads last month
143
GGUF
Model size
8B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FalconSuzuki/Llama-3.1-Swallow-8B-Instruct-v0.5-Q4_K_M-GGUF