⚑ PocketWeights: Qwen2.5-0.5B-Instruct (GGUF)

Heavy models, made light.

This repository provides high-quality, optimized GGUF quantizations of the ultra-lightweight Qwen2.5 0.5B architecture. By fusing the instruct and base models, this build aims to retain peak conversational alignment while achieving an extreme micro-footprint suitable for background tasks and edge hardware.


πŸ“– Model Details

  • Architecture: Qwen2.5 (0.5 Billion Parameters)
  • Lineage: Fused from Qwen/Qwen2.5-0.5B-Instruct and Qwen/Qwen2.5-0.5B
  • Quantization Engine: llama.cpp
  • Target Hardware: Raspberry Pi, older smartphones, IoT edge devices, micro-controllers, and fast background agent execution.

πŸ“¦ Available Files & Formats

File Name Quant Type Precision Recommended Use
model-Q4_K_M.gguf Q4_K_M 4-bit Medium Maximum speed and minimal size (< 500 MB RAM).
model-Q6_K.gguf Q6_K 6-bit High fidelity with low perplexity loss.
model-Q8_0.gguf Q8_0 8-bit Near-lossless precision (Highest accuracy).

Note: You can rename these files locally to Qwen2.5-0.5B-Q4_K_M.gguf after downloading if preferred.


πŸš€ Quick Start Guide

Run with Ollama

You can stream and run these weights directly from Hugging Face using Ollama:

# Recommended 4-bit quantization
ollama run hf.co/PocketWeights/PocketWeights-Qwen2.5-0.5B-GGUF:model-Q4_K_M

# Maximum 8-bit quality
ollama run hf.co/PocketWeights/PocketWeights-Qwen2.5-0.5B-GGUF:model-Q8_0

Run with llama.cpp

./llama-cli -m model-Q4_K_M.gguf -p "You are a helpful assistant." -cnv

🀝 Support the PocketWeights Mission

I build, verify, and maintain these quantization pipelines to provide high-quality, unrestricted, and hardware-friendly models to the open-source community for free.

Running conversion setups, cloud instances, and storage requires ongoing resources. If these weights have saved you time, compute overhead, or API bills, please consider supporting the project with a small tip!

β˜• Donation Options

Ko-fi: ko-fi.com/iamvishalnarayan

Web3 / Crypto (Polygon / ETH):

0x4FC189bf839A89259dd28DE8cD97883c49e15615

Tip: Sending via the Polygon network keeps transfer gas fees below $0.01!

Downloads last month
91
GGUF
Model size
0.5B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for PocketWeights/PocketWeights-Qwen2.5-0.5B-GGUF

Quantized
(116)
this model