error Add 4-bit / 8-bit quantization support (bitsandbytes) for 8GB VRAM GPUs
#14
by netforcetech - opened
Hi! The current model requires too much VRAM. On an RTX 3070 Ti (8GB VRAM), it overflows into the shared system RAM, making generation extremely slow (around 26 seconds per token).
Could you please add --load-in-4bit or --load-in-8bit flags using bitsandbytes to the start.py script so users with 8GB cards can run it completely in VRAM? Thanks!