File size: 2,140 Bytes
ab0079a
 
 
d20c5aa
 
 
 
 
ab0079a
 
d20c5aa
ab0079a
d20c5aa
ab0079a
d20c5aa
ab0079a
d20c5aa
ab0079a
d20c5aa
 
 
ab0079a
d20c5aa
 
 
 
 
 
ab0079a
d20c5aa
ab0079a
d20c5aa
ab0079a
d20c5aa
 
 
ab0079a
d20c5aa
 
 
ab0079a
d20c5aa
 
 
 
 
 
 
ab0079a
d20c5aa
ab0079a
 
d20c5aa
ab0079a
d20c5aa
 
 
 
 
 
 
ab0079a
d20c5aa
ab0079a
d20c5aa
 
 
ab0079a
d20c5aa
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
---
license: gemma
tags:
  - bigsmall
  - compression
  - lossless
  - gemma
  - google
---

# Gemma 2 9B Instruct (BigSmall compressed)

**17.2 GB -> 11.2 GB (BF16). Lossless. Zero inference overhead. Any hardware.**

Compressed with [BigSmall](https://github.com/wpferrell/Bigsmall) -- decompresses once at load time, runs at full native speed. Every weight is bit-identical to the original.

## Quick start

`ash
pip install bigsmall
`

`python
import bigsmall
bigsmall.install_hook()
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("wpferrell/gemma-2-9b-it-bigsmall")
`

## Streaming loader -- run on any hardware

BigSmall's streaming loader decompresses one layer at a time directly into VRAM. Peak memory is one layer -- not the whole model. A 4 GB GPU can run Mistral 7B losslessly.

`python
from bigsmall import StreamingLoader
from transformers import AutoModelForCausalLM

with StreamingLoader("wpferrell/gemma-2-9b-it-bigsmall", device="cuda") as loader:
    model = loader.load_model(AutoModelForCausalLM)
`

| Your GPU | Models you can run |
|----------|--------------------|
| 2 GB | Small models, GPT-2, Gemma 270M |
| 4 GB | Mistral 7B, Llama 3.1 8B, Gemma 2B, Llama 3.2 3B |
| 8 GB | Qwen 2.5 14B, Gemma 2 9B |
| 24 GB | Llama 70B, Qwen 72B, DeepSeek V4-Flash |
| CPU only | Everything -- slower but full quality |

BigSmall is the only lossless compression tool with a streaming loader. DFloat11 and ZipNN load the full model into memory.


## Why BigSmall vs DFloat11

| | BigSmall | DFloat11 |
|--|--|--|
| Inference overhead | **None** | ~2x at batch=1 |
| Hardware | **CPU, Apple Silicon, AMD, any GPU** | CUDA only |
| FP32 support | **Yes** | No |
| Fine-tuning safe | **Yes** | No |
| Streaming loader | **Yes -- peak RAM < 2 GB** | No |

## Compression stats

| Original | Compressed | Ratio | Format | Verified |
|----------|------------|-------|--------|---------|
| 17.2 GB | 11.2 GB | 65.1% | BF16 | md5 every tensor |

- GitHub: [wpferrell/Bigsmall](https://github.com/wpferrell/Bigsmall)
- All models: [huggingface.co/wpferrell](https://huggingface.co/wpferrell)