Download README.md from wpferrell/gemma-2-9b-it-bigsmall: direct link, hf CLI and curl.
- Browser
- Download file 2.14 kB
-
https://huggingface.co/wpferrell/gemma-2-9b-it-bigsmall/resolve/3d382b36e5cb40837b026d1c056c30d3ffc090dc/README.md
- Command line
-
hf download hf://wpferrell/gemma-2-9b-it-bigsmall@3d382b36e5cb40837b026d1c056c30d3ffc090dc/README.md
-
curl -L -o README.md https://huggingface.co/wpferrell/gemma-2-9b-it-bigsmall/resolve/3d382b36e5cb40837b026d1c056c30d3ffc090dc/README.md
license: gemma
tags:
- bigsmall
- compression
- lossless
- gemma
- google
Gemma 2 9B Instruct (BigSmall compressed)
17.2 GB -> 11.2 GB (BF16). Lossless. Zero inference overhead. Any hardware.
Compressed with BigSmall -- decompresses once at load time, runs at full native speed. Every weight is bit-identical to the original.
Quick start
ash pip install bigsmall
python import bigsmall bigsmall.install_hook() from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("wpferrell/gemma-2-9b-it-bigsmall")
Streaming loader -- run on any hardware
BigSmall's streaming loader decompresses one layer at a time directly into VRAM. Peak memory is one layer -- not the whole model. A 4 GB GPU can run Mistral 7B losslessly.
`python from bigsmall import StreamingLoader from transformers import AutoModelForCausalLM
with StreamingLoader("wpferrell/gemma-2-9b-it-bigsmall", device="cuda") as loader: model = loader.load_model(AutoModelForCausalLM) `
| Your GPU | Models you can run |
|---|---|
| 2 GB | Small models, GPT-2, Gemma 270M |
| 4 GB | Mistral 7B, Llama 3.1 8B, Gemma 2B, Llama 3.2 3B |
| 8 GB | Qwen 2.5 14B, Gemma 2 9B |
| 24 GB | Llama 70B, Qwen 72B, DeepSeek V4-Flash |
| CPU only | Everything -- slower but full quality |
BigSmall is the only lossless compression tool with a streaming loader. DFloat11 and ZipNN load the full model into memory.
Why BigSmall vs DFloat11
| BigSmall | DFloat11 | |
|---|---|---|
| Inference overhead | None | ~2x at batch=1 |
| Hardware | CPU, Apple Silicon, AMD, any GPU | CUDA only |
| FP32 support | Yes | No |
| Fine-tuning safe | Yes | No |
| Streaming loader | Yes -- peak RAM < 2 GB | No |
Compression stats
| Original | Compressed | Ratio | Format | Verified |
|---|---|---|---|---|
| 17.2 GB | 11.2 GB | 65.1% | BF16 | md5 every tensor |
- GitHub: wpferrell/Bigsmall
- All models: huggingface.co/wpferrell