Text-to-Image
Transformers
English
SupraDiT
feature-extraction
small
supra
image
flux
img
t2i
from scratch
custom_code
Instructions to use SupraLabs/Supra2-IMG with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SupraLabs/Supra2-IMG with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("SupraLabs/Supra2-IMG", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from SupraLabs/Supra2-IMG: direct link, hf CLI and curl.
- Browser
- Download file 4.22 kB
-
https://huggingface.co/SupraLabs/Supra2-IMG/resolve/main/README.md
- Command line
-
hf download hf://SupraLabs/Supra2-IMG/README.md
-
curl -L -o README.md https://huggingface.co/SupraLabs/Supra2-IMG/resolve/main/README.md
4.22 kB
| license: apache-2.0 | |
| viewer: false | |
| datasets: | |
| - LucasFang/FLUX-Reason-6M | |
| language: | |
| - en | |
| pipeline_tag: text-to-image | |
| library_name: transformers | |
| tags: | |
| - small | |
| - supra | |
| - image | |
| - flux | |
| - img | |
| - t2i | |
| - from scratch | |
| <h1 align="center">Supra2-IMG</h1> | |
| <p align="center"> | |
| Text-To-Image • 100M Parameters • SOTA quality | |
| </p> | |
|  | |
| **Supra2-IMG** is a tiny 100M parameters text-to-image (T2I) model that has been trained from scratch on high-quality synthetic data and delivers state-of-the-art image quality for its size. | |
| --- | |
| ## Samples | |
|  | |
| --- | |
| ## Model | |
| ### About the model | |
| The model is a tiny diffusion transformer (DiT) with ~105M parameters. | |
| - Pipeline: text-to-image | |
| - Parameter Count: 104.1M | |
| - Encoder: frozen [Flan-T5-Base](https://huggingface.co/google/flan-t5-base) | |
| - VAE: [SD-VAE-FT-MSE](https://huggingface.co/stabilityai/sd-vae-ft-mse) | |
| - Image resolution: 256² | |
| - Latent size: 32² | |
| - Patch: 2 | |
| - Context length (Flan-T5): 128 tokens | |
| ### Model config | |
| - `D_MODEL`: 576 | |
| - `DEPTH`: 14 | |
| - `N_HEADS`: 9 | |
| - `HEAD_DIM`: 64 | |
| - `MLP_RATIO`: 4.0 | |
| - `D_CTX`: 768 | |
| - `VAE_SCALE`: 0.18215 | |
| --- | |
| ## Training | |
| ### Dataset | |
| The model was trained for **10 epochs** on the full [LucasFang/FLUX-Reason-6M](https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M) dataset. | |
| ### Data preparation | |
| All train data images were downloaded as parquets + metadata and prepared by first chosing the prompt.<br> | |
| This was done in the following order (each next prompt is a fallback for the previous prompt): `caption_composition` → `caption_entity` → `caption_text` → `caption_style` → `caption_imaginative`.<br> | |
| That way, we ensured only using the highest quality data for pretraining the model. | |
| ### Exact image count | |
| **5.6M** images | |
| ### Epochs count | |
| **10** epochs | |
| ### Hardware | |
| The training ran on a single Nvidia H100 SXM 80GB Runpod Pod for 9 hours (incl. data preparation) with a 2.5TB disk. | |
| --- | |
| ## How to run the model | |
| First, run: | |
| ```bash | |
| # Create project directory | |
| mkdir Supra2-IMG | |
| cd Supra2-IMG | |
| # Download the inference script | |
| wget https://huggingface.co/SupraLabs/Supra2-IMG/resolve/main/inference.py | |
| ``` | |
| Then, you can generate images by running: | |
| ```bash | |
| python inference.py --prompt "a sea jellyfish floating in the pitch-black ocean depths" --seed 0 --cfg 3.0 --steps 50 --n 1 --out jellyfish.png | |
| ``` | |
| ### Recommended settings for sampling | |
| - `--seed`: 0 | |
| - `--cfg`: 3.0 | |
| - `--steps`: 50 | |
| ### Expected Output | |
| The script will output something like: | |
| ```bash | |
| === Supra2-IMG inference === | |
| [device] ... | |
| [ckpt] found ./model_final_ema.pt | |
| [model] building SupraDiT ... | |
| [model] 104.1M parameters | |
| [model] loading weights from ./model_final_ema.pt ... | |
| [model] weights loaded in 0.7s | |
| [text] ctx_len=128 | |
| [text] loading tokenizer + google/flan-t5-base ... | |
| ... | |
| [text] prompt tokens=15 n=1 seed=0 cfg=3.0 steps=50 | |
| [cfg] using stored unconditional embeddings | |
| [sample] Euler flow, 50 steps ... | |
| Generating: ... | |
| [sample] denoising done in ...s | |
| [vae] decoding latents ... | |
| [done] saved 1 image(s) -> ...png | |
| ``` | |
| ### Possible output | |
|  | |
| ## Acknowledgments | |
| - Thank you, FLUX-Reason-6M team, for proving the [LucasFang/FLUX-Reason-6M](https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M) dataset. Without **your** work, this would never be possible! | |
| - Thanks to the creator of [HobbyLM-Image](https://huggingface.co/rootxhacker/HobbyLM-Image) for inspiration | |
| - Thanks to Runpod for providing the compute! | |
| - Also a big Thank You To Stability AI and Google for providing SD-VAE and Flan-T5 | |
| - Thanks to all the other people inspiring us to make great open-source models for the community! | |
| ## What's next | |
| We will keep improving Supra2-IMG, maybe for a next-gen like Supra2.5-IMG, and we will share our progress and findings on the way to the best open-source T2I model 🤗<br> | |
| Please give us a like and a follow on Hugging Face if you want to support our work! | |