Text-to-Image
Transformers
English
SupraDiT
feature-extraction
small
supra
image
flux
img
t2i
from scratch
custom_code
Instructions to use SupraLabs/Supra2-IMG with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SupraLabs/Supra2-IMG with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("SupraLabs/Supra2-IMG", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 4,219 Bytes
716594d 58e5238 716594d 58e5238 cb60685 8849cac cb60685 8849cac 10dec6e 8849cac b22ffe6 8849cac 59aacb2 7f1eebe 59aacb2 8849cac 7f1eebe 8849cac 10dec6e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 | ---
license: apache-2.0
viewer: false
datasets:
- LucasFang/FLUX-Reason-6M
language:
- en
pipeline_tag: text-to-image
library_name: transformers
tags:
- small
- supra
- image
- flux
- img
- t2i
- from scratch
---
<h1 align="center">Supra2-IMG</h1>
<p align="center">
Text-To-Image • 100M Parameters • SOTA quality
</p>

**Supra2-IMG** is a tiny 100M parameters text-to-image (T2I) model that has been trained from scratch on high-quality synthetic data and delivers state-of-the-art image quality for its size.
---
## Samples

---
## Model
### About the model
The model is a tiny diffusion transformer (DiT) with ~105M parameters.
- Pipeline: text-to-image
- Parameter Count: 104.1M
- Encoder: frozen [Flan-T5-Base](https://huggingface.co/google/flan-t5-base)
- VAE: [SD-VAE-FT-MSE](https://huggingface.co/stabilityai/sd-vae-ft-mse)
- Image resolution: 256²
- Latent size: 32²
- Patch: 2
- Context length (Flan-T5): 128 tokens
### Model config
- `D_MODEL`: 576
- `DEPTH`: 14
- `N_HEADS`: 9
- `HEAD_DIM`: 64
- `MLP_RATIO`: 4.0
- `D_CTX`: 768
- `VAE_SCALE`: 0.18215
---
## Training
### Dataset
The model was trained for **10 epochs** on the full [LucasFang/FLUX-Reason-6M](https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M) dataset.
### Data preparation
All train data images were downloaded as parquets + metadata and prepared by first chosing the prompt.<br>
This was done in the following order (each next prompt is a fallback for the previous prompt): `caption_composition` → `caption_entity` → `caption_text` → `caption_style` → `caption_imaginative`.<br>
That way, we ensured only using the highest quality data for pretraining the model.
### Exact image count
**5.6M** images
### Epochs count
**10** epochs
### Hardware
The training ran on a single Nvidia H100 SXM 80GB Runpod Pod for 9 hours (incl. data preparation) with a 2.5TB disk.
---
## How to run the model
First, run:
```bash
# Create project directory
mkdir Supra2-IMG
cd Supra2-IMG
# Download the inference script
wget https://huggingface.co/SupraLabs/Supra2-IMG/resolve/main/inference.py
```
Then, you can generate images by running:
```bash
python inference.py --prompt "a sea jellyfish floating in the pitch-black ocean depths" --seed 0 --cfg 3.0 --steps 50 --n 1 --out jellyfish.png
```
### Recommended settings for sampling
- `--seed`: 0
- `--cfg`: 3.0
- `--steps`: 50
### Expected Output
The script will output something like:
```bash
=== Supra2-IMG inference ===
[device] ...
[ckpt] found ./model_final_ema.pt
[model] building SupraDiT ...
[model] 104.1M parameters
[model] loading weights from ./model_final_ema.pt ...
[model] weights loaded in 0.7s
[text] ctx_len=128
[text] loading tokenizer + google/flan-t5-base ...
...
[text] prompt tokens=15 n=1 seed=0 cfg=3.0 steps=50
[cfg] using stored unconditional embeddings
[sample] Euler flow, 50 steps ...
Generating: ...
[sample] denoising done in ...s
[vae] decoding latents ...
[done] saved 1 image(s) -> ...png
```
### Possible output

## Acknowledgments
- Thank you, FLUX-Reason-6M team, for proving the [LucasFang/FLUX-Reason-6M](https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M) dataset. Without **your** work, this would never be possible!
- Thanks to the creator of [HobbyLM-Image](https://huggingface.co/rootxhacker/HobbyLM-Image) for inspiration
- Thanks to Runpod for providing the compute!
- Also a big Thank You To Stability AI and Google for providing SD-VAE and Flan-T5
- Thanks to all the other people inspiring us to make great open-source models for the community!
## What's next
We will keep improving Supra2-IMG, maybe for a next-gen like Supra2.5-IMG, and we will share our progress and findings on the way to the best open-source T2I model 🤗<br>
Please give us a like and a follow on Hugging Face if you want to support our work!
|