File size: 4,219 Bytes
716594d
 
58e5238
 
 
 
 
 
 
 
 
 
 
 
 
 
 
716594d
58e5238
 
 
 
 
 
 
 
cb60685
 
 
8849cac
 
cb60685
 
 
8849cac
 
 
 
 
 
 
 
 
 
10dec6e
8849cac
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b22ffe6
8849cac
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
59aacb2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7f1eebe
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
59aacb2
8849cac
7f1eebe
 
 
 
8849cac
 
 
 
 
 
 
 
 
 
 
10dec6e
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
---
license: apache-2.0
viewer: false
datasets:
- LucasFang/FLUX-Reason-6M
language:
- en
pipeline_tag: text-to-image
library_name: transformers
tags:
- small
- supra
- image
- flux
- img
- t2i
- from scratch
---

<h1 align="center">Supra2-IMG</h1>

<p align="center">
  Text-To-Image • 100M Parameters • SOTA quality
</p>

![supra2-IMG](https://cdn-uploads.huggingface.co/production/uploads/697f2832c2c5e4daa93cece7/dhbB5CYrGYCnbPdOUHFnj.png)

**Supra2-IMG** is a tiny 100M parameters text-to-image (T2I) model that has been trained from scratch on high-quality synthetic data and delivers state-of-the-art image quality for its size.

---

## Samples

![Samples Supra2-IMG](https://cdn-uploads.huggingface.co/production/uploads/697f2832c2c5e4daa93cece7/HrnJrCBuXsbqoKVF1UdlM.png)

---

## Model

### About the model

The model is a tiny diffusion transformer (DiT) with ~105M parameters.

- Pipeline: text-to-image
- Parameter Count: 104.1M
- Encoder: frozen [Flan-T5-Base](https://huggingface.co/google/flan-t5-base)
- VAE: [SD-VAE-FT-MSE](https://huggingface.co/stabilityai/sd-vae-ft-mse)
- Image resolution: 256²
- Latent size: 32²
- Patch: 2
- Context length (Flan-T5): 128 tokens

### Model config

- `D_MODEL`: 576
- `DEPTH`: 14
- `N_HEADS`: 9
- `HEAD_DIM`: 64
- `MLP_RATIO`: 4.0
- `D_CTX`: 768
- `VAE_SCALE`: 0.18215

---

## Training

### Dataset

The model was trained for **10 epochs** on the full [LucasFang/FLUX-Reason-6M](https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M) dataset.

### Data preparation

All train data images were downloaded as parquets + metadata and prepared by first chosing the prompt.<br>
This was done in the following order (each next prompt is a fallback for the previous prompt): `caption_composition` → `caption_entity` → `caption_text` → `caption_style` → `caption_imaginative`.<br>
That way, we ensured only using the highest quality data for pretraining the model.

### Exact image count

**5.6M** images

### Epochs count

**10** epochs

### Hardware

The training ran on a single Nvidia H100 SXM 80GB Runpod Pod for 9 hours (incl. data preparation) with a 2.5TB disk.

---

## How to run the model

First, run:
```bash
# Create project directory
mkdir Supra2-IMG
cd Supra2-IMG

# Download the inference script
wget https://huggingface.co/SupraLabs/Supra2-IMG/resolve/main/inference.py
```

Then, you can generate images by running:
```bash
python inference.py --prompt "a sea jellyfish floating in the pitch-black ocean depths"  --seed 0  --cfg 3.0  --steps 50  --n 1  --out jellyfish.png
```

### Recommended settings for sampling

- `--seed`: 0
- `--cfg`: 3.0
- `--steps`: 50

### Expected Output

The script will output something like:
```bash
=== Supra2-IMG inference ===
[device] ...
[ckpt] found ./model_final_ema.pt
[model] building SupraDiT ...
[model] 104.1M parameters
[model] loading weights from ./model_final_ema.pt ...
[model] weights loaded in 0.7s
[text] ctx_len=128
[text] loading tokenizer + google/flan-t5-base ...
...
[text] prompt tokens=15  n=1  seed=0  cfg=3.0  steps=50
[cfg] using stored unconditional embeddings
[sample] Euler flow, 50 steps ...
Generating: ...
[sample] denoising done in ...s
[vae] decoding latents ...
[done] saved 1 image(s) -> ...png
```

### Possible output

![image](https://cdn-uploads.huggingface.co/production/uploads/697f2832c2c5e4daa93cece7/6dolZG9Wvvnu_ltsRxlRq.png)

## Acknowledgments

- Thank you, FLUX-Reason-6M team, for proving the [LucasFang/FLUX-Reason-6M](https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M) dataset. Without **your** work, this would never be possible!
- Thanks to the creator of [HobbyLM-Image](https://huggingface.co/rootxhacker/HobbyLM-Image) for inspiration
- Thanks to Runpod for providing the compute!
- Also a big Thank You To Stability AI and Google for providing SD-VAE and Flan-T5
- Thanks to all the other people inspiring us to make great open-source models for the community!

## What's next

We will keep improving Supra2-IMG, maybe for a next-gen like Supra2.5-IMG, and we will share our progress and findings on the way to the best open-source T2I model 🤗<br>
Please give us a like and a follow on Hugging Face if you want to support our work!