--- license: apache-2.0 viewer: false datasets: - LucasFang/FLUX-Reason-6M language: - en pipeline_tag: text-to-image library_name: transformers tags: - small - supra - image - flux - img - t2i - from scratch ---

Supra2-IMG

Text-To-Image • 100M Parameters • SOTA quality

![supra2-IMG](https://cdn-uploads.huggingface.co/production/uploads/697f2832c2c5e4daa93cece7/dhbB5CYrGYCnbPdOUHFnj.png) **Supra2-IMG** is a tiny 100M parameters text-to-image (T2I) model that has been trained from scratch on high-quality synthetic data and delivers state-of-the-art image quality for its size. --- ## Samples ![Samples Supra2-IMG](https://cdn-uploads.huggingface.co/production/uploads/697f2832c2c5e4daa93cece7/HrnJrCBuXsbqoKVF1UdlM.png) --- ## Model ### About the model The model is a tiny diffusion transformer (DiT) with ~105M parameters. - Pipeline: text-to-image - Parameter Count: 104.1M - Encoder: frozen [Flan-T5-Base](https://huggingface.co/google/flan-t5-base) - VAE: [SD-VAE-FT-MSE](https://huggingface.co/stabilityai/sd-vae-ft-mse) - Image resolution: 256² - Latent size: 32² - Patch: 2 - Context length (Flan-T5): 128 tokens ### Model config - `D_MODEL`: 576 - `DEPTH`: 14 - `N_HEADS`: 9 - `HEAD_DIM`: 64 - `MLP_RATIO`: 4.0 - `D_CTX`: 768 - `VAE_SCALE`: 0.18215 --- ## Training ### Dataset The model was trained for **10 epochs** on the full [LucasFang/FLUX-Reason-6M](https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M) dataset. ### Data preparation All train data images were downloaded as parquets + metadata and prepared by first chosing the prompt.
This was done in the following order (each next prompt is a fallback for the previous prompt): `caption_composition` → `caption_entity` → `caption_text` → `caption_style` → `caption_imaginative`.
That way, we ensured only using the highest quality data for pretraining the model. ### Exact image count **5.6M** images ### Epochs count **10** epochs ### Hardware The training ran on a single Nvidia H100 SXM 80GB Runpod Pod for 9 hours (incl. data preparation) with a 2.5TB disk. --- ## How to run the model First, run: ```bash # Create project directory mkdir Supra2-IMG cd Supra2-IMG # Download the inference script wget https://huggingface.co/SupraLabs/Supra2-IMG/resolve/main/inference.py ``` Then, you can generate images by running: ```bash python inference.py --prompt "a sea jellyfish floating in the pitch-black ocean depths" --seed 0 --cfg 3.0 --steps 50 --n 1 --out jellyfish.png ``` ### Recommended settings for sampling - `--seed`: 0 - `--cfg`: 3.0 - `--steps`: 50 ### Expected Output The script will output something like: ```bash === Supra2-IMG inference === [device] ... [ckpt] found ./model_final_ema.pt [model] building SupraDiT ... [model] 104.1M parameters [model] loading weights from ./model_final_ema.pt ... [model] weights loaded in 0.7s [text] ctx_len=128 [text] loading tokenizer + google/flan-t5-base ... ... [text] prompt tokens=15 n=1 seed=0 cfg=3.0 steps=50 [cfg] using stored unconditional embeddings [sample] Euler flow, 50 steps ... Generating: ... [sample] denoising done in ...s [vae] decoding latents ... [done] saved 1 image(s) -> ...png ``` ### Possible output ![image](https://cdn-uploads.huggingface.co/production/uploads/697f2832c2c5e4daa93cece7/6dolZG9Wvvnu_ltsRxlRq.png) ## Acknowledgments - Thank you, FLUX-Reason-6M team, for proving the [LucasFang/FLUX-Reason-6M](https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M) dataset. Without **your** work, this would never be possible! - Thanks to the creator of [HobbyLM-Image](https://huggingface.co/rootxhacker/HobbyLM-Image) for inspiration - Thanks to Runpod for providing the compute! - Also a big Thank You To Stability AI and Google for providing SD-VAE and Flan-T5 - Thanks to all the other people inspiring us to make great open-source models for the community! ## What's next We will keep improving Supra2-IMG, maybe for a next-gen like Supra2.5-IMG, and we will share our progress and findings on the way to the best open-source T2I model 🤗
Please give us a like and a follow on Hugging Face if you want to support our work!