Instructions to use austin-k-wang/SanaSprint1.6B-RWTD-GenEval with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use austin-k-wang/SanaSprint1.6B-RWTD-GenEval with PEFT:
Task type is invalid.
- Sana
How to use austin-k-wang/SanaSprint1.6B-RWTD-GenEval with Sana:
# Load the model and infer image from text import torch from app.sana_pipeline import SanaPipeline from torchvision.utils import save_image sana = SanaPipeline("configs/sana_config/1024ms/Sana_1600M_img1024.yaml") sana.from_pretrained("hf://austin-k-wang/SanaSprint1.6B-RWTD-GenEval") image = sana( prompt='a cyberpunk cat with a neon sign that says "Sana"', height=1024, width=1024, guidance_scale=5.0, pag_guidance_scale=2.0, num_inference_steps=18, ) - Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.base_model_name_or_path" must be a string
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
SANA-Sprint 1.6B RWTD GenEval LoRA
This repository contains a Reward-Weighted Transport Distillation (RWTD) LoRA adapter for SANA-Sprint 1.6B 1024px.
This adapter is obtained by post-training SANA Sprint 1.6B with the method presented in the paper Aligning One-Step Generative Models with Reward-Weighted Transport Distillation. Only the adapter is included. The base model, text encoder, and VAE are downloaded separately from their respective repositories.
Evaluation
The adapter achieves an official GenEval overall score of 0.80 rounded.
Evaluation used all 553 official GenEval prompts with four images per prompt.
- Position: 0.73
- Counting: 0.61
- Color attribution: 0.60
- Single object: 1.00
- Colors: 0.90
- Two objects: 0.95
Evaluation settings:
- Resolution: 1024 × 1024
- Inference steps: 1
- Guidance scale: 4.5
- Maximum timestep: 1.5708
- Official GenEval seeds
- Precision: bfloat16
Installation
This adapter targets the native SanaMSCM implementation from
NVlabs/Sana. It is not directly compatible
with DiffusionPipeline.load_lora_weights().
git clone https://github.com/NVlabs/Sana.git
cd Sana
./environment_setup.sh sana
conda activate sana
pip install "peft==0.18.0" huggingface_hub
Inference
Run the following from the root of the cloned NVlabs/Sana repository. A
CUDA-capable GPU is required.
import torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from torchvision.utils import save_image
from app.sana_sprint_pipeline import SanaSprintPipeline
BASE_MODEL_ID = "Efficient-Large-Model/Sana_Sprint_1.6B_1024px"
ADAPTER_ID = "austin-k-wang/SanaSprint1.6B-RWTD-GenEval"
CONFIG_PATH = (
"configs/sana_sprint_config/1024ms/"
"SanaSprint_1600M_1024px_allqknorm_bf16_scm_ladd.yaml"
)
if not torch.cuda.is_available():
raise RuntimeError("A CUDA-capable GPU is required.")
device = torch.device("cuda:0")
base_checkpoint = hf_hub_download(
repo_id=BASE_MODEL_ID,
filename="checkpoints/Sana_Sprint_1.6B_1024px.pth",
)
pipeline = SanaSprintPipeline(CONFIG_PATH, device=device)
pipeline.config.max_timesteps = 1.5708
pipeline.from_pretrained(base_checkpoint)
pipeline.model = PeftModel.from_pretrained(
pipeline.model,
ADAPTER_ID,
is_trainable=False,
)
pipeline.model = pipeline.model.merge_and_unload()
pipeline.eval().requires_grad_(False)
generator = torch.Generator(device=device).manual_seed(42)
image = pipeline(
prompt="a photo of a bench",
height=1024,
width=1024,
guidance_scale=4.5,
num_inference_steps=1,
generator=generator,
)
save_image(
image,
"sana_sprint_rwtd.png",
normalize=True,
value_range=(-1, 1),
)
Limitations
This adapter inherits the limitations and biases of SANA-Sprint 1.6B. It does not guarantee correct object counts, spatial relationships, colors, text, or anatomy. Results may change with different seeds, timesteps, guidance scales, or dependency versions.
License
The adapter is released under the Apache License 2.0. Use remains subject to the licenses and terms of the base model and supporting components.
Citation
@misc{xie2024sana,
title={Sana: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer},
author={Enze Xie and Junsong Chen and Junyu Chen and Han Cai and Haotian Tang and Yujun Lin and Zhekai Zhang and Muyang Li and Ligeng Zhu and Yao Lu and Song Han},
year={2024},
eprint={2410.10629},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2410.10629}
}
@misc{chen2025sanasprint,
title={SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation},
author={Junsong Chen and Shuchen Xue and Yuyang Zhao and Jincheng Yu and Sayak Paul and Junyu Chen and Han Cai and Enze Xie and Song Han},
year={2025},
eprint={2503.09641},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.09641}
}
The RWTD paper citation will be added once its details are made public.
Acknowledgments
This adapter builds on the SANA and SANA-Sprint models and code released by NVIDIA and the SANA team.
- Downloads last month
- 20
Model tree for austin-k-wang/SanaSprint1.6B-RWTD-GenEval
Unable to build the model tree, the base model loops to the model itself. Learn more.