Text-to-Image
PEFT
Safetensors
Sana
sana-sprint
lora
rwtd
geneval

Configuration Parsing Warning:In adapter_config.json: "peft.base_model_name_or_path" must be a string

Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

SANA-Sprint 1.6B RWTD GenEval LoRA

This repository contains a Reward-Weighted Transport Distillation (RWTD) LoRA adapter for SANA-Sprint 1.6B 1024px.

This adapter is obtained by post-training SANA Sprint 1.6B with the method presented in the paper Aligning One-Step Generative Models with Reward-Weighted Transport Distillation. Only the adapter is included. The base model, text encoder, and VAE are downloaded separately from their respective repositories.

Evaluation

The adapter achieves an official GenEval overall score of 0.80 rounded.

Evaluation used all 553 official GenEval prompts with four images per prompt.

  • Position: 0.73
  • Counting: 0.61
  • Color attribution: 0.60
  • Single object: 1.00
  • Colors: 0.90
  • Two objects: 0.95

Evaluation settings:

  • Resolution: 1024 × 1024
  • Inference steps: 1
  • Guidance scale: 4.5
  • Maximum timestep: 1.5708
  • Official GenEval seeds
  • Precision: bfloat16

Installation

This adapter targets the native SanaMSCM implementation from NVlabs/Sana. It is not directly compatible with DiffusionPipeline.load_lora_weights().

git clone https://github.com/NVlabs/Sana.git
cd Sana
./environment_setup.sh sana
conda activate sana
pip install "peft==0.18.0" huggingface_hub

Inference

Run the following from the root of the cloned NVlabs/Sana repository. A CUDA-capable GPU is required.

import torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from torchvision.utils import save_image

from app.sana_sprint_pipeline import SanaSprintPipeline


BASE_MODEL_ID = "Efficient-Large-Model/Sana_Sprint_1.6B_1024px"
ADAPTER_ID = "austin-k-wang/SanaSprint1.6B-RWTD-GenEval"
CONFIG_PATH = (
    "configs/sana_sprint_config/1024ms/"
    "SanaSprint_1600M_1024px_allqknorm_bf16_scm_ladd.yaml"
)

if not torch.cuda.is_available():
    raise RuntimeError("A CUDA-capable GPU is required.")

device = torch.device("cuda:0")

base_checkpoint = hf_hub_download(
    repo_id=BASE_MODEL_ID,
    filename="checkpoints/Sana_Sprint_1.6B_1024px.pth",
)

pipeline = SanaSprintPipeline(CONFIG_PATH, device=device)
pipeline.config.max_timesteps = 1.5708
pipeline.from_pretrained(base_checkpoint)

pipeline.model = PeftModel.from_pretrained(
    pipeline.model,
    ADAPTER_ID,
    is_trainable=False,
)
pipeline.model = pipeline.model.merge_and_unload()
pipeline.eval().requires_grad_(False)

generator = torch.Generator(device=device).manual_seed(42)

image = pipeline(
    prompt="a photo of a bench",
    height=1024,
    width=1024,
    guidance_scale=4.5,
    num_inference_steps=1,
    generator=generator,
)

save_image(
    image,
    "sana_sprint_rwtd.png",
    normalize=True,
    value_range=(-1, 1),
)

Limitations

This adapter inherits the limitations and biases of SANA-Sprint 1.6B. It does not guarantee correct object counts, spatial relationships, colors, text, or anatomy. Results may change with different seeds, timesteps, guidance scales, or dependency versions.

License

The adapter is released under the Apache License 2.0. Use remains subject to the licenses and terms of the base model and supporting components.

Citation

@misc{xie2024sana,
  title={Sana: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer},
  author={Enze Xie and Junsong Chen and Junyu Chen and Han Cai and Haotian Tang and Yujun Lin and Zhekai Zhang and Muyang Li and Ligeng Zhu and Yao Lu and Song Han},
  year={2024},
  eprint={2410.10629},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2410.10629}
}

@misc{chen2025sanasprint,
  title={SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation},
  author={Junsong Chen and Shuchen Xue and Yuyang Zhao and Jincheng Yu and Sayak Paul and Junyu Chen and Han Cai and Enze Xie and Song Han},
  year={2025},
  eprint={2503.09641},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2503.09641}
}

The RWTD paper citation will be added once its details are made public.

Acknowledgments

This adapter builds on the SANA and SANA-Sprint models and code released by NVIDIA and the SANA team.

Downloads last month
20
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for austin-k-wang/SanaSprint1.6B-RWTD-GenEval

Unable to build the model tree, the base model loops to the model itself. Learn more.

Collection including austin-k-wang/SanaSprint1.6B-RWTD-GenEval

Papers for austin-k-wang/SanaSprint1.6B-RWTD-GenEval