--- base_model: Wan-AI/Wan2.1-T2V-1.3B-Diffusers library_name: peft license: apache-2.0 pipeline_tag: text-to-video tags: - text-to-video - wan - lora - peft - flow-grpo - solace - reinforcement-learning --- # Wan2.1-T2V-1.3B-SOLACE LoRA adapter from **SOLACE** (**S**elf-c**O**nfidence reward for a**L**igning text-to-im**A**ge models via **C**onfidenc**E** optimization), CVPR 2026. SOLACE applied to the **Wan2.1-T2V-1.3B** text-to-video model, using the model's own denoising confidence as an intrinsic reward (no external reward model at training time). - **Base model:** [`Wan-AI/Wan2.1-T2V-1.3B-Diffusers`](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B-Diffusers) - **Method:** SOLACE intrinsic self-confidence reward (built on Flow-GRPO) - **Code:** https://github.com/wookiekim/SOLACE - **Adapter type:** PEFT LoRA (rank 32) on the Wan 3D transformer ## Usage ```python import torch from diffusers import WanPipeline from diffusers.utils import export_to_video from peft import PeftModel model_id = "Wan-AI/Wan2.1-T2V-1.3B-Diffusers" lora_ckpt_path = "wookiekim/Wan2.1-T2V-1.3B-SOLACE" device = "cuda" pipe = WanPipeline.from_pretrained(model_id, torch_dtype=torch.bfloat16) pipe.transformer = PeftModel.from_pretrained(pipe.transformer, lora_ckpt_path) pipe.transformer = pipe.transformer.merge_and_unload() pipe = pipe.to(device) frames = pipe( "a cat walking across a sunlit kitchen floor", height=480, width=832, num_frames=81, num_inference_steps=50, guidance_scale=5.0, ).frames[0] export_to_video(frames, "solace_wan.mp4", fps=16) ``` ## Citation ```bibtex @inproceedings{kim2026solace, title={Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards}, author={Kim, Wookyoung and others}, booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year={2026} } ``` ## Acknowledgments This work builds upon [Flow-GRPO](https://github.com/yifan123/flow_grpo) by Jie Liu et al.