MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis Paper • 2509.14784 • Published Sep 18, 2025 • 1
MOVA: Towards Scalable and Synchronized Video-Audio Generation Paper • 2602.08794 • Published Feb 9 • 159