nari-labs/Dia-1.6B-0626
Text-to-Speech β’ 2B β’ Updated β’ 13.8k β’ 134
Control 3D models using gestures and voice
Audio-Driven Multi-Person Conversational Video Generation
edit images with Kontext and LoRAs
Hand-controlled arpeggiator, drum machine, and visualizer
Chat with an AI that understands text, images, audio, and video
Kontext multi image composition on FLUX[dev]
LightGlue demo
Convert web content to JSON using a custom schema
OmniGen2: Unified Image Understanding and Generation.
Generate or edit images using text prompts