--- title: Lyra-2 Explorable Scene emoji: 🌀 colorFrom: blue colorTo: purple sdk: docker app_port: 7860 pinned: false suggested_hardware: a100-large suggested_storage: large sleep_time: 7200 license: other license_name: nvidia-source-code-license-for-lyra-2-0 license_link: https://huggingface.co/nvidia/Lyra-2.0/blob/main/LICENSE hf_oauth: false startup_duration_timeout: 2h models: - nvidia/Lyra-2.0 tags: - gaussian-splatting - 3d-reconstruction - video-diffusion - text-to-3d - lyra - nvidia --- # Lyra-2 — Image to Explorable 3D Scene Upload a single image and a short caption; this Space produces an exploration video and a walkable Gaussian-splat scene of what's around your viewpoint. Powered by [`nvidia/Lyra-2.0`](https://huggingface.co/nvidia/Lyra-2.0). ## What you get | Artifact | Description | |---|---| | `mp4` video | Lyra-2's generated camera trajectory through the scene. | | `reconstructed_scene.ply` | Standard binary Gaussian-splat PLY. Works with any GS viewer. | ## Runtime - **Cold-boot takes ~60–65 min** before the UI is available — the DCP checkpoint load alone is ~60 min on A100 80GB. Space status stays `APP_STARTING` during that time; this is covered by `startup_duration_timeout: 2h` below. - **Once warm: ~13 min per request** on A100 80GB (DMD distillation always on). - The Lyra-2 graph + DA3 stay resident in GPU memory between requests — the 60-min warmup is amortized across every user the Space serves until it scales to zero. - Queue is serial (concurrency 1) — wait for the request ahead of yours. - Scale-to-zero after 2h idle (next request pays the warmup again). ## Walk the scene on macOS — *coming soon* The downloaded `.ply` is a standard GS format today; you can drop it into any GS viewer (Supersplat, antimatter15/splat, etc.). A dedicated **first-person, walkable** macOS viewer (WASD + mouse-look) is in the works — stay tuned. ## Source & attribution - **Model**: [`nvidia/Lyra-2.0`](https://huggingface.co/nvidia/Lyra-2.0) — weights released under NVIDIA Source Code License. - **Inference code**: [`nv-tlabs/lyra`](https://github.com/nv-tlabs/lyra) (Lyra-2 subfolder). - **This Space**: wraps the upstream inference pipeline in a Gradio UI; no model modifications. ## Tips for good results - **Composition matters**. Scenes with visible depth cues (e.g. hallways, foreground objects, parallax between planes) reconstruct better than flat frontal shots. - **Captions guide the diffusion**. Describe the setting, mood, and any specific elements you want preserved. 1–2 sentences is enough. - **Hallucinations are expected** in regions the input image can't see. Lyra-2 fills unseen areas plausibly but not faithfully. - **DMD vs quality**: fast mode occasionally produces repetitive textures. Disable DMD for production-grade results at the cost of ~10× wall-clock.