cezar-hapiko's picture
Load model synchronously before Gradio launch (fix HF gcTimeout kill)
405bc15
|
Raw History Blame Contribute Delete
3.45 kB
metadata
title: Lyra-2 Explorable Scene
emoji: πŸŒ€
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
suggested_hardware: a100-large
suggested_storage: large
sleep_time: 7200
license: other
license_name: nvidia-source-code-license-for-lyra-2-0
license_link: https://huggingface.co/nvidia/Lyra-2.0/blob/main/LICENSE
hf_oauth: false
startup_duration_timeout: 2h
models:
  - nvidia/Lyra-2.0
tags:
  - gaussian-splatting
  - 3d-reconstruction
  - video-diffusion
  - text-to-3d
  - lyra
  - nvidia

Lyra-2 β€” Image to Explorable 3D Scene

Upload a single image and a short caption; this Space produces an exploration video and a walkable Gaussian-splat scene of what's around your viewpoint. Powered by nvidia/Lyra-2.0.

What you get

Artifact Description
mp4 video Lyra-2's generated camera trajectory through the scene.
reconstructed_scene.ply Standard binary Gaussian-splat PLY. Works with any GS viewer.

Runtime

  • Cold-boot takes ~60–65 min before the UI is available β€” the DCP checkpoint load alone is ~60 min on A100 80GB. Space status stays APP_STARTING during that time; this is covered by startup_duration_timeout: 2h below.
  • Once warm: ~13 min per request on A100 80GB (DMD distillation always on).
  • The Lyra-2 graph + DA3 stay resident in GPU memory between requests β€” the 60-min warmup is amortized across every user the Space serves until it scales to zero.
  • Queue is serial (concurrency 1) β€” wait for the request ahead of yours.
  • Scale-to-zero after 2h idle (next request pays the warmup again).

Walk the scene on macOS β€” coming soon

The downloaded .ply is a standard GS format today; you can drop it into any GS viewer (Supersplat, antimatter15/splat, etc.).

A dedicated first-person, walkable macOS viewer (WASD + mouse-look) is in the works β€” stay tuned.

Source & attribution

  • Model: nvidia/Lyra-2.0 β€” weights released under NVIDIA Source Code License.
  • Inference code: nv-tlabs/lyra (Lyra-2 subfolder).
  • This Space: wraps the upstream inference pipeline in a Gradio UI; no model modifications.

Tips for good results

  • Composition matters. Scenes with visible depth cues (e.g. hallways, foreground objects, parallax between planes) reconstruct better than flat frontal shots.
  • Captions guide the diffusion. Describe the setting, mood, and any specific elements you want preserved. 1–2 sentences is enough.
  • Hallucinations are expected in regions the input image can't see. Lyra-2 fills unseen areas plausibly but not faithfully.
  • DMD vs quality: fast mode occasionally produces repetitive textures. Disable DMD for production-grade results at the cost of ~10Γ— wall-clock.