Spaces:
Running on Zero
Download README.md from rerun/4danyone-rerun: direct link, hf CLI and curl.
- Browser
- Download file 5.21 kB
-
https://huggingface.co/spaces/rerun/4danyone-rerun/resolve/main/README.md
- Command line
-
hf download hf://spaces/rerun/4danyone-rerun/README.md
-
curl -L -o README.md https://huggingface.co/spaces/rerun/4danyone-rerun/resolve/main/README.md
A newer version of the Gradio SDK is available: 6.29.1
title: 4DAnyone × Rerun
emoji: 🕺
colorFrom: indigo
colorTo: pink
sdk: gradio
app_file: app.py
python_version: 3.11
pinned: false
short_description: One monocular clip of a person becomes six novel views
4DAnyone × Rerun
Upload a clip of a person. GVHMR recovers the SMPL-X motion, a Wan 2.2 diffusion transformer generates six synchronized views around the subject, and every phase streams into the embedded Rerun viewer while it happens.
What you watch
The run is one four-link Gradio event chain, and the viewer's layout switches with it:
- Motion — the source clip beside its detections. GVHMR's tracked boxes and
ViTPose keypoints appear on the
frametimeline as each stage completes, and the posed SMPL-X body lands at the end. - Diffusion — a 2×3 grid of per-view previews on a
diffusion_steptimeline. Turbo mode takes four steps; after each one the clean latent estimate is decoded through TAEW2.2 at eight frames per view. Every preview is stamped on bothdiffusion_stepandframe. - Result — the camera rig in the canonical human world, the six generated
MP4s in a grid, and the
frametimeline playing on a loop.
Files
app.py— the Hugging Face entrypoint. Standard library only; it installs the pinned Pixi CLI when absent and runs the lockedapptask.fdanyone_app.py— the whole application: validation, the ZeroGPU callbacks, the Rerun choreography, and the Gradio interface.download_assets.py— every CPU-side download, idempotent, run before Gradio starts.sync_vendor.shandPROVENANCE.md— howfdanyone/got here and from where.pixi.tomlandpixi.lock— the complete CUDA 13, PyTorch 2.12.1, Rerun 0.36.1, Gradio 6.20.0 environment, locked forlinux-64andlinux-aarch64.
Streaming mechanics
prepare_run and generate_run are single blocking calls that report progress
through synchronous hooks. The @spaces.GPU callback therefore runs each
pipeline call on a worker thread and yields BinaryStream.read() as the hooks
fill an explicit RecordingStream. Yielding from inside a hook is impossible,
and yielding only after the call returns would leave the viewer blank for
minutes. The recording is an explicit object rather than a thread-local one,
which is what makes cross-thread logging safe.
spaces.GPU runs its callback in a forked child process, and that shapes the
rest. A RecordingStream cannot cross the fork — the SDK refuses to flush one
whose pid has changed — so every link opens its own stream under the run's
token and the viewer merges same-token streams into one recording. Run state
cannot cross it either: the chain passes a picklable RunSpec, and motion and
generation share one link because PreparedRun holds decoded frames and a live
barrier that no pickle can carry into another worker. Inside that child the
pipeline runs with inline_workers=True, because a grandchild of its own would
not hold the GPU allocation.
Fixed policy
- Turbo mode: four denoising steps, FP8 DiT, the pinned Wan2.2 Turbo-LoRA, the TAEW2.2 tiny decoder, and NVDEC skeleton decode.
- Six views on one ring pitched 15° above the subject, spanning 360°. Six or fewer targets skip the RCP proposal stage entirely.
- The frozen 121-frame, 1280×704 contract. Inputs shorter than 121 frames from the chosen start time are rejected on CPU, before any GPU work.
- The fixed prompt ships as an exported embedding, so the 11 GB UMT5-XXL encoder is never downloaded and never loaded.
Run it locally
On a CUDA 13 host with HF_TOKEN set for the private asset mirror:
CONDA_OVERRIDE_CUDA=13.0 pixi run --frozen app
FDANYONE_MODEL_DIR (default models) and FDANYONE_DATA_DIR (default data)
move the weights and the scratch tree. pixi run test runs the unit tests for
the pure helpers.
Compatibility notes
gradio==6.20.0,gradio-rerun==0.36.1, andrerun-sdk==0.36.1belong together. Gradio 6.21 through 6.26 carry a frontend regression that remounts custom components on every output update, which stops the Rerun WebViewer mid-stream.transformersis 5.x, not the source repository's 4.57. BiRefNet is loaded throughAutoModelForImageSegmentationwithtrust_remote_code, and its custom code only needsPretrainedConfigandPreTrainedModel. Transformers 4.x capshuggingface_hubbelow 1.0, which Gradio 6.20 forbids.- Start the app through
pixi run, never by calling the environment'spythondirectly. Turbo mode compiles the DiT with Triton, and conda-forge's Triton finds a Blackwell-capableptxasthrough$CONDA_PREFIX, which only the activation sets. Without it the compile fails withCannot find ptxas-blackwell. - SageAttention is built from source: no wheel exists for sm_120 or sm_121.
TORCH_CUDA_ARCH_LIST=12.0is set forlinux-64only, because a Space build host has no GPU to detect;linux-aarch64detects its GB10 natively.
Credits
4DAnyone by Ant Research. GVHMR by ZJU3DV. TAEHV by Ollin Boer Bohan.