Spaces:
Running on Zero
Running on Zero
|
Download README.md from rerun/4danyone-rerun: direct link, hf CLI and curl.
- Browser
- Download file 5.21 kB
-
https://huggingface.co/spaces/rerun/4danyone-rerun/resolve/main/README.md
- Command line
-
hf download hf://spaces/rerun/4danyone-rerun/README.md
-
curl -L -o README.md https://huggingface.co/spaces/rerun/4danyone-rerun/resolve/main/README.md
5.21 kB
| title: 4DAnyone × Rerun | |
| emoji: 🕺 | |
| colorFrom: indigo | |
| colorTo: pink | |
| sdk: gradio | |
| app_file: app.py | |
| python_version: 3.11 | |
| pinned: false | |
| short_description: One monocular clip of a person becomes six novel views | |
| # 4DAnyone × Rerun | |
| Upload a clip of a person. GVHMR recovers the SMPL-X motion, a Wan 2.2 diffusion | |
| transformer generates six synchronized views around the subject, and every phase | |
| streams into the embedded Rerun viewer while it happens. | |
| ## What you watch | |
| The run is one four-link Gradio event chain, and the viewer's layout switches | |
| with it: | |
| 1. **Motion** — the source clip beside its detections. GVHMR's tracked boxes and | |
| ViTPose keypoints appear on the `frame` timeline as each stage completes, and | |
| the posed SMPL-X body lands at the end. | |
| 2. **Diffusion** — a 2×3 grid of per-view previews on a `diffusion_step` | |
| timeline. Turbo mode takes four steps; after each one the clean latent | |
| estimate is decoded through TAEW2.2 at eight frames per view. Every preview | |
| is stamped on both `diffusion_step` and `frame`. | |
| 3. **Result** — the camera rig in the canonical human world, the six generated | |
| MP4s in a grid, and the `frame` timeline playing on a loop. | |
| ## Files | |
| - `app.py` — the Hugging Face entrypoint. Standard library only; it installs the | |
| pinned Pixi CLI when absent and runs the locked `app` task. | |
| - `fdanyone_app.py` — the whole application: validation, the ZeroGPU callbacks, | |
| the Rerun choreography, and the Gradio interface. | |
| - `download_assets.py` — every CPU-side download, idempotent, run before Gradio | |
| starts. | |
| - `sync_vendor.sh` and `PROVENANCE.md` — how `fdanyone/` got here and from where. | |
| - `pixi.toml` and `pixi.lock` — the complete CUDA 13, PyTorch 2.12.1, Rerun | |
| 0.36.1, Gradio 6.20.0 environment, locked for `linux-64` and `linux-aarch64`. | |
| ## Streaming mechanics | |
| `prepare_run` and `generate_run` are single blocking calls that report progress | |
| through synchronous hooks. The `@spaces.GPU` callback therefore runs each | |
| pipeline call on a worker thread and yields `BinaryStream.read()` as the hooks | |
| fill an explicit `RecordingStream`. Yielding from inside a hook is impossible, | |
| and yielding only after the call returns would leave the viewer blank for | |
| minutes. The recording is an explicit object rather than a thread-local one, | |
| which is what makes cross-thread logging safe. | |
| `spaces.GPU` runs its callback in a forked child process, and that shapes the | |
| rest. A `RecordingStream` cannot cross the fork — the SDK refuses to flush one | |
| whose pid has changed — so every link opens its own stream under the run's | |
| token and the viewer merges same-token streams into one recording. Run state | |
| cannot cross it either: the chain passes a picklable `RunSpec`, and motion and | |
| generation share one link because `PreparedRun` holds decoded frames and a live | |
| barrier that no pickle can carry into another worker. Inside that child the | |
| pipeline runs with `inline_workers=True`, because a grandchild of its own would | |
| not hold the GPU allocation. | |
| ## Fixed policy | |
| - Turbo mode: four denoising steps, FP8 DiT, the pinned Wan2.2 Turbo-LoRA, the | |
| TAEW2.2 tiny decoder, and NVDEC skeleton decode. | |
| - Six views on one ring pitched 15° above the subject, spanning 360°. Six or | |
| fewer targets skip the RCP proposal stage entirely. | |
| - The frozen 121-frame, 1280×704 contract. Inputs shorter than 121 frames from | |
| the chosen start time are rejected on CPU, before any GPU work. | |
| - The fixed prompt ships as an exported embedding, so the 11 GB UMT5-XXL encoder | |
| is never downloaded and never loaded. | |
| ## Run it locally | |
| On a CUDA 13 host with `HF_TOKEN` set for the private asset mirror: | |
| ```bash | |
| CONDA_OVERRIDE_CUDA=13.0 pixi run --frozen app | |
| ``` | |
| `FDANYONE_MODEL_DIR` (default `models`) and `FDANYONE_DATA_DIR` (default `data`) | |
| move the weights and the scratch tree. `pixi run test` runs the unit tests for | |
| the pure helpers. | |
| ## Compatibility notes | |
| - `gradio==6.20.0`, `gradio-rerun==0.36.1`, and `rerun-sdk==0.36.1` belong | |
| together. Gradio 6.21 through 6.26 carry a frontend regression that remounts | |
| custom components on every output update, which stops the Rerun WebViewer | |
| mid-stream. | |
| - `transformers` is 5.x, not the source repository's 4.57. BiRefNet is loaded | |
| through `AutoModelForImageSegmentation` with `trust_remote_code`, and its | |
| custom code only needs `PretrainedConfig` and `PreTrainedModel`. Transformers | |
| 4.x caps `huggingface_hub` below 1.0, which Gradio 6.20 forbids. | |
| - Start the app through `pixi run`, never by calling the environment's `python` | |
| directly. Turbo mode compiles the DiT with Triton, and conda-forge's Triton | |
| finds a Blackwell-capable `ptxas` through `$CONDA_PREFIX`, which only the | |
| activation sets. Without it the compile fails with `Cannot find | |
| ptxas-blackwell`. | |
| - SageAttention is built from source: no wheel exists for sm_120 or sm_121. | |
| `TORCH_CUDA_ARCH_LIST=12.0` is set for `linux-64` only, because a Space build | |
| host has no GPU to detect; `linux-aarch64` detects its GB10 natively. | |
| ## Credits | |
| 4DAnyone by [Ant Research](https://github.com/ant-research/4DAnyone). | |
| [GVHMR](https://github.com/zju3dv/GVHMR) by ZJU3DV. | |
| [TAEHV](https://github.com/madebyollin/taehv) by Ollin Boer Bohan. | |