Spaces:
Running on Zero
Running on Zero
|
Download README.md from hugging-apps/zing-0-5-world-model-aoti-compile: direct link, hf CLI and curl.
- Browser
- Download file 2.54 kB
-
https://huggingface.co/spaces/hugging-apps/zing-0-5-world-model-aoti-compile/resolve/main/README.md
- Command line
-
hf download hf://spaces/hugging-apps/zing-0-5-world-model-aoti-compile/README.md
-
curl -L -o README.md https://huggingface.co/spaces/hugging-apps/zing-0-5-world-model-aoti-compile/resolve/main/README.md
2.54 kB
| title: Zing 0.5 AoTI Compiler | |
| emoji: ⚙️ | |
| colorFrom: yellow | |
| colorTo: gray | |
| sdk: gradio | |
| sdk_version: 6.26.0 | |
| app_file: app.py | |
| pinned: false | |
| license: apache-2.0 | |
| python_version: "3.12" | |
| startup_duration_timeout: 1h | |
| short_description: torch.export + AOTInductor compiler for Zing-0.5 | |
| models: | |
| - seedleap/zing-0.5 | |
| # Zing-0.5 — ahead-of-time compiler | |
| Companion Space to | |
| [`hugging-apps/zing-0-5-world-model`](https://huggingface.co/spaces/hugging-apps/zing-0-5-world-model). | |
| The demo spends essentially all of its GPU time inside 30 identical | |
| `WanAttentionBlock`s — five forward passes per generated block (four DMD denoising | |
| steps plus the KV-cache commit), so 150 block executions for every 16 frames it | |
| streams. This Space compiles that block **once, ahead of time**, and publishes the | |
| artifact so the demo never has to compile anything at runtime. | |
| ## What it does | |
| 1. Loads the exact same vendored `zing_v0_5` package and sliding-window config | |
| (`local_attn_size=33`, `sink_size=5`) as the demo. | |
| 2. Runs a short real rollout with `spaces.aoti_capture()` wrapped around | |
| `generator.blocks[0]`, so the block is exported against genuine tensors rather | |
| than hand-made ones. The capture deliberately lands on the *second* generated | |
| block: block 0 runs with a cold (empty) KV cache and is kept on the eager path | |
| in the demo, so the compiled graph only ever sees a non-empty history. | |
| 3. `torch.export.export(..., strict=False)` with three dynamic axes — | |
| `sequence` (tokens per block, i.e. resolution), `history` (KV-cache length, | |
| which grows until the sliding window caps it) and `context` (prompt length, | |
| which changes on every mid-session rewrite). | |
| 4. `spaces.aoti_compile_and_save()` → AOTInductor archive. | |
| 5. Uploads it to [`hugging-apps/zing-0-5-world-model-aoti`](https://huggingface.co/hugging-apps/zing-0-5-world-model-aoti) | |
| as `WanAttentionBlock/package.pt2`. | |
| That path is the `{block class name}/package.pt2` layout the `spaces` package | |
| expects — the same convention as `zerogpu-aoti/FLUX.1` and `zerogpu-aoti/Wan2` — | |
| so the demo picks it up with one line: | |
| ```python | |
| spaces.aoti_blocks_load(PIPELINE.generator, "hugging-apps/zing-0-5-world-model-aoti") | |
| ``` | |
| Weights are **not** baked into the archive: `spaces.aoti_patch` hands each block its | |
| own `state_dict()` at call time, so a single `package.pt2` serves all 30 layers. | |
| ## Requirements | |
| Needs an `HF_TOKEN` secret with write access to the artifact repo. Compilation runs | |
| inside a single ZeroGPU allocation (`@spaces.GPU(duration=1500)`). | |