--- title: Zing 0.5 AoTI Compiler emoji: ⚙️ colorFrom: yellow colorTo: gray sdk: gradio sdk_version: 6.26.0 app_file: app.py pinned: false license: apache-2.0 python_version: "3.12" startup_duration_timeout: 1h short_description: torch.export + AOTInductor compiler for Zing-0.5 models: - seedleap/zing-0.5 --- # Zing-0.5 — ahead-of-time compiler Companion Space to [`hugging-apps/zing-0-5-world-model`](https://huggingface.co/spaces/hugging-apps/zing-0-5-world-model). The demo spends essentially all of its GPU time inside 30 identical `WanAttentionBlock`s — five forward passes per generated block (four DMD denoising steps plus the KV-cache commit), so 150 block executions for every 16 frames it streams. This Space compiles that block **once, ahead of time**, and publishes the artifact so the demo never has to compile anything at runtime. ## What it does 1. Loads the exact same vendored `zing_v0_5` package and sliding-window config (`local_attn_size=33`, `sink_size=5`) as the demo. 2. Runs a short real rollout with `spaces.aoti_capture()` wrapped around `generator.blocks[0]`, so the block is exported against genuine tensors rather than hand-made ones. The capture deliberately lands on the *second* generated block: block 0 runs with a cold (empty) KV cache and is kept on the eager path in the demo, so the compiled graph only ever sees a non-empty history. 3. `torch.export.export(..., strict=False)` with three dynamic axes — `sequence` (tokens per block, i.e. resolution), `history` (KV-cache length, which grows until the sliding window caps it) and `context` (prompt length, which changes on every mid-session rewrite). 4. `spaces.aoti_compile_and_save()` → AOTInductor archive. 5. Uploads it to [`hugging-apps/zing-0-5-world-model-aoti`](https://huggingface.co/hugging-apps/zing-0-5-world-model-aoti) as `WanAttentionBlock/package.pt2`. That path is the `{block class name}/package.pt2` layout the `spaces` package expects — the same convention as `zerogpu-aoti/FLUX.1` and `zerogpu-aoti/Wan2` — so the demo picks it up with one line: ```python spaces.aoti_blocks_load(PIPELINE.generator, "hugging-apps/zing-0-5-world-model-aoti") ``` Weights are **not** baked into the archive: `spaces.aoti_patch` hands each block its own `state_dict()` at call time, so a single `package.pt2` serves all 30 layers. ## Requirements Needs an `HF_TOKEN` secret with write access to the artifact repo. Compilation runs inside a single ZeroGPU allocation (`@spaces.GPU(duration=1500)`).