Spaces:
Running on Zero
Download README.md from hugging-apps/zing-0-5-world-model-aoti-compile: direct link, hf CLI and curl.
- Browser
- Download file 2.54 kB
-
https://huggingface.co/spaces/hugging-apps/zing-0-5-world-model-aoti-compile/resolve/main/README.md
- Command line
-
hf download hf://spaces/hugging-apps/zing-0-5-world-model-aoti-compile/README.md
-
curl -L -o README.md https://huggingface.co/spaces/hugging-apps/zing-0-5-world-model-aoti-compile/resolve/main/README.md
A newer version of the Gradio SDK is available: 6.29.1
title: Zing 0.5 AoTI Compiler
emoji: ⚙️
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
pinned: false
license: apache-2.0
python_version: '3.12'
startup_duration_timeout: 1h
short_description: torch.export + AOTInductor compiler for Zing-0.5
models:
- seedleap/zing-0.5
Zing-0.5 — ahead-of-time compiler
Companion Space to
hugging-apps/zing-0-5-world-model.
The demo spends essentially all of its GPU time inside 30 identical
WanAttentionBlocks — five forward passes per generated block (four DMD denoising
steps plus the KV-cache commit), so 150 block executions for every 16 frames it
streams. This Space compiles that block once, ahead of time, and publishes the
artifact so the demo never has to compile anything at runtime.
What it does
- Loads the exact same vendored
zing_v0_5package and sliding-window config (local_attn_size=33,sink_size=5) as the demo. - Runs a short real rollout with
spaces.aoti_capture()wrapped aroundgenerator.blocks[0], so the block is exported against genuine tensors rather than hand-made ones. The capture deliberately lands on the second generated block: block 0 runs with a cold (empty) KV cache and is kept on the eager path in the demo, so the compiled graph only ever sees a non-empty history. torch.export.export(..., strict=False)with three dynamic axes —sequence(tokens per block, i.e. resolution),history(KV-cache length, which grows until the sliding window caps it) andcontext(prompt length, which changes on every mid-session rewrite).spaces.aoti_compile_and_save()→ AOTInductor archive.- Uploads it to
hugging-apps/zing-0-5-world-model-aotiasWanAttentionBlock/package.pt2.
That path is the {block class name}/package.pt2 layout the spaces package
expects — the same convention as zerogpu-aoti/FLUX.1 and zerogpu-aoti/Wan2 —
so the demo picks it up with one line:
spaces.aoti_blocks_load(PIPELINE.generator, "hugging-apps/zing-0-5-world-model-aoti")
Weights are not baked into the archive: spaces.aoti_patch hands each block its
own state_dict() at call time, so a single package.pt2 serves all 30 layers.
Requirements
Needs an HF_TOKEN secret with write access to the artifact repo. Compilation runs
inside a single ZeroGPU allocation (@spaces.GPU(duration=1500)).