multimodalart's picture
multimodalart HF Staff
AoTI compiler
52e0fb7 verified
|
Raw History Blame Contribute Delete
2.54 kB
---
title: Zing 0.5 AoTI Compiler
emoji: ⚙️
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
pinned: false
license: apache-2.0
python_version: "3.12"
startup_duration_timeout: 1h
short_description: torch.export + AOTInductor compiler for Zing-0.5
models:
- seedleap/zing-0.5
---
# Zing-0.5 — ahead-of-time compiler
Companion Space to
[`hugging-apps/zing-0-5-world-model`](https://huggingface.co/spaces/hugging-apps/zing-0-5-world-model).
The demo spends essentially all of its GPU time inside 30 identical
`WanAttentionBlock`s — five forward passes per generated block (four DMD denoising
steps plus the KV-cache commit), so 150 block executions for every 16 frames it
streams. This Space compiles that block **once, ahead of time**, and publishes the
artifact so the demo never has to compile anything at runtime.
## What it does
1. Loads the exact same vendored `zing_v0_5` package and sliding-window config
(`local_attn_size=33`, `sink_size=5`) as the demo.
2. Runs a short real rollout with `spaces.aoti_capture()` wrapped around
`generator.blocks[0]`, so the block is exported against genuine tensors rather
than hand-made ones. The capture deliberately lands on the *second* generated
block: block 0 runs with a cold (empty) KV cache and is kept on the eager path
in the demo, so the compiled graph only ever sees a non-empty history.
3. `torch.export.export(..., strict=False)` with three dynamic axes —
`sequence` (tokens per block, i.e. resolution), `history` (KV-cache length,
which grows until the sliding window caps it) and `context` (prompt length,
which changes on every mid-session rewrite).
4. `spaces.aoti_compile_and_save()` → AOTInductor archive.
5. Uploads it to [`hugging-apps/zing-0-5-world-model-aoti`](https://huggingface.co/hugging-apps/zing-0-5-world-model-aoti)
as `WanAttentionBlock/package.pt2`.
That path is the `{block class name}/package.pt2` layout the `spaces` package
expects — the same convention as `zerogpu-aoti/FLUX.1` and `zerogpu-aoti/Wan2` —
so the demo picks it up with one line:
```python
spaces.aoti_blocks_load(PIPELINE.generator, "hugging-apps/zing-0-5-world-model-aoti")
```
Weights are **not** baked into the archive: `spaces.aoti_patch` hands each block its
own `state_dict()` at call time, so a single `package.pt2` serves all 30 layers.
## Requirements
Needs an `HF_TOKEN` secret with write access to the artifact repo. Compilation runs
inside a single ZeroGPU allocation (`@spaces.GPU(duration=1500)`).