Spaces:
Running on Zero
Running on Zero
File size: 2,544 Bytes
8ca16d2 52e0fb7 8ca16d2 52e0fb7 8ca16d2 52e0fb7 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 | ---
title: Zing 0.5 AoTI Compiler
emoji: ⚙️
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
pinned: false
license: apache-2.0
python_version: "3.12"
startup_duration_timeout: 1h
short_description: torch.export + AOTInductor compiler for Zing-0.5
models:
- seedleap/zing-0.5
---
# Zing-0.5 — ahead-of-time compiler
Companion Space to
[`hugging-apps/zing-0-5-world-model`](https://huggingface.co/spaces/hugging-apps/zing-0-5-world-model).
The demo spends essentially all of its GPU time inside 30 identical
`WanAttentionBlock`s — five forward passes per generated block (four DMD denoising
steps plus the KV-cache commit), so 150 block executions for every 16 frames it
streams. This Space compiles that block **once, ahead of time**, and publishes the
artifact so the demo never has to compile anything at runtime.
## What it does
1. Loads the exact same vendored `zing_v0_5` package and sliding-window config
(`local_attn_size=33`, `sink_size=5`) as the demo.
2. Runs a short real rollout with `spaces.aoti_capture()` wrapped around
`generator.blocks[0]`, so the block is exported against genuine tensors rather
than hand-made ones. The capture deliberately lands on the *second* generated
block: block 0 runs with a cold (empty) KV cache and is kept on the eager path
in the demo, so the compiled graph only ever sees a non-empty history.
3. `torch.export.export(..., strict=False)` with three dynamic axes —
`sequence` (tokens per block, i.e. resolution), `history` (KV-cache length,
which grows until the sliding window caps it) and `context` (prompt length,
which changes on every mid-session rewrite).
4. `spaces.aoti_compile_and_save()` → AOTInductor archive.
5. Uploads it to [`hugging-apps/zing-0-5-world-model-aoti`](https://huggingface.co/hugging-apps/zing-0-5-world-model-aoti)
as `WanAttentionBlock/package.pt2`.
That path is the `{block class name}/package.pt2` layout the `spaces` package
expects — the same convention as `zerogpu-aoti/FLUX.1` and `zerogpu-aoti/Wan2` —
so the demo picks it up with one line:
```python
spaces.aoti_blocks_load(PIPELINE.generator, "hugging-apps/zing-0-5-world-model-aoti")
```
Weights are **not** baked into the archive: `spaces.aoti_patch` hands each block its
own `state_dict()` at call time, so a single `package.pt2` serves all 30 layers.
## Requirements
Needs an `HF_TOKEN` secret with write access to the artifact repo. Compilation runs
inside a single ZeroGPU allocation (`@spaces.GPU(duration=1500)`).
|