File size: 2,544 Bytes
8ca16d2
52e0fb7
 
 
 
8ca16d2
 
 
 
52e0fb7
 
 
 
 
 
8ca16d2
 
52e0fb7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
---
title: Zing 0.5 AoTI Compiler
emoji: ⚙️
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
pinned: false
license: apache-2.0
python_version: "3.12"
startup_duration_timeout: 1h
short_description: torch.export + AOTInductor compiler for Zing-0.5
models:
  - seedleap/zing-0.5
---

# Zing-0.5 — ahead-of-time compiler

Companion Space to
[`hugging-apps/zing-0-5-world-model`](https://huggingface.co/spaces/hugging-apps/zing-0-5-world-model).

The demo spends essentially all of its GPU time inside 30 identical
`WanAttentionBlock`s — five forward passes per generated block (four DMD denoising
steps plus the KV-cache commit), so 150 block executions for every 16 frames it
streams. This Space compiles that block **once, ahead of time**, and publishes the
artifact so the demo never has to compile anything at runtime.

## What it does

1. Loads the exact same vendored `zing_v0_5` package and sliding-window config
   (`local_attn_size=33`, `sink_size=5`) as the demo.
2. Runs a short real rollout with `spaces.aoti_capture()` wrapped around
   `generator.blocks[0]`, so the block is exported against genuine tensors rather
   than hand-made ones. The capture deliberately lands on the *second* generated
   block: block 0 runs with a cold (empty) KV cache and is kept on the eager path
   in the demo, so the compiled graph only ever sees a non-empty history.
3. `torch.export.export(..., strict=False)` with three dynamic axes —
   `sequence` (tokens per block, i.e. resolution), `history` (KV-cache length,
   which grows until the sliding window caps it) and `context` (prompt length,
   which changes on every mid-session rewrite).
4. `spaces.aoti_compile_and_save()` → AOTInductor archive.
5. Uploads it to [`hugging-apps/zing-0-5-world-model-aoti`](https://huggingface.co/hugging-apps/zing-0-5-world-model-aoti)
   as `WanAttentionBlock/package.pt2`.

That path is the `{block class name}/package.pt2` layout the `spaces` package
expects — the same convention as `zerogpu-aoti/FLUX.1` and `zerogpu-aoti/Wan2` —
so the demo picks it up with one line:

```python
spaces.aoti_blocks_load(PIPELINE.generator, "hugging-apps/zing-0-5-world-model-aoti")
```

Weights are **not** baked into the archive: `spaces.aoti_patch` hands each block its
own `state_dict()` at call time, so a single `package.pt2` serves all 30 layers.

## Requirements

Needs an `HF_TOKEN` secret with write access to the artifact repo. Compilation runs
inside a single ZeroGPU allocation (`@spaces.GPU(duration=1500)`).