multimodalart's picture
multimodalart HF Staff
AoTI compiler
52e0fb7 verified
|
Raw History Blame Contribute Delete
2.54 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: Zing 0.5 AoTI Compiler
emoji: ⚙️
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
pinned: false
license: apache-2.0
python_version: '3.12'
startup_duration_timeout: 1h
short_description: torch.export + AOTInductor compiler for Zing-0.5
models:
  - seedleap/zing-0.5

Zing-0.5 — ahead-of-time compiler

Companion Space to hugging-apps/zing-0-5-world-model.

The demo spends essentially all of its GPU time inside 30 identical WanAttentionBlocks — five forward passes per generated block (four DMD denoising steps plus the KV-cache commit), so 150 block executions for every 16 frames it streams. This Space compiles that block once, ahead of time, and publishes the artifact so the demo never has to compile anything at runtime.

What it does

  1. Loads the exact same vendored zing_v0_5 package and sliding-window config (local_attn_size=33, sink_size=5) as the demo.
  2. Runs a short real rollout with spaces.aoti_capture() wrapped around generator.blocks[0], so the block is exported against genuine tensors rather than hand-made ones. The capture deliberately lands on the second generated block: block 0 runs with a cold (empty) KV cache and is kept on the eager path in the demo, so the compiled graph only ever sees a non-empty history.
  3. torch.export.export(..., strict=False) with three dynamic axes — sequence (tokens per block, i.e. resolution), history (KV-cache length, which grows until the sliding window caps it) and context (prompt length, which changes on every mid-session rewrite).
  4. spaces.aoti_compile_and_save() → AOTInductor archive.
  5. Uploads it to hugging-apps/zing-0-5-world-model-aoti as WanAttentionBlock/package.pt2.

That path is the {block class name}/package.pt2 layout the spaces package expects — the same convention as zerogpu-aoti/FLUX.1 and zerogpu-aoti/Wan2 — so the demo picks it up with one line:

spaces.aoti_blocks_load(PIPELINE.generator, "hugging-apps/zing-0-5-world-model-aoti")

Weights are not baked into the archive: spaces.aoti_patch hands each block its own state_dict() at call time, so a single package.pt2 serves all 30 layers.

Requirements

Needs an HF_TOKEN secret with write access to the artifact repo. Compilation runs inside a single ZeroGPU allocation (@spaces.GPU(duration=1500)).