model-glue-nova-qwen21-experimental / reports /nova-qwen21-interface-20260920.md
ntc-ai's picture
Archive Nova Qwen bridge checkpoints and measured combined sampling
1f167a1 verified
|
Raw History Blame Contribute Delete
7.14 kB
# Nova Anime AM v5 ↔ Qwen-Image-2.1: interface and full-codec controls
The exact user-selected Nova checkpoint and Qwen-Image-2.1 now load and generate
through their native interfaces. Both full-codec conversion directions and both
starting domains in shared sampling work. **No learned Nova/Qwen bridge has been
trained yet.** These measurements establish the pretrained baseline and expose
what the future compact bridge must preserve.
## Pinned models and weight checks
Nova is [Civitai version 3338179](https://civitai.com/models/2604424/nova-anime-am?modelVersionId=3338179),
named v5.0:2.9B. Its SHA-256 is
`fdbbfbc3386dfb66eba4d2b47106377e7348a8b24b10f77e3f9584444949371e`.
All 2,786,883,584 transformer and 134,663,680 conditioner parameters loaded
strictly. This version needs **40 transformer layers**, and uses its own text
conditioner. The conditioner's nonpersistent rotary buffer stays float32.
[Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) is pinned to
`b3179ad355be050328e483a9dfdd9e60cd62adfa`. All 297 visual-transformer and 750
text-encoder tensor keys and shapes match their native architectures, without
missing or extra keys. All checkpoint tensors in those components are BF16.
All six transformer/text weight shards also pass SHA-256 checks against the
pinned Hub revision's LFS metadata; the [checksum record](nova-qwen21-weight-integrity-20260920.json)
is committed with this report.
The visual transformer has 7,115,124,736 parameters; the text/vision encoder has
8,767,123,696. The text encoder is released before loading the visual transformer
to avoid retaining both large components in host RAM.
The [pair configuration](../../configs/bridge-nova-qwen21-20260920.json) pins both
codecs and their weight checksums. Nova's codec uses 16-channel, 8Γ— RGB latents;
Qwen 2.1 uses 64-channel, 16Γ— RGBA latents. Per-channel normalization is applied
once at the sampler boundary. RGB gains opaque alpha; RGBA is composited over
white when entering Nova. Transparency is therefore not preserved by an RGB
round trip.
Both codec adapters match native BF16 encoder and decoder forwards at 512px.
Both ordinary full-codec bridges repeat exactly without advancing global RNG.
Native Nova and Qwen denoiser first-call parity passed on both development prompts.
The Qwen adapter explicitly strips leading conditioning tokens from the native
output. A regression test covers that behavior, packing, masks and timestep
precision.
## Finished conversion
On eight fixed SD-generated development images, LPIPS compares each converted
image with the **source codec's decoded image**. Timing is batch-one bridge
latency on RTX A6000, excluding final recipient decoding and LPIPS. It is the
median of seven repeats on the first case, not a latency distribution over cases.
| Direction | Mean conversion LPIPS | Mean round-trip LPIPS | Bridge ms |
|---|---:|---:|---:|
| Nova codec β†’ Qwen 2.1 codec | 0.020476 | 0.026145 | 125.80 |
| Qwen 2.1 codec β†’ Nova codec | 0.037181 | 0.033791 | 298.01 |
The subsequent native-generation probe used two fixed prompts, a ceramic teapot
and an autumn mountain stream, with 40 denoising calls at 512px. Native finished
conversion LPIPS was 0.001676 / 0.004972 for Nova β†’ Qwen, and 0.005421 / 0.098531
for Qwen β†’ Nova. The detailed photographic landscape loses texture through the
older Nova/Anima codec. These two examples are not a broad quality estimate.
## Switching and repeatability
Each starting domain used a shared flow grid with shift 3, either staying in one
model, switching after 20 calls, or switching after calls 10 and 30. The latter
two policies each allocate 20 calls to each model. Clean estimates pass through
the full codecs; noise coordinates transfer by exact pixel shuffle/unshuffle.
The noise dimensionality is equal, but this does not align the models' denoising
fields. Sampling is separate from the deterministic glue operation.
All **12 shared trajectories repeated bitwise**, with unchanged CPU/CUDA global
RNG. All four prompt/domain native first-call checks passed.
| Starting model | One-switch LPIPS vs shared single-model control, two prompts | Switch-and-return LPIPS, two prompts |
|---|---|---|
| Nova | 0.259918 / 0.513853 | 0.362673 / 0.604508 |
| Qwen 2.1 | 0.320215 / 0.698308 | 0.528536 / 0.766178 |
Visual inspection found recognizable requested subjects, substantial style and
composition changes, and some softened detail after switching back. These LPIPS
values measure departure from a control, not a generation-quality ranking.
The native versus shared single-model controls also differ: Nova 0.009478 /
0.010177 and Qwen 0.014215 / 0.321615. Native schedules and BF16 scheduler states
differ from the shared grid and float32 state. The initial FP32 noise is supplied
to both controls; native Qwen rounds it to BF16. Do not attribute the native/shared
differences entirely to glue.
## Evidence and reproduction
Source commit `8533c1e` contains the sampling probe and regression tests.
The isolated runtime and setup commands are in [the recipe](../image-bridge-recipe.md).
The running SD/Sana numerical trainer and shared Music packages were not changed.
```bash
CUDA_VISIBLE_DEVICES= PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
.venv/bin/python -m scripts.audit_qwen21_weight_headers \
--verify-weights --out artifacts/my-qwen-weight-audit.json
CUDA_VISIBLE_DEVICES=1 PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
.venv/bin/python -m scripts.probe_nova_qwen21_sampling \
--out artifacts/my-nova-qwen-probe --max-cases 2 --steps 40 \
--gallery-root artifacts/runs/my-live-run
```
Use an idle assigned GPU. The actual probes ran inside GPU 1's idle window after
the first forward trainer finished. The new-pair work used **762.80 allocated GPU
seconds (0.212 GPU-hours)**, including the first failed token-layout smoke run.
That run and its error log are preserved. This is a separate ledger from the
SD/Sana training campaign.
Exact evidence is under `artifacts/bridge-transfer-20260920/`: `SUMMARY.json`,
`inspection/codec-interface-gpu512.json`, `inspection/qwen21-weight-headers.json`,
`codec-transfer-512/report.json` and `sampling-development-512/report.json`.
All four completed prompt/domain galleries publish automatically on port 8779.
Desktop/mobile checks found all five images loaded, no JavaScript errors and no
horizontal overflow. The reserved test split remains unopened.
The learned-transfer pilot is now queued behind the entire SD/Sana three-seed
matrix and its qualification/export plan. Source commit `c131d2f` implements
native training pairs, both directional movable/fixed controls, validation-selected
compact exports and matched full-codec switching comparisons. Waiter PID 1168603
is recorded in `artifacts/runs/bridge-nova-qwen21-20260920/launch.json`.
No learned Nova/Qwen GPU work has started and no learned-quality result is claimed.
The first queued stage checks the actual pretrained prefix before generating data;
44 CPU tests passed from a clean source archive. See [the recipe](../image-bridge-recipe.md).