model-glue-nova-qwen21-experimental / reports /nova-qwen21-interface-20260920.md
ntc-ai's picture
Archive Nova Qwen bridge checkpoints and measured combined sampling
1f167a1 verified
|
Raw
History Blame Contribute Delete
7.14 kB

Nova Anime AM v5 ↔ Qwen-Image-2.1: interface and full-codec controls

The exact user-selected Nova checkpoint and Qwen-Image-2.1 now load and generate through their native interfaces. Both full-codec conversion directions and both starting domains in shared sampling work. No learned Nova/Qwen bridge has been trained yet. These measurements establish the pretrained baseline and expose what the future compact bridge must preserve.

Pinned models and weight checks

Nova is Civitai version 3338179, named v5.0:2.9B. Its SHA-256 is fdbbfbc3386dfb66eba4d2b47106377e7348a8b24b10f77e3f9584444949371e. All 2,786,883,584 transformer and 134,663,680 conditioner parameters loaded strictly. This version needs 40 transformer layers, and uses its own text conditioner. The conditioner's nonpersistent rotary buffer stays float32.

Qwen-Image-2.1 is pinned to b3179ad355be050328e483a9dfdd9e60cd62adfa. All 297 visual-transformer and 750 text-encoder tensor keys and shapes match their native architectures, without missing or extra keys. All checkpoint tensors in those components are BF16. All six transformer/text weight shards also pass SHA-256 checks against the pinned Hub revision's LFS metadata; the checksum record is committed with this report. The visual transformer has 7,115,124,736 parameters; the text/vision encoder has 8,767,123,696. The text encoder is released before loading the visual transformer to avoid retaining both large components in host RAM.

The pair configuration pins both codecs and their weight checksums. Nova's codec uses 16-channel, 8× RGB latents; Qwen 2.1 uses 64-channel, 16× RGBA latents. Per-channel normalization is applied once at the sampler boundary. RGB gains opaque alpha; RGBA is composited over white when entering Nova. Transparency is therefore not preserved by an RGB round trip.

Both codec adapters match native BF16 encoder and decoder forwards at 512px. Both ordinary full-codec bridges repeat exactly without advancing global RNG. Native Nova and Qwen denoiser first-call parity passed on both development prompts. The Qwen adapter explicitly strips leading conditioning tokens from the native output. A regression test covers that behavior, packing, masks and timestep precision.

Finished conversion

On eight fixed SD-generated development images, LPIPS compares each converted image with the source codec's decoded image. Timing is batch-one bridge latency on RTX A6000, excluding final recipient decoding and LPIPS. It is the median of seven repeats on the first case, not a latency distribution over cases.

Direction Mean conversion LPIPS Mean round-trip LPIPS Bridge ms
Nova codec → Qwen 2.1 codec 0.020476 0.026145 125.80
Qwen 2.1 codec → Nova codec 0.037181 0.033791 298.01

The subsequent native-generation probe used two fixed prompts, a ceramic teapot and an autumn mountain stream, with 40 denoising calls at 512px. Native finished conversion LPIPS was 0.001676 / 0.004972 for Nova → Qwen, and 0.005421 / 0.098531 for Qwen → Nova. The detailed photographic landscape loses texture through the older Nova/Anima codec. These two examples are not a broad quality estimate.

Switching and repeatability

Each starting domain used a shared flow grid with shift 3, either staying in one model, switching after 20 calls, or switching after calls 10 and 30. The latter two policies each allocate 20 calls to each model. Clean estimates pass through the full codecs; noise coordinates transfer by exact pixel shuffle/unshuffle. The noise dimensionality is equal, but this does not align the models' denoising fields. Sampling is separate from the deterministic glue operation.

All 12 shared trajectories repeated bitwise, with unchanged CPU/CUDA global RNG. All four prompt/domain native first-call checks passed.

Starting model One-switch LPIPS vs shared single-model control, two prompts Switch-and-return LPIPS, two prompts
Nova 0.259918 / 0.513853 0.362673 / 0.604508
Qwen 2.1 0.320215 / 0.698308 0.528536 / 0.766178

Visual inspection found recognizable requested subjects, substantial style and composition changes, and some softened detail after switching back. These LPIPS values measure departure from a control, not a generation-quality ranking. The native versus shared single-model controls also differ: Nova 0.009478 / 0.010177 and Qwen 0.014215 / 0.321615. Native schedules and BF16 scheduler states differ from the shared grid and float32 state. The initial FP32 noise is supplied to both controls; native Qwen rounds it to BF16. Do not attribute the native/shared differences entirely to glue.

Evidence and reproduction

Source commit 8533c1e contains the sampling probe and regression tests. The isolated runtime and setup commands are in the recipe. The running SD/Sana numerical trainer and shared Music packages were not changed.

CUDA_VISIBLE_DEVICES= PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
  .venv/bin/python -m scripts.audit_qwen21_weight_headers \
  --verify-weights --out artifacts/my-qwen-weight-audit.json
CUDA_VISIBLE_DEVICES=1 PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
  .venv/bin/python -m scripts.probe_nova_qwen21_sampling \
  --out artifacts/my-nova-qwen-probe --max-cases 2 --steps 40 \
  --gallery-root artifacts/runs/my-live-run

Use an idle assigned GPU. The actual probes ran inside GPU 1's idle window after the first forward trainer finished. The new-pair work used 762.80 allocated GPU seconds (0.212 GPU-hours), including the first failed token-layout smoke run. That run and its error log are preserved. This is a separate ledger from the SD/Sana training campaign.

Exact evidence is under artifacts/bridge-transfer-20260920/: SUMMARY.json, inspection/codec-interface-gpu512.json, inspection/qwen21-weight-headers.json, codec-transfer-512/report.json and sampling-development-512/report.json. All four completed prompt/domain galleries publish automatically on port 8779. Desktop/mobile checks found all five images loaded, no JavaScript errors and no horizontal overflow. The reserved test split remains unopened.

The learned-transfer pilot is now queued behind the entire SD/Sana three-seed matrix and its qualification/export plan. Source commit c131d2f implements native training pairs, both directional movable/fixed controls, validation-selected compact exports and matched full-codec switching comparisons. Waiter PID 1168603 is recorded in artifacts/runs/bridge-nova-qwen21-20260920/launch.json. No learned Nova/Qwen GPU work has started and no learned-quality result is claimed. The first queued stage checks the actual pretrained prefix before generating data; 44 CPU tests passed from a clean source archive. See the recipe.