Nova Anime AM v5 ↔ Qwen-Image-2.1: interface and full-codec controls
The exact user-selected Nova checkpoint and Qwen-Image-2.1 now load and generate through their native interfaces. Both full-codec conversion directions and both starting domains in shared sampling work. No learned Nova/Qwen bridge has been trained yet. These measurements establish the pretrained baseline and expose what the future compact bridge must preserve.
Pinned models and weight checks
Nova is Civitai version 3338179,
named v5.0:2.9B. Its SHA-256 is
fdbbfbc3386dfb66eba4d2b47106377e7348a8b24b10f77e3f9584444949371e.
All 2,786,883,584 transformer and 134,663,680 conditioner parameters loaded
strictly. This version needs 40 transformer layers, and uses its own text
conditioner. The conditioner's nonpersistent rotary buffer stays float32.
Qwen-Image-2.1 is pinned to
b3179ad355be050328e483a9dfdd9e60cd62adfa. All 297 visual-transformer and 750
text-encoder tensor keys and shapes match their native architectures, without
missing or extra keys. All checkpoint tensors in those components are BF16.
All six transformer/text weight shards also pass SHA-256 checks against the
pinned Hub revision's LFS metadata; the checksum record
is committed with this report.
The visual transformer has 7,115,124,736 parameters; the text/vision encoder has
8,767,123,696. The text encoder is released before loading the visual transformer
to avoid retaining both large components in host RAM.
The pair configuration pins both codecs and their weight checksums. Nova's codec uses 16-channel, 8× RGB latents; Qwen 2.1 uses 64-channel, 16× RGBA latents. Per-channel normalization is applied once at the sampler boundary. RGB gains opaque alpha; RGBA is composited over white when entering Nova. Transparency is therefore not preserved by an RGB round trip.
Both codec adapters match native BF16 encoder and decoder forwards at 512px. Both ordinary full-codec bridges repeat exactly without advancing global RNG. Native Nova and Qwen denoiser first-call parity passed on both development prompts. The Qwen adapter explicitly strips leading conditioning tokens from the native output. A regression test covers that behavior, packing, masks and timestep precision.
Finished conversion
On eight fixed SD-generated development images, LPIPS compares each converted image with the source codec's decoded image. Timing is batch-one bridge latency on RTX A6000, excluding final recipient decoding and LPIPS. It is the median of seven repeats on the first case, not a latency distribution over cases.
| Direction | Mean conversion LPIPS | Mean round-trip LPIPS | Bridge ms |
|---|---|---|---|
| Nova codec → Qwen 2.1 codec | 0.020476 | 0.026145 | 125.80 |
| Qwen 2.1 codec → Nova codec | 0.037181 | 0.033791 | 298.01 |
The subsequent native-generation probe used two fixed prompts, a ceramic teapot and an autumn mountain stream, with 40 denoising calls at 512px. Native finished conversion LPIPS was 0.001676 / 0.004972 for Nova → Qwen, and 0.005421 / 0.098531 for Qwen → Nova. The detailed photographic landscape loses texture through the older Nova/Anima codec. These two examples are not a broad quality estimate.
Switching and repeatability
Each starting domain used a shared flow grid with shift 3, either staying in one model, switching after 20 calls, or switching after calls 10 and 30. The latter two policies each allocate 20 calls to each model. Clean estimates pass through the full codecs; noise coordinates transfer by exact pixel shuffle/unshuffle. The noise dimensionality is equal, but this does not align the models' denoising fields. Sampling is separate from the deterministic glue operation.
All 12 shared trajectories repeated bitwise, with unchanged CPU/CUDA global RNG. All four prompt/domain native first-call checks passed.
| Starting model | One-switch LPIPS vs shared single-model control, two prompts | Switch-and-return LPIPS, two prompts |
|---|---|---|
| Nova | 0.259918 / 0.513853 | 0.362673 / 0.604508 |
| Qwen 2.1 | 0.320215 / 0.698308 | 0.528536 / 0.766178 |
Visual inspection found recognizable requested subjects, substantial style and composition changes, and some softened detail after switching back. These LPIPS values measure departure from a control, not a generation-quality ranking. The native versus shared single-model controls also differ: Nova 0.009478 / 0.010177 and Qwen 0.014215 / 0.321615. Native schedules and BF16 scheduler states differ from the shared grid and float32 state. The initial FP32 noise is supplied to both controls; native Qwen rounds it to BF16. Do not attribute the native/shared differences entirely to glue.
Evidence and reproduction
Source commit 8533c1e contains the sampling probe and regression tests.
The isolated runtime and setup commands are in the recipe.
The running SD/Sana numerical trainer and shared Music packages were not changed.
CUDA_VISIBLE_DEVICES= PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
.venv/bin/python -m scripts.audit_qwen21_weight_headers \
--verify-weights --out artifacts/my-qwen-weight-audit.json
CUDA_VISIBLE_DEVICES=1 PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
.venv/bin/python -m scripts.probe_nova_qwen21_sampling \
--out artifacts/my-nova-qwen-probe --max-cases 2 --steps 40 \
--gallery-root artifacts/runs/my-live-run
Use an idle assigned GPU. The actual probes ran inside GPU 1's idle window after the first forward trainer finished. The new-pair work used 762.80 allocated GPU seconds (0.212 GPU-hours), including the first failed token-layout smoke run. That run and its error log are preserved. This is a separate ledger from the SD/Sana training campaign.
Exact evidence is under artifacts/bridge-transfer-20260920/: SUMMARY.json,
inspection/codec-interface-gpu512.json, inspection/qwen21-weight-headers.json,
codec-transfer-512/report.json and sampling-development-512/report.json.
All four completed prompt/domain galleries publish automatically on port 8779.
Desktop/mobile checks found all five images loaded, no JavaScript errors and no
horizontal overflow. The reserved test split remains unopened.
The learned-transfer pilot is now queued behind the entire SD/Sana three-seed
matrix and its qualification/export plan. Source commit c131d2f implements
native training pairs, both directional movable/fixed controls, validation-selected
compact exports and matched full-codec switching comparisons. Waiter PID 1168603
is recorded in artifacts/runs/bridge-nova-qwen21-20260920/launch.json.
No learned Nova/Qwen GPU work has started and no learned-quality result is claimed.
The first queued stage checks the actual pretrained prefix before generating data;
44 CPU tests passed from a clean source archive. See the recipe.