# Fresh start: image-model bridges without pixel-space target preparation ## User intent The user questioned the current bridge approach after seeing combined character sampling. They asked to publish the current work, stop training, clear the conversation context, and investigate an approach that does not leave latent space. **Do not resume the archived decoder/encoder-teacher training or data producer. No successor training job has been launched.** Start with this short file and `AGENTS.md`; the long historical handoff is optional reference, not the starting plan. Design the new experiment from the user's latent-only constraint. In particular, do not silently reuse decode→encode targets or pixel losses. Define how aligned targets or another justified objective can be obtained. Independently generated same-prompt images are not automatically aligned latent targets. Keep held-out quality evaluation separate from target preparation and optimization; clarify its permitted decoding scope if necessary. ## Pair and useful reusable code - Source: Nova Anime AM v5.0:2.9B, Civitai model 2604424 / version 3338179, SHA-256 `fdbbfbc3386dfb66eba4d2b47106377e7348a8b24b10f77e3f9584444949371e`. - Recipient: Qwen-Image-2.1, revision `b3179ad355be050328e483a9dfdd9e60cd62adfa`. - Exact interfaces: `configs/bridge-nova-qwen21-20260920.json`. - Raw latent layouts at 512px: Nova `B×16×64×64`, Qwen `B×64×32×32`. Equal coordinate counts allow an exact pixel-shuffle noise permutation; that does not establish semantic alignment. - `model_glue/nova_qwen_transfer.py` loads the exact pretrained networks and native prompt conditioning. `image_codec_interface.py` records normalization. `image_flow_handoff.py` implements the existing shared flow trajectory. These interfaces can be reused without assuming the old objective is correct. The old bridge's inference was already a latent-to-latent network forward pass. Its supervised targets came from source decoding followed by recipient encoding. It uses frozen source decoder prefixes and recipient encoder heads. Its errors and visible changes motivate investigation; they do not prove a fundamental incompatibility between the pretrained models. ## Archived, stopped, and recoverable Public [gallery](https://huggingface.co/spaces/ntc-ai/model-glue-nova-qwen21) and [weights/source/results](https://huggingface.co/ntc-ai/model-glue-nova-qwen21-experimental). GitHub `255BITS/model-glue` is synced and remains private. Public inference and experiment source are included as `source.zip` on Hugging Face. The growing16x trainers stopped with full recovery at **120,964 forward** and **97,924 reverse** local updates (absolute 193,719 / 154,968). Full model, EMA, optimizers, RNG and data snapshots are retained under `artifacts/runs/bridge-nova-qwen21-growing16x-20260921/movable/{forward,reverse}/recovery.pt`. The old target producer stopped after **1,854 new prompts**, or **2,110 total** of 4,096 planned; completed latent pairs remain under `artifacts/runs/bridge-nova-qwen21-data16x-20260921`. The raw supervisors report `failed` because intentional interruption raises after saving; the authoritative user-stop verification is `artifacts/releases/nova-qwen21-20260921/stop-record.json`. Character samples use the separately pinned **90k / 78k** pair. Validation-best EMA remains forward local step 0 / LPIPS 0.1020884, reverse step 90k / 0.0875336. No final test or matched fixed-cloud comparison was run. No particle advantage or general-purpose bridge qualification is established. ## Workspace constraints Use `.venv/bin/python`; isolate Qwen dependencies with `PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src`. Both GPUs 0 and 1 are assigned; prefix GPU commands with `CUDA_VISIBLE_DEVICES`. Preserve unrelated Music/Lumen/ntc-image-studio processes and preexisting dirty files. Do not upgrade the shared environment, delete old artifacts, or commit generated weights/data. Existing local dashboard remains on port 8779. Inference must remain an ordinary forward pass. Keep train/validation/test separate, measure held-out output quality, and use explicitly adult characters when people are needed. Follow current `AGENTS.md`.