model-glue-nova-qwen21-experimental / latent-only-next-session.md
ntc-ai's picture
Record verified training shutdown and final archive snapshot
053ffd6 verified
|
Raw History Blame Contribute Delete
4.24 kB
# Fresh start: image-model bridges without pixel-space target preparation
## User intent
The user questioned the current bridge approach after seeing combined character
sampling. They asked to publish the current work, stop training, clear the
conversation context, and investigate an approach that does not leave latent
space. **Do not resume the archived decoder/encoder-teacher training or data
producer. No successor training job has been launched.**
Start with this short file and `AGENTS.md`; the long historical handoff is optional
reference, not the starting plan. Design the new experiment from the user's
latent-only constraint. In particular, do not silently reuse decode→encode
targets or pixel losses. Define how aligned targets or another justified
objective can be obtained. Independently generated same-prompt images are not
automatically aligned latent targets. Keep held-out quality evaluation separate
from target preparation and optimization; clarify its permitted decoding scope
if necessary.
## Pair and useful reusable code
- Source: Nova Anime AM v5.0:2.9B, Civitai model 2604424 / version 3338179,
SHA-256 `fdbbfbc3386dfb66eba4d2b47106377e7348a8b24b10f77e3f9584444949371e`.
- Recipient: Qwen-Image-2.1, revision `b3179ad355be050328e483a9dfdd9e60cd62adfa`.
- Exact interfaces: `configs/bridge-nova-qwen21-20260920.json`.
- Raw latent layouts at 512px: Nova `B×16×64×64`, Qwen `B×64×32×32`.
Equal coordinate counts allow an exact pixel-shuffle noise permutation;
that does not establish semantic alignment.
- `model_glue/nova_qwen_transfer.py` loads the exact pretrained networks and
native prompt conditioning. `image_codec_interface.py` records normalization.
`image_flow_handoff.py` implements the existing shared flow trajectory.
These interfaces can be reused without assuming the old objective is correct.
The old bridge's inference was already a latent-to-latent network forward pass.
Its supervised targets came from source decoding followed by recipient encoding.
It uses frozen source decoder prefixes and recipient encoder heads. Its errors
and visible changes motivate investigation; they do not prove a fundamental
incompatibility between the pretrained models.
## Archived, stopped, and recoverable
Public [gallery](https://huggingface.co/spaces/ntc-ai/model-glue-nova-qwen21) and
[weights/source/results](https://huggingface.co/ntc-ai/model-glue-nova-qwen21-experimental).
GitHub `255BITS/model-glue` is synced and remains private. Public inference and
experiment source are included as `source.zip` on Hugging Face.
The growing16x trainers stopped with full recovery at **120,964 forward** and
**97,924 reverse** local updates (absolute 193,719 / 154,968). Full model, EMA,
optimizers, RNG and data snapshots are retained under
`artifacts/runs/bridge-nova-qwen21-growing16x-20260921/movable/{forward,reverse}/recovery.pt`.
The old target producer stopped after **1,854 new prompts**, or **2,110 total**
of 4,096 planned; completed latent pairs remain under
`artifacts/runs/bridge-nova-qwen21-data16x-20260921`.
The raw supervisors report `failed` because intentional interruption raises
after saving; the authoritative user-stop verification is
`artifacts/releases/nova-qwen21-20260921/stop-record.json`.
Character samples use the separately pinned **90k / 78k** pair. Validation-best
EMA remains forward local step 0 / LPIPS 0.1020884, reverse step 90k / 0.0875336.
No final test or matched fixed-cloud comparison was run. No particle advantage
or general-purpose bridge qualification is established.
## Workspace constraints
Use `.venv/bin/python`; isolate Qwen dependencies with
`PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src`.
Both GPUs 0 and 1 are assigned; prefix GPU commands with `CUDA_VISIBLE_DEVICES`.
Preserve unrelated Music/Lumen/ntc-image-studio processes and preexisting dirty
files. Do not upgrade the shared environment, delete old artifacts, or commit
generated weights/data. Existing local dashboard remains on port 8779.
Inference must remain an ordinary forward pass. Keep train/validation/test
separate, measure held-out output quality, and use explicitly adult characters
when people are needed. Follow current `AGENTS.md`.