|
Download latent-only-next-session.md from ntc-ai/model-glue-nova-qwen21-experimental: direct link, hf CLI and curl.
- Browser
- Download file 4.24 kB
-
https://huggingface.co/ntc-ai/model-glue-nova-qwen21-experimental/resolve/main/latent-only-next-session.md
- Command line
-
hf download hf://ntc-ai/model-glue-nova-qwen21-experimental/latent-only-next-session.md
-
curl -L -o latent-only-next-session.md https://huggingface.co/ntc-ai/model-glue-nova-qwen21-experimental/resolve/main/latent-only-next-session.md
4.24 kB
| # Fresh start: image-model bridges without pixel-space target preparation | |
| ## User intent | |
| The user questioned the current bridge approach after seeing combined character | |
| sampling. They asked to publish the current work, stop training, clear the | |
| conversation context, and investigate an approach that does not leave latent | |
| space. **Do not resume the archived decoder/encoder-teacher training or data | |
| producer. No successor training job has been launched.** | |
| Start with this short file and `AGENTS.md`; the long historical handoff is optional | |
| reference, not the starting plan. Design the new experiment from the user's | |
| latent-only constraint. In particular, do not silently reuse decode→encode | |
| targets or pixel losses. Define how aligned targets or another justified | |
| objective can be obtained. Independently generated same-prompt images are not | |
| automatically aligned latent targets. Keep held-out quality evaluation separate | |
| from target preparation and optimization; clarify its permitted decoding scope | |
| if necessary. | |
| ## Pair and useful reusable code | |
| - Source: Nova Anime AM v5.0:2.9B, Civitai model 2604424 / version 3338179, | |
| SHA-256 `fdbbfbc3386dfb66eba4d2b47106377e7348a8b24b10f77e3f9584444949371e`. | |
| - Recipient: Qwen-Image-2.1, revision `b3179ad355be050328e483a9dfdd9e60cd62adfa`. | |
| - Exact interfaces: `configs/bridge-nova-qwen21-20260920.json`. | |
| - Raw latent layouts at 512px: Nova `B×16×64×64`, Qwen `B×64×32×32`. | |
| Equal coordinate counts allow an exact pixel-shuffle noise permutation; | |
| that does not establish semantic alignment. | |
| - `model_glue/nova_qwen_transfer.py` loads the exact pretrained networks and | |
| native prompt conditioning. `image_codec_interface.py` records normalization. | |
| `image_flow_handoff.py` implements the existing shared flow trajectory. | |
| These interfaces can be reused without assuming the old objective is correct. | |
| The old bridge's inference was already a latent-to-latent network forward pass. | |
| Its supervised targets came from source decoding followed by recipient encoding. | |
| It uses frozen source decoder prefixes and recipient encoder heads. Its errors | |
| and visible changes motivate investigation; they do not prove a fundamental | |
| incompatibility between the pretrained models. | |
| ## Archived, stopped, and recoverable | |
| Public [gallery](https://huggingface.co/spaces/ntc-ai/model-glue-nova-qwen21) and | |
| [weights/source/results](https://huggingface.co/ntc-ai/model-glue-nova-qwen21-experimental). | |
| GitHub `255BITS/model-glue` is synced and remains private. Public inference and | |
| experiment source are included as `source.zip` on Hugging Face. | |
| The growing16x trainers stopped with full recovery at **120,964 forward** and | |
| **97,924 reverse** local updates (absolute 193,719 / 154,968). Full model, EMA, | |
| optimizers, RNG and data snapshots are retained under | |
| `artifacts/runs/bridge-nova-qwen21-growing16x-20260921/movable/{forward,reverse}/recovery.pt`. | |
| The old target producer stopped after **1,854 new prompts**, or **2,110 total** | |
| of 4,096 planned; completed latent pairs remain under | |
| `artifacts/runs/bridge-nova-qwen21-data16x-20260921`. | |
| The raw supervisors report `failed` because intentional interruption raises | |
| after saving; the authoritative user-stop verification is | |
| `artifacts/releases/nova-qwen21-20260921/stop-record.json`. | |
| Character samples use the separately pinned **90k / 78k** pair. Validation-best | |
| EMA remains forward local step 0 / LPIPS 0.1020884, reverse step 90k / 0.0875336. | |
| No final test or matched fixed-cloud comparison was run. No particle advantage | |
| or general-purpose bridge qualification is established. | |
| ## Workspace constraints | |
| Use `.venv/bin/python`; isolate Qwen dependencies with | |
| `PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src`. | |
| Both GPUs 0 and 1 are assigned; prefix GPU commands with `CUDA_VISIBLE_DEVICES`. | |
| Preserve unrelated Music/Lumen/ntc-image-studio processes and preexisting dirty | |
| files. Do not upgrade the shared environment, delete old artifacts, or commit | |
| generated weights/data. Existing local dashboard remains on port 8779. | |
| Inference must remain an ordinary forward pass. Keep train/validation/test | |
| separate, measure held-out output quality, and use explicitly adult characters | |
| when people are needed. Follow current `AGENTS.md`. | |