model-glue-nova-qwen21-experimental / reports /nova-qwen21-data16x-20260921.md
ntc-ai's picture
Archive Nova Qwen bridge checkpoints and measured combined sampling
1f167a1 verified
|
Raw History Blame Contribute Delete
6 kB

Background 16x prompt expansion

The 256-prompt Nova/Qwen campaign is now stopped at the user's request while this data is prepared. Recovery is saved at forward 72,755 / reverse 57,044; best validation checkpoints and optimizer/RNG states are preserved. Fixed controls and qualification did not run. A separate, resumable dataset adds 3,840 distinct prompts sampled from the prompt column of DiffusionDB, for 4,096 total training prompts. The source revision is fb620fbe49fa4420e0734bd9c0df11f51176b61f; metadata SHA256 is eecd341187bc91c07f5994ad0660d40228ea025616fd57a509bef8323677c68f. The dataset card declares CC0-1.0. No source images or image scores are read.

The user explicitly requested people and Anima's anime specialization. New prompt quotas are:

Stratum New prompts
Anime/manga/cel-shaded adult people 480
Other-style adult people 480
Anime/manga/cel-shaded nonhuman scenes 1,440
Other-style nonhuman scenes 1,440

Thus half the additions carry anime/manga/Ghibli/cel-shading intent, and one quarter depict adult people. People prompts explicitly require every person to be aged thirty or older. Text-only exclusions reject young-person references and explicit content. Selection does not inspect or rank generated results. Nonhuman subjects include objects, animals, landscapes, architecture, abstract patterns, and space scenes. The pinned 2M-row source has 1,956 / 42,628 / 4,862 / 128,491 eligible unique content stems in these four strata respectively. This is coverage of prompt intent; no claim is made that the models always follow a prompt or that these categories describe observed training outputs.

Sampling uses seed 42021 and SHA256 ordering, with whitespace/case normalization, first-comma-clause deduplication, and exclusions for exact and near-copies of inherited/reserved prompts. Existing validation shards remain byte-identical; the test split is unopened. Per-prompt seeds begin at 940000. The committed config records source row indices and hashes for reproduction.

Each new prompt produces native trajectories in both models with eight denoising steps and clean-estimate captures at 1, 3, 5, 7 plus finished step 8. Targets use the authorized full decoder→encoder teacher, scalar batch size one, transient pixels only. No training images are saved or reviewed. Both directions must be committed together; shards contain only source/target latents and noise fraction. Completion will produce 20,736 training rows per direction (1,536 inherited

  • 19,200 new) and the same 176 validation rows. This is 16x prompt coverage and 13.5x latent rows, since the retained older trajectories have more captures.

The initial worker used GPU 0 with a 0.25 active/sleep wall-time ratio. Once bridge training was stopped, it was restarted with no pacing (ratio 1.0). The 0.50 per-process CUDA allocator cap remains, and it waits for 26 GiB free before model loading. The ratio is voluntary pacing between operations, not a hardware partition. Prompt contexts reside on CPU and only the current trajectory's contexts move to CUDA. Nova is temporarily offloaded while the 8B Qwen text encoder is resident. The native numerical path is unchanged.

A real GPU test regenerated an inherited prompt using the CPU-context path: source, target, and noise fraction were bitwise identical in both directions. Peak allocated memory was 22,603,312,640 bytes; peak reserved was 23,154,655,232 bytes while the reverse trainer remained active. This is a single-case implementation check, not a throughput prediction for the corpus. Evidence: artifacts/verification/nova-qwen21-data16x-20260921/cpu-context-gpu-check.json. The 28 focused CPU tests pass, including age exclusions, quota selection, resume integrity, context restoration after failure, and pacing.

Reproduce selection and launch from the repository root:

CUDA_VISIBLE_DEVICES= .venv/bin/python -m scripts.prepare_nova_prompt_expansion
CUDA_VISIBLE_DEVICES=0 PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
  .venv/bin/python -m scripts.run_bridge_coverage \
  --root artifacts/runs/bridge-nova-qwen21-data16x-fullspeed-20260921 \
  --plan configs/bridge-nova-qwen21-data16x-fullspeed-queue-20260921.json \
  --no-time-cap \
  --prior-ledger artifacts/runs/bridge-nova-qwen21-data16x-20260921/compute-ledger.json \
  --allow-stopped-prior

The active supervisor is 2280383 and worker is 2280432. Check compute-ledger.json, launch.json, and generate-data16x.log in the fullspeed root before starting anything. Dataset progress.json and manifest.json remain in the original data root; active-worker.json there points to this supervisor. Do not duplicate a live worker. manifest.json is published only when all additions are complete. After a stopped worker, resume generation with the exact job arguments in the plan; preserve the old supervisor ledger rather than overwrite it.

The stopped 300k campaign keeps its original pinned dataset, and is incomplete. No automatic training restart is queued. This expansion is prepared for the next training phase. Do not hot-swap its manifest, or put the larger frozen-feature cache entirely on GPU: its expected size exceeds available training headroom. A future expanded-data training phase must explicitly select a bounded or CPU feature-cache strategy.

The provisional estimate after removing pacing is 7–9 hours remaining, using the earlier 128-prompt mean of 6.085 seconds per prompt for both directions (3,840 additions × 6.085 seconds = 6.49 hours), plus text encoding, loading and larger-manifest overhead. This is an extrapolation; the larger corpus was still in prompt encoding when quoted and had no measured native-case rate yet. The old paced worker completed no new cases; its partial unpublished embeddings were re-encoded on the unpaced restart. All durable inherited data is preserved.