UniVideo CI-VID full-finetune checkpoint β step 1750
Public migration checkpoint for the CI-VID interleaved text/video run:
- Training run: https://wandb.ai/diffusionrl-dyh/univideo-civid/runs/w2292csr
- Code and runbook: https://github.com/karrykkk/UniVideo-CI-VID-finetune/tree/civid-handoff-2026-09-20-v6
- Dataset: https://huggingface.co/datasets/karrykkk/CI-VID-SC-4072-repro
This repository is a research artifact. The upstream model and CI-VID dataset licenses continue to apply; publication here does not broaden those terms.
Training state
global step: 1750 / 10000
world size: 16 (2 nodes Γ 8 ranks)
FSDP: HYBRID_SHARD, shard=8, replicate=2
trainable transformer: full 12.9B MMDiT + qwen_project_in
Qwen adaptation: LoRA rank 16
VAE: frozen
precision: BF16
resolution budget: 160Γ288 area, 24 FPS, at most 129 frames
Total checkpoint size is approximately 132.35 GB.
Files
checkpoint-0001750/
βββ complete
βββ transformer.pt
βββ training_state.pt
βββ optimizer.00000-of-00016.pt ... optimizer.00015-of-00016.pt
βββ mllm_lora/
βββ adapter_config.json
βββ adapter_model.safetensors
βββ README.md
transformer.pt + mllm_lora/ are sufficient for inference. Exact optimizer
continuation additionally requires all 16 optimizer shards and the same
world-size-16 topology.
Download
hf download karrykkk/UniVideo-CI-VID-fullft-step1750 \
--revision step-1750 \
--local-dir /shared/checkpoints/UniVideo-CI-VID-fullft-step1750
Exact resume
export CIVID_RESUME_FROM=/shared/checkpoints/UniVideo-CI-VID-fullft-step1750/checkpoint-0001750
export CIVID_MAX_TRAIN_STEPS=10000
export WANDB_RUN_ID=w2292csr # omit this to create a new W&B run
# Run on both nodes with NNODES=2, NPROC_PER_NODE=8 and NODE_RANK=0/1.
bash scripts/run_civid_distributed.sh configs/train_civid_10k_fullft.yaml
The checkpoint preserves model, optimizer, sampler position, and global step, but not all Python/NumPy/CUDA RNG states; resumed training is not guaranteed to be bitwise identical to an uninterrupted run.