infosave commited on
Commit
aabea2e
·
verified ·
1 Parent(s): db5fb4b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +35 -1
README.md CHANGED
@@ -170,7 +170,41 @@ then three deterministic steps that refine what the upscale invented.
170
  distilled — 8 is it exactly, other counts land on sigmas the model never saw
171
  and usually soften the frame. Detail comes from resolution and `--two-stage`.
172
 
173
- ### Measured
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
174
 
175
  49 frames at 24 fps, container on local storage:
176
 
 
170
  distilled — 8 is it exactly, other counts land on sigmas the model never saw
171
  and usually soften the frame. Detail comes from resolution and `--two-stage`.
172
 
173
+ #
174
+ ### LoRA adapters, and multi-subject references
175
+
176
+ ```sh
177
+ cortiq ltx-video --model $M --lora adapter.safetensors --lora-strength 0.8 \
178
+ --prompt "…" --out clip.y4m
179
+ ```
180
+
181
+ q4tp weights cannot absorb a low-rank update without dequantizing the whole
182
+ DiT, so the branch runs beside them — `y = x·Wᵀ + s·(x·Aᵀ)·Bᵀ`, on every path
183
+ including the fused Metal q/k/v submission. Rank 128 against a 4096×4096
184
+ projection is about 6% more arithmetic and no memory beyond the file — but
185
+ measured on an M4 a 384-token step goes 8.6 s to about 22 s, because that
186
+ arithmetic is host-side f32 beside a device-side 4-bit GEMM and does not
187
+ overlap it. The flops are cheap; the placement is what costs.
188
+
189
+ An adapter that also carries a `reference_slot_embedding` takes reference
190
+ stills, which is how the multi-subject adapters work:
191
+
192
+ ```sh
193
+ cortiq ltx-video --model $M --lora msr.safetensors \
194
+ --ref a.ppm --ref b.ppm --ref c.ppm \
195
+ --prompt "Image 1: … Image 2: … Image 3: …" --out clip.y4m
196
+ ```
197
+
198
+ Each still is held for 25 or 33 pixel frames (`--ref-frames`, whichever the
199
+ adapter was trained on), encoded by the same video VAE the render uses, given
200
+ its slot's learned per-channel bias on the latent, and placed at a negative
201
+ frame offset — slot 1 furthest back. Those tokens ride in the same sequence,
202
+ frozen, and are cropped off the result.
203
+
204
+ They cost sequence length: three references at 384×256 add 1152 tokens beside
205
+ 384 of clip. The stills must already be the render's size.
206
+
207
+ ## Measured
208
 
209
  49 frames at 24 fps, container on local storage:
210