infosave commited on
Commit
7f4b1b5
Β·
verified Β·
1 Parent(s): a5dafbc

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +36 -3
README.md CHANGED
@@ -187,6 +187,37 @@ cortiq ltx-render --model $M --context context.safetensors \
187
  cortiq ltx-decode --model $M --latent latent.safetensors --out-dir frames/
188
  ```
189
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
190
  ## What is inside
191
 
192
  | component | weights | in the file | codec |
@@ -282,9 +313,11 @@ machine, no Python, no GPU. `--quant` picks the codec for the big planes
282
  48-block audio-video transformer, the sampler, the latent upscaler, the
283
  video VAE, the audio VAE with its BigVGAN vocoder and bandwidth extension,
284
  and the duration head.
285
- * ⏳ **Image and video conditioning, LoRAs and the IC-LoRA upscaler** are next.
286
- The weights for conditioning are in this file (the video VAE's encoder half);
287
- the LoRAs and the IC-LoRA upscaler are separate releases.
 
 
288
 
289
  Everything above is honest about what it is: a 4-bit repack. The reference at
290
  bf16 is the quality ceiling, and the codec's cost was measured stage by stage
 
187
  cortiq ltx-decode --model $M --latent latent.safetensors --out-dir frames/
188
  ```
189
 
190
+ ## Every mode the model has
191
+
192
+ LTX-2.5 is one network with two streams, and a *mode* is simply which parts
193
+ you hold fixed. Conditioning is encoded into the model's own latent space and
194
+ frozen there β€” the sampler gets a timestep of zero for those tokens and
195
+ leaves them alone β€” so all of this is one command with different inputs.
196
+
197
+ | mode | how |
198
+ |---|---|
199
+ | text β†’ video + sound | `--prompt "…" --out-audio track.wav` |
200
+ | text β†’ video | the same, without `--out-audio` |
201
+ | text β†’ sound | the same, keeping only the wav |
202
+ | image + text β†’ video (+ sound) | `--image still.ppm` |
203
+ | video β†’ video | `--video frames/` |
204
+ | video β†’ sound | `--video frames/ --video-to-audio` |
205
+ | sound β†’ video | `--audio-in track.wav` |
206
+ | sound β†’ sound | `--audio-in track.wav --out-audio out.wav` |
207
+ | image + sound β†’ video | `--image still.ppm --audio-in track.wav` |
208
+
209
+ ```bash
210
+ # a still into a shot, with its soundtrack
211
+ ffmpeg -i photo.jpg -vf scale=384:256 -pix_fmt rgb24 still.ppm
212
+ cortiq ltx-video --model $M --image still.ppm \
213
+ --prompt "the camera pushes in slowly as the light shifts" \
214
+ --height 256 --width 384 --frames 49 --out-dir out/ --out-audio out.wav
215
+ ```
216
+
217
+ Image conditioning runs through the video VAE's **encoder**, audio
218
+ conditioning through the audio VAE's and a log-mel front end β€” both in this
219
+ same file, along with everything else.
220
+
221
  ## What is inside
222
 
223
  | component | weights | in the file | codec |
 
313
  48-block audio-video transformer, the sampler, the latent upscaler, the
314
  video VAE, the audio VAE with its BigVGAN vocoder and bandwidth extension,
315
  and the duration head.
316
+ * βœ… **Every conditioning mode**: image-to-video, video-to-video,
317
+ video-to-audio, audio-to-video, audio-to-audio and the image+audio pairs β€”
318
+ all from this one file, because both VAE encoders are in it.
319
+ * ⏳ **LoRAs and the IC-LoRA upscaler** are separate releases and not packed
320
+ here yet.
321
 
322
  Everything above is honest about what it is: a 4-bit repack. The reference at
323
  bf16 is the quality ceiling, and the codec's cost was measured stage by stage