Text-to-Video
cortiq
Rust
cmf
video
image-to-video
image-text-to-video
video-to-video
video-to-audio
audio-to-video
text-to-audio
audio-to-audio
any-to-any
text-to-audio-video
ltx-video
ltx-2.5
4-bit precision
Instructions to use infosave/LTX-2.5-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/LTX-2.5-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/LTX-2.5-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -187,6 +187,37 @@ cortiq ltx-render --model $M --context context.safetensors \
|
|
| 187 |
cortiq ltx-decode --model $M --latent latent.safetensors --out-dir frames/
|
| 188 |
```
|
| 189 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 190 |
## What is inside
|
| 191 |
|
| 192 |
| component | weights | in the file | codec |
|
|
@@ -282,9 +313,11 @@ machine, no Python, no GPU. `--quant` picks the codec for the big planes
|
|
| 282 |
48-block audio-video transformer, the sampler, the latent upscaler, the
|
| 283 |
video VAE, the audio VAE with its BigVGAN vocoder and bandwidth extension,
|
| 284 |
and the duration head.
|
| 285 |
-
*
|
| 286 |
-
|
| 287 |
-
|
|
|
|
|
|
|
| 288 |
|
| 289 |
Everything above is honest about what it is: a 4-bit repack. The reference at
|
| 290 |
bf16 is the quality ceiling, and the codec's cost was measured stage by stage
|
|
|
|
| 187 |
cortiq ltx-decode --model $M --latent latent.safetensors --out-dir frames/
|
| 188 |
```
|
| 189 |
|
| 190 |
+
## Every mode the model has
|
| 191 |
+
|
| 192 |
+
LTX-2.5 is one network with two streams, and a *mode* is simply which parts
|
| 193 |
+
you hold fixed. Conditioning is encoded into the model's own latent space and
|
| 194 |
+
frozen there β the sampler gets a timestep of zero for those tokens and
|
| 195 |
+
leaves them alone β so all of this is one command with different inputs.
|
| 196 |
+
|
| 197 |
+
| mode | how |
|
| 198 |
+
|---|---|
|
| 199 |
+
| text β video + sound | `--prompt "β¦" --out-audio track.wav` |
|
| 200 |
+
| text β video | the same, without `--out-audio` |
|
| 201 |
+
| text β sound | the same, keeping only the wav |
|
| 202 |
+
| image + text β video (+ sound) | `--image still.ppm` |
|
| 203 |
+
| video β video | `--video frames/` |
|
| 204 |
+
| video β sound | `--video frames/ --video-to-audio` |
|
| 205 |
+
| sound β video | `--audio-in track.wav` |
|
| 206 |
+
| sound β sound | `--audio-in track.wav --out-audio out.wav` |
|
| 207 |
+
| image + sound β video | `--image still.ppm --audio-in track.wav` |
|
| 208 |
+
|
| 209 |
+
```bash
|
| 210 |
+
# a still into a shot, with its soundtrack
|
| 211 |
+
ffmpeg -i photo.jpg -vf scale=384:256 -pix_fmt rgb24 still.ppm
|
| 212 |
+
cortiq ltx-video --model $M --image still.ppm \
|
| 213 |
+
--prompt "the camera pushes in slowly as the light shifts" \
|
| 214 |
+
--height 256 --width 384 --frames 49 --out-dir out/ --out-audio out.wav
|
| 215 |
+
```
|
| 216 |
+
|
| 217 |
+
Image conditioning runs through the video VAE's **encoder**, audio
|
| 218 |
+
conditioning through the audio VAE's and a log-mel front end β both in this
|
| 219 |
+
same file, along with everything else.
|
| 220 |
+
|
| 221 |
## What is inside
|
| 222 |
|
| 223 |
| component | weights | in the file | codec |
|
|
|
|
| 313 |
48-block audio-video transformer, the sampler, the latent upscaler, the
|
| 314 |
video VAE, the audio VAE with its BigVGAN vocoder and bandwidth extension,
|
| 315 |
and the duration head.
|
| 316 |
+
* β
**Every conditioning mode**: image-to-video, video-to-video,
|
| 317 |
+
video-to-audio, audio-to-video, audio-to-audio and the image+audio pairs β
|
| 318 |
+
all from this one file, because both VAE encoders are in it.
|
| 319 |
+
* β³ **LoRAs and the IC-LoRA upscaler** are separate releases and not packed
|
| 320 |
+
here yet.
|
| 321 |
|
| 322 |
Everything above is honest about what it is: a 4-bit repack. The reference at
|
| 323 |
bf16 is the quality ceiling, and the codec's cost was measured stage by stage
|