Text-to-Video
cortiq
Rust
cmf
video
image-to-video
image-text-to-video
video-to-video
video-to-audio
audio-to-video
text-to-audio
audio-to-audio
any-to-any
text-to-audio-video
ltx-video
ltx-2.5
4-bit precision
Instructions to use infosave/LTX-2.5-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/LTX-2.5-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/LTX-2.5-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -170,7 +170,41 @@ then three deterministic steps that refine what the upscale invented.
|
|
| 170 |
distilled — 8 is it exactly, other counts land on sigmas the model never saw
|
| 171 |
and usually soften the frame. Detail comes from resolution and `--two-stage`.
|
| 172 |
|
| 173 |
-
#
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 174 |
|
| 175 |
49 frames at 24 fps, container on local storage:
|
| 176 |
|
|
|
|
| 170 |
distilled — 8 is it exactly, other counts land on sigmas the model never saw
|
| 171 |
and usually soften the frame. Detail comes from resolution and `--two-stage`.
|
| 172 |
|
| 173 |
+
#
|
| 174 |
+
### LoRA adapters, and multi-subject references
|
| 175 |
+
|
| 176 |
+
```sh
|
| 177 |
+
cortiq ltx-video --model $M --lora adapter.safetensors --lora-strength 0.8 \
|
| 178 |
+
--prompt "…" --out clip.y4m
|
| 179 |
+
```
|
| 180 |
+
|
| 181 |
+
q4tp weights cannot absorb a low-rank update without dequantizing the whole
|
| 182 |
+
DiT, so the branch runs beside them — `y = x·Wᵀ + s·(x·Aᵀ)·Bᵀ`, on every path
|
| 183 |
+
including the fused Metal q/k/v submission. Rank 128 against a 4096×4096
|
| 184 |
+
projection is about 6% more arithmetic and no memory beyond the file — but
|
| 185 |
+
measured on an M4 a 384-token step goes 8.6 s to about 22 s, because that
|
| 186 |
+
arithmetic is host-side f32 beside a device-side 4-bit GEMM and does not
|
| 187 |
+
overlap it. The flops are cheap; the placement is what costs.
|
| 188 |
+
|
| 189 |
+
An adapter that also carries a `reference_slot_embedding` takes reference
|
| 190 |
+
stills, which is how the multi-subject adapters work:
|
| 191 |
+
|
| 192 |
+
```sh
|
| 193 |
+
cortiq ltx-video --model $M --lora msr.safetensors \
|
| 194 |
+
--ref a.ppm --ref b.ppm --ref c.ppm \
|
| 195 |
+
--prompt "Image 1: … Image 2: … Image 3: …" --out clip.y4m
|
| 196 |
+
```
|
| 197 |
+
|
| 198 |
+
Each still is held for 25 or 33 pixel frames (`--ref-frames`, whichever the
|
| 199 |
+
adapter was trained on), encoded by the same video VAE the render uses, given
|
| 200 |
+
its slot's learned per-channel bias on the latent, and placed at a negative
|
| 201 |
+
frame offset — slot 1 furthest back. Those tokens ride in the same sequence,
|
| 202 |
+
frozen, and are cropped off the result.
|
| 203 |
+
|
| 204 |
+
They cost sequence length: three references at 384×256 add 1152 tokens beside
|
| 205 |
+
384 of clip. The stills must already be the render's size.
|
| 206 |
+
|
| 207 |
+
## Measured
|
| 208 |
|
| 209 |
49 frames at 24 fps, container on local storage:
|
| 210 |
|