Text-to-Video
cortiq
Rust
cmf
video
image-to-video
image-text-to-video
video-to-video
video-to-audio
audio-to-video
text-to-audio
audio-to-audio
any-to-any
text-to-audio-video
ltx-video
ltx-2.5
4-bit precision
Instructions to use infosave/LTX-2.5-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/LTX-2.5-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/LTX-2.5-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -29,6 +29,11 @@ tags:
|
|
| 29 |
<img src="assets/glass.gif" width="49%" alt="Molten glass blown into a bulb over an orange furnace">
|
| 30 |
</p>
|
| 31 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
**Every frame above was produced by `cortiq`** β a single Rust binary with no
|
| 33 |
PyTorch, no diffusers, no CUDA toolkit and no Python anywhere in the process β
|
| 34 |
reading one memory-mapped [CMF](https://github.com/infosave2007/cmf) file.
|
|
@@ -114,8 +119,26 @@ ffmpeg -i corgi.mp4 -vf "fps=12,scale=384:-1:flags=lanczos,split[s0][s1];\
|
|
| 114 |
[YUV4MPEG2](https://wiki.multimedia.cx/index.php/YUV4MPEG2) stream instead,
|
| 115 |
which every tool reads β so the renderer needs no video encoder of its own.
|
| 116 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 117 |
### Higher resolution
|
| 118 |
|
|
|
|
|
|
|
| 119 |
```bash
|
| 120 |
cortiq ltx-video --model $M --two-stage \
|
| 121 |
--height 512 --width 768 --frames 49 --seed 42 \
|
|
@@ -131,18 +154,20 @@ frame count `8k + 1` (its temporal stride plus the standalone first frame).
|
|
| 131 |
|
| 132 |
### Measured
|
| 133 |
|
| 134 |
-
|
| 135 |
|
| 136 |
-
| stage | 384Γ256 | 768Γ512
|
| 137 |
-
|---|---|---|
|
| 138 |
-
| prompt encode (Gemma-4 12 B + connectors) |
|
| 139 |
-
| denoise | 8 Γ
|
| 140 |
-
| latent upscale | β |
|
| 141 |
-
|
|
|
|
|
|
|
|
| 142 |
|
| 143 |
-
|
| 144 |
-
|
| 145 |
-
|
| 146 |
|
| 147 |
## The stages, separately
|
| 148 |
|
|
@@ -252,15 +277,14 @@ machine, no Python, no GPU. `--quant` picks the codec for the big planes
|
|
| 252 |
|
| 253 |
## Status
|
| 254 |
|
| 255 |
-
* β
**Text β video runs end to end on the Rust engine**: the
|
| 256 |
-
encoder, the aggregate projections, the connectors, the
|
| 257 |
-
transformer, the sampler, the latent upscaler
|
| 258 |
-
|
| 259 |
-
|
| 260 |
-
|
| 261 |
-
are
|
| 262 |
-
|
| 263 |
-
duration head** are in the file but not yet wired into the CLI.
|
| 264 |
|
| 265 |
Everything above is honest about what it is: a 4-bit repack. The reference at
|
| 266 |
bf16 is the quality ceiling, and the codec's cost was measured stage by stage
|
|
|
|
| 29 |
<img src="assets/glass.gif" width="49%" alt="Molten glass blown into a bulb over an orange furnace">
|
| 30 |
</p>
|
| 31 |
|
| 32 |
+
> **The clips above are silent GIFs. The videos are not.** The same 48 blocks
|
| 33 |
+
> denoise the soundtrack alongside the picture β hear it in
|
| 34 |
+
> [`examples/`](./tree/main/examples): six mp4s with audio, their raw 48 kHz
|
| 35 |
+
> stereo wavs, and the exact command that made each one.
|
| 36 |
+
|
| 37 |
**Every frame above was produced by `cortiq`** β a single Rust binary with no
|
| 38 |
PyTorch, no diffusers, no CUDA toolkit and no Python anywhere in the process β
|
| 39 |
reading one memory-mapped [CMF](https://github.com/infosave2007/cmf) file.
|
|
|
|
| 119 |
[YUV4MPEG2](https://wiki.multimedia.cx/index.php/YUV4MPEG2) stream instead,
|
| 120 |
which every tool reads β so the renderer needs no video encoder of its own.
|
| 121 |
|
| 122 |
+
### Sound
|
| 123 |
+
|
| 124 |
+
```bash
|
| 125 |
+
cortiq ltx-video --model $M --prompt "β¦" \
|
| 126 |
+
--height 256 --width 384 --frames 49 --seed 3 \
|
| 127 |
+
--out-dir frames/ --out-audio track.wav
|
| 128 |
+
|
| 129 |
+
ffmpeg -framerate 24 -i frames/frame_%04d.ppm -i track.wav \
|
| 130 |
+
-pix_fmt yuv420p -c:v libx264 -crf 18 -c:a aac -b:a 192k -shortest out.mp4
|
| 131 |
+
```
|
| 132 |
+
|
| 133 |
+
The transformer has been denoising the soundtrack in the same blocks as the
|
| 134 |
+
picture the whole time; `--out-audio` decodes it β the spectrogram VAE, then
|
| 135 |
+
BigVGAN v2, then a bandwidth extender that lifts 16 kHz to 48 kHz stereo.
|
| 136 |
+
Eight seconds of work behind minutes of denoising.
|
| 137 |
+
|
| 138 |
### Higher resolution
|
| 139 |
|
| 140 |
+
<p align="center"><img src="assets/hq-still.png" width="70%" alt="768x512, two-stage"></p>
|
| 141 |
+
|
| 142 |
```bash
|
| 143 |
cortiq ltx-video --model $M --two-stage \
|
| 144 |
--height 512 --width 768 --frames 49 --seed 42 \
|
|
|
|
| 154 |
|
| 155 |
### Measured
|
| 156 |
|
| 157 |
+
49 frames at 24 fps, container on local storage:
|
| 158 |
|
| 159 |
+
| stage | RTX 5090, 384Γ256 | RTX 5090, 768Γ512 `--two-stage` | **M4 MacBook, 24 GB**, 384Γ256 |
|
| 160 |
+
|---|---|---|---|
|
| 161 |
+
| prompt encode (Gemma-4 12 B + connectors) | 26 s | 26 s | 34 s |
|
| 162 |
+
| denoise | 8 Γ 19 s | 8 Γ 19 s + 3 Γ 193 s | 8 Γ 23 s |
|
| 163 |
+
| latent upscale | β | 12 s | β |
|
| 164 |
+
| audio VAE + vocoder | 8 s | 8 s | 9 s |
|
| 165 |
+
| video VAE | 50 s | 580 s | 26 s |
|
| 166 |
+
| **total** | **3 min** | **17 min** | **4 min** |
|
| 167 |
|
| 168 |
+
A 21 B video model, its 12 B prompt encoder and both VAEs, rendering a clip
|
| 169 |
+
with sound on a laptop with 24 GB of unified memory β because nothing is ever
|
| 170 |
+
loaded, only mapped, and the pipeline touches one component at a time.
|
| 171 |
|
| 172 |
## The stages, separately
|
| 173 |
|
|
|
|
| 277 |
|
| 278 |
## Status
|
| 279 |
|
| 280 |
+
* β
**Text β video *and sound* runs end to end on the Rust engine**: the
|
| 281 |
+
Gemma-4 prompt encoder, the aggregate projections, the connectors, the
|
| 282 |
+
48-block audio-video transformer, the sampler, the latent upscaler, the
|
| 283 |
+
video VAE, the audio VAE with its BigVGAN vocoder and bandwidth extension,
|
| 284 |
+
and the duration head.
|
| 285 |
+
* β³ **Image and video conditioning, LoRAs and the IC-LoRA upscaler** are next.
|
| 286 |
+
The weights for conditioning are in this file (the video VAE's encoder half);
|
| 287 |
+
the LoRAs and the IC-LoRA upscaler are separate releases.
|
|
|
|
| 288 |
|
| 289 |
Everything above is honest about what it is: a 4-bit repack. The reference at
|
| 290 |
bf16 is the quality ceiling, and the codec's cost was measured stage by stage
|