Text-to-Video
cortiq
Rust
cmf
video
image-to-video
image-text-to-video
video-to-video
video-to-audio
audio-to-video
text-to-audio
audio-to-audio
any-to-any
text-to-audio-video
ltx-video
ltx-2.5
4-bit precision
Instructions to use infosave/LTX-2.5-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/LTX-2.5-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/LTX-2.5-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -12,6 +12,15 @@ tags:
|
|
| 12 |
- cortiq
|
| 13 |
- video
|
| 14 |
- text-to-video
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
- ltx-video
|
| 16 |
- ltx-2.5
|
| 17 |
- rust
|
|
@@ -38,6 +47,14 @@ tags:
|
|
| 38 |
PyTorch, no diffusers, no CUDA toolkit and no Python anywhere in the process β
|
| 39 |
reading one memory-mapped [CMF](https://github.com/infosave2007/cmf) file.
|
| 40 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
[LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5) renders video **and its
|
| 42 |
soundtrack** from one prompt: a 21 B audio-video diffusion transformer that
|
| 43 |
denoises picture and sound in the same 48 blocks, a Gemma-4 12 B prompt
|
|
|
|
| 12 |
- cortiq
|
| 13 |
- video
|
| 14 |
- text-to-video
|
| 15 |
+
- image-to-video
|
| 16 |
+
- image-text-to-video
|
| 17 |
+
- video-to-video
|
| 18 |
+
- video-to-audio
|
| 19 |
+
- audio-to-video
|
| 20 |
+
- text-to-audio
|
| 21 |
+
- audio-to-audio
|
| 22 |
+
- any-to-any
|
| 23 |
+
- text-to-audio-video
|
| 24 |
- ltx-video
|
| 25 |
- ltx-2.5
|
| 26 |
- rust
|
|
|
|
| 47 |
PyTorch, no diffusers, no CUDA toolkit and no Python anywhere in the process β
|
| 48 |
reading one memory-mapped [CMF](https://github.com/infosave2007/cmf) file.
|
| 49 |
|
| 50 |
+
**All nine modes run from this one file**: text β video, text β sound,
|
| 51 |
+
text β video + sound, image + text β video, video β video, video β sound,
|
| 52 |
+
sound β video, sound β sound, and image + sound β video. Both VAE encoders
|
| 53 |
+
are packed alongside the decoders, so conditioning needs nothing else β see
|
| 54 |
+
[the table below](#every-mode-the-model-has). The `pipeline_tag` says
|
| 55 |
+
`text-to-video` because that is the one tag Hugging Face lets a model carry
|
| 56 |
+
and it is where people look for this; the rest are in `tags`.
|
| 57 |
+
|
| 58 |
[LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5) renders video **and its
|
| 59 |
soundtrack** from one prompt: a 21 B audio-video diffusion transformer that
|
| 60 |
denoises picture and sound in the same 48 blocks, a Gemma-4 12 B prompt
|