Text-to-Video
cortiq
Rust
cmf
video
image-to-video
image-text-to-video
video-to-video
video-to-audio
audio-to-video
text-to-audio
audio-to-audio
any-to-any
text-to-audio-video
ltx-video
ltx-2.5
4-bit precision
Instructions to use infosave/LTX-2.5-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/LTX-2.5-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/LTX-2.5-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -6,83 +6,199 @@ license_link: https://github.com/Lightricks/LTX-2/blob/main/LICENSE.md
|
|
| 6 |
base_model:
|
| 7 |
- Lightricks/LTX-2.5
|
| 8 |
base_model_relation: quantized
|
| 9 |
-
pipeline_tag:
|
| 10 |
tags:
|
| 11 |
- cmf
|
| 12 |
- cortiq
|
| 13 |
- video
|
| 14 |
-
-
|
| 15 |
- ltx-video
|
| 16 |
- ltx-2.5
|
|
|
|
| 17 |
- 4-bit
|
| 18 |
---
|
| 19 |
|
| 20 |
-
# LTX-2.5 — the whole pipeline in one
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
|
| 22 |
[LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5) renders video **and its
|
| 23 |
soundtrack** from one prompt: a 21 B audio-video diffusion transformer that
|
| 24 |
denoises picture and sound in the same 48 blocks, a Gemma-4 12 B prompt
|
| 25 |
encoder, a 3-D video VAE, an audio VAE, two latent upscalers and a duration
|
| 26 |
head. The reference checkout is **71.35 GB across six safetensors** plus a
|
| 27 |
-
|
| 28 |
|
| 29 |
-
Here it is **one
|
| 30 |
-
|
| 31 |
-
inside it — packed by `cortiq`, a Rust binary with no ML framework
|
| 32 |
-
underneath.
|
| 33 |
|
| 34 |
| | reference | this file |
|
| 35 |
|---|---|---|
|
| 36 |
| files | 6 safetensors + configs + tokenizer | **1** |
|
| 37 |
-
| bytes | 71.35 GB | **
|
| 38 |
| weights | 35.65 B | 35.65 B — all of them |
|
| 39 |
| loader | diffusers / ComfyUI + PyTorch | `mmap` |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
```
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 48 |
```
|
| 49 |
|
| 50 |
## What is inside
|
| 51 |
|
| 52 |
| component | weights | in the file | codec |
|
| 53 |
|---|---|---|---|
|
| 54 |
-
| `dit.*` — LTX-2.5 22B audio-video DiT (distilled) | 21.004 B | 10.
|
| 55 |
-
| `te.*` — Gemma-4 12B prompt encoder
|
| 56 |
| `vvae.*` — video VAE (3-D conv, encoder + decoder) | 0.726 B | 1.35 GiB | f16 |
|
| 57 |
| `avae.*` — audio VAE | 0.182 B | 0.34 GiB | f16 |
|
| 58 |
| `ups.*` — latent spatial upscaler ×2 | 0.498 B | 0.93 GiB | f16 |
|
| 59 |
| `upt.*` — latent temporal upscaler ×2 | 0.131 B | 0.24 GiB | f16 |
|
| 60 |
| `dhead.*` — duration head | 1.9 M | 3.8 MB | f16 |
|
| 61 |
-
|
|
| 62 |
| tokenizer (`tokenizer_json`, 32 MB) | — | VOCAB section | raw |
|
| 63 |
|
| 64 |
-
|
| 65 |
-
transformer is the same shape and packs the same way (`--dit …dev…`); the
|
| 66 |
-
LoRAs, the int8/nvfp4 builds and the IC-LoRA upscaler are not in this file.
|
| 67 |
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
|
| 72 |
* **2-D planes of at least 2²⁰ weights → q4tp**, 4.16 bits with a per-row
|
| 73 |
-
scale ladder.
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
* **
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
|
|
|
|
|
|
|
|
|
| 86 |
|
| 87 |
## The architecture it carries
|
| 88 |
|
|
@@ -90,67 +206,65 @@ Four bits is not applied by fiat. `cortiq ltx-pack` decides per tensor:
|
|
| 90 |
file as `ltx.config_json`):
|
| 91 |
|
| 92 |
* **48 blocks**, video stream 4096 (32 heads × 128), **audio stream 2048**
|
| 93 |
-
(32 × 64), joint audio↔video cross-attention with adaLN-gated fusion
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
-
`
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
*
|
| 103 |
-
|
| 104 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 105 |
|
| 106 |
## Packing it yourself
|
| 107 |
|
| 108 |
-
Three passes,
|
| 109 |
-
stand this was packed on
|
| 110 |
|
| 111 |
```bash
|
| 112 |
-
|
| 113 |
-
cortiq ltx-pack --out p1.cmf
|
| 114 |
-
--dit ltx-2.5-22b-distilled-transformer-bf16.safetensors
|
| 115 |
-
|
| 116 |
-
# 2 — the Gemma-4 12B encoder on top (--in carries pass 1 byte for byte)
|
| 117 |
-
cortiq ltx-pack --out p2.cmf --in p1.cmf \
|
| 118 |
-
--te gemma4-12b-with-proj-ltx-2.5-bf16.safetensors
|
| 119 |
-
|
| 120 |
-
# 3 — the VAEs, the upscalers and the duration head
|
| 121 |
cortiq ltx-pack --out ltx25-q4tp.cmf --in p2.cmf \
|
| 122 |
--video-vae ltx-2.5-video-vae-conv-bf16.safetensors \
|
| 123 |
--audio-vae ltx-2.5-audio-vae-bf16.safetensors \
|
| 124 |
--spatial-upscaler ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
|
| 125 |
--temporal-upscaler ltx-2.5-latent-temporal-upscaler-x2-bf16-1.0.safetensors \
|
| 126 |
--duration-head ltx-2.5-duration-head-bf16.safetensors
|
| 127 |
-
|
| 128 |
cortiq verify ltx25-q4tp.cmf && cortiq info ltx25-q4tp.cmf
|
| 129 |
```
|
| 130 |
|
| 131 |
-
Measured on a 32-core pod:
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
| 135 |
-
## Status
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
|
| 139 |
-
|
| 140 |
-
|
| 141 |
-
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
|
| 145 |
-
|
| 146 |
-
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
|
| 150 |
-
|
| 151 |
-
Until that lands this is a **format artifact**: the pipeline in one
|
| 152 |
-
verifiable container, 3.4× smaller, for anyone who wants to read LTX-2.5's
|
| 153 |
-
weights without a PyTorch stack — or to watch the port land against it.
|
| 154 |
|
| 155 |
## Provenance
|
| 156 |
|
|
|
|
| 6 |
base_model:
|
| 7 |
- Lightricks/LTX-2.5
|
| 8 |
base_model_relation: quantized
|
| 9 |
+
pipeline_tag: text-to-video
|
| 10 |
tags:
|
| 11 |
- cmf
|
| 12 |
- cortiq
|
| 13 |
- video
|
| 14 |
+
- text-to-video
|
| 15 |
- ltx-video
|
| 16 |
- ltx-2.5
|
| 17 |
+
- rust
|
| 18 |
- 4-bit
|
| 19 |
---
|
| 20 |
|
| 21 |
+
# LTX-2.5 — the whole pipeline in one 22 GB file, rendered by Rust
|
| 22 |
+
|
| 23 |
+
<p align="center">
|
| 24 |
+
<img src="assets/corgi.gif" width="49%" alt="A corgi in a chef hat flips a pancake in a sunlit kitchen">
|
| 25 |
+
<img src="assets/neon.gif" width="49%" alt="Neon rain on a Tokyo side street at night">
|
| 26 |
+
</p>
|
| 27 |
+
<p align="center">
|
| 28 |
+
<img src="assets/whale.gif" width="49%" alt="A humpback whale glides through a shaft of sunlight">
|
| 29 |
+
<img src="assets/glass.gif" width="49%" alt="Molten glass blown into a bulb over an orange furnace">
|
| 30 |
+
</p>
|
| 31 |
+
|
| 32 |
+
**Every frame above was produced by `cortiq`** — a single Rust binary with no
|
| 33 |
+
PyTorch, no diffusers, no CUDA toolkit and no Python anywhere in the process —
|
| 34 |
+
reading one memory-mapped [CMF](https://github.com/infosave2007/cmf) file.
|
| 35 |
|
| 36 |
[LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5) renders video **and its
|
| 37 |
soundtrack** from one prompt: a 21 B audio-video diffusion transformer that
|
| 38 |
denoises picture and sound in the same 48 blocks, a Gemma-4 12 B prompt
|
| 39 |
encoder, a 3-D video VAE, an audio VAE, two latent upscalers and a duration
|
| 40 |
head. The reference checkout is **71.35 GB across six safetensors** plus a
|
| 41 |
+
PyTorch stack.
|
| 42 |
|
| 43 |
+
Here it is **one file of 22.07 GB** — every component, the Gemma-4 tokenizer
|
| 44 |
+
and every config inside it.
|
|
|
|
|
|
|
| 45 |
|
| 46 |
| | reference | this file |
|
| 47 |
|---|---|---|
|
| 48 |
| files | 6 safetensors + configs + tokenizer | **1** |
|
| 49 |
+
| bytes | 71.35 GB | **22.07 GB** (3.2× smaller) |
|
| 50 |
| weights | 35.65 B | 35.65 B — all of them |
|
| 51 |
| loader | diffusers / ComfyUI + PyTorch | `mmap` |
|
| 52 |
+
| renderer | Python | **one Rust binary** |
|
| 53 |
+
|
| 54 |
+
## Quick start
|
| 55 |
+
|
| 56 |
+
```bash
|
| 57 |
+
# 1 — the runtime (Rust 1.85+; nothing else)
|
| 58 |
+
cargo install cortiq-cli
|
| 59 |
+
|
| 60 |
+
# 2 — the model
|
| 61 |
+
hf download infosave/LTX-2.5-cmf ltx25-q4tp.cmf --local-dir .
|
| 62 |
+
cortiq verify ltx25-q4tp.cmf # every tensor is hashed in the directory
|
| 63 |
+
|
| 64 |
+
# 3 — a video
|
| 65 |
+
cortiq ltx-video --model ltx25-q4tp.cmf \
|
| 66 |
+
--prompt "A corgi in a chef hat flips a pancake in a sunlit kitchen. \
|
| 67 |
+
Warm morning light, static camera." \
|
| 68 |
+
--height 256 --width 384 --frames 49 --fps 24 --seed 42 \
|
| 69 |
+
--out corgi.y4m
|
| 70 |
+
|
| 71 |
+
ffmpeg -i corgi.y4m -pix_fmt yuv420p corgi.mp4
|
| 72 |
+
```
|
| 73 |
+
|
| 74 |
+
That is the whole thing: prompt in, frames out, one process, one file. The
|
| 75 |
+
GPU is found at run time — Vulkan on Linux and Windows, Metal on Apple
|
| 76 |
+
silicon — and everything falls back to the CPU when there is none.
|
| 77 |
+
|
| 78 |
+
> **Keep the file on local storage.** It is memory-mapped, so every weight is
|
| 79 |
+
> a page fault. On a network filesystem (NFS, MooseFS, a rented pod's
|
| 80 |
+
> `/workspace` volume) that is a network round trip per weight and the
|
| 81 |
+
> process will sit at 1 % CPU looking hung. Copy it to a local disk — or to
|
| 82 |
+
> `/dev/shm` if you have the RAM.
|
| 83 |
+
|
| 84 |
+
### The examples above, exactly
|
| 85 |
+
|
| 86 |
+
```bash
|
| 87 |
+
M=ltx25-q4tp.cmf
|
| 88 |
+
cortiq ltx-video --model $M --seed 42 --height 256 --width 384 --frames 49 \
|
| 89 |
+
--out-dir corgi/ --prompt \
|
| 90 |
+
"A corgi in a chef hat flips a pancake in a sunlit kitchen. Warm morning light, static camera."
|
| 91 |
+
|
| 92 |
+
cortiq ltx-video --model $M --seed 7 --height 256 --width 384 --frames 49 \
|
| 93 |
+
--out-dir neon/ --prompt \
|
| 94 |
+
"Neon rain on a Tokyo side street at night, a lone figure with a translucent umbrella \
|
| 95 |
+
walks past ramen shop signs, reflections rippling in the puddles, slow dolly."
|
| 96 |
+
|
| 97 |
+
cortiq ltx-video --model $M --seed 11 --height 256 --width 384 --frames 49 \
|
| 98 |
+
--out-dir whale/ --prompt \
|
| 99 |
+
"A humpback whale glides through a shaft of sunlight in deep blue water, plankton \
|
| 100 |
+
drifting like dust, the camera rises with it toward the surface."
|
| 101 |
+
|
| 102 |
+
cortiq ltx-video --model $M --seed 23 --height 256 --width 384 --frames 49 \
|
| 103 |
+
--out-dir glass/ --prompt \
|
| 104 |
+
"Molten glass is blown into a bulb over an orange furnace, the glowing gather \
|
| 105 |
+
stretching and rotating, sparks drifting in the dark workshop."
|
| 106 |
+
|
| 107 |
+
# frames → mp4 → gif
|
| 108 |
+
ffmpeg -framerate 24 -i corgi/frame_%04d.ppm -pix_fmt yuv420p -crf 18 corgi.mp4
|
| 109 |
+
ffmpeg -i corgi.mp4 -vf "fps=12,scale=384:-1:flags=lanczos,split[s0][s1];\
|
| 110 |
+
[s0]palettegen[p];[s1][p]paletteuse" corgi.gif
|
| 111 |
+
```
|
| 112 |
+
|
| 113 |
+
`--out-dir` writes `frame_0000.ppm …`; `--out file.y4m` writes one
|
| 114 |
+
[YUV4MPEG2](https://wiki.multimedia.cx/index.php/YUV4MPEG2) stream instead,
|
| 115 |
+
which every tool reads — so the renderer needs no video encoder of its own.
|
| 116 |
+
|
| 117 |
+
### Higher resolution
|
| 118 |
|
| 119 |
+
```bash
|
| 120 |
+
cortiq ltx-video --model $M --two-stage \
|
| 121 |
+
--height 512 --width 768 --frames 49 --seed 42 \
|
| 122 |
+
--prompt "…" --out hq.y4m
|
| 123 |
```
|
| 124 |
+
|
| 125 |
+
`--two-stage` samples the way the distilled model was trained: eight
|
| 126 |
+
ancestral Euler steps at half resolution, the learned latent upscaler ×2,
|
| 127 |
+
then three deterministic steps that refine what the upscale invented.
|
| 128 |
+
|
| 129 |
+
Resolution must be a multiple of 32 (the video VAE's spatial stride) and the
|
| 130 |
+
frame count `8k + 1` (its temporal stride plus the standalone first frame).
|
| 131 |
+
|
| 132 |
+
### Measured
|
| 133 |
+
|
| 134 |
+
RTX 5090, `/dev/shm`, 49 frames at 24 fps:
|
| 135 |
+
|
| 136 |
+
| stage | 384×256 | 768×512 (`--two-stage`) |
|
| 137 |
+
|---|---|---|
|
| 138 |
+
| prompt encode (Gemma-4 12 B + connectors) | 28 s | 28 s |
|
| 139 |
+
| denoise | 8 × 30 s | 8 × 30 s + 3 × 120 s |
|
| 140 |
+
| latent upscale | — | 25 s |
|
| 141 |
+
| video VAE | 50 s | 200 s |
|
| 142 |
+
|
| 143 |
+
Nothing here is tuned yet — the transformer runs one 48-block forward per
|
| 144 |
+
step against `mmap`ped 4-bit weights, and the VAE is a straight
|
| 145 |
+
im2col + GEMM.
|
| 146 |
+
|
| 147 |
+
## The stages, separately
|
| 148 |
+
|
| 149 |
+
Each stage is its own command. That is how the port was gated: every one of
|
| 150 |
+
them can be run against a dump of the reference implementation's own
|
| 151 |
+
activations and will report the first place it diverges.
|
| 152 |
+
|
| 153 |
+
```bash
|
| 154 |
+
# prompt → the two context tensors the transformer cross-attends to
|
| 155 |
+
cortiq ltx-encode --model $M --prompt "…" --out context.safetensors
|
| 156 |
+
|
| 157 |
+
# context → latent → frames
|
| 158 |
+
cortiq ltx-render --model $M --context context.safetensors \
|
| 159 |
+
--height 256 --width 384 --frames 49 --out-latent latent.safetensors --out out.y4m
|
| 160 |
+
|
| 161 |
+
# latent → frames, the 3-D convolutional decoder alone
|
| 162 |
+
cortiq ltx-decode --model $M --latent latent.safetensors --out-dir frames/
|
| 163 |
```
|
| 164 |
|
| 165 |
## What is inside
|
| 166 |
|
| 167 |
| component | weights | in the file | codec |
|
| 168 |
|---|---|---|---|
|
| 169 |
+
| `dit.*` — LTX-2.5 22B audio-video DiT (distilled) | 21.004 B | 10.84 GiB | q4tp + exact adaLN |
|
| 170 |
+
| `te.*` — Gemma-4 12B prompt encoder, aggregates, vision tower | 13.116 B | 6.87 GiB | q4tp + q8 embeddings |
|
| 171 |
| `vvae.*` — video VAE (3-D conv, encoder + decoder) | 0.726 B | 1.35 GiB | f16 |
|
| 172 |
| `avae.*` — audio VAE | 0.182 B | 0.34 GiB | f16 |
|
| 173 |
| `ups.*` — latent spatial upscaler ×2 | 0.498 B | 0.93 GiB | f16 |
|
| 174 |
| `upt.*` — latent temporal upscaler ×2 | 0.131 B | 0.24 GiB | f16 |
|
| 175 |
| `dhead.*` — duration head | 1.9 M | 3.8 MB | f16 |
|
| 176 |
+
| configs, HF assets | — | 55 KB | raw |
|
| 177 |
| tokenizer (`tokenizer_json`, 32 MB) | — | VOCAB section | raw |
|
| 178 |
|
| 179 |
+
### What the codec does per tensor, and why
|
|
|
|
|
|
|
| 180 |
|
| 181 |
+
Four bits is not applied by fiat. `cortiq ltx-pack` decides per tensor, and
|
| 182 |
+
two of those decisions were made by measuring against the reference rather
|
| 183 |
+
than by taste:
|
| 184 |
|
| 185 |
* **2-D planes of at least 2²⁰ weights → q4tp**, 4.16 bits with a per-row
|
| 186 |
+
scale ladder. Every projection in the transformer and in the encoder.
|
| 187 |
+
* **The adaLN-single stacks stay exact.** Their output is not a residual —
|
| 188 |
+
it is the scale and the shift applied to every token in every block, so an
|
| 189 |
+
error there is multiplied into the whole stream instead of being averaged
|
| 190 |
+
away by anything downstream. Quantized, they put 3.6·10⁻² of relative
|
| 191 |
+
error into the very first normalization of block 0; exact, 5.9·10⁻³.
|
| 192 |
+
0.56 GB.
|
| 193 |
+
* **The token table stays 8-bit.** It *is* the residual stream at layer zero
|
| 194 |
+
and it carries through forty-eight residual additions. q4tp put 11 % into
|
| 195 |
+
every hidden state the prompt encoder produced; q8 puts 0.5 % there, for
|
| 196 |
+
0.5 GB.
|
| 197 |
+
* **The adaLN tables, the connector's learnable registers and the VAE's
|
| 198 |
+
`per_channel_statistics` stay exact** — 19 MB in total, read once a step,
|
| 199 |
+
modulating everything.
|
| 200 |
+
* **Convolutions stay f16.** Both VAEs and both upscalers are convolutional,
|
| 201 |
+
and the decoder is what the eye actually sees.
|
| 202 |
|
| 203 |
## The architecture it carries
|
| 204 |
|
|
|
|
| 206 |
file as `ltx.config_json`):
|
| 207 |
|
| 208 |
* **48 blocks**, video stream 4096 (32 heads × 128), **audio stream 2048**
|
| 209 |
+
(32 × 64), joint audio↔video cross-attention with adaLN-gated fusion that
|
| 210 |
+
reads the *pre-fusion* state of both streams, so the order the two
|
| 211 |
+
directions run in cannot bias the result.
|
| 212 |
+
* Per block: self-attention, cross-attention to the prompt with its own
|
| 213 |
+
adaLN pair on the query *and* on the prompt's keys and values, **RMS
|
| 214 |
+
q/k-norm across the whole inner dimension**, gated attention
|
| 215 |
+
(`2·sigmoid` per head), and a gelu-approximate feed-forward — all
|
| 216 |
+
modulated from per-block `[9, 4096]` / `[9, 2048]` tables.
|
| 217 |
+
* **Split 3-D RoPE** over (seconds, pixel row, pixel column) evaluated at the
|
| 218 |
+
*middle* of each patch's bounds, θ = 10000, with the causal correction that
|
| 219 |
+
gives the first latent frame one pixel frame where every later one gets
|
| 220 |
+
eight. The audio stream shares the time axis in seconds, which is what lets
|
| 221 |
+
the two cross-attend positionally.
|
| 222 |
+
* **The prompt encoder** is Gemma-4 12 B — forty sliding-window layers at head
|
| 223 |
+
256 and eight full-attention layers at head 512 whose value projection *is*
|
| 224 |
+
the key projection — and the features are not its last hidden state: all
|
| 225 |
+
**forty-nine** layer outputs are RMS-normalized per token per layer,
|
| 226 |
+
concatenated to 188160 numbers and projected once to 4096 (video) and once
|
| 227 |
+
to 2048 (audio).
|
| 228 |
+
* **Embeddings connectors** — 8 gated-attention blocks each for video and
|
| 229 |
+
audio, with **128 learnable registers** that replace every padded position,
|
| 230 |
+
which is why the transformer needs no prompt mask at all.
|
| 231 |
|
| 232 |
## Packing it yourself
|
| 233 |
|
| 234 |
+
Three passes, each one able to delete its source before the next lands — the
|
| 235 |
+
stand this was packed on had a 50 GB disk quota and the sources are 71 GB:
|
| 236 |
|
| 237 |
```bash
|
| 238 |
+
cortiq ltx-pack --out p1.cmf --dit ltx-2.5-22b-distilled-transformer-bf16.safetensors
|
| 239 |
+
cortiq ltx-pack --out p2.cmf --in p1.cmf --te gemma4-12b-with-proj-ltx-2.5-bf16.safetensors
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 240 |
cortiq ltx-pack --out ltx25-q4tp.cmf --in p2.cmf \
|
| 241 |
--video-vae ltx-2.5-video-vae-conv-bf16.safetensors \
|
| 242 |
--audio-vae ltx-2.5-audio-vae-bf16.safetensors \
|
| 243 |
--spatial-upscaler ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
|
| 244 |
--temporal-upscaler ltx-2.5-latent-temporal-upscaler-x2-bf16-1.0.safetensors \
|
| 245 |
--duration-head ltx-2.5-duration-head-bf16.safetensors
|
|
|
|
| 246 |
cortiq verify ltx25-q4tp.cmf && cortiq info ltx25-q4tp.cmf
|
| 247 |
```
|
| 248 |
|
| 249 |
+
Measured on a 32-core pod: **five minutes** for 71 GB of bf16, single
|
| 250 |
+
machine, no Python, no GPU. `--quant` picks the codec for the big planes
|
| 251 |
+
(`q4tp`, `q8`, `f16`, `f32`), `--vae-quant` the one for convolutions.
|
| 252 |
+
|
| 253 |
+
## Status
|
| 254 |
+
|
| 255 |
+
* ✅ **Text → video runs end to end on the Rust engine**: the Gemma-4 prompt
|
| 256 |
+
encoder, the aggregate projections, the connectors, the 48-block audio-video
|
| 257 |
+
transformer, the sampler, the latent upscaler and the video VAE.
|
| 258 |
+
* ⏳ **Sound is generated but not yet decoded.** The transformer denoises the
|
| 259 |
+
audio latent in the same blocks as the picture, and it is in the output
|
| 260 |
+
latent — the audio VAE and its vocoder are the next port, so the clips here
|
| 261 |
+
are silent.
|
| 262 |
+
* ⏳ **Image and video conditioning, LoRAs, the IC-LoRA upscaler and the
|
| 263 |
+
duration head** are in the file but not yet wired into the CLI.
|
| 264 |
+
|
| 265 |
+
Everything above is honest about what it is: a 4-bit repack. The reference at
|
| 266 |
+
bf16 is the quality ceiling, and the codec's cost was measured stage by stage
|
| 267 |
+
rather than assumed — see the numbers in the codec section.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 268 |
|
| 269 |
## Provenance
|
| 270 |
|