File size: 20,334 Bytes
05fb95e
 
 
 
 
 
 
 
80bd061
05fb95e
 
 
 
80bd061
405d5e7
 
 
 
 
 
 
 
 
05fb95e
 
80bd061
05fb95e
 
 
80bd061
 
 
 
 
 
 
 
 
 
 
e7e3021
 
 
 
 
80bd061
 
 
05fb95e
405d5e7
 
 
 
 
 
 
 
05fb95e
 
 
 
 
80bd061
05fb95e
80bd061
 
05fb95e
 
 
 
80bd061
05fb95e
 
80bd061
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e7e3021
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
80bd061
05fb95e
e7e3021
 
80bd061
 
 
 
05fb95e
80bd061
 
 
 
 
db5fb4b
 
 
80bd061
aabea2e
 
 
 
 
 
 
 
 
 
 
7571997
 
 
 
 
 
 
 
 
 
 
 
aabea2e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
80bd061
e7e3021
80bd061
e7e3021
 
6d66336
8cc48df
e7e3021
8cc48df
 
 
d509cca
8cc48df
d509cca
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
80bd061
8cc48df
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e7e3021
 
6d66336
 
 
80bd061
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
05fb95e
 
7f4b1b5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
05fb95e
 
 
 
80bd061
 
05fb95e
 
 
 
 
80bd061
05fb95e
 
80bd061
05fb95e
80bd061
 
 
05fb95e
 
80bd061
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
05fb95e
 
 
 
 
 
 
80bd061
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
05fb95e
 
 
80bd061
 
05fb95e
 
80bd061
 
05fb95e
 
 
 
 
 
 
 
 
80bd061
 
 
 
 
 
e7e3021
 
 
 
 
7f4b1b5
 
 
 
 
80bd061
 
 
 
05fb95e
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
---
library_name: cortiq
license: other
license_name: ltx-2-community-license-agreement
license_link: https://github.com/Lightricks/LTX-2/blob/main/LICENSE.md
base_model:
- Lightricks/LTX-2.5
base_model_relation: quantized
pipeline_tag: text-to-video
tags:
- cmf
- cortiq
- video
- text-to-video
- image-to-video
- image-text-to-video
- video-to-video
- video-to-audio
- audio-to-video
- text-to-audio
- audio-to-audio
- any-to-any
- text-to-audio-video
- ltx-video
- ltx-2.5
- rust
- 4-bit
---

# LTX-2.5 β€” the whole pipeline in one 22 GB file, rendered by Rust

<p align="center">
  <img src="assets/corgi.gif" width="49%" alt="A corgi in a chef hat flips a pancake in a sunlit kitchen">
  <img src="assets/neon.gif" width="49%" alt="Neon rain on a Tokyo side street at night">
</p>
<p align="center">
  <img src="assets/whale.gif" width="49%" alt="A humpback whale glides through a shaft of sunlight">
  <img src="assets/glass.gif" width="49%" alt="Molten glass blown into a bulb over an orange furnace">
</p>

> **The clips above are silent GIFs. The videos are not.** The same 48 blocks
> denoise the soundtrack alongside the picture β€” hear it in
> [`examples/`](./tree/main/examples): six mp4s with audio, their raw 48 kHz
> stereo wavs, and the exact command that made each one.

**Every frame above was produced by `cortiq`** β€” a single Rust binary with no
PyTorch, no diffusers, no CUDA toolkit and no Python anywhere in the process β€”
reading one memory-mapped [CMF](https://github.com/infosave2007/cmf) file.

**All nine modes run from this one file**: text β†’ video, text β†’ sound,
text β†’ video + sound, image + text β†’ video, video β†’ video, video β†’ sound,
sound β†’ video, sound β†’ sound, and image + sound β†’ video. Both VAE encoders
are packed alongside the decoders, so conditioning needs nothing else β€” see
[the table below](#every-mode-the-model-has). The `pipeline_tag` says
`text-to-video` because that is the one tag Hugging Face lets a model carry
and it is where people look for this; the rest are in `tags`.

[LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5) renders video **and its
soundtrack** from one prompt: a 21 B audio-video diffusion transformer that
denoises picture and sound in the same 48 blocks, a Gemma-4 12 B prompt
encoder, a 3-D video VAE, an audio VAE, two latent upscalers and a duration
head. The reference checkout is **71.35 GB across six safetensors** plus a
PyTorch stack.

Here it is **one file of 22.07 GB** β€” every component, the Gemma-4 tokenizer
and every config inside it.

| | reference | this file |
|---|---|---|
| files | 6 safetensors + configs + tokenizer | **1** |
| bytes | 71.35 GB | **22.07 GB** (3.2Γ— smaller) |
| weights | 35.65 B | 35.65 B β€” all of them |
| loader | diffusers / ComfyUI + PyTorch | `mmap` |
| renderer | Python | **one Rust binary** |

## Quick start

```bash
# 1 β€” the runtime (Rust 1.85+; nothing else)
cargo install cortiq-cli

# 2 β€” the model
hf download infosave/LTX-2.5-cmf ltx25-q4tp.cmf --local-dir .
cortiq verify ltx25-q4tp.cmf        # every tensor is hashed in the directory

# 3 β€” a video
cortiq ltx-video --model ltx25-q4tp.cmf \
  --prompt "A corgi in a chef hat flips a pancake in a sunlit kitchen. \
Warm morning light, static camera." \
  --height 256 --width 384 --frames 49 --fps 24 --seed 42 \
  --out corgi.y4m

ffmpeg -i corgi.y4m -pix_fmt yuv420p corgi.mp4
```

That is the whole thing: prompt in, frames out, one process, one file. The
GPU is found at run time β€” Vulkan on Linux and Windows, Metal on Apple
silicon β€” and everything falls back to the CPU when there is none.

> **Keep the file on local storage.** It is memory-mapped, so every weight is
> a page fault. On a network filesystem (NFS, MooseFS, a rented pod's
> `/workspace` volume) that is a network round trip per weight and the
> process will sit at 1 % CPU looking hung. Copy it to a local disk β€” or to
> `/dev/shm` if you have the RAM.

### The examples above, exactly

```bash
M=ltx25-q4tp.cmf
cortiq ltx-video --model $M --seed 42 --height 256 --width 384 --frames 49 \
  --out-dir corgi/ --prompt \
  "A corgi in a chef hat flips a pancake in a sunlit kitchen. Warm morning light, static camera."

cortiq ltx-video --model $M --seed 7 --height 256 --width 384 --frames 49 \
  --out-dir neon/ --prompt \
  "Neon rain on a Tokyo side street at night, a lone figure with a translucent umbrella \
walks past ramen shop signs, reflections rippling in the puddles, slow dolly."

cortiq ltx-video --model $M --seed 11 --height 256 --width 384 --frames 49 \
  --out-dir whale/ --prompt \
  "A humpback whale glides through a shaft of sunlight in deep blue water, plankton \
drifting like dust, the camera rises with it toward the surface."

cortiq ltx-video --model $M --seed 23 --height 256 --width 384 --frames 49 \
  --out-dir glass/ --prompt \
  "Molten glass is blown into a bulb over an orange furnace, the glowing gather \
stretching and rotating, sparks drifting in the dark workshop."

# frames β†’ mp4 β†’ gif
ffmpeg -framerate 24 -i corgi/frame_%04d.ppm -pix_fmt yuv420p -crf 18 corgi.mp4
ffmpeg -i corgi.mp4 -vf "fps=12,scale=384:-1:flags=lanczos,split[s0][s1];\
[s0]palettegen[p];[s1][p]paletteuse" corgi.gif
```

`--out-dir` writes `frame_0000.ppm …`; `--out file.y4m` writes one
[YUV4MPEG2](https://wiki.multimedia.cx/index.php/YUV4MPEG2) stream instead,
which every tool reads β€” so the renderer needs no video encoder of its own.

### Sound

```bash
cortiq ltx-video --model $M --prompt "…" \
  --height 256 --width 384 --frames 49 --seed 3 \
  --out-dir frames/ --out-audio track.wav

ffmpeg -framerate 24 -i frames/frame_%04d.ppm -i track.wav \
  -pix_fmt yuv420p -c:v libx264 -crf 18 -c:a aac -b:a 192k -shortest out.mp4
```

The transformer has been denoising the soundtrack in the same blocks as the
picture the whole time; `--out-audio` decodes it β€” the spectrogram VAE, then
BigVGAN v2, then a bandwidth extender that lifts 16 kHz to 48 kHz stereo.
Eight seconds of work behind minutes of denoising.

### Higher resolution

<p align="center"><img src="assets/hq-still.png" width="70%" alt="768x512, two-stage"></p>

```bash
cortiq ltx-video --model $M --two-stage \
  --height 512 --width 768 --frames 49 --seed 42 \
  --prompt "…" --out hq.y4m
```

`--two-stage` samples the way the distilled model was trained: eight
ancestral Euler steps at half resolution, the learned latent upscaler Γ—2,
then three deterministic steps that refine what the upscale invented.

`--steps N` / `--steps2 N` resample that schedule. The shipped ladder is
distilled β€” 8 is it exactly, other counts land on sigmas the model never saw
and usually soften the frame. Detail comes from resolution and `--two-stage`.

#
### LoRA adapters, and multi-subject references

```sh
cortiq ltx-video --model $M --lora adapter.safetensors --lora-strength 0.8 \
  --prompt "…" --out clip.y4m
```

q4tp weights cannot absorb a low-rank update without dequantizing the whole
DiT, so the branch runs beside them β€” `y = xΒ·Wα΅€ + sΒ·(xΒ·Aα΅€)Β·Bα΅€`, on every path
including the fused Metal q/k/v submission. Rank 128 against a 4096Γ—4096
projection is about 6% more arithmetic and no memory beyond the file. On
Metal the branch rides in the base GEMM's own submission, reading the
activation already uploaded and accumulating into the output already
written, so it costs no transfer: a 384-token step goes 8.6 s to 10.1 s on
an M4.

Three naming conventions are read as they ship β€” `diffusion_model.…`
(ComfyUI single-file), `base_model.model.…` (PEFT) and the bare module path β€”
in F32, F16 or BF16, with `lora_A`/`lora_B` or `lora_down`/`lora_up` spelling.
`CMF_LORA_PROBE=1` prints each branch's measured contribution and
`CMF_LORA_ROUTE=<r>` switches off the ones below `r`; both are described in
[docs/LORA.md](https://github.com/infosave2007/cmf/blob/master/docs/LORA.md).

An adapter that also carries a `reference_slot_embedding` takes reference
stills, which is how the multi-subject adapters work:

```sh
cortiq ltx-video --model $M --lora msr.safetensors \
  --ref a.ppm --ref b.ppm --ref c.ppm \
  --prompt "Image 1: … Image 2: … Image 3: …" --out clip.y4m
```

Each still is held for 25 or 33 pixel frames (`--ref-frames`, whichever the
adapter was trained on), encoded by the same video VAE the render uses, given
its slot's learned per-channel bias on the latent, and placed at a negative
frame offset β€” slot 1 furthest back. Those tokens ride in the same sequence,
frozen, and are cropped off the result.

They cost sequence length: three references at 384Γ—256 add 1152 tokens beside
384 of clip. The stills must already be the render's size.

## Measured

49 frames at 24 fps, container on local storage:

| stage | RTX 5090, 384Γ—256 | RTX 5090, 768Γ—512 `--two-stage` | **M4 MacBook, 24 GB**, 384Γ—256 |
|---|---|---|---|
| prompt encode (Gemma-4 12 B + connectors) | 26 s | 26 s | 32 s, then cached |
| denoise | 8 Γ— 19 s | 8 Γ— 19 s + 3 Γ— 70 s | 8 Γ— 13 s |
| latent upscale | β€” | 12 s | β€” |
| audio VAE + vocoder | 8 s | 8 s | 8.6 s |
| video VAE | 50 s | 200 s | 24 s |
| **total** | **3 min** | **10 min** | **2.3 min** |

The Mac number is the interesting one, and it took three rounds to get there.

First, keeping the device at all. A 22 GB container does not fit in a single
Metal buffer, so it is mapped as two overlapping windows β€” and the driver
accounts its working set by buffer length, not by unique pages. With both
windows on its books it evicts and re-wires between commits, and a 190 ms
matmul takes 2.7 s. So the windows are built on first use, the prompt encoder
parks the device for its phase (its weights live in the window the denoising
loop never touches), and the denoising loop takes the per-op probe out of the
picture β€” forty-eight identical blocks with the device warm throughout is the
opposite of what a probe that alternates arms can measure.

Then, three things found by profiling rather than guessing, worth another
quarter of every step: the feed-forward's gelu ran in f64 on one thread (half
a billion values a step), the Metal path scanned every activation buffer
scalar-with-a-branch to check it fits in half (2.7 billion floats a step, one
thread), and independent projections each paid their own ~1.3 ms
command-buffer completion. Same arithmetic β€” a render at the same seed before
and after matches at 42.6 dB, which is the last-bit difference between f32 and
f64 amplified by eight sampling steps, not a change in what the model draws.

Third, giving the memory back. The container is 20.5 GiB and the machine has
24 GB, so holding all of it resident leaves nothing for the render and macOS
answers with the compressor: at 384Γ—256Γ—25 the steps used to climb through a
run, 12.4 s to 13.1 s with a 26 s spike, and none of that was arithmetic. But
the pipeline
touches one component at a time and never comes back β€” the prompt encoder is
6.8 GiB read once, the DiT is 10.8 GiB finished before either VAE opens. Both
are handed back to the system the moment they stop being read, and because
they are clean file-backed pages, anything that wants them again just refaults.
At that size the steps now hold 8.5–8.7 s flat and the stage goes 117.5 s β†’
72.8 s; the 49-frame row above is the same change measured at the size the
table quotes.

On device-vs-host: `CMF_MM_AB=1` runs both arms of every eligible q4tp GEMM
back to back on the same data inside one call, which is the only comparison a
laptop that drifts between runs can be trusted to give. Over a whole render the
Metal kernel is **2.01Γ— the host** β€” 2.17Γ— on 4096Γ—16384, 2.21Γ— on 16384Γ—4096,
1.81Γ— on 4096Γ—4096 β€” and the two arms disagree by at most 9e-4 relative.

A 21 B video model, its 12 B prompt encoder and both VAEs, rendering a clip
with sound on a laptop with 24 GB of unified memory β€” because nothing is ever
loaded, only mapped, and the pipeline touches one component at a time. The
encoded prompt is cached, so a second take on the same text starts at the
first denoising step.

## The stages, separately

Each stage is its own command. That is how the port was gated: every one of
them can be run against a dump of the reference implementation's own
activations and will report the first place it diverges.

```bash
# prompt β†’ the two context tensors the transformer cross-attends to
cortiq ltx-encode --model $M --prompt "…" --out context.safetensors

# context β†’ latent β†’ frames
cortiq ltx-render --model $M --context context.safetensors \
  --height 256 --width 384 --frames 49 --out-latent latent.safetensors --out out.y4m

# latent β†’ frames, the 3-D convolutional decoder alone
cortiq ltx-decode --model $M --latent latent.safetensors --out-dir frames/
```

## Every mode the model has

LTX-2.5 is one network with two streams, and a *mode* is simply which parts
you hold fixed. Conditioning is encoded into the model's own latent space and
frozen there β€” the sampler gets a timestep of zero for those tokens and
leaves them alone β€” so all of this is one command with different inputs.

| mode | how |
|---|---|
| text β†’ video + sound | `--prompt "…" --out-audio track.wav` |
| text β†’ video | the same, without `--out-audio` |
| text β†’ sound | the same, keeping only the wav |
| image + text β†’ video (+ sound) | `--image still.ppm` |
| video β†’ video | `--video frames/` |
| video β†’ sound | `--video frames/ --video-to-audio` |
| sound β†’ video | `--audio-in track.wav` |
| sound β†’ sound | `--audio-in track.wav --out-audio out.wav` |
| image + sound β†’ video | `--image still.ppm --audio-in track.wav` |

```bash
# a still into a shot, with its soundtrack
ffmpeg -i photo.jpg -vf scale=384:256 -pix_fmt rgb24 still.ppm
cortiq ltx-video --model $M --image still.ppm \
  --prompt "the camera pushes in slowly as the light shifts" \
  --height 256 --width 384 --frames 49 --out-dir out/ --out-audio out.wav
```

Image conditioning runs through the video VAE's **encoder**, audio
conditioning through the audio VAE's and a log-mel front end β€” both in this
same file, along with everything else.

## What is inside

| component | weights | in the file | codec |
|---|---|---|---|
| `dit.*` β€” LTX-2.5 22B audio-video DiT (distilled) | 21.004 B | 10.84 GiB | q4tp + exact adaLN |
| `te.*` β€” Gemma-4 12B prompt encoder, aggregates, vision tower | 13.116 B | 6.87 GiB | q4tp + q8 embeddings |
| `vvae.*` β€” video VAE (3-D conv, encoder + decoder) | 0.726 B | 1.35 GiB | f16 |
| `avae.*` β€” audio VAE | 0.182 B | 0.34 GiB | f16 |
| `ups.*` β€” latent spatial upscaler Γ—2 | 0.498 B | 0.93 GiB | f16 |
| `upt.*` β€” latent temporal upscaler Γ—2 | 0.131 B | 0.24 GiB | f16 |
| `dhead.*` β€” duration head | 1.9 M | 3.8 MB | f16 |
| configs, HF assets | β€” | 55 KB | raw |
| tokenizer (`tokenizer_json`, 32 MB) | β€” | VOCAB section | raw |

### What the codec does per tensor, and why

Four bits is not applied by fiat. `cortiq ltx-pack` decides per tensor, and
two of those decisions were made by measuring against the reference rather
than by taste:

* **2-D planes of at least 2²⁰ weights β†’ q4tp**, 4.16 bits with a per-row
  scale ladder. Every projection in the transformer and in the encoder.
* **The adaLN-single stacks stay exact.** Their output is not a residual β€”
  it is the scale and the shift applied to every token in every block, so an
  error there is multiplied into the whole stream instead of being averaged
  away by anything downstream. Quantized, they put 3.6·10⁻² of relative
  error into the very first normalization of block 0; exact, 5.9·10⁻³.
  0.56 GB.
* **The token table stays 8-bit.** It *is* the residual stream at layer zero
  and it carries through forty-eight residual additions. q4tp put 11 % into
  every hidden state the prompt encoder produced; q8 puts 0.5 % there, for
  0.5 GB.
* **The adaLN tables, the connector's learnable registers and the VAE's
  `per_channel_statistics` stay exact** β€” 19 MB in total, read once a step,
  modulating everything.
* **Convolutions stay f16.** Both VAEs and both upscalers are convolutional,
  and the decoder is what the eye actually sees.

## The architecture it carries

`AVTransformer3DModel`, from the release's own config (kept verbatim in the
file as `ltx.config_json`):

* **48 blocks**, video stream 4096 (32 heads Γ— 128), **audio stream 2048**
  (32 Γ— 64), joint audio↔video cross-attention with adaLN-gated fusion that
  reads the *pre-fusion* state of both streams, so the order the two
  directions run in cannot bias the result.
* Per block: self-attention, cross-attention to the prompt with its own
  adaLN pair on the query *and* on the prompt's keys and values, **RMS
  q/k-norm across the whole inner dimension**, gated attention
  (`2Β·sigmoid` per head), and a gelu-approximate feed-forward β€” all
  modulated from per-block `[9, 4096]` / `[9, 2048]` tables.
* **Split 3-D RoPE** over (seconds, pixel row, pixel column) evaluated at the
  *middle* of each patch's bounds, ΞΈ = 10000, with the causal correction that
  gives the first latent frame one pixel frame where every later one gets
  eight. The audio stream shares the time axis in seconds, which is what lets
  the two cross-attend positionally.
* **The prompt encoder** is Gemma-4 12 B β€” forty sliding-window layers at head
  256 and eight full-attention layers at head 512 whose value projection *is*
  the key projection β€” and the features are not its last hidden state: all
  **forty-nine** layer outputs are RMS-normalized per token per layer,
  concatenated to 188160 numbers and projected once to 4096 (video) and once
  to 2048 (audio).
* **Embeddings connectors** β€” 8 gated-attention blocks each for video and
  audio, with **128 learnable registers** that replace every padded position,
  which is why the transformer needs no prompt mask at all.

## Packing it yourself

Three passes, each one able to delete its source before the next lands β€” the
stand this was packed on had a 50 GB disk quota and the sources are 71 GB:

```bash
cortiq ltx-pack --out p1.cmf --dit ltx-2.5-22b-distilled-transformer-bf16.safetensors
cortiq ltx-pack --out p2.cmf --in p1.cmf --te gemma4-12b-with-proj-ltx-2.5-bf16.safetensors
cortiq ltx-pack --out ltx25-q4tp.cmf --in p2.cmf \
  --video-vae ltx-2.5-video-vae-conv-bf16.safetensors \
  --audio-vae ltx-2.5-audio-vae-bf16.safetensors \
  --spatial-upscaler  ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
  --temporal-upscaler ltx-2.5-latent-temporal-upscaler-x2-bf16-1.0.safetensors \
  --duration-head     ltx-2.5-duration-head-bf16.safetensors
cortiq verify ltx25-q4tp.cmf && cortiq info ltx25-q4tp.cmf
```

Measured on a 32-core pod: **five minutes** for 71 GB of bf16, single
machine, no Python, no GPU. `--quant` picks the codec for the big planes
(`q4tp`, `q8`, `f16`, `f32`), `--vae-quant` the one for convolutions.

## Status

* βœ… **Text β†’ video *and sound* runs end to end on the Rust engine**: the
  Gemma-4 prompt encoder, the aggregate projections, the connectors, the
  48-block audio-video transformer, the sampler, the latent upscaler, the
  video VAE, the audio VAE with its BigVGAN vocoder and bandwidth extension,
  and the duration head.
* βœ… **Every conditioning mode**: image-to-video, video-to-video,
  video-to-audio, audio-to-video, audio-to-audio and the image+audio pairs β€”
  all from this one file, because both VAE encoders are in it.
* ⏳ **LoRAs and the IC-LoRA upscaler** are separate releases and not packed
  here yet.

Everything above is honest about what it is: a 4-bit repack. The reference at
bf16 is the quality ceiling, and the codec's cost was measured stage by stage
rather than assumed β€” see the numbers in the codec section.

## Provenance

The weights are Lightricks' LTX-2.5 release and remain under the
[LTX-2.x Community License](https://github.com/Lightricks/LTX-2/blob/main/LICENSE.md)
β€” the licence text ships inside the source checkpoints and applies to this
repack unchanged. The CMF container format and the `cortiq` runtime are
Apache-2.0 ([repository](https://github.com/infosave2007/cmf), `PATENTS.md`).

No weight was altered: the pack is a codec change and a container change.
Every tensor's bytes are hashed in the directory, so `cortiq verify` proves
the file is the one that was written.