infosave commited on
Commit
80bd061
·
verified ·
1 Parent(s): b2bbbdd

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +198 -84
README.md CHANGED
@@ -6,83 +6,199 @@ license_link: https://github.com/Lightricks/LTX-2/blob/main/LICENSE.md
6
  base_model:
7
  - Lightricks/LTX-2.5
8
  base_model_relation: quantized
9
- pipeline_tag: image-to-video
10
  tags:
11
  - cmf
12
  - cortiq
13
  - video
14
- - audio
15
  - ltx-video
16
  - ltx-2.5
 
17
  - 4-bit
18
  ---
19
 
20
- # LTX-2.5 — the whole pipeline in one 21 GB file
 
 
 
 
 
 
 
 
 
 
 
 
 
21
 
22
  [LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5) renders video **and its
23
  soundtrack** from one prompt: a 21 B audio-video diffusion transformer that
24
  denoises picture and sound in the same 48 blocks, a Gemma-4 12 B prompt
25
  encoder, a 3-D video VAE, an audio VAE, two latent upscalers and a duration
26
  head. The reference checkout is **71.35 GB across six safetensors** plus a
27
- ComfyUI install.
28
 
29
- Here it is **one memory-mapped [CMF](https://github.com/infosave2007/cmf) file
30
- of 20.99 GB** — every component, the Gemma-4 tokenizer and every config
31
- inside it — packed by `cortiq`, a Rust binary with no ML framework
32
- underneath.
33
 
34
  | | reference | this file |
35
  |---|---|---|
36
  | files | 6 safetensors + configs + tokenizer | **1** |
37
- | bytes | 71.35 GB | **20.99 GB** (3.4× smaller) |
38
  | weights | 35.65 B | 35.65 B — all of them |
39
  | loader | diffusers / ComfyUI + PyTorch | `mmap` |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
40
 
 
 
 
 
41
  ```
42
- Arch: ltx-2.5-av
43
- Layers: 48 (48 full) Hidden: 4096 Heads: 32 (KV: 32)
44
- Vocab: 262144 Params: 35.65B Tensors: 6693
45
- Tokenizer: embedded
46
- ✓ envelope, sections, tensor directory (6693 tensors)
47
- ✓ all tensor hashes match
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
48
  ```
49
 
50
  ## What is inside
51
 
52
  | component | weights | in the file | codec |
53
  |---|---|---|---|
54
- | `dit.*` — LTX-2.5 22B audio-video DiT (distilled) | 21.004 B | 10.28 GiB | q4tp |
55
- | `te.*` — Gemma-4 12B prompt encoder + aggregate projections + vision tower | 13.116 B | 6.37 GiB | q4tp |
56
  | `vvae.*` — video VAE (3-D conv, encoder + decoder) | 0.726 B | 1.35 GiB | f16 |
57
  | `avae.*` — audio VAE | 0.182 B | 0.34 GiB | f16 |
58
  | `ups.*` — latent spatial upscaler ×2 | 0.498 B | 0.93 GiB | f16 |
59
  | `upt.*` — latent temporal upscaler ×2 | 0.131 B | 0.24 GiB | f16 |
60
  | `dhead.*` — duration head | 1.9 M | 3.8 MB | f16 |
61
- | `ltx.config_json`, `te.asset.*` | — | 55 KB | raw |
62
  | tokenizer (`tokenizer_json`, 32 MB) | — | VOCAB section | raw |
63
 
64
- The DiT is the **distilled** transformer — the short-schedule one. The `dev`
65
- transformer is the same shape and packs the same way (`--dit …dev…`); the
66
- LoRAs, the int8/nvfp4 builds and the IC-LoRA upscaler are not in this file.
67
 
68
- ## What the codec does per tensor
69
-
70
- Four bits is not applied by fiat. `cortiq ltx-pack` decides per tensor:
71
 
72
  * **2-D planes of at least 2²⁰ weights → q4tp**, 4.16 bits with a per-row
73
- scale ladder. That is every projection in the DiT and in the encoder —
74
- 20.9 B of the 21.0 B, and 13.1 B of the 13.1 B.
75
- * **The adaLN tables stay exact.** `scale_shift_table`, the per-block
76
- prompt/audio tables, the connector's learnable registers and the VAE's
77
- `per_channel_statistics` ship as F32 in the release, are read once a step
78
- and modulate everything downstream. 19 MB in total — quantizing them
79
- would move every block's normalization to save nothing.
80
- * **Convolutions stay f16.** Both VAEs are 3-D/2-D convolutions; the decoder
81
- is what the eye sees. `--vae-quant q4tp` folds kernels to
82
- `[out, in·k·k·k]` and quantizes them anyway — offered, not the default.
83
- * **Small 2-D planes, norms and biases**: exact when the source is F32,
84
- f16 otherwise. The attention gates (`to_gate`, `[32, 4096]`), the patchify
85
- projection and `proj_out` are tiny and read at full sequence length.
 
 
 
86
 
87
  ## The architecture it carries
88
 
@@ -90,67 +206,65 @@ Four bits is not applied by fiat. `cortiq ltx-pack` decides per tensor:
90
  file as `ltx.config_json`):
91
 
92
  * **48 blocks**, video stream 4096 (32 heads × 128), **audio stream 2048**
93
- (32 × 64), joint audio↔video cross-attention with adaLN-gated fusion
94
- (`av_ca_a2v_gate`, `av_ca_v2a_gate`).
95
- * Per block: self-attention + cross-attention to the prompt, **RMS q/k-norm**,
96
- gated attention output (`to_gate`), gelu-approximate feed-forward without
97
- bias, and adaLN modulation from per-block `[9, 4096]` / `[9, 2048]` tables.
98
- * **3-D RoPE** over (frames, height, width), θ = 10000, max positions
99
- `[20, 2048, 2048]`, causal temporal positioning.
100
- * **Embeddings connectors** — 8 transformer layers each for video and audio
101
- with **128 learnable registers**, sitting between the encoder and the DiT.
102
- * The prompt encoder projects the **concatenation of all 49 Gemma-4 layer
103
- outputs** (`[4096, 188160]` for video, `[2048, 188160]` for audio), not the
104
- last hidden state — the file carries both aggregates.
 
 
 
 
 
 
 
 
 
 
105
 
106
  ## Packing it yourself
107
 
108
- Three passes, and each one deletes its source before the next lands — the
109
- stand this was packed on has a 50 GB disk quota and the sources are 71 GB:
110
 
111
  ```bash
112
- # 1 — the 22B transformer (42.0 GB → 11.0 GB), then delete it
113
- cortiq ltx-pack --out p1.cmf \
114
- --dit ltx-2.5-22b-distilled-transformer-bf16.safetensors
115
-
116
- # 2 — the Gemma-4 12B encoder on top (--in carries pass 1 byte for byte)
117
- cortiq ltx-pack --out p2.cmf --in p1.cmf \
118
- --te gemma4-12b-with-proj-ltx-2.5-bf16.safetensors
119
-
120
- # 3 — the VAEs, the upscalers and the duration head
121
  cortiq ltx-pack --out ltx25-q4tp.cmf --in p2.cmf \
122
  --video-vae ltx-2.5-video-vae-conv-bf16.safetensors \
123
  --audio-vae ltx-2.5-audio-vae-bf16.safetensors \
124
  --spatial-upscaler ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
125
  --temporal-upscaler ltx-2.5-latent-temporal-upscaler-x2-bf16-1.0.safetensors \
126
  --duration-head ltx-2.5-duration-head-bf16.safetensors
127
-
128
  cortiq verify ltx25-q4tp.cmf && cortiq info ltx25-q4tp.cmf
129
  ```
130
 
131
- Measured on a 32-core pod: 154 s for the transformer, 123 s for the encoder,
132
- 24 s for the rest — **five minutes** for 71 GB of bf16, single machine, no
133
- Python, no GPU.
134
-
135
- ## Status: the container is done, the renderer is not
136
-
137
- Be precise about what you are downloading:
138
-
139
- * ✅ **The file is complete and verified** — all 35.65 B weights, the
140
- tokenizer, every config, `cortiq verify` clean, and it opens through the
141
- same `mmap` path as every other CMF model.
142
- * ✅ **`cortiq info` / `verify` / `dequant` work on it** — any tensor can be
143
- read back to raw f32 for a reference comparison, which is how the port
144
- will be gated.
145
- * ⏳ **`cortiq` cannot render LTX-2.5 yet.** The runtime speaks
146
- [MiniMax-H3](https://huggingface.co/infosave/MiniMax-H3-Turbo-cmf) and
147
- MiniMax-Music-3 today; LTX-2.5's `AVTransformer3DModel` — the audio���video
148
- gated fusion, the connectors with their learnable registers, the 3-D RoPE
149
- and the LTX VAEs — is a separate port, in progress against this file.
150
-
151
- Until that lands this is a **format artifact**: the pipeline in one
152
- verifiable container, 3.4× smaller, for anyone who wants to read LTX-2.5's
153
- weights without a PyTorch stack — or to watch the port land against it.
154
 
155
  ## Provenance
156
 
 
6
  base_model:
7
  - Lightricks/LTX-2.5
8
  base_model_relation: quantized
9
+ pipeline_tag: text-to-video
10
  tags:
11
  - cmf
12
  - cortiq
13
  - video
14
+ - text-to-video
15
  - ltx-video
16
  - ltx-2.5
17
+ - rust
18
  - 4-bit
19
  ---
20
 
21
+ # LTX-2.5 — the whole pipeline in one 22 GB file, rendered by Rust
22
+
23
+ <p align="center">
24
+ <img src="assets/corgi.gif" width="49%" alt="A corgi in a chef hat flips a pancake in a sunlit kitchen">
25
+ <img src="assets/neon.gif" width="49%" alt="Neon rain on a Tokyo side street at night">
26
+ </p>
27
+ <p align="center">
28
+ <img src="assets/whale.gif" width="49%" alt="A humpback whale glides through a shaft of sunlight">
29
+ <img src="assets/glass.gif" width="49%" alt="Molten glass blown into a bulb over an orange furnace">
30
+ </p>
31
+
32
+ **Every frame above was produced by `cortiq`** — a single Rust binary with no
33
+ PyTorch, no diffusers, no CUDA toolkit and no Python anywhere in the process —
34
+ reading one memory-mapped [CMF](https://github.com/infosave2007/cmf) file.
35
 
36
  [LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5) renders video **and its
37
  soundtrack** from one prompt: a 21 B audio-video diffusion transformer that
38
  denoises picture and sound in the same 48 blocks, a Gemma-4 12 B prompt
39
  encoder, a 3-D video VAE, an audio VAE, two latent upscalers and a duration
40
  head. The reference checkout is **71.35 GB across six safetensors** plus a
41
+ PyTorch stack.
42
 
43
+ Here it is **one file of 22.07 GB** — every component, the Gemma-4 tokenizer
44
+ and every config inside it.
 
 
45
 
46
  | | reference | this file |
47
  |---|---|---|
48
  | files | 6 safetensors + configs + tokenizer | **1** |
49
+ | bytes | 71.35 GB | **22.07 GB** (3.2× smaller) |
50
  | weights | 35.65 B | 35.65 B — all of them |
51
  | loader | diffusers / ComfyUI + PyTorch | `mmap` |
52
+ | renderer | Python | **one Rust binary** |
53
+
54
+ ## Quick start
55
+
56
+ ```bash
57
+ # 1 — the runtime (Rust 1.85+; nothing else)
58
+ cargo install cortiq-cli
59
+
60
+ # 2 — the model
61
+ hf download infosave/LTX-2.5-cmf ltx25-q4tp.cmf --local-dir .
62
+ cortiq verify ltx25-q4tp.cmf # every tensor is hashed in the directory
63
+
64
+ # 3 — a video
65
+ cortiq ltx-video --model ltx25-q4tp.cmf \
66
+ --prompt "A corgi in a chef hat flips a pancake in a sunlit kitchen. \
67
+ Warm morning light, static camera." \
68
+ --height 256 --width 384 --frames 49 --fps 24 --seed 42 \
69
+ --out corgi.y4m
70
+
71
+ ffmpeg -i corgi.y4m -pix_fmt yuv420p corgi.mp4
72
+ ```
73
+
74
+ That is the whole thing: prompt in, frames out, one process, one file. The
75
+ GPU is found at run time — Vulkan on Linux and Windows, Metal on Apple
76
+ silicon — and everything falls back to the CPU when there is none.
77
+
78
+ > **Keep the file on local storage.** It is memory-mapped, so every weight is
79
+ > a page fault. On a network filesystem (NFS, MooseFS, a rented pod's
80
+ > `/workspace` volume) that is a network round trip per weight and the
81
+ > process will sit at 1 % CPU looking hung. Copy it to a local disk — or to
82
+ > `/dev/shm` if you have the RAM.
83
+
84
+ ### The examples above, exactly
85
+
86
+ ```bash
87
+ M=ltx25-q4tp.cmf
88
+ cortiq ltx-video --model $M --seed 42 --height 256 --width 384 --frames 49 \
89
+ --out-dir corgi/ --prompt \
90
+ "A corgi in a chef hat flips a pancake in a sunlit kitchen. Warm morning light, static camera."
91
+
92
+ cortiq ltx-video --model $M --seed 7 --height 256 --width 384 --frames 49 \
93
+ --out-dir neon/ --prompt \
94
+ "Neon rain on a Tokyo side street at night, a lone figure with a translucent umbrella \
95
+ walks past ramen shop signs, reflections rippling in the puddles, slow dolly."
96
+
97
+ cortiq ltx-video --model $M --seed 11 --height 256 --width 384 --frames 49 \
98
+ --out-dir whale/ --prompt \
99
+ "A humpback whale glides through a shaft of sunlight in deep blue water, plankton \
100
+ drifting like dust, the camera rises with it toward the surface."
101
+
102
+ cortiq ltx-video --model $M --seed 23 --height 256 --width 384 --frames 49 \
103
+ --out-dir glass/ --prompt \
104
+ "Molten glass is blown into a bulb over an orange furnace, the glowing gather \
105
+ stretching and rotating, sparks drifting in the dark workshop."
106
+
107
+ # frames → mp4 → gif
108
+ ffmpeg -framerate 24 -i corgi/frame_%04d.ppm -pix_fmt yuv420p -crf 18 corgi.mp4
109
+ ffmpeg -i corgi.mp4 -vf "fps=12,scale=384:-1:flags=lanczos,split[s0][s1];\
110
+ [s0]palettegen[p];[s1][p]paletteuse" corgi.gif
111
+ ```
112
+
113
+ `--out-dir` writes `frame_0000.ppm …`; `--out file.y4m` writes one
114
+ [YUV4MPEG2](https://wiki.multimedia.cx/index.php/YUV4MPEG2) stream instead,
115
+ which every tool reads — so the renderer needs no video encoder of its own.
116
+
117
+ ### Higher resolution
118
 
119
+ ```bash
120
+ cortiq ltx-video --model $M --two-stage \
121
+ --height 512 --width 768 --frames 49 --seed 42 \
122
+ --prompt "…" --out hq.y4m
123
  ```
124
+
125
+ `--two-stage` samples the way the distilled model was trained: eight
126
+ ancestral Euler steps at half resolution, the learned latent upscaler ×2,
127
+ then three deterministic steps that refine what the upscale invented.
128
+
129
+ Resolution must be a multiple of 32 (the video VAE's spatial stride) and the
130
+ frame count `8k + 1` (its temporal stride plus the standalone first frame).
131
+
132
+ ### Measured
133
+
134
+ RTX 5090, `/dev/shm`, 49 frames at 24 fps:
135
+
136
+ | stage | 384×256 | 768×512 (`--two-stage`) |
137
+ |---|---|---|
138
+ | prompt encode (Gemma-4 12 B + connectors) | 28 s | 28 s |
139
+ | denoise | 8 × 30 s | 8 × 30 s + 3 × 120 s |
140
+ | latent upscale | — | 25 s |
141
+ | video VAE | 50 s | 200 s |
142
+
143
+ Nothing here is tuned yet — the transformer runs one 48-block forward per
144
+ step against `mmap`ped 4-bit weights, and the VAE is a straight
145
+ im2col + GEMM.
146
+
147
+ ## The stages, separately
148
+
149
+ Each stage is its own command. That is how the port was gated: every one of
150
+ them can be run against a dump of the reference implementation's own
151
+ activations and will report the first place it diverges.
152
+
153
+ ```bash
154
+ # prompt → the two context tensors the transformer cross-attends to
155
+ cortiq ltx-encode --model $M --prompt "…" --out context.safetensors
156
+
157
+ # context → latent → frames
158
+ cortiq ltx-render --model $M --context context.safetensors \
159
+ --height 256 --width 384 --frames 49 --out-latent latent.safetensors --out out.y4m
160
+
161
+ # latent → frames, the 3-D convolutional decoder alone
162
+ cortiq ltx-decode --model $M --latent latent.safetensors --out-dir frames/
163
  ```
164
 
165
  ## What is inside
166
 
167
  | component | weights | in the file | codec |
168
  |---|---|---|---|
169
+ | `dit.*` — LTX-2.5 22B audio-video DiT (distilled) | 21.004 B | 10.84 GiB | q4tp + exact adaLN |
170
+ | `te.*` — Gemma-4 12B prompt encoder, aggregates, vision tower | 13.116 B | 6.87 GiB | q4tp + q8 embeddings |
171
  | `vvae.*` — video VAE (3-D conv, encoder + decoder) | 0.726 B | 1.35 GiB | f16 |
172
  | `avae.*` — audio VAE | 0.182 B | 0.34 GiB | f16 |
173
  | `ups.*` — latent spatial upscaler ×2 | 0.498 B | 0.93 GiB | f16 |
174
  | `upt.*` — latent temporal upscaler ×2 | 0.131 B | 0.24 GiB | f16 |
175
  | `dhead.*` — duration head | 1.9 M | 3.8 MB | f16 |
176
+ | configs, HF assets | — | 55 KB | raw |
177
  | tokenizer (`tokenizer_json`, 32 MB) | — | VOCAB section | raw |
178
 
179
+ ### What the codec does per tensor, and why
 
 
180
 
181
+ Four bits is not applied by fiat. `cortiq ltx-pack` decides per tensor, and
182
+ two of those decisions were made by measuring against the reference rather
183
+ than by taste:
184
 
185
  * **2-D planes of at least 2²⁰ weights → q4tp**, 4.16 bits with a per-row
186
+ scale ladder. Every projection in the transformer and in the encoder.
187
+ * **The adaLN-single stacks stay exact.** Their output is not a residual —
188
+ it is the scale and the shift applied to every token in every block, so an
189
+ error there is multiplied into the whole stream instead of being averaged
190
+ away by anything downstream. Quantized, they put 3.6·10⁻² of relative
191
+ error into the very first normalization of block 0; exact, 5.9·10⁻³.
192
+ 0.56 GB.
193
+ * **The token table stays 8-bit.** It *is* the residual stream at layer zero
194
+ and it carries through forty-eight residual additions. q4tp put 11 % into
195
+ every hidden state the prompt encoder produced; q8 puts 0.5 % there, for
196
+ 0.5 GB.
197
+ * **The adaLN tables, the connector's learnable registers and the VAE's
198
+ `per_channel_statistics` stay exact** — 19 MB in total, read once a step,
199
+ modulating everything.
200
+ * **Convolutions stay f16.** Both VAEs and both upscalers are convolutional,
201
+ and the decoder is what the eye actually sees.
202
 
203
  ## The architecture it carries
204
 
 
206
  file as `ltx.config_json`):
207
 
208
  * **48 blocks**, video stream 4096 (32 heads × 128), **audio stream 2048**
209
+ (32 × 64), joint audio↔video cross-attention with adaLN-gated fusion that
210
+ reads the *pre-fusion* state of both streams, so the order the two
211
+ directions run in cannot bias the result.
212
+ * Per block: self-attention, cross-attention to the prompt with its own
213
+ adaLN pair on the query *and* on the prompt's keys and values, **RMS
214
+ q/k-norm across the whole inner dimension**, gated attention
215
+ (`2·sigmoid` per head), and a gelu-approximate feed-forward — all
216
+ modulated from per-block `[9, 4096]` / `[9, 2048]` tables.
217
+ * **Split 3-D RoPE** over (seconds, pixel row, pixel column) evaluated at the
218
+ *middle* of each patch's bounds, θ = 10000, with the causal correction that
219
+ gives the first latent frame one pixel frame where every later one gets
220
+ eight. The audio stream shares the time axis in seconds, which is what lets
221
+ the two cross-attend positionally.
222
+ * **The prompt encoder** is Gemma-4 12 B — forty sliding-window layers at head
223
+ 256 and eight full-attention layers at head 512 whose value projection *is*
224
+ the key projection — and the features are not its last hidden state: all
225
+ **forty-nine** layer outputs are RMS-normalized per token per layer,
226
+ concatenated to 188160 numbers and projected once to 4096 (video) and once
227
+ to 2048 (audio).
228
+ * **Embeddings connectors** — 8 gated-attention blocks each for video and
229
+ audio, with **128 learnable registers** that replace every padded position,
230
+ which is why the transformer needs no prompt mask at all.
231
 
232
  ## Packing it yourself
233
 
234
+ Three passes, each one able to delete its source before the next lands — the
235
+ stand this was packed on had a 50 GB disk quota and the sources are 71 GB:
236
 
237
  ```bash
238
+ cortiq ltx-pack --out p1.cmf --dit ltx-2.5-22b-distilled-transformer-bf16.safetensors
239
+ cortiq ltx-pack --out p2.cmf --in p1.cmf --te gemma4-12b-with-proj-ltx-2.5-bf16.safetensors
 
 
 
 
 
 
 
240
  cortiq ltx-pack --out ltx25-q4tp.cmf --in p2.cmf \
241
  --video-vae ltx-2.5-video-vae-conv-bf16.safetensors \
242
  --audio-vae ltx-2.5-audio-vae-bf16.safetensors \
243
  --spatial-upscaler ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
244
  --temporal-upscaler ltx-2.5-latent-temporal-upscaler-x2-bf16-1.0.safetensors \
245
  --duration-head ltx-2.5-duration-head-bf16.safetensors
 
246
  cortiq verify ltx25-q4tp.cmf && cortiq info ltx25-q4tp.cmf
247
  ```
248
 
249
+ Measured on a 32-core pod: **five minutes** for 71 GB of bf16, single
250
+ machine, no Python, no GPU. `--quant` picks the codec for the big planes
251
+ (`q4tp`, `q8`, `f16`, `f32`), `--vae-quant` the one for convolutions.
252
+
253
+ ## Status
254
+
255
+ * ✅ **Text → video runs end to end on the Rust engine**: the Gemma-4 prompt
256
+ encoder, the aggregate projections, the connectors, the 48-block audio-video
257
+ transformer, the sampler, the latent upscaler and the video VAE.
258
+ * ⏳ **Sound is generated but not yet decoded.** The transformer denoises the
259
+ audio latent in the same blocks as the picture, and it is in the output
260
+ latent — the audio VAE and its vocoder are the next port, so the clips here
261
+ are silent.
262
+ * ⏳ **Image and video conditioning, LoRAs, the IC-LoRA upscaler and the
263
+ duration head** are in the file but not yet wired into the CLI.
264
+
265
+ Everything above is honest about what it is: a 4-bit repack. The reference at
266
+ bf16 is the quality ceiling, and the codec's cost was measured stage by stage
267
+ rather than assumed — see the numbers in the codec section.
 
 
 
 
268
 
269
  ## Provenance
270