infosave commited on
Commit
e7e3021
Β·
verified Β·
1 Parent(s): 5ba9a32

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +43 -19
README.md CHANGED
@@ -29,6 +29,11 @@ tags:
29
  <img src="assets/glass.gif" width="49%" alt="Molten glass blown into a bulb over an orange furnace">
30
  </p>
31
 
 
 
 
 
 
32
  **Every frame above was produced by `cortiq`** β€” a single Rust binary with no
33
  PyTorch, no diffusers, no CUDA toolkit and no Python anywhere in the process β€”
34
  reading one memory-mapped [CMF](https://github.com/infosave2007/cmf) file.
@@ -114,8 +119,26 @@ ffmpeg -i corgi.mp4 -vf "fps=12,scale=384:-1:flags=lanczos,split[s0][s1];\
114
  [YUV4MPEG2](https://wiki.multimedia.cx/index.php/YUV4MPEG2) stream instead,
115
  which every tool reads β€” so the renderer needs no video encoder of its own.
116
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
117
  ### Higher resolution
118
 
 
 
119
  ```bash
120
  cortiq ltx-video --model $M --two-stage \
121
  --height 512 --width 768 --frames 49 --seed 42 \
@@ -131,18 +154,20 @@ frame count `8k + 1` (its temporal stride plus the standalone first frame).
131
 
132
  ### Measured
133
 
134
- RTX 5090, `/dev/shm`, 49 frames at 24 fps:
135
 
136
- | stage | 384Γ—256 | 768Γ—512 (`--two-stage`) |
137
- |---|---|---|
138
- | prompt encode (Gemma-4 12 B + connectors) | 28 s | 28 s |
139
- | denoise | 8 Γ— 30 s | 8 Γ— 30 s + 3 Γ— 120 s |
140
- | latent upscale | β€” | 25 s |
141
- | video VAE | 50 s | 200 s |
 
 
142
 
143
- Nothing here is tuned yet β€” the transformer runs one 48-block forward per
144
- step against `mmap`ped 4-bit weights, and the VAE is a straight
145
- im2col + GEMM.
146
 
147
  ## The stages, separately
148
 
@@ -252,15 +277,14 @@ machine, no Python, no GPU. `--quant` picks the codec for the big planes
252
 
253
  ## Status
254
 
255
- * βœ… **Text β†’ video runs end to end on the Rust engine**: the Gemma-4 prompt
256
- encoder, the aggregate projections, the connectors, the 48-block audio-video
257
- transformer, the sampler, the latent upscaler and the video VAE.
258
- * ⏳ **Sound is generated but not yet decoded.** The transformer denoises the
259
- audio latent in the same blocks as the picture, and it is in the output
260
- latent β€” the audio VAE and its vocoder are the next port, so the clips here
261
- are silent.
262
- * ⏳ **Image and video conditioning, LoRAs, the IC-LoRA upscaler and the
263
- duration head** are in the file but not yet wired into the CLI.
264
 
265
  Everything above is honest about what it is: a 4-bit repack. The reference at
266
  bf16 is the quality ceiling, and the codec's cost was measured stage by stage
 
29
  <img src="assets/glass.gif" width="49%" alt="Molten glass blown into a bulb over an orange furnace">
30
  </p>
31
 
32
+ > **The clips above are silent GIFs. The videos are not.** The same 48 blocks
33
+ > denoise the soundtrack alongside the picture β€” hear it in
34
+ > [`examples/`](./tree/main/examples): six mp4s with audio, their raw 48 kHz
35
+ > stereo wavs, and the exact command that made each one.
36
+
37
  **Every frame above was produced by `cortiq`** β€” a single Rust binary with no
38
  PyTorch, no diffusers, no CUDA toolkit and no Python anywhere in the process β€”
39
  reading one memory-mapped [CMF](https://github.com/infosave2007/cmf) file.
 
119
  [YUV4MPEG2](https://wiki.multimedia.cx/index.php/YUV4MPEG2) stream instead,
120
  which every tool reads β€” so the renderer needs no video encoder of its own.
121
 
122
+ ### Sound
123
+
124
+ ```bash
125
+ cortiq ltx-video --model $M --prompt "…" \
126
+ --height 256 --width 384 --frames 49 --seed 3 \
127
+ --out-dir frames/ --out-audio track.wav
128
+
129
+ ffmpeg -framerate 24 -i frames/frame_%04d.ppm -i track.wav \
130
+ -pix_fmt yuv420p -c:v libx264 -crf 18 -c:a aac -b:a 192k -shortest out.mp4
131
+ ```
132
+
133
+ The transformer has been denoising the soundtrack in the same blocks as the
134
+ picture the whole time; `--out-audio` decodes it β€” the spectrogram VAE, then
135
+ BigVGAN v2, then a bandwidth extender that lifts 16 kHz to 48 kHz stereo.
136
+ Eight seconds of work behind minutes of denoising.
137
+
138
  ### Higher resolution
139
 
140
+ <p align="center"><img src="assets/hq-still.png" width="70%" alt="768x512, two-stage"></p>
141
+
142
  ```bash
143
  cortiq ltx-video --model $M --two-stage \
144
  --height 512 --width 768 --frames 49 --seed 42 \
 
154
 
155
  ### Measured
156
 
157
+ 49 frames at 24 fps, container on local storage:
158
 
159
+ | stage | RTX 5090, 384Γ—256 | RTX 5090, 768Γ—512 `--two-stage` | **M4 MacBook, 24 GB**, 384Γ—256 |
160
+ |---|---|---|---|
161
+ | prompt encode (Gemma-4 12 B + connectors) | 26 s | 26 s | 34 s |
162
+ | denoise | 8 Γ— 19 s | 8 Γ— 19 s + 3 Γ— 193 s | 8 Γ— 23 s |
163
+ | latent upscale | β€” | 12 s | β€” |
164
+ | audio VAE + vocoder | 8 s | 8 s | 9 s |
165
+ | video VAE | 50 s | 580 s | 26 s |
166
+ | **total** | **3 min** | **17 min** | **4 min** |
167
 
168
+ A 21 B video model, its 12 B prompt encoder and both VAEs, rendering a clip
169
+ with sound on a laptop with 24 GB of unified memory β€” because nothing is ever
170
+ loaded, only mapped, and the pipeline touches one component at a time.
171
 
172
  ## The stages, separately
173
 
 
277
 
278
  ## Status
279
 
280
+ * βœ… **Text β†’ video *and sound* runs end to end on the Rust engine**: the
281
+ Gemma-4 prompt encoder, the aggregate projections, the connectors, the
282
+ 48-block audio-video transformer, the sampler, the latent upscaler, the
283
+ video VAE, the audio VAE with its BigVGAN vocoder and bandwidth extension,
284
+ and the duration head.
285
+ * ⏳ **Image and video conditioning, LoRAs and the IC-LoRA upscaler** are next.
286
+ The weights for conditioning are in this file (the video VAE's encoder half);
287
+ the LoRAs and the IC-LoRA upscaler are separate releases.
 
288
 
289
  Everything above is honest about what it is: a 4-bit repack. The reference at
290
  bf16 is the quality ceiling, and the codec's cost was measured stage by stage