infosave commited on
Commit
37d210a
Β·
verified Β·
1 Parent(s): 0f34ddc

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +118 -31
README.md CHANGED
@@ -58,56 +58,143 @@ and reference images, videos and audio (`ref2va`); those paths are not ported
58
  and the vision tower is not packed. What is here is `t2va`: prompt in, video
59
  and audio out.
60
 
61
- ## Getting a file
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
 
63
  ```bash
64
- huggingface-cli download infosave/MiniMax-H3-Turbo-cmf \
65
- --include 'parts-q4tp/part_*' --local-dir .
66
- cat parts-q4tp/part_* > mmh3-turbo-q4tp.cmf
67
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
68
  cortiq animate mmh3-turbo-q4tp.cmf \
69
  --prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
70
- --width 512 --height 288 --frames 39 --out corgi.avi
 
 
 
 
 
 
 
 
 
 
 
 
 
 
71
  ```
72
 
73
- `cortiq animate` writes an MJPEG+PCM AVI and the stereo track beside it as
74
- `.wav`. There is no ffmpeg in the loop and nothing to install: the JPEG encoder
75
- and the RIFF muxer are part of the binary, because a pipeline that ends in a
76
- shell-out to a 20 MB dependency is not a pipeline you can ship.
 
 
 
 
 
 
77
 
78
- Width and height are multiples of 32. Frame counts snap **up** to the model's
79
- 17k+5 grid (5, 22, 39, 56, … 124), which is the grid the temporal VAE and the
80
- DiT's frame-span pattern agree on. 4 steps is what the LoRA is trained for;
81
- more still helps a little.
 
 
 
 
 
 
 
82
 
83
  ## What it costs to run
84
 
85
  The file is memory-mapped, so plan on RAM at least its size or every step
86
  touches non-resident pages.
87
 
88
- Measured on 48 CPU cores, 4 steps:
 
89
 
90
- | | tokens in the pack | denoise (4 steps) | decode | total |
91
- |---|---|---|---|---|
92
- | 256Γ—160, 22 frames (0.9 s) | 375 | 45.0 s | 10.7 s | **55.7 s** |
93
- | 512Γ—288, 39 frames (1.6 s) | 1 879 | 220.8 s | 150.0 s | **371.3 s** |
94
 
95
- Nearly all of the decode is the video VAE β€” the vocoder is 4 s of the 150.
96
 
97
- The packed sequence is `[text | audio | video]` and everything attends to
98
- everything, so cost grows with the token count and then with its square. A
 
99
  512Γ—288 second is five times the tokens of a 256Γ—160 one.
100
 
101
- **The GPU path is off, on purpose.** `cortiq`'s wide-GEMM arm on wgpu is
102
- measured three times faster than the host here and it is also wrong: on an
103
- RTX PRO 6000 Blackwell the DiT's first step disagrees with the CPU (video
104
- velocity rms 1.32 against 1.73, audio 0.16 against 1.01) and the second step
105
- returns NaN β€” a flat grey frame and a clipped waveform. The op probe cannot
106
- notice: it only times the two arms, and it runs the bad one for real while it
107
- is deciding. So `cortiq animate` pins `CMF_GPU=0` before the backend comes up,
108
- which is early enough to be sure. `CMF_MMH3_GPU=1` opts back in if you want to
109
- watch it fail. Fixing that kernel is the next thing worth doing to this model β€”
110
- it is the whole difference between minutes and hours.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
111
 
112
  ## What the conversion did
113
 
 
58
  and the vision tower is not packed. What is here is `t2va`: prompt in, video
59
  and audio out.
60
 
61
+ ## Running it
62
+
63
+ ### 1. Get the runtime
64
+
65
+ `cortiq` is one Rust binary. Either install it β€”
66
+
67
+ ```bash
68
+ cargo install cortiq-cli # needs Rust 1.85+; brings the GPU backend
69
+ ```
70
+
71
+ β€” or take a prebuilt archive from the
72
+ [latest release](https://github.com/infosave2007/cmf/releases/latest)
73
+ (Linux x86-64, macOS on Apple Silicon and Intel, Windows x86-64 and ARM64;
74
+ each ships a `.sha256`). Nothing else is required: no Python, no PyTorch, no
75
+ CUDA toolkit, no ffmpeg.
76
+
77
+ Check it took:
78
 
79
  ```bash
80
+ cortiq --version
81
+ ```
 
82
 
83
+ ### 2. Get the weights
84
+
85
+ One file, 23.5 GB.
86
+
87
+ ```bash
88
+ pip install -U "huggingface_hub[cli]" # only to fetch the file
89
+ hf download infosave/MiniMax-H3-Turbo-cmf mmh3-turbo-q4tp.cmf --local-dir .
90
+ ```
91
+
92
+ Confirm it arrived whole β€” the container carries a hash per tensor:
93
+
94
+ ```bash
95
+ cortiq verify mmh3-turbo-q4tp.cmf # β†’ βœ“ all tensor hashes match
96
+ cortiq info mmh3-turbo-q4tp.cmf # β†’ arch, layers, 47.83B params
97
+ ```
98
+
99
+ ### 3. Render
100
+
101
+ ```bash
102
  cortiq animate mmh3-turbo-q4tp.cmf \
103
  --prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
104
+ --width 512 --height 288 --frames 39 --steps 4 --seed 42 \
105
+ --out corgi.avi
106
+ ```
107
+
108
+ That writes `corgi.avi` β€” MJPEG video with PCM stereo, playable in VLC, mpv,
109
+ QuickTime and Windows Media Player β€” and `corgi.wav` beside it. The JPEG
110
+ encoder and the RIFF muxer are inside the binary: a pipeline that ends in a
111
+ shell-out to a 20 MB dependency is not a pipeline you can ship. If you want an
112
+ mp4 for a browser, remux it yourself; the model never needs one.
113
+
114
+ **On a GPU.** The device path is opt-in for this model while its kernels earn
115
+ their keep (see below):
116
+
117
+ ```bash
118
+ CMF_MMH3_GPU=1 cortiq animate mmh3-turbo-q4tp.cmf --prompt "…" --out corgi.avi
119
  ```
120
 
121
+ ### Options that matter
122
+
123
+ | flag | default | what it does |
124
+ |---|---|---|
125
+ | `--width` / `--height` | 512 Γ— 288 | multiples of 32. The trained short edge is 768; below ~256 the model drifts off-distribution |
126
+ | `--frames` | 39 | at 24 fps, snapped **up** to the model's 17k+5 grid: 5, 22, 39, 56, … 124. 124 β‰ˆ 5 s, and 124–362 is the validated range |
127
+ | `--steps` | 4 | what the Turbo LoRA is trained for. More still helps a little |
128
+ | `--seed` | 42 | same seed, same prompt, same size β†’ the same clip, byte for byte |
129
+ | `--quality` | 92 | JPEG quality of the AVI's frames |
130
+ | `--stock-sampler` | off | integrate the audio on the video's clock, as a single-schedule sampler does. Wrong at 4 steps β€” it is here to hear how wrong |
131
 
132
+ | environment | what it does |
133
+ |---|---|
134
+ | `CMF_MMH3_GPU=1` | opt into the device path |
135
+ | `CMF_THREADS=n` | cap the worker pool (defaults to the machine's cores) |
136
+ | `CMF_ANIM_PROF=1` | per-step rms of both latent streams and both velocities |
137
+
138
+ ### What it needs
139
+
140
+ RAM at least the file's size β€” 24 GB β€” or every step faults on non-resident
141
+ pages; the weights are memory-mapped, not read. Disk: 24 GB. A GPU is optional
142
+ and wants ~14 GB of VRAM for the DiT's planes. No network access at run time.
143
 
144
  ## What it costs to run
145
 
146
  The file is memory-mapped, so plan on RAM at least its size or every step
147
  touches non-resident pages.
148
 
149
+ 512Γ—288, 39 frames, 4 steps, one machine β€” 48 CPU cores and one RTX PRO 6000
150
+ Blackwell:
151
 
152
+ | | denoise | decode | total |
153
+ |---|---|---|---|
154
+ | host | 198.2 s | 147.8 s | **346.5 s** |
155
+ | `CMF_MMH3_GPU=1` | 96.8 s | 74.7 s | **172.0 s** |
156
 
157
+ and the smaller size, on the device: 256Γ—160 over 22 frames in **29.2 s**.
158
 
159
+ Nearly all of the decode is the video VAE β€” the vocoder is 4 s of it. The
160
+ packed sequence is `[text | audio | video]` and everything attends to
161
+ everything, so cost grows with the token count and then with its square: a
162
  512Γ—288 second is five times the tokens of a 256Γ—160 one.
163
 
164
+ **A free 2Γ— on the decoder, if you want it.** The video VAE decodes in
165
+ 256-pixel tiles, always, and grows the OVERLAP rather than the tile count β€”
166
+ so a 288-pixel edge is covered by two 256-pixel tiles overlapping by 224, and
167
+ you pay for 512 rows to get 288. An edge of exactly 256 is one tile. 512Γ—256
168
+ therefore decodes three tiles where 512Γ—288 decodes six, for 89% of the
169
+ pixels. The schedule is the reference's and this port reproduces it exactly;
170
+ picking an edge that lands on it is free.
171
+
172
+ **Host and device do not agree to the last bit, and neither is wrong.** The
173
+ host arm quantizes activations to int8 (`CMF_SDOT`) where the device
174
+ dequantizes to f32, so the two renders differ by a few per cent in latent rms
175
+ and visibly in fine texture. Set `CMF_SDOT=0` on both sides to compare
176
+ arithmetic instead of that approximation.
177
+
178
+ **Why the device is opt-in.** Getting it right took three fixes, and one
179
+ thing is still held back.
180
+
181
+ The engine's blocked f32 GEMM cached its weight-side device buffer **by
182
+ pointer address**. Every batched attention allocates one k/v scratch pair per
183
+ call and refills it per head β€” same address, different matrix β€” so head 0's
184
+ keys came back for every head, on the GPU only, silently. It is keyed on a
185
+ content fingerprint now. The same GEMM also took every job over 4 M MACs on
186
+ sight with no CPU arm to lose to, which on this model's decoder was three
187
+ times *slower* than the host it displaced; it goes through the same
188
+ measure-don't-assume probe as every other op class now, and on this stack the
189
+ probe hands that work back (0.24 ms device against 0.13 host) while sending
190
+ the weight GEMMs to the card (25.8 ms against 92.0).
191
+
192
+ Still held: **the cooperative-matrix kernel runs this model out of f16 range.**
193
+ At 256Γ—160 the render is correct; at 512Γ—288 the audio stream goes NaN on the
194
+ second sampling step and the video follows. Bisected β€” `CMF_BAKE_GPU=0` does
195
+ not help, `CMF_COOP=0` does β€” so `cortiq animate` pins `CMF_COOP=0`. Tensor
196
+ cores are worth having here; the kernel needs to carry a scale before it can
197
+ carry these activations. That is the next real speedup in this model.
198
 
199
  ## What the conversion did
200