infosave commited on
Commit
41b5099
·
verified ·
1 Parent(s): 9e08d5a

card: field notes from discussions #1/#2/#4 — 24 GB frame budgets, keyframe VAE encode on CPU, kill disarmed during prompt encode (0.5.80), P-core QoS

Browse files
Files changed (1) hide show
  1. README.md +67 -0
README.md CHANGED
@@ -196,6 +196,7 @@ binary both splits and replicates across cards — see
196
  | `CMF_GPU_PROBE=0` | pin the op arbitration (already the default for `animate`, so a seed reproduces) |
197
  | `CMF_THREADS=n` | cap the worker pool (defaults to the machine's cores) |
198
  | `CMF_ANIM_PROF=1` | per-step rms of both latent streams and both velocities |
 
199
 
200
  ### What it needs
201
 
@@ -208,6 +209,72 @@ peaks at 15.1 GB of VRAM — a whole-run maximum, polled, not a snapshot. That i
208
  what makes it the one to reach for on a 20 GB card: below that the prompt
209
  encoder pages, and on this workload paging costs more than any kernel.
210
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
211
  ## Making it smaller
212
 
213
  Half this file is the PROMPT ENCODER — 12.2 GB of Qwen3-VL against the DiT that
 
196
  | `CMF_GPU_PROBE=0` | pin the op arbitration (already the default for `animate`, so a seed reproduces) |
197
  | `CMF_THREADS=n` | cap the worker pool (defaults to the machine's cores) |
198
  | `CMF_ANIM_PROF=1` | per-step rms of both latent streams and both velocities |
199
+ | `CMF_MM_KILL=0` | never fall back to the host on a slow device op. The engine treats three consecutive over-budget GEMMs as "another process owns the card" and finishes the run on the CPU; a weight paging in from disk is exempt, but on a machine where the file exceeds RAM the first steps can be slow for reasons that are not contention (see the field notes) |
200
 
201
  ### What it needs
202
 
 
209
  what makes it the one to reach for on a 20 GB card: below that the prompt
210
  encoder pages, and on this workload paging costs more than any kernel.
211
 
212
+ ### Field notes: 24 GB Macs and 20 GB cards
213
+
214
+ Two users measured what this card could not, and the engine changed on
215
+ their reports (HF discussions #1, #2 and #4). The numbers, as they sent them:
216
+
217
+ | machine | file | render | result |
218
+ |---|---|---|---|
219
+ | RTX 3080 20 GB, Windows | `mmh3-turbo-q4tp` (25.2 GB) | 512×288×39, 4 steps | **1218 s** — the encoder did not fit the 19 GB budget and paged every step |
220
+ | RTX 3080 20 GB, Windows | `mmh3-turbo-clipproj4b-q4tp` (13.2 GB) | same | **105 s** — resident; denoise 62.8 s, video VAE 30.6 s |
221
+ | Mac mini M4 24 GB | `mmh3-turbo-clipproj4b-q4tp` | 512×256×39, 4 steps | **174 s** — denoise 122 s, video VAE 41 s |
222
+ | Mac mini M4 24 GB | `mmh3-turbo-clipproj4b-fl2va-q4tp` (14.5 GB) | 22 frames from a keyframe | **92 s**, no swap |
223
+ | Mac mini M4 24 GB | `mmh3-turbo-fl2va-q4tp` (25.7 GB), 0.5.79 | 10 / 40 frames | **48 s / 140 s per denoise step** on the GPU, 20 GB resident, 0.3 GB swap |
224
+ | Mac mini M4 24 GB | same, 150 frames | | 8 GB of swap and a sawtooth — the activation cache no longer fits |
225
+ | Mac mini M4 24 GB | `mmh3-turbo-clipproj4b-fl2va-v2-q4tp` (14.5 GB) | 448×768, 50 frames | **136 s per denoise step**, faces hold through the clip; 90 frames at that size is the paging threshold |
226
+ | Mac mini M4 24 GB | `mmh3-turbo-fl2va-q4tp` (25.7 GB) | 512×256 | safe to ~70–90 frames; **768×448 caps at 40–50** (90 frames: 22.2 GB RAM + 1.1 GB swap, the GPU stalls) |
227
+ | Mac mini M4 24 GB | any file, `--first-frame` | | the video-VAE **encoder** of the keyframe ran ~100–140 s on the CPU (`encode 0/1 (140.8s)`); text-to-video skips it (0.1 s) |
228
+
229
+ What follows from them, if you are on such a machine:
230
+
231
+ - **On a 24 GB Mac the 25.7 GB file works after 0.5.79, for clips of
232
+ ≤ 40 frames.** The prompt encoder's pages are released after the text
233
+ encode; what remains — DiT, VAE decoders and the activation cache — fits
234
+ up to ~40 frames. Past that the cache spills to swap. For longer clips
235
+ chain 40-frame chunks (`--last-frame` of one render becomes
236
+ `--first-frame` of the next) or use the `clipproj4b` files, which leave
237
+ the room.
238
+ - **The Metal driver budgets by buffer size, not resident pages.** The
239
+ weight arena's overlapping windows read as ~27 GB to the driver even
240
+ after the encoder is released, so the first ops after the encode page
241
+ from the SSD and can take seconds. Before 0.5.80 a single such op tripped
242
+ the contention kill and the rest of the run walked the CPU (>60 s a
243
+ step); the kill now needs three consecutive strikes, exempts weights
244
+ that were not resident, and `CMF_MM_KILL=0` turns it off. If a run still
245
+ says `device contended, CPU for the rest of the process` on a machine
246
+ nobody else is using, that variable is the answer.
247
+ - **The kill is disarmed during the prompt encode (0.5.80).** The
248
+ encoder is a one-shot pass over 12 GB of weights; on a machine the file
249
+ does not fit it streams from disk, and its GEMMs run over any budget
250
+ for reasons that are not contention — the report that had to gut
251
+ `mm_kill` in the source was on exactly that. The kill now arms only
252
+ when the denoise loop starts, and three strikes are still needed.
253
+ - **The keyframe encode is CPU work today.** With `--first-frame` the
254
+ video-VAE encoder runs the reference picture on the host (~100–140 s
255
+ on an M4 for a 448×768 frame; 0.1 s without a keyframe). 0.5.80 puts
256
+ the pool's workers on the performance cores (they were landing on the
257
+ E-cores with the P-cores asleep — that report's `asitop`), which
258
+ shortens it; a device path for the encoder's 3-D convolutions is the
259
+ real fix and is on the list.
260
+ - **Frame budgets on 24 GB, from that user's sweep:** with the 25.7 GB
261
+ file 512×256 is safe to ~70–90 frames and 768×448 caps at 40–50; with
262
+ the 14.5 GB v2 file 448×768 at 50 frames is the sweet spot (136 s per
263
+ step) and 90 frames is where paging starts. Past those the machine
264
+ swaps and the GPU stalls to nothing — better to chain clips than to
265
+ push the frame count.
266
+ - **The draft → final workflow.** Block the shot on `clipproj4b-fl2va-v2`
267
+ (90 s on the Mac; v2 holds faces through a clip — "25 GB-level face
268
+ consistency at 14 GB speed", that user's words), then pay the full
269
+ encoder once for the final take. Identity of *specific* real people is
270
+ still the full encoder's territory.
271
+ - **`--height 256` instead of 288** halves the video-VAE decode on every
272
+ machine (three 256-pixel tiles instead of six).
273
+ - **Voices are prompt space.** There is no reference-audio input — H3
274
+ conditions on text and keyframes only — but speaker identity, timbre,
275
+ pace and emotion respond to stage directions in the prompt, and a fixed
276
+ seed keeps the same actor across takes.
277
+
278
  ## Making it smaller
279
 
280
  Half this file is the PROMPT ENCODER — 12.2 GB of Qwen3-VL against the DiT that