infosave commited on
Commit
79fdefb
Β·
verified Β·
1 Parent(s): 7097d8d

ClipProj section: scope the timings to the card they were taken on

Browse files
Files changed (1) hide show
  1. README.md +23 -16
README.md CHANGED
@@ -63,7 +63,7 @@ not ported.
63
  |---|---|---|
64
  | `mmh3-turbo-fl2va-q4tp.cmf` | 23.94 GB | **use this one** β€” text-to-video AND keyframe-to-video |
65
  | `mmh3-turbo-q4tp.cmf` | 23.47 GB | the same without the vision tower: text-to-video only |
66
- | `mmh3-turbo-clipproj4b-q4tp.cmf` | **13.16 GB** | text-to-video with a 4B stand-in prompt encoder. Peaks at 15.1 GB of VRAM, so a 20 GB card holds the whole run; 2x faster, and still four bits everywhere. Costs some set dressing β€” see below |
67
  | `mmh3-turbo-fl2va-q2tp.cmf` | 18.74 GB | two bits on the gate/up planes. Smaller, faster, and it stops following the prompt β€” kept for anyone who wants to push on it, not for rendering. See below |
68
 
69
  ## Keyframe to video
@@ -280,20 +280,27 @@ that the mechanism is wired right.
280
 
281
  ### What it costs
282
 
283
- Same prompt, same seed, same four steps, one RTX 3090, both files under
284
- identical settings:
 
285
 
286
- | | encode | denoise | video VAE | audio VAE | **total** |
287
- |---|---|---|---|---|---|
288
- | `mmh3-turbo-q4tp` | 49.2 s | 116.4 s | 101.9 s | 6.7 s | **281 s** |
289
- | `mmh3-turbo-clipproj4b-q4tp` | **1.9 s** | 84.5 s | 45.6 s | 5.9 s | **140 s** |
290
-
291
- Encoding is 26x faster, which is the part ClipProj actually replaced. The rest
292
- is residency: 13.16 GB leaves the decoder room on a 24 GB card where 23.47 GB
293
- does not, and the video VAE more than halves without a line of its own
294
- changing. On the 20 GB card this variant exists for, that gap is the whole
295
- point β€” a reader of this repo measured 20 minutes for one 512x288 clip on an
296
- RTX 3080 20 GB, paging the encoder in and out of a ~19 GB weight budget.
 
 
 
 
 
 
297
 
298
  The parity probe takes the device arm on both files at the same `rel rms
299
  4.65e-3`, which is its own small proof that the DiT came through untouched.
@@ -319,8 +326,8 @@ plain ground. A 0.92 cosine is close, not equal, and where it is not equal is
319
  the scene's furniture rather than its subject or its action.
320
 
321
  So the ordering is not "smaller is worse". The two-bit file answers a different
322
- question; this one answers the right question with a plainer set, for 56% of
323
- the size and half the time. If you have the VRAM for the 32B encoder, use it β€”
324
  its framing is richer. If you are paging, this is the better trade, and unlike
325
  the two-bit build it is a trade rather than a loss.
326
 
 
63
  |---|---|---|
64
  | `mmh3-turbo-fl2va-q4tp.cmf` | 23.94 GB | **use this one** β€” text-to-video AND keyframe-to-video |
65
  | `mmh3-turbo-q4tp.cmf` | 23.47 GB | the same without the vision tower: text-to-video only |
66
+ | `mmh3-turbo-clipproj4b-q4tp.cmf` | **13.16 GB** | text-to-video with a 4B stand-in prompt encoder. Peaks at 15.1 GB of VRAM, so a 20 GB card holds the whole run instead of paging it; still four bits everywhere. Costs some set dressing β€” see below |
67
  | `mmh3-turbo-fl2va-q2tp.cmf` | 18.74 GB | two bits on the gate/up planes. Smaller, faster, and it stops following the prompt β€” kept for anyone who wants to push on it, not for rendering. See below |
68
 
69
  ## Keyframe to video
 
280
 
281
  ### What it costs
282
 
283
+ The one number that transfers between machines is the **encode stage**, because
284
+ that is the only stage ClipProj replaced. Same prompt, seed and four steps,
285
+ 512x288x39, measured inside one run:
286
 
287
+ | | prompt encode |
288
+ |---|---|
289
+ | `mmh3-turbo-q4tp` β€” 32B tapped at 50 | 49.2 s |
290
+ | `mmh3-turbo-clipproj4b-q4tp` β€” 4B tapped at 24 | **1.9 s** |
291
+
292
+ Everything after that is residency, and residency is a property of YOUR card,
293
+ not of this file. On the 24 GB RTX 3090 these were taken on, the whole render
294
+ came out roughly half the time of the 23.47 GB file and the video VAE decode
295
+ more than halved β€” but totals on that box drifted between 140 s and 210 s for
296
+ one unchanged configuration as the card warmed, so treat the ratio as a
297
+ direction and measure your own.
298
+
299
+ The point is the 20 GB card this variant exists for, where the difference is
300
+ not a ratio but a cliff: a reader of this repo measured **20 minutes** for one
301
+ 512x288 clip on an RTX 3080 20 GB, paging the 23.47 GB file through a ~19 GB
302
+ weight budget. At 13.16 GB, with a polled whole-run peak of 15.1 GB of VRAM,
303
+ there is nothing left to page.
304
 
305
  The parity probe takes the device arm on both files at the same `rel rms
306
  4.65e-3`, which is its own small proof that the DiT came through untouched.
 
326
  the scene's furniture rather than its subject or its action.
327
 
328
  So the ordering is not "smaller is worse". The two-bit file answers a different
329
+ question; this one answers the right question with a plainer set, at 56% of
330
+ the size. If you have the VRAM for the 32B encoder, use it β€”
331
  its framing is richer. If you are paging, this is the better trade, and unlike
332
  the two-bit build it is a trade rather than a loss.
333