Instructions to use infosave/MiniMax-H3-Turbo-cmf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- cortiq
How to use infosave/MiniMax-H3-Turbo-cmf with cortiq:
# one Rust binary, no additional dependencies cargo install cortiq-cli # or a prebuilt binary from github.com/infosave2007/cmf/releases hf download infosave/MiniMax-H3-Turbo-cmf --include "*.cmf" --local-dir . ls *.cmf # some repos ship more than one quantization
cortiq animate FILE.cmf --prompt "a corgi in a chef hat flipping a pancake" --out clip.avi
- Notebooks
- Google Colab
- Kaggle
ClipProj section: scope the timings to the card they were taken on
Browse files
README.md
CHANGED
|
@@ -63,7 +63,7 @@ not ported.
|
|
| 63 |
|---|---|---|
|
| 64 |
| `mmh3-turbo-fl2va-q4tp.cmf` | 23.94 GB | **use this one** β text-to-video AND keyframe-to-video |
|
| 65 |
| `mmh3-turbo-q4tp.cmf` | 23.47 GB | the same without the vision tower: text-to-video only |
|
| 66 |
-
| `mmh3-turbo-clipproj4b-q4tp.cmf` | **13.16 GB** | text-to-video with a 4B stand-in prompt encoder. Peaks at 15.1 GB of VRAM, so a 20 GB card holds the whole run
|
| 67 |
| `mmh3-turbo-fl2va-q2tp.cmf` | 18.74 GB | two bits on the gate/up planes. Smaller, faster, and it stops following the prompt β kept for anyone who wants to push on it, not for rendering. See below |
|
| 68 |
|
| 69 |
## Keyframe to video
|
|
@@ -280,20 +280,27 @@ that the mechanism is wired right.
|
|
| 280 |
|
| 281 |
### What it costs
|
| 282 |
|
| 283 |
-
|
| 284 |
-
|
|
|
|
| 285 |
|
| 286 |
-
| | encode |
|
| 287 |
-
|---|---|
|
| 288 |
-
| `mmh3-turbo-q4tp`
|
| 289 |
-
| `mmh3-turbo-clipproj4b-q4tp`
|
| 290 |
-
|
| 291 |
-
|
| 292 |
-
|
| 293 |
-
|
| 294 |
-
|
| 295 |
-
|
| 296 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 297 |
|
| 298 |
The parity probe takes the device arm on both files at the same `rel rms
|
| 299 |
4.65e-3`, which is its own small proof that the DiT came through untouched.
|
|
@@ -319,8 +326,8 @@ plain ground. A 0.92 cosine is close, not equal, and where it is not equal is
|
|
| 319 |
the scene's furniture rather than its subject or its action.
|
| 320 |
|
| 321 |
So the ordering is not "smaller is worse". The two-bit file answers a different
|
| 322 |
-
question; this one answers the right question with a plainer set,
|
| 323 |
-
the size
|
| 324 |
its framing is richer. If you are paging, this is the better trade, and unlike
|
| 325 |
the two-bit build it is a trade rather than a loss.
|
| 326 |
|
|
|
|
| 63 |
|---|---|---|
|
| 64 |
| `mmh3-turbo-fl2va-q4tp.cmf` | 23.94 GB | **use this one** β text-to-video AND keyframe-to-video |
|
| 65 |
| `mmh3-turbo-q4tp.cmf` | 23.47 GB | the same without the vision tower: text-to-video only |
|
| 66 |
+
| `mmh3-turbo-clipproj4b-q4tp.cmf` | **13.16 GB** | text-to-video with a 4B stand-in prompt encoder. Peaks at 15.1 GB of VRAM, so a 20 GB card holds the whole run instead of paging it; still four bits everywhere. Costs some set dressing β see below |
|
| 67 |
| `mmh3-turbo-fl2va-q2tp.cmf` | 18.74 GB | two bits on the gate/up planes. Smaller, faster, and it stops following the prompt β kept for anyone who wants to push on it, not for rendering. See below |
|
| 68 |
|
| 69 |
## Keyframe to video
|
|
|
|
| 280 |
|
| 281 |
### What it costs
|
| 282 |
|
| 283 |
+
The one number that transfers between machines is the **encode stage**, because
|
| 284 |
+
that is the only stage ClipProj replaced. Same prompt, seed and four steps,
|
| 285 |
+
512x288x39, measured inside one run:
|
| 286 |
|
| 287 |
+
| | prompt encode |
|
| 288 |
+
|---|---|
|
| 289 |
+
| `mmh3-turbo-q4tp` β 32B tapped at 50 | 49.2 s |
|
| 290 |
+
| `mmh3-turbo-clipproj4b-q4tp` β 4B tapped at 24 | **1.9 s** |
|
| 291 |
+
|
| 292 |
+
Everything after that is residency, and residency is a property of YOUR card,
|
| 293 |
+
not of this file. On the 24 GB RTX 3090 these were taken on, the whole render
|
| 294 |
+
came out roughly half the time of the 23.47 GB file and the video VAE decode
|
| 295 |
+
more than halved β but totals on that box drifted between 140 s and 210 s for
|
| 296 |
+
one unchanged configuration as the card warmed, so treat the ratio as a
|
| 297 |
+
direction and measure your own.
|
| 298 |
+
|
| 299 |
+
The point is the 20 GB card this variant exists for, where the difference is
|
| 300 |
+
not a ratio but a cliff: a reader of this repo measured **20 minutes** for one
|
| 301 |
+
512x288 clip on an RTX 3080 20 GB, paging the 23.47 GB file through a ~19 GB
|
| 302 |
+
weight budget. At 13.16 GB, with a polled whole-run peak of 15.1 GB of VRAM,
|
| 303 |
+
there is nothing left to page.
|
| 304 |
|
| 305 |
The parity probe takes the device arm on both files at the same `rel rms
|
| 306 |
4.65e-3`, which is its own small proof that the DiT came through untouched.
|
|
|
|
| 326 |
the scene's furniture rather than its subject or its action.
|
| 327 |
|
| 328 |
So the ordering is not "smaller is worse". The two-bit file answers a different
|
| 329 |
+
question; this one answers the right question with a plainer set, at 56% of
|
| 330 |
+
the size. If you have the VRAM for the 32B encoder, use it β
|
| 331 |
its framing is richer. If you are paging, this is the better trade, and unlike
|
| 332 |
the two-bit build it is a trade rather than a loss.
|
| 333 |
|