infosave commited on
Commit
7419377
·
verified ·
1 Parent(s): 70d3fcb

the q8_2f build's host-path cost, measured

Browse files
Files changed (1) hide show
  1. README.md +30 -1
README.md CHANGED
@@ -56,7 +56,7 @@ One transformer denoises picture and sound together in one packed sequence.
56
  | `mmh3-turbo-q4tp.cmf` | 23.47 GB | 24 GB+ | no | same, without the vision tower |
57
  | **`mmh3-turbo-clipproj4b-q4tp.cmf`** | **13.16 GB** | **16–20 GB** | no | **the small one**, peaks at 15.1 GB |
58
  | **`mmh3-turbo-clipproj4b-fl2va-q4tp.cmf`** | **14.48 GB** | **16–24 GB** | **yes** | the small one WITH start/end frames — the 4B vision tower and the VAE encoder join the compact build |
59
- | **`mmh3-turbo-clipproj4b-fl2va-v2-q8_2f.cmf`** | **26.90 GB** | **32 GB+** | **yes** | **eight bits**: the two-field int8, `w = q·row[o]·col[i]`. The step up from four bits for a machine with the memory |
60
  | `mmh3-turbo-fl2va-q2tp.cmf` | 18.74 GB | — | yes | don't render with this |
61
 
62
  Text-to-video everywhere; the `fl2va` files also take a first and/or last frame
@@ -615,6 +615,35 @@ Two more, on the runtime side: the latent upscaler published for H3
615
  (a 345 M-parameter 3-D conv net) is not ported, and there is no fused
616
  device kernel for adapter branches on this model — see below.
617
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
618
  ## The latent upscaler
619
 
620
  Render small, resize the **latent**, decode once:
 
56
  | `mmh3-turbo-q4tp.cmf` | 23.47 GB | 24 GB+ | no | same, without the vision tower |
57
  | **`mmh3-turbo-clipproj4b-q4tp.cmf`** | **13.16 GB** | **16–20 GB** | no | **the small one**, peaks at 15.1 GB |
58
  | **`mmh3-turbo-clipproj4b-fl2va-q4tp.cmf`** | **14.48 GB** | **16–24 GB** | **yes** | the small one WITH start/end frames — the 4B vision tower and the VAE encoder join the compact build |
59
+ | **`mmh3-turbo-clipproj4b-fl2va-v2-q8_2f.cmf`** | **26.90 GB** | **32 GB+** | **yes** | **eight bits**: the two-field int8, `w = q·row[o]·col[i]`. More weight fidelity — but see the caveat below, it renders on the **host** today |
60
  | `mmh3-turbo-fl2va-q2tp.cmf` | 18.74 GB | — | yes | don't render with this |
61
 
62
  Text-to-video everywhere; the `fl2va` files also take a first and/or last frame
 
615
  (a 345 M-parameter 3-D conv net) is not ported, and there is no fused
616
  device kernel for adapter branches on this model — see below.
617
 
618
+ ## The eight-bit build, and what it costs today
619
+
620
+ `mmh3-turbo-clipproj4b-fl2va-v2-q8_2f.cmf` (26.90 GB, 26.21 B parameters,
621
+ `cortiq verify` clean) packs the DiT as **`q8_2f`** — the two-field int8,
622
+ `w = q·row[o]·col[i]`: eight bits with a second scale field along the *input*
623
+ axis, which is where an activation-outlier channel shows up from the weight
624
+ side. As a codec it is strictly more faithful than the four-bit ladder.
625
+
626
+ **The caveat, measured rather than guessed: it renders on the host.** The
627
+ engine says so itself at startup —
628
+
629
+ ```
630
+ mmh3 GPU parity probe: no q4tp qkv tensor — host path
631
+ ```
632
+
633
+ — because every device path here reads the four-bit tiled layout: the fused
634
+ qkv/attention/output kernel, the packed FFN, all of it. `q8_2f` has a matvec
635
+ on both backends and no *matmat*, which is what a video DiT needs. Measured on
636
+ an RTX 5090, 512×288, 22 frames, four steps: **357.3 s** (denoise 198.8,
637
+ video VAE 151.6) with the card idle and sixteen cores busy, against **60.2 s**
638
+ for the q4tp file at 39 frames on the same class of card.
639
+
640
+ Until that kernel lands: take this file to compare codecs or to render
641
+ overnight on a big-memory machine, and take the q4tp file when you want the
642
+ GPU.
643
+
644
+ That is the honest state of it, and the missing piece is one kernel, not a
645
+ redesign.
646
+
647
  ## The latent upscaler
648
 
649
  Render small, resize the **latent**, decode once: