infosave commited on
Commit
e14c9c9
Β·
verified Β·
1 Parent(s): 921dce2

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +12 -6
README.md CHANGED
@@ -58,8 +58,10 @@ give it a first and/or last frame and it continues from there. The release's
58
  third path β€” `ref2va`, conditioning on reference images, clips and audio β€” is
59
  not ported.
60
 
61
- One file: `mmh3-turbo-fl2va-q4tp.cmf`, 23.94 GB, every projection at
62
- four bits. A two-bit build exists and is **not** published β€” see below.
 
 
63
 
64
  ## Keyframe to video
65
 
@@ -166,15 +168,16 @@ RAM at least the file's size β€” 24 GB β€” or every step faults on non-resident
166
  pages; the weights are memory-mapped, not read. Disk: 24 GB. A GPU is optional
167
  and wants ~14 GB of VRAM for the DiT's planes. No network access at run time.
168
 
169
- ## Two bits was tried, and it is not shipped
170
 
171
  The obvious next cut is the DeepSeek-V4 policy: gate/up at two bits,
172
  everything else at four. It builds β€” 23.94 GB down to **18.74**, and
173
  with a device kernel of its own it renders *faster* than the four-bit
174
  file (217.3 s against 258.4, because there is less weight to move).
175
- `cortiq verify` passes.
176
 
177
- It also stops following the prompt.
 
178
 
179
  | | |
180
  |---|---|
@@ -193,7 +196,10 @@ PROMPT ENCODER, and the policy put two bits on its gate/up planes along
193
  with the DiT's β€” so the loss falls on the part that decides what the
194
  clip is about, not on the part that draws it. A two-bit build confined
195
  to the DiT would save ~2.9 GB instead of 5.2 and is the version worth
196
- measuring next. Until someone does, four bits is what ships.
 
 
 
197
 
198
  ## What it costs to run
199
 
 
58
  third path β€” `ref2va`, conditioning on reference images, clips and audio β€” is
59
  not ported.
60
 
61
+ | file | size | |
62
+ |---|---|---|
63
+ | `mmh3-turbo-fl2va-q4tp.cmf` | 23.94 GB | **use this one** |
64
+ | `mmh3-turbo-fl2va-q2tp.cmf` | 18.74 GB | two bits on the gate/up planes. Smaller, faster, and it stops following the prompt β€” kept for anyone who wants to push on it, not for rendering. See below |
65
 
66
  ## Keyframe to video
67
 
 
168
  pages; the weights are memory-mapped, not read. Disk: 24 GB. A GPU is optional
169
  and wants ~14 GB of VRAM for the DiT's planes. No network access at run time.
170
 
171
+ ## Two bits: smaller, faster, and answering a different question
172
 
173
  The obvious next cut is the DeepSeek-V4 policy: gate/up at two bits,
174
  everything else at four. It builds β€” 23.94 GB down to **18.74**, and
175
  with a device kernel of its own it renders *faster* than the four-bit
176
  file (217.3 s against 258.4, because there is less weight to move).
177
+ `cortiq verify` passes. The file is here.
178
 
179
+ It also stops following the prompt, which is why it is not the one to
180
+ reach for.
181
 
182
  | | |
183
  |---|---|
 
196
  with the DiT's β€” so the loss falls on the part that decides what the
197
  clip is about, not on the part that draws it. A two-bit build confined
198
  to the DiT would save ~2.9 GB instead of 5.2 and is the version worth
199
+ measuring next β€” the packer's policy is one predicate,
200
+ `is_wide_plane`, if you want to try it. The file above is published so
201
+ that experiment starts from something rather than nothing; four bits is
202
+ what to render with.
203
 
204
  ## What it costs to run
205