infosave commited on
Commit
8c5652c
·
verified ·
1 Parent(s): ab31ab0

card: point the opening at the 32-step soul render, add the Russian one

Browse files
Files changed (1) hide show
  1. README.md +43 -11
README.md CHANGED
@@ -17,16 +17,27 @@ tags:
17
 
18
  ```bash
19
  cortiq music minimax-music3-q4tp.cmf \
20
- --prompt "bpm is 92, key is E minor. Electric blues rock, gritty slide \
21
- guitar, walking bass, brushed drums, warm analog production." \
 
 
22
  --lyrics "[verse]
23
- I woke up on a dusty road" \
24
- --seconds 15 --steps 8 --seed 42 --out song.wav
 
 
 
25
  ```
26
 
27
- **[Listen to that command's output.](https://huggingface.co/infosave/MiniMax-Music-3-cmf/resolve/main/samples/blues_15s.wav)**
28
  One binary, one file, no Python.
29
 
 
 
 
 
 
 
30
  [MiniMax-Music-3](https://huggingface.co/MiniMaxAI/MiniMax-Music3)
31
  generates music with vocals from a caption and lyrics. This is its
32
  [Comfy-Org repackage](https://huggingface.co/Comfy-Org/MiniMax-Music-3)
@@ -162,12 +173,33 @@ the source:
162
 
163
  ### What it costs
164
 
165
- Measured on an Apple M4, 5 s at 8 steps, 206 s total: AR 0.45 s/frame,
166
- denoise 6.5 s/step over 430 latent frames, vocoder 19 s. The 15-second
167
- sample above took 823 s.
168
 
169
- The vocoder was 75 s before it was handed the thread pool, which is a
170
- quarter of a short render for a one-word change.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
171
 
172
  ## Running it
173
 
@@ -176,7 +208,7 @@ cargo install cortiq-cli # 0.5.74+
176
  hf download infosave/MiniMax-Music-3-cmf minimax-music3-q4tp.cmf --local-dir .
177
  cortiq verify minimax-music3-q4tp.cmf # → ✓ all tensor hashes match
178
  cortiq music minimax-music3-q4tp.cmf --prompt "..." --lyrics "..." \
179
- --seconds 15 --steps 8 --seed 42 --out song.wav
180
  ```
181
 
182
  `--seconds` is a ceiling: the model can stop earlier. Same seed, same
 
17
 
18
  ```bash
19
  cortiq music minimax-music3-q4tp.cmf \
20
+ --prompt "Classic 1960s soul, passionate male tenor with rich vibrato, \
21
+ lush female backing vocals, gospel choir harmonies, vintage Motown and \
22
+ Stax atmosphere, groovy bassline, warm Hammond organ, horn section, \
23
+ analog tape saturation, romantic nighttime mood." \
24
  --lyrics "[verse]
25
+ Baby when the midnight comes around
26
+ I still hear your footsteps on the ground
27
+ [chorus]
28
+ Oh, come back to me" \
29
+ --seconds 20 --steps 32 --seed 7 --out song.wav
30
  ```
31
 
32
+ **[Listen to that command's output.](https://huggingface.co/infosave/MiniMax-Music-3-cmf/resolve/main/samples/soul_20s_32steps.wav)**
33
  One binary, one file, no Python.
34
 
35
+ The lyrics are not English-only — the same command
36
+ [in Russian](https://huggingface.co/infosave/MiniMax-Music-3-cmf/resolve/main/samples/russian_20s_32steps.wav),
37
+ caption and all. And steps buy audible quality: the
38
+ [same soul prompt at 16](https://huggingface.co/infosave/MiniMax-Music-3-cmf/resolve/main/samples/soul_16steps.wav)
39
+ is where the sibilance stops being distracting, 32 is where it settles.
40
+
41
  [MiniMax-Music-3](https://huggingface.co/MiniMaxAI/MiniMax-Music3)
42
  generates music with vocals from a caption and lyrics. This is its
43
  [Comfy-Org repackage](https://huggingface.co/Comfy-Org/MiniMax-Music-3)
 
173
 
174
  ### What it costs
175
 
176
+ Every render prints where its time went. On an RTX 3090, 4 s at 4
177
+ steps:
 
178
 
179
+ | stage | GPU | CPU only |
180
+ |---|---|---|
181
+ | AR | 52.3 s | 44.2 s |
182
+ | denoise | 32.9 s | 36.2 s |
183
+ | vocoder | **18.2 s** | 31.3 s |
184
+ | total | **103.5 s** | 111.8 s |
185
+
186
+ The vocoder was 54.0 s until its convolutions stopped shipping their
187
+ column matrix across the bus — the host built it, transposed it into a
188
+ second buffer of the same size and uploaded that, up to 2.37 GB for a
189
+ 20-second song, when the input it expands from is `k` times smaller.
190
+ It is expanded on the card now.
191
+
192
+ Two things that table will not tell you. The CPU column is a 256-core
193
+ EPYC, so an ordinary machine's fallback is far slower than this and the
194
+ device gap far wider. And the shape that matters for real songs is not
195
+ this one: attention over latent frames is quadratic, so a 20-second
196
+ render at 32 steps spends about **80% of its time in the denoise**,
197
+ around 32-40 s a step. `CMF_MUSIC3_PROF=1` splits a step four ways if
198
+ you want to see it.
199
+
200
+ Earlier, on an Apple M4, 5 s at 8 steps took 206 s: AR 0.45 s/frame,
201
+ denoise 6.5 s/step over 430 latent frames, vocoder 19 s — and the
202
+ vocoder was 75 s before it was handed the thread pool.
203
 
204
  ## Running it
205
 
 
208
  hf download infosave/MiniMax-Music-3-cmf minimax-music3-q4tp.cmf --local-dir .
209
  cortiq verify minimax-music3-q4tp.cmf # → ✓ all tensor hashes match
210
  cortiq music minimax-music3-q4tp.cmf --prompt "..." --lyrics "..." \
211
+ --seconds 20 --steps 32 --seed 7 --out song.wav
212
  ```
213
 
214
  `--seconds` is a ceiling: the model can stop earlier. Same seed, same