alfajgjK commited on
Commit
791f996
·
verified ·
1 Parent(s): 443f0f3

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +68 -0
README.md ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ library_name: onnx
4
+ tags:
5
+ - audio-to-audio
6
+ - audio-source-separation
7
+ - onnx
8
+ - onnxruntime-web
9
+ - webgpu
10
+ - demucs
11
+ ---
12
+
13
+ # htdemucs — ONNX graphs for onnxruntime-web (WebGPU)
14
+
15
+ Hybrid Transformer Demucs, exported to ONNX and prepared for **in-browser**
16
+ inference with [onnxruntime-web](https://onnxruntime.ai/) on the **WebGPU**
17
+ backend. Used by [stemstudio.jp](https://www.stemstudio.jp) for local,
18
+ upload-free stem separation.
19
+
20
+ Derived from [facebookresearch/demucs](https://github.com/facebookresearch/demucs)
21
+ (MIT). See `LICENSE` — the upstream notice travels with these files.
22
+
23
+ ## Files
24
+
25
+ | File | Transfer (gzip) | Inflated | Stems |
26
+ |---|---|---|---|
27
+ | `htdemucs_ft.opt.fp16.istft.onnx.gzbin` | 141,591,461 | 172,746,888 | vocals / instrumental (fine-tuned) |
28
+ | `htdemucs_4s.opt.fp16.istft.onnx.gzbin` | 141,601,329 | 172,746,888 | vocals / drums / bass / other |
29
+ | `htdemucs_6s.opt.fp16.istft.onnx.gzbin` | 114,203,693 | 142,573,723 | + guitar / piano |
30
+
31
+ ## Two things that are easy to get wrong
32
+
33
+ **1. The `.gzbin` extension is deliberate.** These are raw gzip streams stored
34
+ as ordinary files, and the browser inflates them with
35
+ `DecompressionStream('gzip')`. They are *not* named `.gz`, because some hosts
36
+ and dev servers see that extension and add `content-encoding: gzip`, at which
37
+ point the transport decompresses them and the application decompresses them
38
+ again — which fails with `incorrect header check`. Detect gzip by the leading
39
+ magic bytes (`1f 8b`), not by header or extension.
40
+
41
+ The compression is in the file rather than in the transport because model hosts
42
+ generally serve `.onnx` uncompressed regardless of `Accept-Encoding`.
43
+
44
+ **2. Build order matters: optimize *before* fp16, not after.** Running the
45
+ graph optimizer on an fp16 graph overflows folded constants to `inf` and every
46
+ output becomes `NaN`. Worse, `nan > tolerance` is false, so a naive comparison
47
+ reports "max diff 0.000e+00 / OK" on a completely broken model. The order used
48
+ here is: fp32 export → graph optimization → fp16 weights → inverse STFT folded
49
+ into the graph.
50
+
51
+ ## Why the optimized graph is WebGPU-only
52
+
53
+ Measured on one 7.8 s segment, same model, same host:
54
+
55
+ | | wasm heap | wasm infer | WebGPU session build |
56
+ |---|---|---|---|
57
+ | unoptimized | 1,223 MB | 16.8 s | 15.6 s |
58
+ | optimized | **3,251 MB** | 18.1 s | **1.4 s** |
59
+
60
+ On WebGPU, session build gets ~11× faster. On wasm, heap grows 2.7× and
61
+ inference gets *slower*. "Ship the optimized one and everyone gets faster" is
62
+ false — pick per execution provider.
63
+
64
+ ## Training data
65
+
66
+ The upstream weights were trained on MUSDB18-HQ and additional sources. The
67
+ MIT license covers the code and these derived graphs; the dataset carries its
68
+ own terms, which are worth checking for your own use.