Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,68 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
library_name: onnx
|
| 4 |
+
tags:
|
| 5 |
+
- audio-to-audio
|
| 6 |
+
- audio-source-separation
|
| 7 |
+
- onnx
|
| 8 |
+
- onnxruntime-web
|
| 9 |
+
- webgpu
|
| 10 |
+
- demucs
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# htdemucs — ONNX graphs for onnxruntime-web (WebGPU)
|
| 14 |
+
|
| 15 |
+
Hybrid Transformer Demucs, exported to ONNX and prepared for **in-browser**
|
| 16 |
+
inference with [onnxruntime-web](https://onnxruntime.ai/) on the **WebGPU**
|
| 17 |
+
backend. Used by [stemstudio.jp](https://www.stemstudio.jp) for local,
|
| 18 |
+
upload-free stem separation.
|
| 19 |
+
|
| 20 |
+
Derived from [facebookresearch/demucs](https://github.com/facebookresearch/demucs)
|
| 21 |
+
(MIT). See `LICENSE` — the upstream notice travels with these files.
|
| 22 |
+
|
| 23 |
+
## Files
|
| 24 |
+
|
| 25 |
+
| File | Transfer (gzip) | Inflated | Stems |
|
| 26 |
+
|---|---|---|---|
|
| 27 |
+
| `htdemucs_ft.opt.fp16.istft.onnx.gzbin` | 141,591,461 | 172,746,888 | vocals / instrumental (fine-tuned) |
|
| 28 |
+
| `htdemucs_4s.opt.fp16.istft.onnx.gzbin` | 141,601,329 | 172,746,888 | vocals / drums / bass / other |
|
| 29 |
+
| `htdemucs_6s.opt.fp16.istft.onnx.gzbin` | 114,203,693 | 142,573,723 | + guitar / piano |
|
| 30 |
+
|
| 31 |
+
## Two things that are easy to get wrong
|
| 32 |
+
|
| 33 |
+
**1. The `.gzbin` extension is deliberate.** These are raw gzip streams stored
|
| 34 |
+
as ordinary files, and the browser inflates them with
|
| 35 |
+
`DecompressionStream('gzip')`. They are *not* named `.gz`, because some hosts
|
| 36 |
+
and dev servers see that extension and add `content-encoding: gzip`, at which
|
| 37 |
+
point the transport decompresses them and the application decompresses them
|
| 38 |
+
again — which fails with `incorrect header check`. Detect gzip by the leading
|
| 39 |
+
magic bytes (`1f 8b`), not by header or extension.
|
| 40 |
+
|
| 41 |
+
The compression is in the file rather than in the transport because model hosts
|
| 42 |
+
generally serve `.onnx` uncompressed regardless of `Accept-Encoding`.
|
| 43 |
+
|
| 44 |
+
**2. Build order matters: optimize *before* fp16, not after.** Running the
|
| 45 |
+
graph optimizer on an fp16 graph overflows folded constants to `inf` and every
|
| 46 |
+
output becomes `NaN`. Worse, `nan > tolerance` is false, so a naive comparison
|
| 47 |
+
reports "max diff 0.000e+00 / OK" on a completely broken model. The order used
|
| 48 |
+
here is: fp32 export → graph optimization → fp16 weights → inverse STFT folded
|
| 49 |
+
into the graph.
|
| 50 |
+
|
| 51 |
+
## Why the optimized graph is WebGPU-only
|
| 52 |
+
|
| 53 |
+
Measured on one 7.8 s segment, same model, same host:
|
| 54 |
+
|
| 55 |
+
| | wasm heap | wasm infer | WebGPU session build |
|
| 56 |
+
|---|---|---|---|
|
| 57 |
+
| unoptimized | 1,223 MB | 16.8 s | 15.6 s |
|
| 58 |
+
| optimized | **3,251 MB** | 18.1 s | **1.4 s** |
|
| 59 |
+
|
| 60 |
+
On WebGPU, session build gets ~11× faster. On wasm, heap grows 2.7× and
|
| 61 |
+
inference gets *slower*. "Ship the optimized one and everyone gets faster" is
|
| 62 |
+
false — pick per execution provider.
|
| 63 |
+
|
| 64 |
+
## Training data
|
| 65 |
+
|
| 66 |
+
The upstream weights were trained on MUSDB18-HQ and additional sources. The
|
| 67 |
+
MIT license covers the code and these derived graphs; the dataset carries its
|
| 68 |
+
own terms, which are worth checking for your own use.
|