non-intrusive-se / DATA_LICENSES.md
vvwangvv's picture
Add non-intrusive DS reference implementation: code, data, checkpoints, recipe
ea2f14c verified
|
Raw History Blame Contribute Delete
1.4 kB

Bundled data — sources and licenses

This repository bundles audio so the recipe is self-contained. Each dataset is redistributed under its original license, reproduced/attributed below. Because ESC-50 is CC BY-NC, the bundled artifact as a whole is for non-commercial / academic use.

Clean speech — LibriTTS (data/clean/)

  • Subset of LibriTTS: train-clean-100 (training), dev-clean (dev), test-clean (test). Each *.wav ships with its *.normalized.txt transcript.
  • Source: https://www.openslr.org/60/
  • License: CC BY 4.0.
  • Citation:

    H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, Y. Wu. "LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech," Interspeech 2019.

Noise — ESC-50 (data/noise/)

  • Full ESC-50 (2000 clips, 5 folds). Folds 1–3 used for training noise, fold 4 for dev, fold 5 for test (disjoint). meta/esc50.csv retained.
  • Source: https://github.com/karolpiczak/ESC-50
  • License: CC BY-NC 3.0 (non-commercial).
  • Citation:

    K. J. Piczak. "ESC: Dataset for Environmental Sound Classification," ACM Multimedia 2015.

ASR model — Whisper (not bundled)

  • openai/whisper-base.en is downloaded from the Hugging Face Hub at runtime (not redistributed here). License: MIT (OpenAI Whisper).

Code

  • The source code in ds4se/ and scripts/ is MIT-licensed — see LICENSE.