Buckets:
36.5 GB
77 files
Updated 17 days ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| data | 75 items | ||
| .gitattributes | 2.31 kB xet | b6a9e0dd | |
| README.md | 1.87 kB xet | 6e017127 |
This dataset contains the Spanish-English track of the benchmark from ICASSP 2024: Zero Resource Code-Switched Speech Benchmark Using Speech Utterance Pairs for Multiple Spoken Languages.
Though the benchmark is originally designed to assess the semantic and syntactic abilities of the speech foundation models, you can also use this dataset for code-switching ASR.
If you find this dataset helpful, please consider to cite the following paper:
@INPROCEEDINGS{10446737,
author={Huang, Kuan-Po and Yang, Chih-Kai and Fu, Yu-Kuan and Dunbar, Ewan and Lee, Hung-Yi},
booktitle={ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
title={Zero Resource Code-Switched Speech Benchmark Using Speech Utterance Pairs for Multiple Spoken Languages},
year={2024},
volume={},
number={},
pages={10006-10010},
keywords={Speech coding;Benchmark testing;Signal processing;Linguistics;Acoustics;Speech processing;Task analysis;Code-switch;Multilingual;Discrete unit;Zero resource;Self-supervised},
doi={10.1109/ICASSP48485.2024.10446737}}
- Total size
- 36.5 GB
- Files
- 77
- Last updated
- Aug 28
- Pre-warmed CDN
- US EU US EU