LibriMix: An Open-Source Dataset for Generalizable Speech Separation

نشر في Joris Cosentino بتاريخ 2020 في مجال هندسة إلكترونية والبحث باللغة English تحميل البحث

الملخص بالإنكليزية

In recent years, wsj0-2mix has become the reference dataset for single-channel speech separation. Most deep learning-based speech separation models today are benchmarked on it. However, recent studies have shown important performance drops when models trained on wsj0-2mix are evaluated on other, similar datasets. To address this generalization issue, we created LibriMix, an open-source alternative to wsj0-2mix, and to its noisy extension, WHAM!. Based on LibriSpeech, LibriMix consists of two- or three-speaker mixtures combined with ambient noise samples from WHAM!. Using Conv-TasNet, we achieve competitive performance on all LibriM

تحميل البحث