Does Speech Enhancement of Publicly Available Data Help Build Robust Speech Recognition Systems?

Bhavya Ghai; Buvana Ramanan; Klaus Mueller

Conference ProceedingsOPEN ACCESS

Does Speech Enhancement of Publicly Available Data Help Build Robust Speech Recognition Systems?

AAAI 2020 - 34th AAAI Conference on Artificial Intelligence (2020) 13793-13794

ArXiv: 1910.13488

0Citations

9Readers

Abstract

Automatic speech recognition(ASR) systems play a key role in many commercial products including voice assistants. Typically, they require large amounts of high quality speech data for training which gives an undue advantage to large organizations which have tons of private data. We investigated if speech data obtained from publicly available sources can be further enhanced to train better speech recognition models. We begin with noisy/contaminated speech data, apply speech enhancement to produce 'cleaned' version and use both the versions to train the ASR model. We have found that using speech enhancement gives 9.5% better word error rate than training on just the original noisy data and 9% better than training on just the ground truth 'clean' data. It's performance is also comparable to the ideal case scenario when trained on noisy and it's ground truth 'clean' version.

Cite

CITATION STYLE

APA

Ghai, B., Ramanan, B., & Mueller, K. (2020). Does Speech Enhancement of Publicly Available Data Help Build Robust Speech Recognition Systems? In AAAI 2020 - 34th AAAI Conference on Artificial Intelligence (pp. 13793–13794). AAAI press.

Does Speech Enhancement of Publicly Available Data Help Build Robust Speech Recognition Systems?

Abstract

Cite

Register to see more suggestions