Personalizing TTS voices for progressive dysarthria

5Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Amyotrophic lateral sclerosis (ALS) patients experience progressive speech deterioration due to muscle paralysis, leading to eventual loss of verbal communication capability. Text-to-speech synthesis (TTS) is an important technology for speech generating devices, enabling users to communicate using generic electronic voices, but often without the vocal identity of the users. Our work is aimed at personalizing TTS voices for people with ALS induced dysarthria by integrating machine learning and speech processing techniques of voice conversion (VC) and TTS. This is challenging as only small quantities of dysarthric speech are available from individual patients. Our system includes both timbre and prosody conversion for VC, neural TTS to generate TTS speech, and neural feature converter to interface VC and TTS. We collected speech data from 4 ALS target speakers with mild to severe dysarthria. Subjective listening tests showed that on average, our approach improved speech intelligibility by about 72% over the target speakers' speech, the converted voice was 2 to 3 times more similar to ALS targets than to TTS sources, and the converted speech quality was in the MOS scale of fair to good.

Cite

CITATION STYLE

APA

Zhao, Y., Song, M., Yue, Y., & Kuruvilla-Dugdale, M. (2021). Personalizing TTS voices for progressive dysarthria. In BHI 2021 - 2021 IEEE EMBS International Conference on Biomedical and Health Informatics, Proceedings. Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/BHI50953.2021.9508522

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free