Incremental Text-to-Speech Synthesis Using Pseudo Lookahead with Large Pretrained Language Model

16Citations
Citations of this article
33Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

This letter presents an incremental text-to-speech (TTS) method that performs synthesis in small linguistic units while maintaining the naturalness of output speech. Incremental TTS is generally subject to a trade-off between latency and synthetic speech quality. It is challenging to produce high-quality speech with a low-latency setup that does not make much use of an unobserved future sentence (hereafter, 'lookahead'). To resolve this issue, we propose an incremental TTS method that uses a pseudo lookahead generated with a language model to take the future contextual information into account without increasing latency. Our method can be regarded as imitating a human's incremental reading and uses pretrained GPT2, which accounts for the large-scale linguistic knowledge, for the lookahead generation. Evaluation results show that our method 1) achieves higher speech quality than the method taking only observed information into account and 2) achieves a speech quality equivalent to waiting for the future context observation.

Cite

CITATION STYLE

APA

Saeki, T., Takamichi, S., & Saruwatari, H. (2021). Incremental Text-to-Speech Synthesis Using Pseudo Lookahead with Large Pretrained Language Model. IEEE Signal Processing Letters, 28, 857–861. https://doi.org/10.1109/LSP.2021.3073869

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free