Neural End-To-End Speech Translation Leveraged by ASR Posterior Distribution

N/ACitations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

End-to-end speech translation (ST) directly renders source language speech to the target language without intermediate automatic speech recognition (ASR) output as in a cascade approach. End-to-end ST avoids error propagation from intermediate ASR results. Although re- cent attempts have applied multi-task learning using an auxiliary task of ASR to improve ST performance, they use cross-entropy loss to one-hot references in the ASR task, and the trained ST models do not consider pos- sible ASR confusion. In this study, we propose a novel multi-task learning framework for end-to-end STs leveraged by ASR-based loss against pos- terior distributions obtained using a pre-trained ASR model called ASR posterior-based loss (ASR-PBL). The ASR-PBL method, which enables a ST model to reflect possible ASR confusion among competing hypotheses with similar pronunciations, can be applied to one of the strong multi-task ST baseline models with Hybrid CTC/Attention ASR task loss. In our experiments on the Fisher Spanish-to-English corpus, the proposed method demonstrated better BLEU results than the baseline that used standard CE loss.

Cite

CITATION STYLE

APA

Ko, Y., Sudoh, K., Sakti, S., & Nakamura, S. (2024). Neural End-To-End Speech Translation Leveraged by ASR Posterior Distribution. IEICE Transactions on Information and Systems, E107.D(10), 1322–1331. https://doi.org/10.1587/transinf.2023EDP7249

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free