Abstract
End-to-end speech translation (ST) directly renders source language speech to the target language without intermediate automatic speech recognition (ASR) output as in a cascade approach. End-to-end ST avoids error propagation from intermediate ASR results. Although re- cent attempts have applied multi-task learning using an auxiliary task of ASR to improve ST performance, they use cross-entropy loss to one-hot references in the ASR task, and the trained ST models do not consider pos- sible ASR confusion. In this study, we propose a novel multi-task learning framework for end-to-end STs leveraged by ASR-based loss against pos- terior distributions obtained using a pre-trained ASR model called ASR posterior-based loss (ASR-PBL). The ASR-PBL method, which enables a ST model to reflect possible ASR confusion among competing hypotheses with similar pronunciations, can be applied to one of the strong multi-task ST baseline models with Hybrid CTC/Attention ASR task loss. In our experiments on the Fisher Spanish-to-English corpus, the proposed method demonstrated better BLEU results than the baseline that used standard CE loss.
Author supplied keywords
Cite
CITATION STYLE
Ko, Y., Sudoh, K., Sakti, S., & Nakamura, S. (2024). Neural End-To-End Speech Translation Leveraged by ASR Posterior Distribution. IEICE Transactions on Information and Systems, E107.D(10), 1322–1331. https://doi.org/10.1587/transinf.2023EDP7249
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.