Going Beyond Your Expectations in Latency Metrics for Simultaneous Speech Translation

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Current evaluation practices in Simultaneous Speech Translation (SimulST) systems typically involve segmenting the input audio and corresponding translations, calculating quality and latency metrics for each segment, and averaging the results. Although this approach may provide a reliable estimation of translation quality, it can lead to misleading values of latency metrics due to an inherent assumption that average latency values are good enough estimators of SimulST systems' response time. However, our detailed analysis of latency evaluations for state-of-the-art SimulST systems demonstrates that latency distributions are often skewed and subject to extreme variations. As a result, the mean in latency metrics fails to capture these anomalies, potentially masking the lack of robustness in some systems and metrics. In this paper, a thorough analysis of the results of systems submitted to recent editions of the IWSLT simultaneous track is provided to support our hypothesis and alternative ways to report latency metrics are proposed in order to provide a better understanding of SimulST systems' latency.

Cite

CITATION STYLE

APA

Iranzo-Sánchez, J., Iranzo-Sánchez, J., Giménez, A., & Civera, J. (2025). Going Beyond Your Expectations in Latency Metrics for Simultaneous Speech Translation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 18205–18228). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.937

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free