Abstract
Simultaneous speech translation (SimulST) systems must balance translation quality with response time, making latency measurement crucial for evaluating their real-world performance. However, there has been a longstanding belief that current metrics yield unrealistically high latency measurements in unsegmented streaming settings. In this paper, we investigate this phenomenon, revealing its root cause in a fundamental misconception underlying existing latency evaluation approaches. We demonstrate that this issue affects not only streaming but also segment-level latency evaluation across different metrics. Furthermore, we propose a modification to correctly measure computation-aware latency for SimulST systems, addressing the limitations present in existing metrics.
Cite
CITATION STYLE
Xu, X., Xu, W., Ouyang, S., & Li, L. (2025). CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation. In 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Proceedings of the Conference Findings, NAACL 2025 (pp. 7077–7082). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-naacl.393
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.