End-to-End Active Speaker Detection

Juan León Alcázar; Moritz Cordes; Chen Zhao; Bernard Ghanem

Conference Proceedings

End-to-End Active Speaker Detection

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2022) 13697 LNCS 126-143

DOI: 10.1007/978-3-031-19836-6_8

6Citations

19Readers

Get full text

Abstract

Recent advances in the Active Speaker Detection (ASD) problem build upon a two-stage process: feature extraction and spatio-temporal context aggregation. In this paper, we propose an end-to-end ASD workflow where feature learning and contextual predictions are jointly learned. Our end-to-end trainable network simultaneously learns multi-modal embeddings and aggregates spatio-temporal context. This results in more suitable feature representations and improved performance in the ASD task. We also introduce interleaved graph neural network (iGNN) blocks, which split the message passing according to the main sources of context in the ASD problem. Experiments show that the aggregated features from the iGNN blocks are more suitable for ASD, resulting in state-of-the art performance. Finally, we design a weakly-supervised strategy, which demonstrates that the ASD problem can also be approached by utilizing audiovisual data but relying exclusively on audio annotations. We achieve this by modelling the direct relationship between the audio signal and the possible sound sources (speakers), as well as introducing a contrastive loss.

Cite

CITATION STYLE

APA

Alcázar, J. L., Cordes, M., Zhao, C., & Ghanem, B. (2022). End-to-End Active Speaker Detection. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 13697 LNCS, pp. 126–143). Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/978-3-031-19836-6_8

End-to-End Active Speaker Detection

Abstract

Cite

Register to see more suggestions