A Transformer-based System for Action Spotting in Soccer Videos

22Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Action Spotting in the broadcast soccer game is important to understand salient actions and video summary applications. In this paper, we propose an efficient transformer-based system for action spotting in soccer videos. We first use the multi-scale vision transformer to extract features from the videos. Then we adopt a sliding window strategy to further utilize temporal features and enhanced temporal understanding. Finally, the features are input to NetVLAD++ model to obtain the final results. Our model can learn a hierarchy of robust representations and perform well in the Action Spotting Task of SoccerNet Challenge 2022. Our method achieves excellent results and outperforms the baseline and previous published works.

Cite

CITATION STYLE

APA

Zhu, H., Liang, J., Lin, C., Zhang, J., & Hu, J. (2022). A Transformer-based System for Action Spotting in Soccer Videos. In MMSports 2022 - Proceedings of the 5th International ACM Workshop on Multimedia Content Analysis in Sports (pp. 103–109). Association for Computing Machinery, Inc. https://doi.org/10.1145/3552437.3555693

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free