Abstract
Understanding the multi-scale visual information in a video is essential for Video Question Answering (VideoQA). Therefore, we propose a novel Multi-Scale Progressive Attention Network (MSPAN) to achieve relational reasoning between cross-scale video information. We construct clips of different lengths to represent different scales of the video. Then, the cliplevel features are aggregated into node features by using max-pool, and a graph is generated for each scale of clips. For cross-scale feature interaction, we design a message passing strategy between adjacent scale graphs, i.e., topdown scale interaction and bottom-up scale interaction. Under the question's guidance of progressive attention, we realize the fusion of all-scale video features. Experimental evaluations on three benchmarks: TGIF-QA, MSVDQA and MSRVTT-QA show our method has achieved state-of-the-art performance.
Cite
CITATION STYLE
Guo, Z., Zhao, J., Jiao, L., Liu, X., & Li, L. (2021). Multi-Scale Progressive Attention Network for Video Question Answering. In ACL-IJCNLP 2021 - 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Proceedings of the Conference (Vol. 2, pp. 973–978). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2021.acl-short.122
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.