Multi-Scale Progressive Attention Network for Video Question Answering

29Citations
Citations of this article
61Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Understanding the multi-scale visual information in a video is essential for Video Question Answering (VideoQA). Therefore, we propose a novel Multi-Scale Progressive Attention Network (MSPAN) to achieve relational reasoning between cross-scale video information. We construct clips of different lengths to represent different scales of the video. Then, the cliplevel features are aggregated into node features by using max-pool, and a graph is generated for each scale of clips. For cross-scale feature interaction, we design a message passing strategy between adjacent scale graphs, i.e., topdown scale interaction and bottom-up scale interaction. Under the question's guidance of progressive attention, we realize the fusion of all-scale video features. Experimental evaluations on three benchmarks: TGIF-QA, MSVDQA and MSRVTT-QA show our method has achieved state-of-the-art performance.

Cite

CITATION STYLE

APA

Guo, Z., Zhao, J., Jiao, L., Liu, X., & Li, L. (2021). Multi-Scale Progressive Attention Network for Video Question Answering. In ACL-IJCNLP 2021 - 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Proceedings of the Conference (Vol. 2, pp. 973–978). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2021.acl-short.122

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free