Query-guided refinement and dynamic spans network for video highlight detection and temporal grounding in online information systems

3Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

With the surge in online video content, finding highlights and key video segments have garnered widespread attention. Given a textual query, video highlight detection (HD) and temporal grounding (TG) aim to predict frame-wise saliency scores from a video while concurrently locating all relevant spans. Despite recent progress in DETR-based works, these methods crudely fuse different inputs in the encoder, which limits effective cross-modal interaction. To solve this challenge, the authors design QD-Net (query-guided refinement and dynamic spans network) tailored for HD&TG. Specifically, they propose a query-guided refinement module to decouple the feature encoding from the interaction process. Furthermore, they present a dynamic span decoder that leverages learnable 2D spans as decoder queries, which accelerates training convergence for TG. On QVHighlights dataset, the proposed QD-Net achieves 61.87 HD-HIT@1 and 61.88 TG-mAP@0.5, yielding a significant improvement of +1.88 and +8.05, respectively, compared to the state-of-The-Art method.

Cite

CITATION STYLE

APA

Xu, Y., Sun, Y., Xie, Z., Zhai, B., Jia, Y., & Du, S. (2023). Query-guided refinement and dynamic spans network for video highlight detection and temporal grounding in online information systems. International Journal on Semantic Web and Information Systems, 19(1). https://doi.org/10.4018/IJSWIS.332768

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free