Abstract
With the surge in online video content, finding highlights and key video segments have garnered widespread attention. Given a textual query, video highlight detection (HD) and temporal grounding (TG) aim to predict frame-wise saliency scores from a video while concurrently locating all relevant spans. Despite recent progress in DETR-based works, these methods crudely fuse different inputs in the encoder, which limits effective cross-modal interaction. To solve this challenge, the authors design QD-Net (query-guided refinement and dynamic spans network) tailored for HD&TG. Specifically, they propose a query-guided refinement module to decouple the feature encoding from the interaction process. Furthermore, they present a dynamic span decoder that leverages learnable 2D spans as decoder queries, which accelerates training convergence for TG. On QVHighlights dataset, the proposed QD-Net achieves 61.87 HD-HIT@1 and 61.88 TG-mAP@0.5, yielding a significant improvement of +1.88 and +8.05, respectively, compared to the state-of-The-Art method.
Author supplied keywords
Cite
CITATION STYLE
Xu, Y., Sun, Y., Xie, Z., Zhai, B., Jia, Y., & Du, S. (2023). Query-guided refinement and dynamic spans network for video highlight detection and temporal grounding in online information systems. International Journal on Semantic Web and Information Systems, 19(1). https://doi.org/10.4018/IJSWIS.332768
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.