Dense video captioning using unsupervised semantic information

4Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We introduce a method to learn unsupervised semantic visual information based on the premise that complex events can be decomposed into simpler events and that these simple events are shared across several complex events. We first employ a clustering method to group representations producing a visual codebook. Then, we learn a dense representation by encoding the co-occurrence probability matrix for the codebook entries. This representation leverages the performance of the dense video captioning task in a scenario with only visual features. For example, we replace the audio signal in the BMT method and produce temporal proposals with comparable performance. Furthermore, we concatenate the visual representation with our descriptor in a vanilla transformer method to achieve state-of-the-art performance in the captioning subtask compared to the methods that explore only visual features, as well as a competitive performance with multi-modal methods. Our code is available at https://github.com/valterlej/dvcusi.

Cite

CITATION STYLE

APA

Estevam, V., Laroca, R., Pedrini, H., & Menotti, D. (2025). Dense video captioning using unsupervised semantic information. Journal of Visual Communication and Image Representation, 107. https://doi.org/10.1016/j.jvcir.2024.104385

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free