Multi-modal object detection via transformer network

9Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.
Get full text

Abstract

According to the fact that single-modal data usually contain limited information, a great deal of effort has been devoted to making use of the complementary information contained in the multi-modal data on various patterns. Thus, this paper is concerned with an object detection method that can fully utilize multi-modal data. First, the method introduces the transformer mechanism to realize the fusion of intra-modal and inter-modal features of different modal data. The aim is to take advantage of the complementarity of data between modalities, which helps to improve the performance of multi-modal object detection. Second, a contrastive loss suitable for contrastive learning is applied. This enables the authors to effectively utilize label information. Extensive experiments are conducted on multiple object detection datasets to demonstrate the effectiveness of our proposed method.

Cite

CITATION STYLE

APA

Liu, W., Wang, H., Gao, Q., & Zhu, Z. (2023). Multi-modal object detection via transformer network. IET Image Processing, 17(12), 3541–3550. https://doi.org/10.1049/ipr2.12884

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free