Image-Text Retrieval via Contrastive Learning with Auxiliary Generative Features and Support-set Regularization

7Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

In this paper, we bridge the heterogeneity gap between different modalities and improve image-text retrieval by taking advantage of auxiliary image-to-text and text-to-image generative features with contrastive learning. Concretely, contrastive learning is devised to narrow the distance between the aligned image-text pairs and push apart the distance between the unaligned pairs from both inter- and intra-modality perspectives with the help of cross-modal retrieval features and auxiliary generative features. In addition, we devise a support-set regularization term to further improve contrastive learning by constraining the distance between each image/text and its corresponding cross-modal support-set information contained in the same semantic category. To evaluate the effectiveness of the proposed method, we conduct experiments on three benchmark datasets (i.e., MIRFLICKR-25K, NUS-WIDE, MS COCO). Experimental results show that our model significantly outperforms the strong baselines for cross-modal image-text retrieval. For reproducibility, we submit the code and data publicly at: \urlhttps: //github.com/Hambaobao/CRCGS.

Cite

CITATION STYLE

APA

Zhang, L., Yang, M., Li, C., & Xu, R. (2022). Image-Text Retrieval via Contrastive Learning with Auxiliary Generative Features and Support-set Regularization. In SIGIR 2022 - Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 1938–1943). Association for Computing Machinery, Inc. https://doi.org/10.1145/3477495.3531783

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free