Enhancing cross-modal retrieval based on modality-specific and embedding spaces

N/ACitations
Citations of this article
8Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

A new approach that drastically improves cross-modal retrieval performance in vision and language (hereinafter referred to as 'vision and language retrieval') is proposed in this paper. Vision and language retrieval takes data of one modality as a query to retrieve relevant data of another modality, and it enables flexible retrieval across different modalities. Most of the existing methods learn optimal embeddings of visual and lingual information to a single common representation space. However, we argue that the forced embedding optimization results in loss of key information for sentences and images. In this paper, we propose an effective utilization of representation spaces in a simple but robust vision and language retrieval method. The proposed method makes use of multiple individual representation spaces through text-to-image and image-to-text models. Experimental results showed that the proposed approach enhances the performance of existing methods that embed visual and lingual information to a single common representation space.

Cite

CITATION STYLE

APA

Yanagi, R., Togo, R., Ogawa, T., & Haseyama, M. (2020). Enhancing cross-modal retrieval based on modality-specific and embedding spaces. IEEE Access, 8, 96777–96786. https://doi.org/10.1109/ACCESS.2020.2995815

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free