Cross-Modal Discrete Representation Learning

24Citations
Citations of this article
104Readers
Mendeley users who have this article in their library.

Abstract

In contrast to recent advances focusing on high-level representation learning across modalities, in this work we present a self-supervised learning framework that is able to learn a representation that captures finer levels of granularity across different modalities such as concepts or events represented by visual objects or spoken words. Our framework relies on a discretized embedding space created via vector quantization that is shared across different modalities. Beyond the shared embedding space, we propose a Cross-Modal Code Matching objective that forces the representations from different views (modalities) to have a similar distribution over the discrete embedding space such that cross-modal objects/actions localization can be performed without direct supervision. We show that the proposed discretized multi-modal fine-grained representation (e.g., pixel/word/frame) can complement high-level summary representations (e.g., video/sentence/waveform) for improved performance on cross-modal retrieval tasks. We also observe that the discretized representation uses individual clusters to represent the same semantic concept across modalities.

Cite

CITATION STYLE

APA

Liu, A. H., Jin, S. Y., Lai, C. I. J., Rouditchenko, A., Oliva, A., & Glass, J. (2022). Cross-Modal Discrete Representation Learning. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 3013–3035). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.acl-long.215

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free