Cross-modal Representation Learning and Relation Reasoning for Bidirectional Adaptive Manipulation

6Citations
Citations of this article
4Readers
Mendeley users who have this article in their library.

Abstract

Since single-modal controllable manipulation typically requires supervision of information from other modalities or cooperation with complex software and experts, this paper addresses the problem of cross-modal adaptive manipulation (CAM). The novel task performs cross-modal semantic alignment from mutual supervision and implements bidirectional exchange of attributes, relations, or objects in parallel, benefiting both modalities while significantly reducing manual effort. We introduce a robust solution for CAM, which includes two essential modules, namely Heterogeneous Representation Learning (HRL) and Cross-modal Relation Reasoning (CRR). The former is designed to perform representation learning for cross-modal semantic alignment on heterogeneous graph nodes. The latter is adopted to identify and exchange the focused attributes, relations, or objects in both modalities. Our method produces pleasing cross-modal outputs on CUB and Visual Genome.

Cite

CITATION STYLE

APA

Li, L., Fan, K., & Yuan, C. (2022). Cross-modal Representation Learning and Relation Reasoning for Bidirectional Adaptive Manipulation. In IJCAI International Joint Conference on Artificial Intelligence (pp. 3222–3228). International Joint Conferences on Artificial Intelligence. https://doi.org/10.24963/ijcai.2022/447

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free