Incorporating visual semantics into sentence representations within a grounded space

18Citations
Citations of this article
106Readers
Mendeley users who have this article in their library.

Abstract

Language grounding is an active field aiming at enriching textual representations with visual information. Generally, textual and visual elements are embedded in the same representation space, which implicitly assumes a one-to-one correspondence between modalities. This hypothesis does not hold when representing words, and becomes problematic when used to learn sentence representations - the focus of this paper - as a visual scene can be described by a wide variety of sentences. To overcome this limitation, we propose to transfer visual information to textual representations by learning an intermediate representation space: the grounded space. We further propose two new complementary objectives ensuring that (1) sentences associated with the same visual content are close in the grounded space and (2) similarities between related elements are preserved across modalities. We show that this model outperforms the previous state-of-the-art on classification and semantic relatedness tasks.

Cite

CITATION STYLE

APA

Bordes, P., Zablocki, É., Soulier, L., Piwowarski, B., & Gallinari, P. (2019). Incorporating visual semantics into sentence representations within a grounded space. In EMNLP-IJCNLP 2019 - 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, Proceedings of the Conference (pp. 696–707). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1064

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free