Illustrative language understanding: Large-scale visual grounding with image search

N/ACitations
Citations of this article
187Readers
Mendeley users who have this article in their library.

Abstract

We introduce Picturebook, a large-scale lookup operation to ground language via 'snapshots' of our physical world accessed through image search. For each word in a vocabulary, we extract the top-k images from Google image search and feed the images through a convolutional network to extract a word embedding. We introduce a multimodal gating function to fuse our Picturebook embeddings with other word representations. We also introduce Inverse Picturebook, a mechanism to map a Picturebook embedding back into words. We experiment and report results across a wide range of tasks: word similarity, natural language inference, semantic relatedness, sentiment/topic classification, image-sentence ranking and machine translation. We also show that gate activations corresponding to Picturebook embeddings are highly correlated to human judgments of concreteness ratings.

Cite

CITATION STYLE

APA

Kiros, J. R., Chan, W., & Hinton, G. E. (2018). Illustrative language understanding: Large-scale visual grounding with image search. In ACL 2018 - 56th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers) (Vol. 1, pp. 922–933). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/p18-1085

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free