Abstract
We introduce Picturebook, a large-scale lookup operation to ground language via 'snapshots' of our physical world accessed through image search. For each word in a vocabulary, we extract the top-k images from Google image search and feed the images through a convolutional network to extract a word embedding. We introduce a multimodal gating function to fuse our Picturebook embeddings with other word representations. We also introduce Inverse Picturebook, a mechanism to map a Picturebook embedding back into words. We experiment and report results across a wide range of tasks: word similarity, natural language inference, semantic relatedness, sentiment/topic classification, image-sentence ranking and machine translation. We also show that gate activations corresponding to Picturebook embeddings are highly correlated to human judgments of concreteness ratings.
Cite
CITATION STYLE
Kiros, J. R., Chan, W., & Hinton, G. E. (2018). Illustrative language understanding: Large-scale visual grounding with image search. In ACL 2018 - 56th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers) (Vol. 1, pp. 922–933). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/p18-1085
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.