Computational Integration of Human Vision and Natural Language through Bitext Alignment

1Citations
Citations of this article
71Readers
Mendeley users who have this article in their library.

Abstract

Multimodal integration of visual and linguistic data is a longstanding but crucial challenge for modeling human understanding. We propose a framework that uses an unsupervised bitext alignment method to integrate visual and linguistic data. We present an empirical study of the various parameters of the framework. Our results exceed baselines using both exact and delayed temporal correspondence. The resulting alignments can be used for image classification and retrieval.

Cite

CITATION STYLE

APA

Vaidyanathan, P., Prud’hommeaux, E., Alm, C. O., & Pelz, J. B. (2015). Computational Integration of Human Vision and Natural Language through Bitext Alignment. In A Workshop of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015 - Workshop on Vision and Language 2015, VL 2015: Vision and Language Meet Cognitive Systems - Proceedings (pp. 4–5). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w15-2802

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free