Abstract
Multimodal integration of visual and linguistic data is a longstanding but crucial challenge for modeling human understanding. We propose a framework that uses an unsupervised bitext alignment method to integrate visual and linguistic data. We present an empirical study of the various parameters of the framework. Our results exceed baselines using both exact and delayed temporal correspondence. The resulting alignments can be used for image classification and retrieval.
Cite
CITATION STYLE
Vaidyanathan, P., Prud’hommeaux, E., Alm, C. O., & Pelz, J. B. (2015). Computational Integration of Human Vision and Natural Language through Bitext Alignment. In A Workshop of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015 - Workshop on Vision and Language 2015, VL 2015: Vision and Language Meet Cognitive Systems - Proceedings (pp. 4–5). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w15-2802
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.