An analysis of action recognition datasets for language and vision tasks

7Citations
Citations of this article
117Readers
Mendeley users who have this article in their library.

Abstract

A large amount of recent research has focused on tasks that combine language and vision, resulting in a proliferation of datasets and methods. One such task is action recognition, whose applications include image annotation, scene understanding and image retrieval. In this survey, we categorize the existing approaches based on how they conceptualize this problem and provide a detailed review of existing datasets, highlighting their diversity as well as advantages and disadvantages. We focus on recently developed datasets which link visual information with linguistic resources and provide a fine-grained syntactic and semantic analysis of actions in images.

Cite

CITATION STYLE

APA

Gella, S., & Keller, F. (2017). An analysis of action recognition datasets for language and vision tasks. In ACL 2017 - 55th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers) (Vol. 2, pp. 64–71). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/P17-2011

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free