Generating image descriptions with gold standard visual inputs: Motivation, evaluation and baselines

8Citations
Citations of this article
69Readers
Mendeley users who have this article in their library.

Abstract

In this paper, we present the task of generating image descriptions with gold standard visual detections as input, rather than directly from an image. This allows the Natural Language Generation community to focus on the text generation process, rather than dealing with the noise and complications arising from the visual detection process. We propose a fine-grained evaluation metric specifically for evaluating the content selection capabilities of image description generation systems. To demonstrate the evaluation metric on the task, several baselines are presented using bounding box information and textual information as priors for content selection. The baselines are evaluated using the proposed metric, showing that the fine-grained metric is useful for evaluating the content selection phase of an image description generation system.

Cite

CITATION STYLE

APA

Wang, J., & Gaizauskas, R. (2015). Generating image descriptions with gold standard visual inputs: Motivation, evaluation and baselines. In ENLG 2015 - Proceedings of the 15th European Workshop on Natural Language Generation (pp. 117–126). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w15-4722

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free