Abstract
In this paper, we present the task of generating image descriptions with gold standard visual detections as input, rather than directly from an image. This allows the Natural Language Generation community to focus on the text generation process, rather than dealing with the noise and complications arising from the visual detection process. We propose a fine-grained evaluation metric specifically for evaluating the content selection capabilities of image description generation systems. To demonstrate the evaluation metric on the task, several baselines are presented using bounding box information and textual information as priors for content selection. The baselines are evaluated using the proposed metric, showing that the fine-grained metric is useful for evaluating the content selection phase of an image description generation system.
Cite
CITATION STYLE
Wang, J., & Gaizauskas, R. (2015). Generating image descriptions with gold standard visual inputs: Motivation, evaluation and baselines. In ENLG 2015 - Proceedings of the 15th European Workshop on Natural Language Generation (pp. 117–126). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w15-4722
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.