Abstract
Detecting discriminative semantic attributes from text which correlate with image features is one of the main challenges of zero-shot learning for fine-grained image classification. Particularly, using full-length encyclopedic articles as textual descriptions has had limited success, one reason being that such documents contain many non-visual or unrelated sentences. We propose a method to automatically extract visually relevant sentences from Wikipedia documents. Our model, based on a convolutional neural network, is robustly tested through ground truth labeling obtained via Amazon Mechanical Turk, achieving 81.73% F1 measure.
Cite
CITATION STYLE
Winn, O., Kidambi, M. K., & Muresan, S. (2016). Detecting visually relevant sentences for fine-grained classification. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 86–91). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w16-3213
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.