Representing sets of instances for visual recognition

2Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

In computer vision, a complex entity such as an image or video is often represented as a set of instance vectors, which are extracted from different parts of that entity. Thus, it is essential to design a representation to encode information in a set of instances robustly. Existing methods such as FV and VLAD are designed based on a generative perspective, and their performances fluctuate when difference types of instance vectors are used (i.e., they are not robust). The proposed D3 method effectively compares two sets as two distributions, and proposes a directional total variation distance (DTVD) to measure their dissimilarity. Furthermore, a robust classifier-based method is proposed to estimate DTVD robustly, and to efficiently represent these sets. D3 is evaluated in action and image recognition tasks. It achieves excellent robustness, accuracy and speed.

Cite

CITATION STYLE

APA

Wu, J., Gao, B. B., & Liu, G. (2016). Representing sets of instances for visual recognition. In 30th AAAI Conference on Artificial Intelligence, AAAI 2016 (pp. 2237–2243). AAAI press. https://doi.org/10.1609/aaai.v30i1.10184

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free