Efficient nearest neighbor search by removing anti-hub

2Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The central research question of the nearest neighbor search is how to reduce the memory cost while maintaining its accuracy. Instead of compressing each vector as is done in the existing methods, we propose a way to subsample unnecessary vectors to save memory. We empirically found that such unnecessary vectors have low hubness scores and thus can be easily identified beforehand. Such points are called anti-hubs in the data mining community. By removing anti-hubs, we achieved a memory-efficient search while preserving accuracy. In million-scale experiments, we showed that any vector compression method improves search accuracy by partial replacement with anti-hub removal under the same memory usage. A billion-scale benchmark showed that our data reduction combined with the best search method achieves higher accuracy under the assumption of fixed memory consumption. For example, our method had a much higher recall@100 (0.53) compared with the existing method (0.23) for the same memory consumption (6GB).

Cite

CITATION STYLE

APA

Tanaka, K., Matsui, Y., & Satoh, S. (2021). Efficient nearest neighbor search by removing anti-hub. In ICMR 2021 - Proceedings of the 2021 International Conference on Multimedia Retrieval (pp. 285–293). Association for Computing Machinery, Inc. https://doi.org/10.1145/3460426.3463622

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free