Abstract
With the ubiquity of sensor-equipped smartphones, it is common to have multimedia documents uploaded to the Internet that have GPS coordinates associated with them. Utilizing such geotags as an additional feature is intuitively appealing for improving the performance of location-aware applications. However, raw GPS coordinates are fine-grained location indicators without any semantic information. Existing methods on geotag semantic encoding mostly extract hand-crafted, application-specific location representations that heavily depend on large-scale supplementary data and thus cannot perform efficiently on mobile devices. In this paper, we present a machine learning based approach, termed GPS2Vec+, which learns rich location representations by capitalizing on the world-wide geotagged images. Once trained, the model has no dependence on the auxiliary data anymore so it encodes geotags highly efficiently by inference. We extract visual and semantic knowledge from image content and user-generated tags, and transfer the information into locations by using geotagged images as a bridge. To adapt to different application domains, we further present an attention-based fusion framework that estimates the importance of the learnt location representations under different contexts for effective feature fusion. Our location representations yield significant performance improvements over the state-of-the-art geotag encoding methods on image classification and venue annotation.
Author supplied keywords
Cite
CITATION STYLE
Yin, Y., Zhang, Y., Liu, Z., Liang, Y., Wang, S., Shah, R. R., & Zimmermann, R. (2021). Learning Multi-context Aware Location Representations from Large-scale Geotagged Images. In MM 2021 - Proceedings of the 29th ACM International Conference on Multimedia (pp. 899–907). Association for Computing Machinery, Inc. https://doi.org/10.1145/3474085.3475268
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.