SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

Zhaoxin Fan; Zhenbo Song; Hongyan Liu; Zhiwu Lu; Jun He; Xiaoyong Du

Conference ProceedingsOPEN ACCESS

SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

Proceedings of the 36th AAAI Conference on Artificial Intelligence, AAAI 2022 (2022) 36 551-560

DOI: 10.1609/aaai.v36i1.19934

72Citations

31Readers

Abstract

Simultaneous Localization and Mapping (SLAM) and Autonomous Driving are becoming increasingly more important in recent years. Point cloud-based large scale place recognition is the spine of them. While many models have been proposed and have achieved acceptable performance by learning short-range local features, they always skip long-range contextual properties. Moreover, the model size also becomes a serious shackle for their wide applications. To overcome these challenges, we propose a super light-weight network model termed SVT-Net. On top of the highly efficient 3D Sparse Convolution (SP-Conv), an Atom-based Sparse Voxel Transformer (ASVT) and a Cluster-based Sparse Voxel Transformer (CSVT) are proposed respectively to learn both short-range local features and long-range contextual features. Consisting of ASVT and CSVT, SVT-Net can achieve state-of-the-art performance in terms of both recognition accuracy and running speed with a super-light model size (0.9M parameters). Meanwhile, for the purpose of further boosting efficiency, we introduce two simplified versions, which also achieve state-of-the-art performance and further reduce the model size to 0.8M and 0.4M respectively.

Cite

CITATION STYLE

APA

Fan, Z., Song, Z., Liu, H., Lu, Z., He, J., & Du, X. (2022). SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition. In Proceedings of the 36th AAAI Conference on Artificial Intelligence, AAAI 2022 (Vol. 36, pp. 551–560). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/aaai.v36i1.19934

SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

Abstract

Cite

Register to see more suggestions