Feature-Level Vehicle-Infrastructure Cooperative Perception with Adaptive Fusion for 3D Object Detection

3Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Highlights: What are the main findings? The proposed feature-level VICP framework consistently outperforms state-of-the-art baselines on the DAIR-V2X-C dataset, achieving higher (Formula presented.) and (Formula presented.). Experiments show that RFR delivers the largest gain, UWF improves robustness via adaptive uncertainty weighting, and CDCA enhances feature calibration. What is the implication of the main finding? Cooperative perception effectively overcomes occlusion and blind-spot limitations of vehicle-centric systems. The proposed model provides a reference for scalable and generalizable deployment of cooperative perception within smart city infrastructure. As vehicle-centric perception struggles with occlusion and dense traffic, vehicle-infrastructure cooperative perception (VICP) offers a viable route to extend sensing coverage and robustness. This study proposes a feature-level VICP framework that fuses vehicle- and roadside-derived visual features via V2X communication. The model integrates four components: regional feature reconstruction (RFR) for transferring region-specific roadside cues, context-driven channel attention (CDCA) for channel recalibration, uncertainty-weighted fusion (UWF) for confidence-guided weighting, and point sampling voxel fusion (PSVF) for efficient alignment. Evaluated on the DAIR-V2X-C benchmark, our method consistently outperforms state-of-the-art feature-level fusion baselines, achieving improved (Formula presented.) and (Formula presented.) (reported settings: 16.31% and 21.49%, respectively). Ablations show RFR provides the largest single-module gain +3.27% (Formula presented.) and +3.85% (Formula presented.), UWF yields substantial robustness gains, and CDCA offers modest calibration benefits. The framework enhances occlusion handling and cross-view detection while reducing dependence on explicit camera calibration, supporting more generalizable cooperative perception.

Cite

CITATION STYLE

APA

Yu, S., Peng, J., Wang, S., Wu, D., & Ma, C. (2025). Feature-Level Vehicle-Infrastructure Cooperative Perception with Adaptive Fusion for 3D Object Detection. Smart Cities, 8(5). https://doi.org/10.3390/smartcities8050171

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free