Abstract
To address pedestrian occlusion and multi-scale issues in crowded scenes, a crowded pedestrian detection algorithm that integrates global context information, spatial information, and pose information is proposed. Firstly, the feature pyramid network (FPN) is enhanced by integrating a dynamic sampling module and a weighted fusion branch, enabling more effective multi-scale feature aggregation and better handling of pedestrian scale variations. Secondly, a global context information module is designed to capture additional potential pedestrian features by fusing global context information and spatial information, which mitigates feature loss due to occlusion. Additionally, an enhanced Transformer-based pose module is constructed to extract pose information as supplementary resources, tackling the challenges of occlusion in crowded pedestrian detection. Since pose estimation is limited by pose estimation algorithms, an improved CA attention is incorporated into the visual information to complement the pose information. Finally, a novel data augmentation method is proposed to simulate occlusion, thereby improving the model's generalization ability. Extensive experimental evaluations on the CrowdHuman and CityPersons benchmarks corroborate the robustness and superior detection capabilities of the proposed algorithm, which consistently surpasses current state-of-the-art approaches in crowded pedestrian detection scenarios.
Author supplied keywords
Cite
CITATION STYLE
Chang, H., Li, M., Huang, S., & Li, T. (2025). GSPNet: A Crowded Pedestrian Detection Network Integrating Global Context, Spatial, and Pose Information. IEEE Access, 13, 149373–149389. https://doi.org/10.1109/ACCESS.2025.3599213
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.