Abstract
Detecting small and occluded objects in complex traffic scenes remains a major challenge for intelligent transportation systems, including the difficulty in balancing high accuracy and low computational cost. To address these challenges, this paper proposes an innovative MSF-YOLOv11n model. With a modular architecture design, we propose three innovative technical frameworks. First, the Adaptive Attention-Enhanced Wavelet Module (AAEWM) performs multi-scale wavelet decomposition combined with attention-guided reconstruction, enabling the model to capture structural information such as edges and textures while suppressing noise interference. Second, the Feature-Enhanced Attention Upsampling Module (FEAUM) integrates lightweight Ghost convolution with Shuffle Attention and employs a hybrid interpolation and convolution strategy, which preserves fine-grained details in the upsampling process. Third, the Multi-Scale Fusion Lightweight Triple Attention Module (MFTAM) employs depthwise separable convolutions and a triple attention mechanism across channel, spatial and positional dimensions, effectively enhancing multi-scale feature interaction and improving small-object localization accuracy. Compared with YOLOv11n, MSF-YOLOv11n improves mAP50-95 by 3.8% and mAP50 by 3.7% on UA-DETRAC, and gains 3.6% in mAP50-95 and 3.9% in mAP50 on DAIR-V2X, all with a computational cost of only 6.5 GFLOPs and an inference speed of 100 FPS on UA-DETRAC. These results demonstrate that MSF-YOLOv11n achieves a strong balance of accuracy and efficiency for real-world traffic perception.
Author supplied keywords
Cite
CITATION STYLE
Xu, S., Sun, Z., Xu, H., Huang, W., Chen, Q., & He, H. (2026). MSF-YOLOv11n: A Multi-Scale Feature Fusion-Based Efficient Detector for Small and Occluded Object Detection in Complex Traffic Scenes. IEEE Access, 14, 81515–81534. https://doi.org/10.1109/ACCESS.2026.3656395
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.