Abstract
Recent advancements in multimodal object detection have predominantly relied on end-to-end training paradigms, which, while effective, demand substantial computational resources and risk feature degradation. To address these challenges, we propose a frozen backbone paradigm, preserving pretrained representations as stable semantic anchors for efficient multimodal fusion. Our approach introduces a lightweight multi-receptive field attention (MRFA) mechanism, enhancing feature interaction and representation diversity without exhaustive retraining. Experiments on the FLIR Aligned and M3FD dataset demonstrate consistent improvements over state-of-the-art end-to-end models, highlighting the potential of pretrained backbones coupled with adaptive attention mechanisms for robust multimodal object detection. The project code is released at https://github.com/LuBingyu11/MRFA.
Author supplied keywords
Cite
CITATION STYLE
Lu, B., Liu, H., & Watanabe, H. (2026). Enhancing RGB-IR object detection: a frozen backbone approach with multi-receptive field attention. Visual Computer, 42(3). https://doi.org/10.1007/s00371-026-04378-1
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.