Abstract
<title>Abstract</title> <p>In recent years, multimodal object detection has garnered significant attention due to its ability to enhance model accuracy and robustness. This is particularly valuable in applications such as autonomous driving, robotics, and UAV-based inspections, where the fusion of visible and infrared thermal imaging enables reliable target recognition under extreme environmental conditions. Although integrating multimodal information improves detection performance, it often comes at the cost of increased computational complexity, which can hinder efficiency. To address this challenge, we propose a lightweight multispectral object detection algorithm. The proposed method incorporates a Dual-feature lightweight backbone, a Skip-Connected channel attention module(SCCA), and a Multi-Scale Spatial Attention Module (MSSA), collectively forming a Detail-Contour Feature Interaction Aggregation Module. This architecture enables effective fusion and detection of information from different modalities, such as infrared and visible images. We validate our approach on several benchmark datasets, including LLVIP, FLIR, and DroneVehicle. Experimental results demonstrate that the proposed method achieves improved detection accuracy while maintaining low computational overhead, exhibiting strong effectiveness and generalization. Furthermore, it outperforms existing mainstream algorithms in terms of detection depth and precision.</p>