Abstract:With the rapid growth in demand for security surveillance, assisted driving, and remote sensing, object detection networks with robust environmental perception and high detection accuracy have become a research focus. However, single-modality image detection technologies face limitations in environmental adaptability, often affected by factors such as lighting conditions, fog, rain, and obstacles like vegetation, leading to information loss and reduced detection accuracy. We propose an object detection network that integrates features from visible light and infrared images—IV-YOLO—to address these challenges. This network is based on YOLOv8 (You Only Look Once v8) and employs a dual-branch fusion structure that leverages the complementary features of infrared and visible light images for target detection. We designed a Bidirectional Pyramid Feature Fusion structure (Bi-Fusion) to effectively integrate multimodal features, reducing errors from feature redundancy and extracting fine-grained features for small object detection. Additionally, we developed a Shuffle-SPP structure that combines channel and spatial attention to enhance the focus on deep features and extract richer information through upsampling. Regarding model optimization, we designed a loss function tailored for multi-scale object detection, accelerating the convergence speed of the network during training. Compared with the current state-of-the-art Dual-YOLO model, IV-YOLO achieves mAP improvements of 2.8%, 1.1%, and 2.2% on the Drone Vehicle, FLIR, and KAIST datasets, respectively. On the Drone Vehicle and FLIR datasets, IV-YOLO has a parameter count of 4.31 M and achieves a frame rate of 203.2 fps, significantly outperforming YOLOv8n (5.92 M parameters, 188.6 fps on the Drone Vehicle dataset) and YOLO-FIR (7.1 M parameters, 83.3 fps on the FLIR dataset), which had previously achieved the best performance on these datasets. This demonstrates that IV-YOLO achieves higher real-time detection performance while maintaining lower parameter complexity, making it highly promising for applications in autonomous driving, public safety, and beyond.

Multimodal Feature Fusion YOLOv5 for RGB-T Object Detection

MMYFnet: Multi-Modality YOLO Fusion Network for Object Detection in Remote Sensing Images

ACDF-YOLO: Attentive and Cross-Differential Fusion Network for Multimodal Remote Sensing Object Detection

Multispectral Object Detection Based on Multilevel Feature Fusion and Dual Feature Modulation

An object detection algorithm based on infrared-visible dual modal feature fusion

DEYOLO: Dual-Feature-Enhancement YOLO for Cross-Modality Object Detection

Unified Information Fusion Network for Multi-Modal RGB-D and RGB-T Salient Object Detection

MMLF: Multi-modal Multi-class Late Fusion for Object Detection with Uncertainty Estimation

IV-YOLO: A Lightweight Dual-Branch Object Detection Network

CMIFDF: A lightweight cross-modal image fusion and weight-sharing object detection network framework

MFFNet: Multi-modal Feature Fusion Network for V-D-T Salient Object Detection

TF-YOLO: A Transformer–Fusion-Based YOLO Detector for Multimodal Pedestrian Detection in Autonomous Driving Scenes

Multi-Dimensional Information Fusion You Only Look Once Network for Suspicious Object Detection in Millimeter Wave Images

A multi‐modal fusion YoLo network for traffic detection

Discriminative unimodal feature selection and fusion for RGB-D salient object detection

RGB-X Object Detection via Scene-Specific Fusion Modules

SMFF-YOLO: A Scale-Adaptive YOLO Algorithm with Multi-Level Feature Fusion for Object Detection in UAV Scenes

MSF-YOLO: A multi-scale features fusion-based method for small object detection

Improving Object Detection in YOLOv8n with the C2f-f Module and Multi-Scale Fusion Reconstruction

Object Detection by Channel and Spatial Exchange for Multimodal Remote Sensing Imagery

A Lightweight YOLO Object Detection Algorithm Based on Bidirectional Multi‐Scale Feature Enhancement