Abstract:Vehicle detection with optical remote sensing images has become widely applied in recent years. However, the following challenges have remained unsolved during remote sensing vehicle target detection. These challenges include the dense and arbitrary angles at which vehicles are distributed and which make it difficult to detect them; the extensive model parameter (Param) that blocks real-time detection; the large differences between larger vehicles in terms of their features, which lead to a reduced detection precision; and the way in which the distribution in vehicle datasets is unbalanced and thus not conducive to training. First, this paper constructs a small dataset of vehicles, MiVehicle. This dataset includes 3000 corresponding infrared and visible image pairs, offering a more balanced distribution. In the infrared part of the dataset, the proportions of different vehicle types are as follows: cars, 48%; buses, 19%; trucks, 15%; freight, cars 10%; and vans, 8%. Second, we choose the rotated box mechanism for detection with the model and we build a new vehicle detector, ML-Det, with a novel multi-scale feature fusion triple cross-criss FPN (TCFPN), which can effectively capture the vehicle features in three different positions with an mAP improvement of 1.97%. Moreover, we propose LKC–INVO, which allows involution to couple the structure of multiple large kernel convolutions, resulting in an mAP increase of 2.86%. We also introduce a novel C2F_ContextGuided module with global context perception, which enhances the perception ability of the model in the global scope and minimizes model Params. Eventually, we propose an assemble–disperse attention module to aggregate local features so as to improve the performance. Overall, ML-Det achieved a 3.22% improvement in accuracy while keeping Params almost unchanged. In the self-built small MiVehicle dataset, we achieved 70.44% on visible images and 79.12% on infrared images with 20.1 GFLOPS, 78.8 FPS, and 7.91 M. Additionally, we trained and tested our model on the following public datasets: UAS-AOD and DOTA. ML-Det was found to be ahead of many other advanced target detection algorithms.

3D-DETNet: a Single Stage Video-Based Vehicle Detector

A Multi-view 3D Vehicle Detection Method Based On Novel 3D Proposal Generation Method

Vehicle Behavior Recognition using Multi-Stream 3D Convolutional Neural Network

Multiclass objects detection algorithm using DarkNet-53 and DenseNet for intelligent vehicles

SuperDet: An Efficient Single-Shot Network for Vehicle Detection in Remote Sensing Images

3D Vehicle Detection Using Multi-Level Fusion From Point Clouds and Images

Multiscale Attention and Feature Decomposition Network for Surveillance Vehicle Detection

Image Guidance Based 3D Vehicle Detection in Traffic Scene.

DALDet: Depth-Aware Learning Based Object Detection for Autonomous Driving

Monocular 3-D Vehicle Detection Using a Cascade Network for Autonomous Driving

Real-Time Vehicle Detection Framework Based on the Fusion of LiDAR and Camera

3D Fully Convolutional Network for Vehicle Detection in Point Cloud

A multi-task Faster R-CNN method for 3D vehicle detection based on a single image

A Multi-Scale Feature Fusion Based Lightweight Vehicle Target Detection Network on Aerial Optical Images

PA3DNet: 3-D Vehicle Detection with Pseudo Shape Segmentation and Adaptive Camera-LiDAR Fusion

3D car-detection based on a Mobile Deep Sensor Fusion Model and real-scene applications

3D Detection for Occluded Vehicles From Point Clouds

DV3-IBi_YOLOv5s: A Lightweight Backbone Network and Multiscale Neck Network Vehicle Detection Algorithm

Sparse Embedded Convolution Based Dual Feature Aggregation 3D Object Detection Network

6DoF-3D: Efficient and accurate 3D object detection using six degrees-of-freedom for autonomous driving

CFENet: an Accurate and Efficient Single-Shot Object Detector for Autonomous Driving.