Abstract:Vehicle detection with optical remote sensing images has become widely applied in recent years. However, the following challenges have remained unsolved during remote sensing vehicle target detection. These challenges include the dense and arbitrary angles at which vehicles are distributed and which make it difficult to detect them; the extensive model parameter (Param) that blocks real-time detection; the large differences between larger vehicles in terms of their features, which lead to a reduced detection precision; and the way in which the distribution in vehicle datasets is unbalanced and thus not conducive to training. First, this paper constructs a small dataset of vehicles, MiVehicle. This dataset includes 3000 corresponding infrared and visible image pairs, offering a more balanced distribution. In the infrared part of the dataset, the proportions of different vehicle types are as follows: cars, 48%; buses, 19%; trucks, 15%; freight, cars 10%; and vans, 8%. Second, we choose the rotated box mechanism for detection with the model and we build a new vehicle detector, ML-Det, with a novel multi-scale feature fusion triple cross-criss FPN (TCFPN), which can effectively capture the vehicle features in three different positions with an mAP improvement of 1.97%. Moreover, we propose LKC–INVO, which allows involution to couple the structure of multiple large kernel convolutions, resulting in an mAP increase of 2.86%. We also introduce a novel C2F_ContextGuided module with global context perception, which enhances the perception ability of the model in the global scope and minimizes model Params. Eventually, we propose an assemble–disperse attention module to aggregate local features so as to improve the performance. Overall, ML-Det achieved a 3.22% improvement in accuracy while keeping Params almost unchanged. In the self-built small MiVehicle dataset, we achieved 70.44% on visible images and 79.12% on infrared images with 20.1 GFLOPS, 78.8 FPS, and 7.91 M. Additionally, we trained and tested our model on the following public datasets: UAS-AOD and DOTA. ML-Det was found to be ahead of many other advanced target detection algorithms.

Transformer-based vehicle detection for surveillance images

A Transformer-Based Object Detector with Coarse-Fine Crossing Representations

Multiclass objects detection algorithm using DarkNet-53 and DenseNet for intelligent vehicles

Dense Vehicle Counting Estimation Via a Synergism Attention Network

A Vehicle Detection Method Based on an Improved U-YOLO Network for High-Resolution Remote-Sensing Images

DS-Trans: A 3D Object Detection Method Based on a Deformable Spatiotemporal Transformer for Autonomous Vehicles

A Fast and Accurate Real-Time Vehicle Detection Method Using Deep Learning for Unconstrained Environments

DAFV: A Unified and Real-Time Framework of Joint Detection and Attributes Recognition for Fast Vehicles

Dense-TNT: Efficient Vehicle Type Classification Neural Network Using Satellite Imagery

A Multi-Scale Feature Fusion Based Lightweight Vehicle Target Detection Network on Aerial Optical Images

TF-YOLO: A Transformer–Fusion-Based YOLO Detector for Multimodal Pedestrian Detection in Autonomous Driving Scenes

Real-time Vehicle Detection under Complex Road Conditions

Fast Automatic Vehicle Detection In Uav Images Using Convolutional Neural Networks

FPGA-Based Vehicle Detection and Tracking Accelerator

A Method Based on Multi-Convolution Layers Joint and Generative Adversarial Networks for Vehicle Detection

Enhancing YOLO for occluded vehicle detection with grouped orthogonal attention and dense object repulsion

Research on dense object detection methods in congested environments of urban streets and roads based on DCYOLO

Research on Microscale Vehicle Logo Detection Based on Real-Time DEtection TRansformer (RT-DETR)

Improved object detection method for unmanned driving based on Transformers

Occlusion-Aware Detection for Internet of Vehicles in Urban Traffic Sensing Systems

CODAN: Counting-driven Attention Network for Vehicle Detection in Congested Scenes