YOLOv8-PoseBoost: Advancements in Multimodal Robot Pose Keypoint Detection

Feng Wang,Gang Wang,Baoli Lu

DOI: https://doi.org/10.3390/electronics13061046

IF: 2.9

2024-03-12

Electronics

Abstract:In the field of multimodal robotics, achieving comprehensive and accurate perception of the surrounding environment is a highly sought-after objective. However, current methods still have limitations in motion keypoint detection, especially in scenarios involving small target detection and complex scenes. To address these challenges, we propose an innovative approach known as YOLOv8-PoseBoost. This method introduces the Channel Attention Module (CBAM) to enhance the network's focus on small targets, thereby increasing sensitivity to small target individuals. Additionally, we employ multiple scale detection heads, enabling the algorithm to comprehensively detect individuals of varying sizes in images. The incorporation of cross-level connectivity channels further enhances the fusion of features between shallow and deep networks, reducing the rate of missed detections for small target individuals. We also introduce a Scale Invariant Intersection over Union (SIoU) redefined bounding box regression localization loss function, which accelerates model training convergence and improves detection accuracy. Through a series of experiments, we validate YOLOv8-PoseBoost's outstanding performance in motion keypoint detection for small targets and complex scenes. This innovative approach provides an effective solution for enhancing the perception and execution capabilities of multimodal robots. It has the potential to drive the development of multimodal robots across various application domains, holding both theoretical and practical significance.

engineering, electrical & electronic,computer science, information systems,physics, applied

What problem does this paper attempt to address?

### What problem does this paper attempt to solve? This paper aims to address two main issues in multimodal robot pose keypoint detection: 1. **Small Object Detection Challenge**: Current methods have limitations in detecting small objects (such as small-sized pedestrians), especially in scenarios requiring fine operations or working in confined spaces. 2. **Keypoint Detection in Complex Scenes**: In dynamic and complex environments (such as multiple object interactions, changing lighting conditions, and occlusions), traditional pose keypoint detection algorithms struggle to achieve ideal performance. To tackle these challenges, the researchers propose the YOLOv8-PoseBoost method. This method enhances the network's ability to detect pose keypoints in small objects and complex scenes through the following innovative strategies: - **Introduction of Channel Attention Module (CBAM)**: Enhances the network's focus on small objects, increasing sensitivity to small-sized pedestrians without significantly increasing computational complexity. - **Multi-Scale Detection Head**: Integrates four different-sized detection heads, enabling the algorithm to comprehensively recognize targets of various sizes in the image. - **Cross-Layer Connection Channels**: Introduces two cross-layer communication channels to enhance feature fusion between shallow and deep networks, reducing the miss rate of small objects. - **Improved Bounding Box Regression Loss Function (SIoU)**: Introduces the Scale-Invariant Intersection over Union (SIoU) loss function, redefining the bounding box regression loss, accelerating training convergence, and improving detection accuracy. A series of experiments validated the outstanding performance of YOLOv8-PoseBoost in pose keypoint detection for small objects and complex scenes. This innovative method provides an effective solution to enhance the perception and execution capabilities of multimodal robots and is expected to promote the development of multimodal robots in various application fields.

YOLOv8-PoseBoost: Advancements in Multimodal Robot Pose Keypoint Detection

The application prospects of robot pose estimation technology: exploring new directions based on YOLOv8-ApexNet

RFA-YOLO-POSE: A Fusion Algorithm for Pose Detection and Object Identification Amidst Complex Crowds

PLOD-YOLO: Premium Lightweight Object Detection for Autonomous Following Robot

An improved YOLOv8 algorithm for small object detection in autonomous driving

SP-YOLOv8s: An Improved YOLOv8s Model for Remote Sensing Image Tiny Object Detection

YOLO Adaptive Developments in Complex Natural Environments for Tiny Object Detection

Mobile Robot Tracking Method Based on Improved YOLOv8 Pedestrian Detection Algorithm

A Two-Stage Monocular Vision Detection Method for 6D Pose Estimation in Multi-Heterogeneous Robot Systems

Small Target-YOLOv5: Enhancing the Algorithm for Small Object Detection in Drone Aerial Imagery Based on YOLOv5

I-YOLO: a novel single-stage framework for small object detection

Improved small-object detection using YOLOv8: A comparative study

YOLO-DRS: A Bioinspired Object Detection Algorithm for Remote Sensing Images Incorporating a Multi-Scale Efficient Lightweight Attention Mechanism

MC-YOLOv5: A Multi-Class Small Object Detection Algorithm

Research on a small target object detection method for aerial photography based on improved YOLOv7

YOLO‐RSFM: An efficient road small object detection method

KSL-POSE: A Real-Time 2D Human Pose Estimation Method Based on Modified YOLOv8-Pose Framework

Improvement and Enhancement of YOLOv5 Small Target Recognition Based on Multi-module Optimization

CF-YOLOX: An Autonomous Driving Detection Model for Multi-Scale Object Detection

MS-YOLO: integration-based multi-subnets neural network for object detection in aerial images

Subtle-YOLOv8: a detection algorithm for tiny and complex targets in UAV aerial imagery