Abstract:Robotic manipulation requires accurate perception of the environment, which poses a significant challenge due to its inherent complexity and constantly changing nature. In this context, RGB image and point-cloud observations are two commonly used modalities in visual-based robotic manipulation, but each of these modalities have their own limitations. Commercial point-cloud observations often suffer from issues like sparse sampling and noisy output due to the limits of the emission-reception imaging principle. On the other hand, RGB images, while rich in texture information, lack essential depth and 3D information crucial for robotic manipulation. To mitigate these challenges, we propose an image-only robotic manipulation framework that leverages an eye-on-hand monocular camera installed on the robot's parallel gripper. By moving with the robot gripper, this camera gains the ability to actively perceive object from multiple perspectives during the manipulation process. This enables the estimation of 6D object poses, which can be utilized for manipulation. While, obtaining images from more and diverse viewpoints typically improves pose estimation, it also increases the manipulation time. To address this trade-off, we employ a reinforcement learning policy to synchronize the manipulation strategy with active perception, achieving a balance between 6D pose accuracy and manipulation efficiency. Our experimental results in both simulated and real-world environments showcase the state-of-the-art effectiveness of our approach. %, which, to the best of our knowledge, is the first to achieve robust real-world robotic manipulation through active pose estimation. We believe that our method will inspire further research on real-world-oriented robotic manipulation.

Bridging the Robot Perception Gap with Mid-Level Vision

Bridging the Robot Perception Gap with Mid-Level Vision

Robustifying Semantic Cognition of Traversability Across Wearable RGB-depth Cameras

Unifying Terrain Awareness Through Real-Time Semantic Segmentation

Real-time 3D Semantic Scene Perception for Egocentric Robots with Binocular Vision

A Universal Semantic-Geometric Representation for Robotic Manipulation

Active Scene Understanding via Online Semantic Reconstruction

Bridging Visual Perception with Contextual Semantics for Understanding Robot Manipulation Tasks

Active Perception with A Monocular Camera for Multiscopic Vision

Robot Manipulation in Salient Vision through Referring Image Segmentation and Geometric Constraints

3D Move to See: Multi-perspective visual servoing for improving object views with semantic segmentation

A Mobile Robot Visual SLAM System With Enhanced Semantics Segmentation

A survey of image semantics-based visual simultaneous localization and mapping: Application-oriented solutions to autonomous navigation of mobile robots

RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation

FusionSense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction

ImageManip: Image-based Robotic Manipulation with Affordance-guided Next View Selection

Efficient Bi-manipulation using RGBD Multi-model Fusion based on Attention Mechanism

Robust Perception-based Visual Simultaneous Localization and Tracking in Dynamic Environments

Indoor Semantic Scene Understanding using Multi-modality Fusion

RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter

Robotic Grasping With Multi-View Image Acquisition and Model-Based Pose Estimation