Abstract:Robotic manipulation requires accurate perception of the environment, which poses a significant challenge due to its inherent complexity and constantly changing nature. In this context, RGB image and point-cloud observations are two commonly used modalities in visual-based robotic manipulation, but each of these modalities have their own limitations. Commercial point-cloud observations often suffer from issues like sparse sampling and noisy output due to the limits of the emission-reception imaging principle. On the other hand, RGB images, while rich in texture information, lack essential depth and 3D information crucial for robotic manipulation. To mitigate these challenges, we propose an image-only robotic manipulation framework that leverages an eye-on-hand monocular camera installed on the robot's parallel gripper. By moving with the robot gripper, this camera gains the ability to actively perceive object from multiple perspectives during the manipulation process. This enables the estimation of 6D object poses, which can be utilized for manipulation. While, obtaining images from more and diverse viewpoints typically improves pose estimation, it also increases the manipulation time. To address this trade-off, we employ a reinforcement learning policy to synchronize the manipulation strategy with active perception, achieving a balance between 6D pose accuracy and manipulation efficiency. Our experimental results in both simulated and real-world environments showcase the state-of-the-art effectiveness of our approach. %, which, to the best of our knowledge, is the first to achieve robust real-world robotic manipulation through active pose estimation. We believe that our method will inspire further research on real-world-oriented robotic manipulation.

Learning Latent Object-Centric Representations for Visual-Based Robot Manipulation

Learning Robot Manipulation Skills from Human Demonstration Videos Using Two-Stream 2-D/3-D Residual Networks with Self-Attention

Dynamics Learning with Object-Centric Interaction Networks for Robot Manipulation

Vision-Based Robotic Object Grasping—A Deep Reinforcement Learning Approach

Vision-Based Categorical Object Pose Estimation and Manipulation.

Learning Robotic Manipulation through Visual Planning and Acting

Manipulate by Seeing: Creating Manipulation Controllers from Pre-Trained Representations

RLAfford: End-to-End Affordance Learning for Robotic Manipulation

RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation

Learning to Imagine Manipulation Goals for Robot Task Planning

Latent Space Planning for Multiobject Manipulation With Environment-Aware Relational Classifiers

Vision-based Manipulation from Single Human Video with Open-World Object Graphs

Latent Space Planning for Multi-Object Manipulation with Environment-Aware Relational Classifiers

LaSeSOM: A Latent and Semantic Representation Framework for Soft Object Manipulation

Human-oriented Representation Learning for Robotic Manipulation

Programmatically Grounded, Compositionally Generalizable Robotic Manipulation

Q-Attention: Enabling Efficient Learning for Vision-Based Robotic Manipulation

CORN: Contact-based Object Representation for Nonprehensile Manipulation of General Unseen Objects

Learning Manipulation by Predicting Interaction

ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation

Reinforcement Learning with Decoupled State Representation for Robot Manipulations