D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement

Yixuan Wang,Mingtong Zhang,Zhuoran Li,Tarik Kelestemur,Katherine Driggs-Campbell,Jiajun Wu,Li Fei-Fei,Yunzhu Li

2024-10-17

Abstract:Scene representation is a crucial design choice in robotic manipulation systems. An ideal representation is expected to be 3D, dynamic, and semantic to meet the demands of diverse manipulation tasks. However, previous works often lack all three properties simultaneously. In this work, we introduce D$^3$Fields -- dynamic 3D descriptor fields. These fields are implicit 3D representations that take in 3D points and output semantic features and instance masks. They can also capture the dynamics of the underlying 3D environments. Specifically, we project arbitrary 3D points in the workspace onto multi-view 2D visual observations and interpolate features derived from visual foundational models. The resulting fused descriptor fields allow for flexible goal specifications using 2D images with varied contexts, styles, and instances. To evaluate the effectiveness of these descriptor fields, we apply our representation to rearrangement tasks in a zero-shot manner. Through extensive evaluation in real worlds and simulations, we demonstrate that D$^3$Fields are effective for zero-shot generalizable rearrangement tasks. We also compare D$^3$Fields with state-of-the-art implicit 3D representations and show significant improvements in effectiveness and efficiency.

Robotics,Computer Vision and Pattern Recognition,Machine Learning

What problem does this paper attempt to address?

The problem that this paper attempts to solve is how to implement a scene representation method with 3D, dynamic and semantic characteristics simultaneously in robotic manipulation tasks to meet diverse manipulation requirements. Specifically, existing works often fail to satisfy these three characteristics at the same time. Some representation methods exist in 3D space but ignore semantic information; others focus on dynamic modeling but only consider 2D data and overlook the role of 3D space; still others are limited to considering the semantic information of object instances and classes. Therefore, this paper proposes D3Fields - a dynamic 3D descriptor field, aiming to overcome these limitations and provide a representation method that can capture 3D geometric structures, dynamic changes and semantic information simultaneously, thus supporting the rearrangement tasks of zero - sample generalization. D3Fields achieves this by mapping any 3D point to multi - view 2D visual observations and interpolating to extract features from them. These fused descriptor fields allow the flexible specification of targets using 2D images with different contexts, styles and instances, and are especially suitable for rearrangement tasks that require zero - sample generalization ability. The paper shows significant improvements in the effectiveness and efficiency of D3Fields through extensive evaluations in real - world and simulated environments, especially in handling zero - sample generalization rearrangement tasks.

D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement

3D-SSD: Learning Hierarchical Features from RGB-D Images for Amodal 3D Object Detection

A Unified Feature Representation and Learning Framework for 3D Shape

GenDP: 3D Semantic Fields for Category-Level Generalizable Diffusion Policy

DM-NeRF: 3D Scene Geometry Decomposition and Manipulation from 2D Images

Diorama: Unleashing Zero-shot Single-view 3D Scene Modeling

DDF-HO: Hand-Held Object Reconstruction via Conditional Directed Distance Field

PACA: Perspective-Aware Cross-Attention Representation for Zero-Shot Scene Rearrangement

Learning Part-aware 3D Representations by Fusing 2D Gaussians and Superquadrics

MSGField: A Unified Scene Representation Integrating Motion, Semantics, and Geometry for Robotic Manipulation

SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation

Neural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping

Probabilistic Directed Distance Fields for Ray-Based Shape Representations

One-Shot Neural Fields for 3D Object Understanding

R3D3: Dense 3D Reconstruction of Dynamic Scenes from Multiple Cameras

Data-Driven 3D Reconstruction of Dressed Humans From Sparse Views

Dynamic Depth Fusion and Transformation for Monocular 3D Object Detection.

A Review and A Robust Framework of Data-Efficient 3D Scene Parsing with Traditional/Learned 3D Descriptors

3D-TAFS: A Training-free Framework for 3D Affordance Segmentation

Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models

Fast and Efficient: Mask Neural Fields for 3D Scene Segmentation