I-PHYRE: Interactive Physical Reasoning

Shiqian Li,Kewen Wu,Chi Zhang,Yixin Zhu

2024-03-25

Abstract:Current evaluation protocols predominantly assess physical reasoning in stationary scenes, creating a gap in evaluating agents' abilities to interact with dynamic events. While contemporary methods allow agents to modify initial scene configurations and observe consequences, they lack the capability to interact with events in real time. To address this, we introduce I-PHYRE, a framework that challenges agents to simultaneously exhibit intuitive physical reasoning, multi-step planning, and in-situ intervention. Here, intuitive physical reasoning refers to a quick, approximate understanding of physics to address complex problems; multi-step denotes the need for extensive sequence planning in I-PHYRE, considering each intervention can significantly alter subsequent choices; and in-situ implies the necessity for timely object manipulation within a scene, where minor timing deviations can result in task failure. We formulate four game splits to scrutinize agents' learning and generalization of essential principles of interactive physical reasoning, fostering learning through interaction with representative scenarios. Our exploration involves three planning strategies and examines several supervised and reinforcement agents' zero-shot generalization proficiency on I-PHYRE. The outcomes highlight a notable gap between existing learning algorithms and human performance, emphasizing the imperative for more research in enhancing agents with interactive physical reasoning capabilities. The environment and baselines will be made publicly available.

Artificial Intelligence,Computer Vision and Pattern Recognition,Machine Learning,Robotics

What problem does this paper attempt to address?

The problem addressed in this paper is that existing evaluation methods mainly focus on physical reasoning in static scenes and fail to adequately assess the ability of agents to interact with dynamic events. Although existing methods allow agents to modify the initial scene configuration and observe the results, they lack real-time interaction with events. To address this issue, the paper introduces the Interactive Physical Reasoning (I-PHYRE) framework, which requires agents to demonstrate intuitive physical reasoning, multi-step planning, and on-site intervention capabilities simultaneously. I-PHYRE tests the learning and generalization of agents on key principles of interactive physical reasoning through four game splits, promoting learning through interaction with representative scenes. The study compares the zero-shot generalization performance of different strategies and reinforcement learning agents, and the results show a significant gap between current learning algorithms and human performance, emphasizing the urgency of research in endowing agents with interactive physical reasoning abilities.

I-PHYRE: Interactive Physical Reasoning

Hi-Phy: A Benchmark for Hierarchical Physical Reasoning

Interactive Visual Reasoning under Uncertainty

NovPhy: A Testbed for Physical Reasoning in Open-world Environments

PhyPlan: Generalizable and Rapid Physical Task Planning with Physics Informed Skill Networks for Robot Manipulators

AGENT: A Benchmark for Core Psychological Reasoning

iCORPP: Interleaved commonsense reasoning and probabilistic planning on robots

Rapid trial-and-error learning with simulation supports flexible tool use and physical reasoning

Benchmarks for Physical Reasoning AI

On the Learning Mechanisms in Physical Reasoning

PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI

Imagined Potential Games: A Framework for Simulating, Learning and Evaluating Interactive Behaviors

IFR-Explore: Learning Inter-object Functional Relationships in 3D Indoor Scenes

Pangu-Agent: A Fine-Tunable Generalist Agent with Structured Reasoning

IR-STP: Enhancing Autonomous Driving With Interaction Reasoning in Spatio-Temporal Planning

Closed Loop Interactive Embodied Reasoning for Robot Manipulation

EscapeBench: Pushing Language Models to Think Outside the Box

SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation

Spatial Reasoning and Planning for Deep Embodied Agents

Physical Reasoning and Object Planning for Household Embodied Agents

Abstract Spatial-Temporal Reasoning Via Probabilistic Abduction and Execution