I-PHYRE: Interactive Physical Reasoning

Shiqian Li,Kewen Wu,Chi Zhang,Yixin Zhu
2024-03-25
Abstract:Current evaluation protocols predominantly assess physical reasoning in stationary scenes, creating a gap in evaluating agents' abilities to interact with dynamic events. While contemporary methods allow agents to modify initial scene configurations and observe consequences, they lack the capability to interact with events in real time. To address this, we introduce I-PHYRE, a framework that challenges agents to simultaneously exhibit intuitive physical reasoning, multi-step planning, and in-situ intervention. Here, intuitive physical reasoning refers to a quick, approximate understanding of physics to address complex problems; multi-step denotes the need for extensive sequence planning in I-PHYRE, considering each intervention can significantly alter subsequent choices; and in-situ implies the necessity for timely object manipulation within a scene, where minor timing deviations can result in task failure. We formulate four game splits to scrutinize agents' learning and generalization of essential principles of interactive physical reasoning, fostering learning through interaction with representative scenarios. Our exploration involves three planning strategies and examines several supervised and reinforcement agents' zero-shot generalization proficiency on I-PHYRE. The outcomes highlight a notable gap between existing learning algorithms and human performance, emphasizing the imperative for more research in enhancing agents with interactive physical reasoning capabilities. The environment and baselines will be made publicly available.
Artificial Intelligence,Computer Vision and Pattern Recognition,Machine Learning,Robotics
What problem does this paper attempt to address?
The problem addressed in this paper is that existing evaluation methods mainly focus on physical reasoning in static scenes and fail to adequately assess the ability of agents to interact with dynamic events. Although existing methods allow agents to modify the initial scene configuration and observe the results, they lack real-time interaction with events. To address this issue, the paper introduces the Interactive Physical Reasoning (I-PHYRE) framework, which requires agents to demonstrate intuitive physical reasoning, multi-step planning, and on-site intervention capabilities simultaneously. I-PHYRE tests the learning and generalization of agents on key principles of interactive physical reasoning through four game splits, promoting learning through interaction with representative scenes. The study compares the zero-shot generalization performance of different strategies and reinforcement learning agents, and the results show a significant gap between current learning algorithms and human performance, emphasizing the urgency of research in endowing agents with interactive physical reasoning abilities.