Abstract:Embodied artificial intelligence (EAI) integrates advanced AI models into physical entities for real-world interaction. The emergence of foundation models as the "brain" of EAI agents for high-level task planning has shown promising results. However, the deployment of these agents in physical environments presents significant safety challenges. For instance, a housekeeping robot lacking sufficient risk awareness might place a metal container in a microwave, potentially causing a fire. To address these critical safety concerns, comprehensive pre-deployment risk assessments are imperative. This study introduces EARBench, a novel framework for automated physical risk assessment in EAI scenarios. EAIRiskBench employs a multi-agent cooperative system that leverages various foundation models to generate safety guidelines, create risk-prone scenarios, make task planning, and evaluate safety systematically. Utilizing this framework, we construct EARDataset, comprising diverse test cases across various domains, encompassing both textual and visual scenarios. Our comprehensive evaluation of state-of-the-art foundation models reveals alarming results: all models exhibit high task risk rates (TRR), with an average of 95.75% across all evaluated models. To address these challenges, we further propose two prompting-based risk mitigation strategies. While these strategies demonstrate some efficacy in reducing TRR, the improvements are limited, still indicating substantial safety concerns. This study provides the first large-scale assessment of physical risk awareness in EAI agents. Our findings underscore the critical need for enhanced safety measures in EAI systems and provide valuable insights for future research directions in developing safer embodied artificial intelligence system. Data and code are available at <a class="link-external link-https" href="https://github.com/zihao-ai/EARBench" rel="external noopener nofollow">this https URL</a>.

EARBench: Towards Evaluating Physical Risk Awareness for Task Planning of Foundation Model-based Embodied AI Agents

SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents

Embodied Red Teaming for Auditing Robotic Foundation Models

EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment

Architecture-level particular risk modeling and analysis for a cyber-physical system with AADL

SafeEmbodAI: a Safety Framework for Mobile Robots in Embodied AI Systems

HAZARD Challenge: Embodied Decision Making in Dynamically Changing Environments

Agent-SafetyBench: Evaluating the Safety of LLM Agents

BBSEA: An Exploration of Brain-Body Synchronization for Embodied Agents

AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models

Alexa Arena: A User-Centric Interactive Platform for Embodied AI

Safe Planner: Empowering Safety Awareness in Large Pre-Trained Models for Robot Task Planning

Learning to Assess Danger from Movies for Cooperative Escape Planning in Hazardous Environments

Multimodal Grounding for Embodied AI via Augmented Reality Headsets for Natural Language Driven Task Planning

Dynamic Risk Assessment and Active Response Strategy of Human-Robot Collaboration Based on Fuzzy Comprehensive Evaluation

MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning

A Survey on Robotics with Foundation Models: toward Embodied AI

Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Need for Developing a Security Robot-Based Risk Management for Emerging Practices in the Workplace Using the Advanced Human-Robot Collaboration Model

Spatial Relationship-Aware Rapid Entire Body Fuzzy Assessment Method for Prevention of Work-Related Musculoskeletal Disorders.