Abstract:The integration of physiological computing into mixed-initiative human-robot interaction systems offers valuable advantages in autonomous task allocation by incorporating real-time features as human state observations into the decision-making system. This approach may alleviate the cognitive load on human operators by intelligently allocating mission tasks between agents. Nevertheless, accommodating a diverse pool of human participants with varying physiological and behavioral measurements presents a substantial challenge. To address this, resorting to a probabilistic framework becomes necessary, given the inherent uncertainty and partial observability on the human's state. Recent research suggests to learn a Partially Observable Markov Decision Process (POMDP) model from a data set of previously collected experiences that can be solved using Offline Reinforcement Learning (ORL) methods. In the present work, we not only highlight the potential of partially observable representations and physiological measurements to improve human operator state estimation and performance, but also enhance the overall mission effectiveness of a human-robot team. Importantly, as the fixed data set may not contain enough information to fully represent complex stochastic processes, we propose a method to incorporate model uncertainty, thus enabling risk-sensitive sequential decision-making. Experiments were conducted with a group of twenty-six human participants within a simulated robot teleoperation environment, yielding empirical evidence of the method's efficacy. The obtained adaptive task allocation policy led to statistically significant higher scores than the one that was used to collect the data set, allowing for generalization across diverse participants also taking into account risk-sensitive metrics.

Proximal Reinforcement Learning: Efficient Off-Policy Evaluation in Partially Observed Markov Decision Processes

Behavior Proximal Policy Optimization

Beyond Reward: Offline Preference-guided Policy Optimization

Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes

Off-policy Evaluation in Infinite-Horizon Reinforcement Learning with Latent Confounders

Offline Risk-sensitive RL with Partial Observability to Enhance Performance in Human-Robot Teaming

Off-Policy Evaluation in Partially Observable Environments

Efficient Online Reinforcement Learning with Offline Data

RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation

Conformal Off-Policy Evaluation in Markov Decision Processes

Deep Offline Reinforcement Learning for Real-world Treatment Optimization Applications

Provably Efficient Reinforcement Learning in Partially Observable Dynamical Systems

Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning

Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency

Provable Offline Preference-Based Reinforcement Learning

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards

Causal Reinforcement Learning using Observational and Interventional Data

Offline Policy Evaluation and Optimization under Confounding

Beyond the Boundaries of Proximal Policy Optimization

Robust Fitted-Q-Evaluation and Iteration under Sequentially Exogenous Unobserved Confounders

Enhancing Sample Efficiency and Exploration in Reinforcement Learning through the Integration of Diffusion Models and Proximal Policy Optimization