Abstract:While prior work has explored whether large language models (LLMs) possess a "theory of mind" (ToM) - the ability to attribute mental states to oneself and others - there has been little work testing whether LLMs can implicitly apply such knowledge to predict behavior, or to judge whether an observed behavior is rational. Such skills are critical for appropriate interaction in social environments. We create a new dataset, SimpleTom, containing concise, diverse stories (e.g., "The can of Pringles has moldy chips in it. Mary picks up the can in the supermarket and walks to the cashier."), each with three questions that test different degrees of ToM reasoning, asking models to predict (a) mental state ("Is Mary aware of the mold?"), (b) behavior ("Will Mary pay for the chips or report the mold?"), and (c) judgment ("Mary paid for the chips. Was that reasonable?"). To our knowledge, SimpleToM is the first dataset to systematically explore downstream reasoning requiring knowledge of mental states in realistic scenarios. Our experimental results are intriguing: While most models can reliably predict mental state on our dataset (a), they often fail to correctly predict the behavior (b), and fare even worse at judging whether given behaviors are reasonable (c), despite being correctly aware of the protagonist's mental state should make such secondary predictions obvious. We further show that we can help models do better at (b) and (c) via interventions such as reminding the model of its earlier mental state answer and mental-state-specific chain-of-thought prompting, raising the action prediction accuracies (e.g., from 49.5% to 93.5% for GPT-4o) and judgment accuracies (e.g., from 15.3% to 94.7% in GPT-4o). While this shows that models can be coaxed to perform well, it requires task-specific interventions, and the natural model performances remain low, a cautionary tale for LLM deployment.

Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models

ToMBench: Benchmarking Theory of Mind in Large Language Models

Probing the Robustness of Theory of Mind in Large Language Models

OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

ToMChallenges: A Principle-Guided Dataset and Diverse Evaluation Tasks for Exploring Theory of Mind

Zero, Finite, and Infinite Belief History of Theory of Mind Reasoning in Large Language Models

Multi-ToM: Evaluating Multilingual Theory of Mind Capabilities in Large Language Models

Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

TimeToM: Temporal Space is the Key to Unlocking the Door of Large Language Models' Theory-of-Mind

Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs. Children Aged 7-10 on Advanced Tests

Theory of Mind for Multi-Agent Collaboration via Large Language Models

SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs

HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

LLM Theory of Mind and Alignment: Opportunities and Risks

Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning

NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding

MindGames: Targeting Theory of Mind in Large Language Models with Dynamic Epistemic Modal Logic

Think Twice: Perspective-Taking Improves Large Language Models' Theory-of-Mind Capabilities

Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?

How FaR Are Large Language Models From Agents with Theory-of-Mind?

Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker