Abstract:While prior work has explored whether large language models (LLMs) possess a "theory of mind" (ToM) - the ability to attribute mental states to oneself and others - there has been little work testing whether LLMs can implicitly apply such knowledge to predict behavior, or to judge whether an observed behavior is rational. Such skills are critical for appropriate interaction in social environments. We create a new dataset, SimpleTom, containing concise, diverse stories (e.g., "The can of Pringles has moldy chips in it. Mary picks up the can in the supermarket and walks to the cashier."), each with three questions that test different degrees of ToM reasoning, asking models to predict (a) mental state ("Is Mary aware of the mold?"), (b) behavior ("Will Mary pay for the chips or report the mold?"), and (c) judgment ("Mary paid for the chips. Was that reasonable?"). To our knowledge, SimpleToM is the first dataset to systematically explore downstream reasoning requiring knowledge of mental states in realistic scenarios. Our experimental results are intriguing: While most models can reliably predict mental state on our dataset (a), they often fail to correctly predict the behavior (b), and fare even worse at judging whether given behaviors are reasonable (c), despite being correctly aware of the protagonist's mental state should make such secondary predictions obvious. We further show that we can help models do better at (b) and (c) via interventions such as reminding the model of its earlier mental state answer and mental-state-specific chain-of-thought prompting, raising the action prediction accuracies (e.g., from 49.5% to 93.5% for GPT-4o) and judgment accuracies (e.g., from 15.3% to 94.7% in GPT-4o). While this shows that models can be coaxed to perform well, it requires task-specific interventions, and the natural model performances remain low, a cautionary tale for LLM deployment.

Do language models have coherent mental models of everyday things?

Can Language Models Understand Physical Concepts?

Do Large Language Models Exhibit Cognitive Dissonance? Studying the Difference Between Revealed Beliefs and Stated Answers

SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs

Evaluating Commonsense in Pre-trained Language Models

Easy Problems That LLMs Get Wrong

Large Language Models as a Tool for Mining Object Knowledge

Language Models (Mostly) Know What They Know

How Well Do Large Language Models Understand Syntax? An Evaluation by Asking Natural Language Questions

Language Models Don't Learn the Physical Manifestation of Language

Language models show human-like content effects on reasoning tasks

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Mental Modeling of Reinforcement Learning Agents by Language Models

Can I understand what I create? Self-Knowledge Evaluation of Large Language Models

Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process

Do Large Language Models Know about Facts?

Can Language Models Be Tricked by Language Illusions? Easier with Syntax, Harder with Semantics

Language Models Benefit from Preparation with Elicited Knowledge

Belief in the Machine: Investigating Epistemological Blind Spots of Language Models

Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology