Abstract:Large language models (LLMs) have significantly advanced the field of artificial intelligence. Yet, evaluating them comprehensively remains challenging. We argue that this is partly due to the predominant focus on performance metrics in most benchmarks. This paper introduces CogBench, a benchmark that includes ten behavioral metrics derived from seven cognitive psychology experiments. This novel approach offers a toolkit for phenotyping LLMs' behavior. We apply CogBench to 35 LLMs, yielding a rich and diverse dataset. We analyze this data using statistical multilevel modeling techniques, accounting for the nested dependencies among fine-tuned versions of specific LLMs. Our study highlights the crucial role of model size and reinforcement learning from human feedback (RLHF) in improving performance and aligning with human behavior. Interestingly, we find that open-source models are less risk-prone than proprietary models and that fine-tuning on code does not necessarily enhance LLMs' behavior. Finally, we explore the effects of prompt-engineering techniques. We discover that chain-of-thought prompting improves probabilistic reasoning, while take-a-step-back prompting fosters model-based behaviors.

What problem does this paper attempt to address?

The paper "CogBench: An Approach to Evaluate Behavior of Large Language Models (LLMs) within the Domain of Cognitive Psychology Experiments" primarily focuses on how to comprehensively evaluate the behavior of large language models. Currently, most benchmark tests place too much emphasis on performance metrics while neglecting an understanding of the underlying behavioral mechanisms. To address this, the researchers propose a new benchmark testing tool called CogBench, which includes ten behavioral metrics from seven cognitive psychology experiments. CogBench tests 35 LLMs and reveals that model size and human feedback reinforcement learning (RLHF) are crucial for performance and consistency with human behavior. Open-source models tend to be more cautious in risk decision-making compared to proprietary models, and fine-tuning for code does not necessarily improve LLM behavior. Additionally, the study explores the impact of prompting engineering techniques and finds that prompting with chain thinking enhances probability reasoning, while "take a step back" prompts promote model baseline behavior. In the paper, the authors emphasize the role of psychology experiments in understanding LLM behavior and provide a comprehensive behavior evaluation through CogBench. They find that although some models excel in specific tasks, they do not perform like humans across many behavioral metrics. This indicates that, in addition to performance metrics, more attention is needed on the behavioral characteristics of models. Through this approach, researchers can gain deeper insights into the behavioral patterns of LLMs and how training and fine-tuning strategies impact these behaviors.

CogBench: a large language model walks into a psychology lab

Benchmarking Cognitive Biases in Large Language Models as Evaluators

CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models

CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy

NLPBench: Evaluating Large Language Models on Solving NLP Problems

CogLM: Tracking Cognitive Development of Large Language Models

Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark

Quantifying AI Psychology: A Psychometrics Benchmark for Large Language Models

TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles

I Think, Therefore I am: Benchmarking Awareness of Large Language Models Using AwareBench

Analysis of hybrid imaging techniques

Human Behavioral Benchmarking: Numeric Magnitude Comparison Effects in Large Language Models

LMentry: A Language Model Benchmark of Elementary Language Tasks

CLR-Bench: Evaluating Large Language Models in College-level Reasoning

Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks

clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents

ML-Bench: Large Language Models Leverage Open-source Libraries for Machine Learning Tasks

Do Large Language Models Exhibit Cognitive Dissonance? Studying the Difference Between Revealed Beliefs and Stated Answers

Turning large language models into cognitive models