Awes, Laws, and Flaws From Today's LLM Research

Adrian de Wynter

2024-08-30

Abstract:We perform a critical examination of the scientific methodology behind contemporary large language model (LLM) research. For this we assess over 2,000 research works based on criteria typical of what is considered good research (e.g. presence of statistical tests and reproducibility) and cross-validate it with arguments that are at the centre of controversy (e.g., claims of emergent behaviour, the use of LLMs as evaluators). We find multiple trends, such as declines in claims of emergent behaviour and ethics disclaimers; the rise of LLMs as evaluators in spite of a lack of consensus from the community about their useability; and an increase of claims of LLM reasoning abilities, typically without leveraging human evaluation. This paper underscores the need for more scrutiny and rigour by and from this field to live up to the fundamentals of a responsible scientific method that is ethical, reproducible, systematic, and open to criticism.

Computation and Language

What problem does this paper attempt to address?

The paper attempts to address the issue of scientific methodology in current large language model (LLM) research. Specifically, the authors evaluated over 2000 related studies based on standards typically considered indicative of good research (such as the presence of statistical tests and reproducibility) and cross-verified with some contentious points (such as claims of emergent behaviors, using LLMs as evaluators, etc.). The authors identified several trends, such as a decrease in claims of emergent behaviors, a decrease in ethical statements, an increase in the use of LLMs as evaluators despite community disagreements on their usability, and an increase in claims of LLM reasoning abilities (often not relying on human evaluation). The paper emphasizes the need for more self-examination and rigor in the field to adhere to the fundamental principles of responsible scientific methodology, namely ethics, reproducibility, systematicity, and open criticism. Summary: - **Issue**: The issue of scientific methodology in current large language model (LLM) research. - **Method**: Evaluation of over 2000 related studies, cross-verified based on standard research quality indicators (such as statistical tests, reproducibility, etc.) and contentious points (such as emergent behaviors, LLMs as evaluators, etc.). - **Findings**: Decrease in claims of emergent behaviors, decrease in ethical statements, increase in the use of LLMs as evaluators, increase in claims of reasoning abilities (not relying on human evaluation). - **Conclusion**: More self-examination and rigor are needed to ensure research adheres to the fundamental principles of responsible scientific methodology.

Awes, Laws, and Flaws From Today's LLM Research

LLMs as Research Tools: A Large Scale Survey of Researchers' Usage and Perceptions

Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility

Caveat Lector: Large Language Models in Legal Practice

Position: Key Claims in LLM Research Have a Long Tail of Footnotes

Breaking the Silence: the Threats of Using LLMs in Software Engineering

Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Evaluation of an LLM in Identifying Logical Fallacies: A Call for Rigor When Adopting LLMs in HCI Research

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs

Objection Overruled! Lay People can Distinguish Large Language Models from Lawyers, but still Favour Advice from an LLM

Easy Problems That LLMs Get Wrong

Exploring Advanced Methodologies in Security Evaluation for LLMs

Are LLMs the Master of All Trades? : Exploring Domain-Agnostic Reasoning Skills of LLMs

Comprehensive Reassessment of Large-Scale Evaluation Outcomes in LLMs: A Multifaceted Statistical Approach

LLMs for science: Usage for code generation and data analysis

Evaluating Large Language Models: A Comprehensive Survey

An Interdisciplinary Outlook on Large Language Models for Scientific Research

Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review

Can LLMs be Fooled? Investigating Vulnerabilities in LLMs