A Novel Nuanced Conversation Evaluation Framework for Large Language Models in Mental Health

Alexander Marrapese,Basem Suleiman,Imdad Ullah,Juno Kim
2024-03-09
Abstract:Understanding the conversation abilities of Large Language Models (LLMs) can help lead to its more cautious and appropriate deployment. This is especially important for safety-critical domains like mental health, where someone's life may depend on the exact wording of a response to an urgent question. In this paper, we propose a novel framework for evaluating the nuanced conversation abilities of LLMs. Within it, we develop a series of quantitative metrics developed from literature on using psychotherapy conversation analysis literature. While we ensure that our framework and metrics are transferable by researchers to relevant adjacent domains, we apply them to the mental health field. We use our framework to evaluate several popular frontier LLMs, including some GPT and Llama models, through a verified mental health dataset. Our results show that GPT4 Turbo can perform significantly more similarly to verified therapists than other selected LLMs. We conduct additional analysis to examine how LLM conversation performance varies across specific mental health topics. Our results indicate that GPT4 Turbo performs well in achieving high correlation with verified therapists in particular topics such as Parenting and Relationships. We believe our contributions will help researchers develop better LLMs that, in turn, will more positively support people's lives.
Computation and Language,Artificial Intelligence,Emerging Technologies
What problem does this paper attempt to address?
This paper aims to address the issue of how to objectively evaluate the nuanced conversational abilities of large language models (LLMs) in the field of mental health. Specifically, the researchers propose a novel evaluation framework to quantitatively analyze the nuanced conversational skills exhibited by LLMs when interacting with humans. The importance of this work lies in ensuring that LLMs can provide appropriate and empathetic support when dealing with sensitive issues such as mental health. The key contributions of the paper include: 1. **Proposing a novel evaluation framework**: This framework focuses on assessing the ability of LLMs to handle nuanced conversations, particularly in the context of mental health issues. By drawing on methods from the literature on psychotherapy dialogue analysis, a series of quantitative evaluation metrics were developed. 2. **Developing multiple quantitative metrics**: These metrics cover various aspects, from emotional content to mimicking dialogue strategies found in psychological counseling literature. Examples include emotional consistency, relative directional emotional change, emotional change within responses, simplicity, effectiveness, readability, element reuse, acceptability, and active listening. 3. **Evaluating cutting-edge LLMs**: The researchers used this framework to evaluate the performance of several popular cutting-edge LLMs (including GPT4 Turbo, Llama 2 7B, and Mistral 7B V1) in handling validated mental health issues. 4. **Analyzing the impact of specific mental health topics**: The study also explored how different mental health topics affect the conversational performance of LLMs, especially on specific themes such as parenting and interpersonal relationships. In summary, this paper provides researchers with a tool to better understand and improve the application of LLMs in the field of mental health by proposing a novel evaluation framework and a series of specific quantitative metrics, thereby offering more positive and supportive services to users.