PsyEval: A Suite of Mental Health Related Tasks for Evaluating Large Language Models

Haoan Jin,Siyuan Chen,Dilawaier Dilixiati,Yewei Jiang,Mengyue Wu,Kenny Q. Zhu
2024-06-03
Abstract:Evaluating Large Language Models (LLMs) in the mental health domain poses distinct challenged from other domains, given the subtle and highly subjective nature of symptoms that exhibit significant variability among individuals. This paper presents PsyEval, the first comprehensive suite of mental health-related tasks for evaluating LLMs. PsyEval encompasses five sub-tasks that evaluate three critical dimensions of mental health. This comprehensive framework is designed to thoroughly assess the unique challenges and intricacies of mental health-related tasks, making PsyEval a highly specialized and valuable tool for evaluating LLM performance in this domain. We evaluate twelve advanced LLMs using PsyEval. Experiment results not only demonstrate significant room for improvement in current LLMs concerning mental health but also unveil potential directions for future model optimization.
Computation and Language
What problem does this paper attempt to address?
The problem that this paper attempts to solve is the unique challenges faced in evaluating large language models (LLMs) in the field of mental health. Specifically, symptoms in the mental health field are subtle and highly subjective, and there are significant differences among different individuals. Therefore, the evaluation of LLMs requires specially designed methods and criteria. The paper proposes a set of mental - health - related tasks named PsyEval, aiming to comprehensively evaluate the performance of LLMs in the mental health field. PsyEval covers five subtasks, which respectively evaluate three key dimensions in the mental health field: knowledge, diagnostic ability, and emotional support ability. Through this evaluation framework, researchers can more accurately identify the current LLMs' strengths and weaknesses in handling mental - health - related tasks and point out the direction for the optimization of future models.