Towards Effective Automatic Evaluation of Generated Reflections for Motivational Interviewing

Rim Helaoui,Daniele Riboni,Diego Reforgiato Recupero,Zixiu "Alex" Wu
DOI: https://doi.org/10.1145/3610661.3616127
2023-10-09
Abstract:Reflection is an essential counselling skill where the therapist communicates their understanding of the client’s words to the client. Recent studies have explored language-model-based reflection generation, but automatic quality evaluation of generated reflections remains under-explored. In this work, we investigate automatic evaluation on one fundamental quality aspect: coherence and context-consistency. We test a range of automatic evaluators/metrics and examine their correlations with expert judgement. We find that large language models (LLMs) as zero-shot evaluators achieve the best performance, while other metrics correlate poorly with expert judgement. We also demonstrate that diverse LLM-as-evaluator configurations need to be explored to find the best setup.
Psychology,Computer Science
What problem does this paper attempt to address?