A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Yining Huang,Keke Tang,Meilian Chen,Boyuan Wang

2024-05-29

Abstract:Since the inception of the Transformer architecture in 2017, Large Language Models (LLMs) such as GPT and BERT have evolved significantly, impacting various industries with their advanced capabilities in language understanding and generation. These models have shown potential to transform the medical field, highlighting the necessity for specialized evaluation frameworks to ensure their effective and ethical deployment. This comprehensive survey delineates the extensive application and requisite evaluation of LLMs within healthcare, emphasizing the critical need for empirical validation to fully exploit their capabilities in enhancing healthcare outcomes. Our survey is structured to provide an in-depth analysis of LLM applications across clinical settings, medical text data processing, research, education, and public health awareness. We begin by exploring the roles of LLMs in various medical applications, detailing their evaluation based on performance in tasks such as clinical diagnosis, medical text data processing, information retrieval, data analysis, and educational content generation. The subsequent sections offer a comprehensive discussion on the evaluation methods and metrics employed, including models, evaluators, and comparative experiments. We further examine the benchmarks and datasets utilized in these evaluations, providing a categorized description of benchmarks for tasks like question answering, summarization, information extraction, bioinformatics, information retrieval and general comprehensive benchmarks. This structure ensures a thorough understanding of how LLMs are assessed for their effectiveness, accuracy, usability, and ethical alignment in the medical domain. ...

Computation and Language

What problem does this paper attempt to address?

### Problems the Paper Aims to Address This paper aims to comprehensively evaluate the application and assessment of large language models (LLMs) in the healthcare industry. Specifically, through detailed analysis, the paper explores the applications of LLMs in various aspects such as clinical settings, medical text data processing, research, education, and public health awareness. It emphasizes the importance of empirical validation of these models to fully leverage their potential in improving healthcare outcomes. The main objectives of the paper include: 1. **Providing in-depth application analysis**: Describing in detail the performance of LLMs in different healthcare application scenarios, such as clinical diagnosis, medical text data processing, information retrieval, data analysis, and educational content generation. 2. **Discussing evaluation methods and metrics**: Introducing various methods and metrics used to evaluate LLMs, including models, evaluators, and comparative experiments. 3. **Benchmarks and datasets**: Summarizing the benchmarks and datasets used in these evaluations, covering tasks such as question answering, summarization, information extraction, bioinformatics, and information retrieval. 4. **Ethics and practicality**: Ensuring that the application of LLMs in the healthcare field is both effective and ethically compliant, providing healthcare professionals, researchers, and policymakers with a comprehensive understanding to responsibly develop and deploy these models. Through these analyses, the paper hopes to provide healthcare professionals with detailed insights into the advantages and limitations of LLMs, guiding the effective implementation and evaluation of these powerful tools while ensuring they achieve their maximum potential while maintaining strict ethical standards.

A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Evaluating large language models in medical applications: a survey

A Survey of Large Language Models in Medicine: Progress, Application, and Challenge

A Survey on Medical Large Language Models: Technology, Application, Trustworthiness, and Future Directions

A Survey on Large Language Models from General Purpose to Medical Applications: Datasets, Methodologies, and Evaluations

A Survey on Evaluation of Large Language Models

A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine

Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review

A Survey on Evaluation of Large Language ModelsJust Accepted

Application Research of Large Language Models in Medicine: Status, Problems, and Future

Critical Care Studies Using Large Language Models Based on Electronic Healthcare Records: A Technical Note

Large language models in healthcare and medical domain: A review

The application of large language models in medicine: A scoping review

A Survey of Large Language Models for Healthcare: from Data, Technology, and Applications to Accountability and Ethics

Large language models in medical and healthcare fields: applications, advances, and challenges

Large Language Models for Medicine: A Survey

Evaluation of large language model performance on the Biomedical Language Understanding and Reasoning Benchmark

Evaluating Large Language Models: A Comprehensive Survey

A Survey for Large Language Models in Biomedicine

Large Language Models in Healthcare: A Comprehensive Benchmark