Large Language Models Fail to Reproduce Level I Recommendations for Breast Radiotherapy

Kathleen Tang,John Han,Shengyang Wu
DOI: https://doi.org/10.1101/2024.11.11.24317123
2024-11-11
Abstract:This study evaluates the reliability of the largest public-facing large language models in providing accurate breast cancer radiotherapy recommendations. We assessed ChatGPT 3.5, ChatGPT 4, ChatGPT 4o, Claude 3.5 Sonnet, and ChatGPT o1 in three common clinical scenarios. The clinical cases are as follows: post-lumpectomy radiotherapy in a 40 year old woman, (2) postmastectomy radiation in a 40 year old woman with 4+ lymph nodes, and (3) postmastectomy radiation in an 80 year old woman with early stage tumor and negative axillary dissection. Each case was designed to be unambiguous with respect to the Level I evidence and clinical guideline-supported approach. The evidence-supported radiation treatments are as follows: (1) Whole breast with boost (2) Regional nodal irradiation (3) Omission of post-operative radiotherapy. Each prompt is presented to each LLM multiple times to ensure reproducibility. Results indicate that the free, public-facing models often fail to provide accurate treatment recommendations, particularly when omission of radiotherapy was the correct course of action. Many recommendations suggested by the LLMs increase morbidity and mortality in patients. Models only accessible through paid subscription (ChatGPT o1 and o1-mini) demonstrated greatly improved accuracy. Some prompt-engineering techniques, rewording and chain-of-reasoning, enhanced the accuracy of the LLMs, while true/false questioning significantly worsened results. While public-facing LLMs show potential for medical applications, their current reliability is unsuitable for clinical decision-making.
Oncology
What problem does this paper attempt to address?