Caution Regarding ChatGPT's Appropriateness and Reliability Regarding Surgery for Wrist Arthritis

Keegan Hones,Emily Krisanda,Harvey Chim
DOI: https://doi.org/10.1177/15589447241265519
2024-07-24
Abstract:Background: Chat Generative Pre-Trained Transformer (ChatGPT), an artificial intelligence (AI) program, is widely used for information compilation. This study sought to analyze the quality and consistency of the information generated by ChatGPT regarding common procedures for wrist arthritis. Methods: 32 standardized questions regarding wrist osteoarthritis and related procedures (4-corner-fusion [4CF], proximal row carpectomy [PRC], resurfacing capitate pyrocarbon implant, wrist denervation, and total wrist arthrodesis and arthroplasty) were presented to the ChatGPT-3.5 interface 3 separate times, without feedback. ChatGPT's answers were evaluated for medical accuracy by 3 reviewers and rated as "appropriate," "appropriate but incomplete," or "inappropriate." Ratings were then converted to numerical values to calculate an intraclass correlation coefficient (ICC). A DISCERN score was used to assess quality, and Flesch-Kincade Grade Level and Flesch Reading Ease Score for readability. Results: 75% of the responses were deemed "appropriate," with 23 questions receiving unanimous appropriate ratings across all responses. The ICC was 0.97 (95% CI [0.46, 0.98]), indicating excellent reliability. DISCERN score was 60 (good). The Flesch-Kincaid Grade Level was 14.6 ± 1.9, and the Flesch Reading Ease Score was 25.3 ± 6.7, implying a college reading level. The information that ChatGPT provided for PRC and total wrist arthrodesis and arthroplasty, appeared to be more reliable than for 4CF and denervation. Conclusion: ChatGPT's reliability and accuracy of information varied across procedures, possibly due to unknown and diverse sources. Furthermore, while some answers were factually correct, many provided generic information across differing questions, limiting usefulness. ChatGPT must be used cautiously and the limitations understood.
What problem does this paper attempt to address?