Abstract:Objectives Artificial intelligence (AI) has shown promise in improving the performance of fetal ultrasound screening in detecting congenital heart disease (CHD). The effect of giving AI advice to human operators has not been studied in this context. Giving additional information about AI model workings, such as confidence scores for AI predictions, may be a way of improving performance further. Our aims were to investigate whether AI advice improved overall diagnostic accuracy (using a single CHD lesion as an exemplar), and to see what, if any, additional information given to clinicians optimized the overall performance of the clinician‐AI team. Methods An AI model was trained to classify a single fetal CHD lesion (atrioventricular septal defect, AVSD), using a retrospective cohort of 121,130 cardiac four chamber images extracted from 173 ultrasound scan videos (98 with normal hearts, 75 with AVSD). A ResNet50 model architecture was used. Temperature scaling of model prediction probability was performed on a validation set, and gradient‐weighted class activation maps (grad‐CAMs) produced. Ten clinicians (two consultant fetal cardiologists, three trainees in pediatric cardiology, and five fetal cardiac sonographers) were recruited from a center of fetal cardiology to participate. Each participant was shown 2000 fetal four chamber images in a random order (1,000 normal and 1,000 AVSD). The dataset was comprised of 500 images, each shown in four conditions: 1) image alone without AI output; 2) image with binary AI classification; 3) image with AI model confidence; 4) image with gradient‐weighted class activation map image overlays. The clinicians were asked to classify each image as normal or AVSD. Results 20,000 image classifications were recorded from 10 clinicians. The AI model alone achieved an accuracy of 0.798 (95% CI 0.760 – 0.832), sensitivity of 0.868 (95% CI 0.834 – 0.902) and specificity of 0.728 (95% CI 0.702 – 0.754, and the clinicians without AI achieved an accuracy of 0.844 (95% CI 0.834 – 0.854), sensitivity of 0.827 (95% CI 0.795 – 0.858) and specificity of 0.861 (95% CI 0.828 – 0.895). Showing a binary (normal or AVSD) AI model output resulted in significant improvement in accuracy to 0.865 (p <0.001). This effect was seen in both experienced and less experienced participants. Giving incorrect AI advice resulted in significant deterioration in overall accuracy from 0.761 to 0.693 (p <0.001), which was driven by an increase in both type I and type II error by the clinicians. This effect was worsened by showing model confidence (accuracy 0.649, p <0.001) or grad‐CAM (accuracy 0.644, p <0.001). Conclusions AI has the potential to improve performance when used in collaboration with clinicians, even if the model performance does not reach expert level. Giving additional information about model workings such as model confidence and class activation map image overlays did not improve overall performance, and actually worsened performance for images where the AI model was incorrect. This article is protected by copyright. All rights reserved.

Enabling faster and more reliable sonographic assessment of gestational age through machine learning

Deep learning to estimate gestational age from fly‐to cineloop videos: A novel approach to ultrasound quality control

AI Estimation of Gestational Age from Blind Ultrasound Sweeps in Low-Resource Settings

Diagnostic Accuracy of an Integrated AI Tool to Estimate Gestational Age From Blind Ultrasound Sweeps

AI supported fetal echocardiography with quality assessment

Whole-examination AI estimation of fetal biometrics from 20-week ultrasound scans

Deep learning fetal ultrasound video model match human observers in biometric measurements

AutoFB: Automating Fetal Biometry Estimation from Standard Ultrasound Planes

Automatic Detection of Standard Planes in Fetal Ultrasound Images based on Convolutional Neural Networks and Ensemble Learning

Artificial Intelligence to Assist in the Screening Fetal Anomaly Ultrasound Scan (PROMETHEUS): A Randomised Controlled Trial

Toward point-of-care ultrasound estimation of fetal gestational age from the trans-cerebellar diameter using CNN-based ultrasound image analysis

Automating the Human Action of First-Trimester Biometry Measurement from Real-World Freehand Ultrasound

Human-level Performance On Automatic Head Biometrics In Fetal Ultrasound Using Fully Convolutional Neural Networks

Advances in the Application of Artificial Intelligence in Fetal Echocardiography

Development and validation of an artificial intelligence assisted prenatal ultrasonography screening system for trainees

Attention-guided deep learning for gestational age prediction using fetal brain MRI

Interaction between clinicians and artificial intelligence to detect fetal atrioventricular septal defects on ultrasound: how can we optimize collaborative performance?

Development and external validation of an ultrasound image-based deep learning model to estimate gestational age in the second and third trimesters of pregnancy using data from Garbh-Ini cohort: a prospective cohort study in North Indian population

Advancing Fetal Ultrasound Diagnostics: Innovative Methodologies for Improved Accuracy in Detecting Down Syndrome

Deep Learning-based Quality Assessment of Clinical Protocol Adherence in Fetal Ultrasound Dating Scans

Towards deep observation: A systematic survey on artificial intelligence techniques to monitor fetus via Ultrasound Images