Abstract:Handling missing data in clinical prognostic studies is an essential yet challenging task. This study aimed to provide a comprehensive assessment of the effectiveness and reliability of different machine learning (ML) imputation methods across various analytical perspectives. Specifically, it focused on three distinct classes of performance metrics used to evaluate ML imputation methods: post-imputation bias of regression estimates, post-imputation predictive accuracy, and substantive model-free metrics. As an illustration, we applied data from a real-world breast cancer survival study. This comprehensive approach aimed to provide a thorough assessment of the effectiveness and reliability of ML imputation methods across various analytical perspectives. A simulated dataset with 30% Missing At Random (MAR) values was used. A number of single imputation (SI) methods - specifically KNN, missMDA, CART, missForest, missRanger, missCforest - and multiple imputation (MI) methods - specifically miceCART and miceRF - were evaluated. The performance metrics used were Gower's distance, estimation bias, empirical standard error, coverage rate, length of confidence interval, predictive accuracy, proportion of falsely classified (PFC), normalized root mean squared error (NRMSE), AUC, and C-index scores. The analysis revealed that in terms of Gower's distance, CART and missForest were the most accurate, while missMDA and CART excelled for binary covariates; missForest and miceCART were superior for continuous covariates. When assessing bias and accuracy in regression estimates, miceCART and miceRF exhibited the least bias. Overall, the various imputation methods demonstrated greater efficiency than complete-case analysis (CCA), with MICE methods providing optimal confidence interval coverage. In terms of predictive accuracy for Cox models, missMDA and missForest had superior AUC and C-index scores. Despite offering better predictive accuracy, the study found that SI methods introduced more bias into the regression coefficients compared to MI methods. This study underlines the importance of selecting appropriate imputation methods based on study goals and data types in time-to-event research. The varying effectiveness of methods across the different performance metrics studied highlights the value of using advanced machine learning algorithms within a multiple imputation framework to enhance research integrity and the robustness of findings.

Addressing Missing Data in a Healthcare Dataset Using an Improved kNN Algorithm

Impact of machine learning-based imputation techniques on medical datasets- a comparative analysis

On the Performance of Imputation Techniques for Missing Values on Healthcare Datasets

Improving Healthcare Prediction of Diabetic Patients Using KNN Imputed Features and Tri-Ensemble Model

Missing Data Imputation for Classification Problems

Extremely missing numerical data in Electronic Health Records for machine learning can be managed through simple imputation methods considering informative missingness: A comparative of solutions in a COVID-19 mortality case study

Systematic Review on Missing Data Imputation Techniques with Machine Learning Algorithms for Healthcare

A novel ranked k-nearest neighbors algorithm for missing data imputation

Improving prediction of cervical cancer using KNN imputer and multi-model ensemble learning

An Innovative Imputation and Classification Approach for Accurate Disease Prediction

An automated approach to predict diabetic patients using KNN imputation and effective data mining techniques

Numerical Data Imputation for Multimodal Data Sets: A Probabilistic Nearest-Neighbor Kernel Density Approach

Integrated ECOD-KNN Algorithm for Missing Values Imputation in Datasets: Outlier Removal

Handling missing values in healthcare data: A systematic review of deep learning-based imputation techniques

Multi-metric comparison of machine learning imputation methods with application to breast cancer survival

Missing data imputation by K nearest neighbours based on grey relational structure and mutual information

Hybrid Missing Value Imputation Algorithm- KLR

An approach to dealing with missing values in heterogeneous data using k-nearest neighbors

An Intelligent Missing Data Imputation Techniques: A Review

Cervical cancer detection using K nearest neighbor imputer and stacked ensemble learningmodel

Imputation techniques on missing values in breast cancer treatment and fertility data