Abstract:Background: Risk prediction models are routinely used to assist in clinical decision making. A small sample size for model development can compromise model performance when the model is applied to new patients. For binary outcomes, the calibration slope (CS) and the mean absolute prediction error (MAPE) are two key measures on which sample size calculations for the development of risk models have been based. CS quantifies the degree of model overfitting while MAPE assesses the accuracy of individual predictions. Methods: Recently, two formulae were proposed to calculate the sample size required, given anticipated features of the development data such as the outcome prevalence and c-statistic, to ensure that the expectation of the CS and MAPE (over repeated samples) in models fitted using MLE will meet prespecified target values. In this article, we use a simulation study to evaluate the performance of these formulae. Results: We found that both formulae work reasonably well when the anticipated model strength is not too high (c-statistic < 0.8), regardless of the outcome prevalence. However, for higher model strengths the CS formula underestimates the sample size substantially. For example, for c-statistic = 0.85 and 0.9, the sample size needed to be increased by at least 50% and 100%, respectively, to meet the target expected CS. On the other hand, the MAPE formula tends to overestimate the sample size for high model strengths. These conclusions were more pronounced for higher prevalence than for lower prevalence. Similar results were drawn when the outcome was time to event with censoring. Given these findings, we propose a simulation-based approach, implemented in the new R package 'samplesizedev', to correctly estimate the sample size even for high model strengths. The software can also calculate the variability in CS and MAPE, thus allowing for assessment of model stability. Conclusions: The calibration and MAPE formulae suggest sample sizes that are generally appropriate for use when the model strength is not too high. However, they tend to be biased for higher model strengths, which are not uncommon in clinical risk prediction studies. On those occasions, our proposed adjustments to the sample size calculations will be relevant.

Evaluation of a decided sample size in machine learning applications

Sample Size Analysis for Machine Learning Clinical Validation Studies

Toward Generalizable Machine Learning Models in Speech, Language, and Hearing Sciences: Estimating Sample Size and Reducing Overfitting

Determining Sample Size in Binary Measurement System

A refined approach for evaluating small datasets via binary classification using machine learning

Small sample size effects in statistical pattern recognition: recommendations for practitioners

An evaluation of sample size requirements for developing risk prediction models with binary outcomes

Machine learning algorithm validation with a limited sample size

Sample size determination for multidimensional parameters and the A-optimal subsampling in a big data linear regression model

How to calculate sample size in animal and human studies

All about sample-size calculations for A/B testing: Novel extensions and practical guide

Sample size requirements are not being considered in studies developing prediction models for binary outcomes: a systematic review

Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning

Sample Size in Natural Language Processing within Healthcare Research

Systematic Bias in Sample Inference and its Effect on Machine Learning

Sample size for developing a prediction model with a binary outcome: targeting precise individual risk estimates to improve clinical decisions and fairness

Sample size calculations for the experimental comparison of multiple algorithms on multiple problem instances

Calculating the sample size required for developing a clinical prediction model

The Effects of Sample Size on Omics Study: from the Perspective of Robustness and Diagnostic Accuracy

Revisiting Sample Size Determination in Natural Language Understanding