Identifying Prediabetes in Canadian Populations Using Machine Learning

Katherine Lu,Paijani Sheth,Zhi Lin Zhou,Kamyar Kazari,Aziz Guergachi,Karim Keshavjee,Mohammad Noaeen,Zahra Shakeri
DOI: https://doi.org/10.1101/2024.02.03.24302301
2024-02-05
Abstract:Prediabetes is a critical health condition characterized by elevated blood glucose levels that fall below the threshold for Type 2 diabetes (T2D) diagnosis. Accurate identification of prediabetes is essential to forestall the progression to T2D among at-risk individuals. This study aims to pinpoint the most effective machine learning (ML) model for prediabetes prediction and to elucidate the key biological variables critical for distinguishing individuals with prediabetes. Utilizing data from the Canadian Primary Care Sentinel Surveillance Network (CPCSSN), our analysis included 6,414 participants identified as either nondiabetic or prediabetic. A rigorous selection process led to the identification of ten variables for the study, informed by literature review, data completeness, and the evaluation of collinearity. Our comparative analysis of seven ML models revealed that the Deep Neural Network (DNN), enhanced with early stop regularization, outshined others by achieving a recall rate of 60%. This model’s performance underscores its potential in effectively identifying prediabetic individuals, showcasing the strategic integration of ML in healthcare. While the model reflects a significant advancement in prediabetes prediction, it also opens avenues for further research to refine prediction accuracy, possibly by integrating novel biological markers or exploring alternative modeling techniques. The results of our work represent a pivotal step forward in the early detection of prediabetes, contributing significantly to preventive healthcare measures and the broader fight against the global epidemic of Type 2 diabetes.
Health Informatics
What problem does this paper attempt to address?
This paper aims to use machine learning (ML) methods to identify prediabetic individuals in Canada and find key biological variables that distinguish prediabetic individuals. The study is based on data from the Canadian Primary Care Sentinel Surveillance Network (CPCSSN) and analyzes 6414 participants, including non-diabetic and prediabetic patients. Through variable selection, feature engineering, and comparative analysis of seven different ML models (including deep neural network, DNN), the DNN model achieved a recall rate of 60% under early stopping regularization conditions, demonstrating excellent performance. The paper points out that accurate identification of prediabetes is crucial as approximately 70% of prediabetic individuals may develop type 2 diabetes. The performance of the DNN model indicates its potential in combining ML technology with healthcare, but also emphasizes the need for further improvement in prediction accuracy through possible methods such as incorporating new biomarkers or exploring alternative modeling techniques. The study also discusses limitations of the model, such as overfitting issues and limitations in predictive ability, and notes that existing variables may be insufficient for sufficient prediction of prediabetes. Additionally, the paper mentions the importance of considering diversity, equity, and inclusion (DEI) in diabetes research, particularly in multicultural contexts like Toronto, but acknowledges that the current data lacks information on race/ethnicity, limiting exploration in this aspect. In conclusion, the paper presents an effective method for identifying prediabetic individuals in Canada through machine learning, while also emphasizing the challenges that future work needs to address in order to improve the accuracy and comprehensiveness of predictions.