Analysis of lung cancer risk factors from medical records in Ethiopia using machine learning

Demeke Endalie,Wondmagegn Taye Abebe
DOI: https://doi.org/10.1371/journal.pdig.0000308
2023-07-21
PLOS Digital Health
Abstract:Cancer is a broad term that refers to a wide range of diseases that can affect any part of the human body. To minimize the number of cancer deaths and to prepare an appropriate health policy on cancer spread mitigation, scientifically supported knowledge of cancer causes is critical. As a result, in this study, we analyzed lung cancer risk factors that lead to a highly severe cancer case using a decision tree-based ranking algorithm. This feature relevance ranking algorithm computes the weight of each feature of the dataset by using split points to improve detection accuracy, and each risk factor is weighted based on the number of observations that occur for it on the decision tree. Coughing of blood, air pollution, and obesity are the most severe lung cancer risk factors out of nine, with a weight of 39%, 21%, and 14%, respectively. We also proposed a machine learning model that uses Extreme Gradient Boosting (XGBoost) to detect lung cancer severity levels in lung cancer patients. We used a dataset of 1000 lung cancer patients and 465 individuals free from lung cancer from Tikur Ambesa (Black Lion) Hospital in Addis Ababa, Ethiopia, to assess the performance of the proposed model. The proposed cancer severity level detection model achieved 98.9%, 99%, and 98.9% accuracy, precision, and recall, respectively, for the testing dataset. The findings can assist governments and non-governmental organizations in making lung cancer-related policy decisions. Lung cancer has become one of the leading causes of mortality in Ethiopia. Lung cancer risk factors vary from place to place since it depends on the people's socio-cultural activities. In this study, we examine lung cancer risk factors from the medical records of lung cancer patients in Addis Ababa, Ethiopia. The data contains the medical records of 872 women and 593 men. The key risk variables for lung cancer in the study area were identified using a decision tree. We discovered that coughing blood is one of the major risk factors for lung cancer, with a weight of 0.39. A feature importance of 0.39 indicates that the feature contributes 39% of the overall decision in the detection model. Furthermore, air pollution and obesity are the most important risk factors for lung cancer, with relevance weights of 0.21 and 0.14, respectively. This implies that these risk factors are causing or indicating most lung cancer cases in the study area. These three factors account for 74% of lung cancer analysis in the study area. Furthermore, we use the XGBoost classifier to detect lung cancer severity levels from risk factors, and the experiment yields a significant detection result.
What problem does this paper attempt to address?