Comparative Machine Learning Framework for Classifying Cancer Incidence in Iraq ‎‎(2018–2020): Evidence from Random Forest, XGBoost, SVM, and Shrinkage-QDA

Authors

  • Omar Fawzi Salih Al-Rawi Department of Statistics and Informatics Techniques, Technical college of Management–Mosul, Northern ‎Technical University, Iraq

DOI:

https://doi.org/10.19139/soic-2310-5070-3795

Keywords:

Neonatal Mortality, Machine Learning, Feature Selection, Demographic Indicators, Predictive ‎Modeling

Abstract

Although the use of machine learning in cancer research is increasing, most studies conducted in Iraq have focused primarily on clinical diagnosis, medical image classification, and survival analysis, with little attention given to the classification of cancer incidence rates at the governorate level, based on unified national datasets. In addition, there is a lack of a comprehensive comparative framework for evaluating various machine learning algorithms under the same analytical conditions. To fill this gap, this study proposes a comparative machine learning framework for the classification of cancer incidence patterns in Iraqi governorates, using the 2018 and 2020 datasets.The analysis is focused on three age-specific cancer incidence variables, which are classified into low, medium, and high levels using a tertile-based classification method. Demographic and health-related predictors were incorporated and further enhanced by spatial feature engineering, including standardized variables, spatial lag effects, and spatial differences, to capture the spatial dependencies between neighboring governorates.The four classification algorithms applied and compared were Shrinkage Quadratic Discriminant Analysis (Shrinkage-QDA), Support Vector Machine (SVM), Random Forest, and Extreme Gradient Boosting (XGBoost). The performance of the models was assessed using accuracy and misclassification rates within a leave-one-out cross-validation (LOOCV) framework, which is suitable for small-sample spatial data.The results indicate a substantial variation in predictive performance across models and variables. Tree-based methods, particularly XGBoost and Random Forest, consistently demonstrated superior classification accuracy and lower misclassification rates, suggesting a greater capacity to model nonlinear relationships and spatial heterogeneity. On the other hand, SVM performed worst, and Shrinkage-QDA yielded moderate results. Temporal analysis showed a decrease in classification performance for one incidence category and an increase for the other categories in 2020. The overall results emphasize the importance of spatial feature engineering combined with machine learning, and provide evidence-based insights for cancer surveillance, spatial epidemiology, and public health planning in Iraq.

Downloads

Published

2026-08-05

How to Cite

Salih Al-Rawi, O. F. (2026). Comparative Machine Learning Framework for Classifying Cancer Incidence in Iraq ‎‎(2018–2020): Evidence from Random Forest, XGBoost, SVM, and Shrinkage-QDA. Statistics, Optimization & Information Computing. https://doi.org/10.19139/soic-2310-5070-3795

Issue

Section

Research Articles