A Statistical Evaluation of Feature Selection Methods for Breast Cancer Classification: A Simulation and Real-Data Study

Authors

  • Hanaa Elgohari Department of Applied Statistics, Faculty of Commerce, Mansoura University, Mansoura, Egypt; Faculty of Business, Horus University, New Damietta, Egypt
  • Mohamed Zakariaa Fouda Department of Applied Statistics, Faculty of Commerce, Mansoura University, Mansoura, Egypt
  • Mohamed Asim Faculty of Business, Horus University, New Damietta, Egypt
  • Omar M. Elzeki Faculty of Computer Science and Engineering, New Mansoura University, New Mansoura, Egypt; Computer Science Department, Faculty of Computers and Information Systems, Mansoura University, Mansoura, Egypt

DOI:

https://doi.org/10.19139/soic-2310-5070-4378

Keywords:

Feature selection, Univariate selection (Select K-Best), Filter methods;, Data leakAage; Model stability

Abstract

Breast cancer remains one of the leading causes of mortality among women worldwide, and early, accuratediagnosis is central to improving survival outcomes. This work develops a statistical evaluation framework for breast cancerclassification that couples machine-learning classifiers with a family of feature selection techniques. Rather than focusingon predictive accuracy alone, we emphasise the statistical behaviour, stability, and robustness of the resulting models undervarying data conditions. Five feature selection strategies—filter methods, univariate selection (Select K-Best), model-basedfeature importance, principal component analysis (PCA), and a hybrid score-combining scheme—are paired with nineclassifiers, including logistic regression, support vector classification, random forests, gradient boosting, and XGBoost. Theevaluation combines a real-data study on the Wisconsin Diagnostic Breast Cancer (WDBC) dataset with a Monte Carlosimulation study that varies sample size, noise level, feature correlation, and the proportion of irrelevant features. In contrastto the original submission, all real-data experiments are performed with feature selection carried out strictly inside eachtraining partition of a 50-split resampling scheme, eliminating the information leakage that produced the previously reportedperfect scores. Under this leakage-free protocol, no classifier attains 100% accuracy; the strongest real-data performanceis obtained by PCA-based reduction paired with logistic regression (accuracy 0.978 ± 0.005, ROC-AUC 0.994 ± 0.002),while the filter, univariate, and hybrid selectors yield identical feature subsets and are statistically indistinguishable fromone another. The simulation study addresses a complementary question: the behaviour of the classifiers themselves undercontrolled conditions, rather than the ranking of selector–classifier pairs on the real data. Now including correlated and nonlinear data-generating designs, it shows that logistic regression maintains the highest mean accuracy and the highest stabilityacross all twelve conditions, whereas ensemble methods remain competitive but slightly more variable. Differences amongselectors are assessed with the Friedman test and Nemenyi and McNemar post-hoc procedures. Overall, the findings confirmthat feature relevance strongly governs model variance and generalisation, and that appropriately regularised linear modelsremain competitive with, and often more stable than, more complex ensembles.

Downloads

Published

2026-07-31

How to Cite

Elgohari, H., Fouda, M., Asim, M., & Elzeki, O. (2026). A Statistical Evaluation of Feature Selection Methods for Breast Cancer Classification: A Simulation and Real-Data Study. Statistics, Optimization & Information Computing, 16(3), 2826–2845. https://doi.org/10.19139/soic-2310-5070-4378

Issue

Section

Research Articles

Categories