A Statistical Evaluation of Feature Selection Methods for Breast Cancer Classification: A Simulation and Real-Data Study
DOI:
https://doi.org/10.19139/soic-2310-5070-4378Keywords:
Feature selection, Univariate selection (Select K-Best), Filter methods;, Data leakAage; Model stabilityAbstract
Breast cancer remains one of the leading causes of mortality among women worldwide, and early, accuratediagnosis is central to improving survival outcomes. This work develops a statistical evaluation framework for breast cancerclassification that couples machine-learning classifiers with a family of feature selection techniques. Rather than focusingon predictive accuracy alone, we emphasise the statistical behaviour, stability, and robustness of the resulting models undervarying data conditions. Five feature selection strategies—filter methods, univariate selection (Select K-Best), model-basedfeature importance, principal component analysis (PCA), and a hybrid score-combining scheme—are paired with nineclassifiers, including logistic regression, support vector classification, random forests, gradient boosting, and XGBoost. Theevaluation combines a real-data study on the Wisconsin Diagnostic Breast Cancer (WDBC) dataset with a Monte Carlosimulation study that varies sample size, noise level, feature correlation, and the proportion of irrelevant features. In contrastto the original submission, all real-data experiments are performed with feature selection carried out strictly inside eachtraining partition of a 50-split resampling scheme, eliminating the information leakage that produced the previously reportedperfect scores. Under this leakage-free protocol, no classifier attains 100% accuracy; the strongest real-data performanceis obtained by PCA-based reduction paired with logistic regression (accuracy 0.978 ± 0.005, ROC-AUC 0.994 ± 0.002),while the filter, univariate, and hybrid selectors yield identical feature subsets and are statistically indistinguishable fromone another. The simulation study addresses a complementary question: the behaviour of the classifiers themselves undercontrolled conditions, rather than the ranking of selector–classifier pairs on the real data. Now including correlated and nonlinear data-generating designs, it shows that logistic regression maintains the highest mean accuracy and the highest stabilityacross all twelve conditions, whereas ensemble methods remain competitive but slightly more variable. Differences amongselectors are assessed with the Friedman test and Nemenyi and McNemar post-hoc procedures. Overall, the findings confirmthat feature relevance strongly governs model variance and generalisation, and that appropriately regularised linear modelsremain competitive with, and often more stable than, more complex ensembles.Downloads
Published
2026-07-31
How to Cite
Elgohari, H., Fouda, M., Asim, M., & Elzeki, O. (2026). A Statistical Evaluation of Feature Selection Methods for Breast Cancer Classification: A Simulation and Real-Data Study. Statistics, Optimization & Information Computing, 16(3), 2826–2845. https://doi.org/10.19139/soic-2310-5070-4378
License
Copyright (c) 2026 Hanaa Elgohari, Mohamed Zakariaa Fouda , Mohamed Asim, Omar M. Elzeki

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).