Benchmarking Attention Mechanisms in ConvNeXt-Tiny for Dermoscopic Skin Lesion Classification Under Stratified and Lesion-Level Leakage-Safe Cross-Validation
DOI:
https://doi.org/10.19139/soic-2310-5070-4134Keywords:
Skin lesion classification, attention mechanisms, ConvNeXt, data leakage, cross-validation, dermoscopy, deep learning, lesion-level data leakageAbstract
Attention mechanisms are widely added to convolutional backbones to improve skin lesion classification, yet their comparative benefit is rarely assessed under evaluation protocols that prevent lesion-level data leakage. We present a controlled benchmark of four canonical attention modules—Channel Attention, Spatial Attention, the Convolutional Block Attention Module (CBAM), and Coordinate Attention—each inserted as a post-backbone plug-in on a shared ConvNeXt-Tiny feature extractor for seven-class dermoscopic classification on the HAM10000 dataset. Every configuration is evaluated twice: under conventional image-level stratified five-fold cross-validation, and under lesion-aware group five-fold cross-validation that ensures all images sharing the same lesion id remain within a single fold. Across all configurations, replacing stratified with group validation reduces the macro-averaged F1-score by an average of 10.82 percentage points, an inflation concentrated mainly in the minority classes—the rarest class loses more than 22 points of F1—and shows an exploratory negative association with class rarity (r = −0.66). The ranking of attention mechanisms is not preserved across protocols: Channel Attention attains the highest observed macro-F1 under image-level stratified validation, whereas Spatial Attention attains the highest observed macro-F1 under lesion-level group validation and is the only attention mechanism with a numerically higher mean than the corresponding attention-free baseline. However, none of the eight fold-level comparisons against the protocol-matched baselines remains statistically significant after Holm correction. Peak GPU-memory requirements are broadly comparable across configurations, although Coordinate Attention requires moderately longer training under lesion-level group validation. These results show that the cross-validation protocol influences reported performance more strongly than the choice of attention mechanism, and that image-level evaluation on HAM10000 can provide optimistic estimates of generalisation to unseen lesions. We recommend lesion-level group partitioning as the preferred protocol for trustworthy evaluation of skin lesion classifiers.Downloads
Published
2026-07-24
How to Cite
Muqtadir, A., Adi, K., & Gunawan, V. (2026). Benchmarking Attention Mechanisms in ConvNeXt-Tiny for Dermoscopic Skin Lesion Classification Under Stratified and Lesion-Level Leakage-Safe Cross-Validation. Statistics, Optimization & Information Computing. https://doi.org/10.19139/soic-2310-5070-4134
License
Copyright (c) 2026 Asfan Muqtadir, Kusworo Adi, Vincensius Gunawan

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).