Top-K Landmark Selection, Optical-Flow Fusion, and BorderlineSMOTE for Subject-Independent Facial Expression Recognition: A Leakage-Aware Cross-Dataset Evaluation

Authors

  • Sri Winarno Faculty of Computer Science, Universitas Dian Nuswantoro, Semarang, Indonesia
  • Farrikh Alzami Faculty of Computer Science, Universitas Dian Nuswantoro, Semarang, Indonesia
  • Dewi Agustini Santoso Faculty of Computer Science, Universitas Dian Nuswantoro, Semarang, Indonesia
  • Muhammad Naufal Information Engineering, Faculty of Computer Science, Universitas Dian Nuswantoro, Semarang, Indonesia
  • Harun Al Azies Information Engineering, Faculty of Computer Science, Universitas Dian Nuswantoro, Semarang, Indonesia
  • Kalaiarasi A/P Sonai Muthu Faculty of Information Science and Technoloty, MNA-R1003, Multimedia University, Jalan Ayer Keroh Lama, 75450, Bukit Beruang, Melaka, Malaysia

DOI:

https://doi.org/10.19139/soic-2310-5070-4519

Keywords:

facial expression recognition, landmark selection, optical flow, class imbalance, BorderlineSMOTE, CNN, GRU, subject-independent evaluation

Abstract

Facial expression recognition (FER) systems increasingly rely on landmark-based and motion-based representations to reduce sensitivity to illumination and identity, yet no published pipeline jointly combines Top-K landmark selection, optical-flow motion fusion, and resampling-based class-imbalance handling under a subject-disjoint evaluation protocol. This study proposes a CNN-GRU architecture that selects informative facial landmarks using an optical-flow-derived importance score, fuses geometric position and motion features, and applies BorderlineSMOTE oversampling to counter class imbalance, evaluated under a leakage-free, subject-disjoint cross-validation protocol on two datasets of contrasting scale: the Extended Cohn-Kanade dataset (CK+, 106 subjects) and the Indonesian Mixed Emotion Dataset (IMED, 15 subjects), each restricted to six basic emotion classes. Across five seeds, ten subject-disjoint folds, five landmark retention levels, three input modalities, and two resampling conditions, Friedman and Wilcoxon signed-rank tests with 95% confidence intervals show that landmark retention of 20 to 30 percent and BorderlineSMOTE significantly improve weighted F1 on CK+ (F1w~=~0.810) and, more weakly, on IMED (F1w~=~0.426), while modality fusion significantly outperforms single-modality input on CK+ but not on IMED. On IMED, the BorderlineSMOTE result is the one finding whose significance depends on the multiple-comparison correction applied: it survives Benjamini-Hochberg FDR correction but not the stricter Bonferroni correction, and is reported with this dependency stated explicitly rather than under a single, correction-agnostic conclusion. A Dataset-by-Factor interaction test shows all three effects, not only modality, are significantly smaller on IMED than on CK+; modality attenuates most severely, to roughly one-eighth of its CK+ size, which is why it alone loses significance within IMED, whereas landmark pruning and resampling retain enough of their larger CK+ margin to stay significant despite comparable proportional shrinkage. Every effect weakens on the smaller dataset, but only the narrowest-margin effect loses significance outright. This specific pattern has not previously been isolated under a controlled, identical-protocol comparison. Given IMED's fifteen-subject pool, its findings are reported as exploratory and dataset-limited rather than confirmatory.

Downloads

Published

2026-10-03

How to Cite

Winarno, S., Alzami, F., Santoso, D. A., Naufal, M., Azies, H. A., & Muthu, K. A. S. (2026). Top-K Landmark Selection, Optical-Flow Fusion, and BorderlineSMOTE for Subject-Independent Facial Expression Recognition: A Leakage-Aware Cross-Dataset Evaluation. Statistics, Optimization & Information Computing. https://doi.org/10.19139/soic-2310-5070-4519

Issue

Section

Research Articles

Categories

Most read articles by the same author(s)