A Categorical Approach to Classifying Poor Households in East Java: Utilizing Support Vector Machine (SVM) for SDG 1

Authors

  • Vita Ratnasari Department of Statistics, Faculty of Science and Data Analytics, Institut Teknologi Sepuluh Nopember, Surabaya 60111, Indonesia
  • Kadek Adi Surya Negara Department of Statistics, Faculty of Science and Data Analytics, Institut Teknologi Sepuluh Nopember, Surabaya 60111, Indonesia
  • Andrea Tri Rian Dani Doctoral Study Program of Mathematics and Natural Sciences, Faculty of Science and Technology, Universitas Airlangga, Surabaya 60115, Indonesia; Statistics Study Program, Department of Mathematics, Faculty of Mathematics and Natural Sciences, Universitas Mulawarman, Samarinda 75119, Indonesia
  • Muhammad Fikry Al-Farizi Department of Statistics, Faculty of Science and Data Analytics, Institut Teknologi Sepuluh Nopember, Surabaya 60111, Indonesia

DOI:

https://doi.org/10.19139/soic-2310-5070-3755

Keywords:

Index Terms-Poverty Classification, Poor Households, Support Vector Machine, Socioeconomic Indicators, SDGs 1

Abstract

Persistent inaccuracies in household poverty classification impede Indonesia’s progress toward Sustainable Development Goal 1 (No Poverty). In March 2023, East Java Province reported the highest absolute number of poor residents nationwide, underscoring the need for more reliable, data-driven targeting mechanisms. This study introduces a novel categorical modeling framework that integrates binary logistic regression, based variable selection with Support Vector Machine (SVM) classification, enhanced by Synthetic Minority Oversampling Technique (SMOTE) to address severe class imbalance. From the 2023 National Socioeconomic Survey (SUSENAS), eleven socioeconomic indicators were identified as statistically significant predictors of poverty status: household size, urban–rural residence, head-of-household employment, literacy, wall material, floor area, sanitation facility, primary drinking-water source, land ownership, motorcycle ownership, and ICT asset ownership. We then constructed linear and Radial Basis Function (RBF) SVM classifiers, each trained on an 80 % SMOTE-augmented dataset. The linear-kernel SVM, (C = 0.1) emerged as the best performer, achieving 78.73 % overall accuracy, 77.07 % sensitivity (poor-household recall), 78.82 % specificity (non-poor recall), a geometric mean of 77.94 %, and an AUC of 77.94 %. Compared to RBF-kernel alternatives, the linear model delivered superior balance between minority- and majority-class discrimination without incurring additional computational complexity. This research offers policymakers a robust, scalable tool for rapid poverty monitoring and targeted intervention design by demonstrating that a rigorously tuned, categorical SVM framework can substantially improve subnational poverty classification. Future work should explore alternative cross-validation schemes and nonlinear kernels to optimize predictive performance further.

Downloads

Published

2026-08-03

How to Cite

Ratnasari, V., Negara, K. A. S., Dani, A. T. R., & Al-Farizi, M. F. (2026). A Categorical Approach to Classifying Poor Households in East Java: Utilizing Support Vector Machine (SVM) for SDG 1. Statistics, Optimization & Information Computing. https://doi.org/10.19139/soic-2310-5070-3755

Issue

Section

Research Articles

Most read articles by the same author(s)