A Systematic Comparison of Cost-Sensitive Methods for Long-Tail Intent Classification in Indonesian Government Complaint Routing

Authors

  • Dewi Agustini Santoso Faculty of Computer Science, Universitas Dian Nuswantoro
  • Heni Indrayani Faculty of Computer Science, Universitas Dian Nuswantoro
  • Karis Widyatmoko Faculty of Computer Science, Universitas Dian Nuswantoro
  • Farrikh Alzami Universitas Dian Nuswantoro
  • Iswahyudi Iswahyudi Dinas Komunikasi dan Informatika Provinsi Jawa Tengah

DOI:

https://doi.org/10.19139/soic-2310-5070-4091

Keywords:

Cost-sensitive learning, Long-tail classification, Indonesian NLP, Government Complaint Routing, XLM-RoBERTa, Focal Loss, Class Imbalance

Abstract

E-government complaint routing systems face extreme class imbalance under which standard transformer fine-tuning leaves most minority agencies without correct routings. This study compares two cost-sensitive loss modification methods, inverse-frequency class weighting and focal loss, for long-tail intent classification in Indonesian government complaint routing. The dataset comprises 56,068 Indonesian-Javanese-English code-mixed complaints across 86 government agencies with an imbalance ratio of 332.86:1, among the most extreme long-tail distributions documented in e-government NLP research. XLM-RoBERTa-base was fine-tuned under three configurations: a CrossEntropyLoss baseline, inverse-frequency class weighting, and focal loss with a focusing parameter of 2.0. Evaluation employed Pareto-based head/tail segmentation, where 33 head classes account for 80% of training data and 53 tail classes cover the remaining 20%. Experiments were repeated across three independent random seeds to assess stability. Class weighting improved tail macro F1 from 0.2048 to 0.3426, a 67.3% gain, with a corresponding 9.4% decline in head macro F1. Focal loss yielded a 63.7% tail gain but incurred a 15.4% head performance reduction. A Friedman test comparing macro F1 across three methods and three seeds yields a test statistic of 6.00 with two degrees of freedom and a p-value of 0.050, a boundary result that reflects perfectly consistent method ranking across seeds and is interpreted as directional evidence rather than conclusive proof. Segment-level Wilcoxon tests on per-class F1 are individually significant in every seed, for both the tail improvement and the head decline, with p-values below 0.001 in all cases, indicating that the non-significant pooled test conceals two opposing within-segment effects. Class weighting demonstrated superior balance, achieving a higher overall macro F1 (0.4441) and higher tail F1 than focal loss (0.4241) under identical training conditions. Spearman rank correlation between class training frequency and F1 improvement is consistently negative across all seeds (mean correlation of -0.585, p-value below 0.001), confirming that lower-frequency agencies benefit systematically from cost-sensitive training. The Pareto-based evaluation convention adopted here provides a reproducible, threshold-free decomposition of imbalanced classifier performance that transfers to other domains with majority/minority class structure.

Downloads

Published

2026-08-08

How to Cite

Santoso, D. A., Indrayani, H., Widyatmoko, K., Alzami, F., & Iswahyudi, I. (2026). A Systematic Comparison of Cost-Sensitive Methods for Long-Tail Intent Classification in Indonesian Government Complaint Routing. Statistics, Optimization & Information Computing. https://doi.org/10.19139/soic-2310-5070-4091

Issue

Section

Research Articles

Categories

Most read articles by the same author(s)