Adaptive Operator Selection in ALNS Using Q-Learning: A Comparative Study of Softmax, UCB, epsilon-Greedy and QTS Strategies for CVRP

Authors

  • Hajar BOUALAMIA Sultan Moulay Slimane University
  • Imad HAFIDI Sultan Moulay Slimane University
  • Abdelmoutalib METRANE Cadi Ayyad University

DOI:

https://doi.org/10.19139/soic-2310-5070-3956

Keywords:

Capacitated Vehicle Routing Problem,, Metaheuristics, Adaptive Large Neighborhood Search, Reinforcement Learning, Q-Learning, Adaptive operator selection

Abstract

This paper presents a comprehensive comparative study of operator selection strategies within a previously proposed Q-learning-based Adaptive Large Neighborhood Search (ALNS) framework for solving the Capacitated Vehicle Routing Problem (CVRP). Unlike classical ALNS approaches, where operators are selected using predefined heuristic rules, the proposed framework dynamically learns effective destroy–repair operator pairs during the search process. In addition, a new Q-Value Thompson Sampling (QTS) action selection strategy is introduced and compared with the conventional Roulette Wheel Selection baseline as well as three widely used reinforcement learning policies, namely ϵ-greedy, Softmax, and Upper Confidence Bound (UCB). The five strategies are evaluated using representative benchmark instances from CVRPLIB under identical experimental conditions over 30 independent runs. The comparative analysis considers solution quality, computational time, convergence behaviour, robustness through boxplot analysis, and statistical significance using Friedman and Wilcoxon signed-rank tests. The results show that the investigated strategies exhibit complementary strengths. These findings provide new insights into the impact of action selection policies on adaptive operator selection and demonstrate that QTS constitutes a robust and competitive alternative within the ALNS framework.

Downloads

Published

2026-07-31

How to Cite

BOUALAMIA, H., HAFIDI, I., & METRANE, A. (2026). Adaptive Operator Selection in ALNS Using Q-Learning: A Comparative Study of Softmax, UCB, epsilon-Greedy and QTS Strategies for CVRP. Statistics, Optimization & Information Computing. https://doi.org/10.19139/soic-2310-5070-3956

Issue

Section

Research Articles

Categories