A meta-learning enhanced deep reinforcement learning approach for generalizing across orienteering problem with time windows

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

The Orienteering Problem with Time Windows (OPTW) is a complex combinatorial optimization problem with applications in logistics, tourist route planning, and emergency services. Traditional methods for solving OPTW, including metaheuristics, often struggle with scalability, adaptability, and generalization to new instances. Recently, deep reinforcement learning (DRL) has shown promise in tackling routing problems. However, existing DRL methods typically rely on non-Markovian state representations and handcrafted masking rules, which limit their adaptability and generalization. This paper presents Meta Pointer Network for OPTW (MetaPNet-OPTW), a meta-learning-enhanced DRL framework that combines a Markovian state formulation with OR-based feasibility rules within a pointer network model. We introduce the Meta-Learning enhanced REINFORCE algorithm, which learns across diverse problem instances and enables rapid adaptation to unseen configurations with minimal fine-tuning. During inference, active search with beam search is used to refine solutions dynamically. Extensive experiments show that MetaPNet-OPTW outperforms existing DRL approaches in efficiency and generalization, and notably improves 20 of 33 best-known solutions on the Gavalas benchmark. We further provide a t-SNE analysis of the learned latent space, enriched with spatio-temporal statistics, which explains why the model excels on Gavalas instances while identifying harder clusters such as r2 and c2. This study contributes a scalable DRL framework for OPTW that not only achieves state-of-the-art performance but also provides new interpretability into benchmark difficulty and model adaptability.

키워드

Orienteering problem with time windowsDeep reinforcement learningPointer networkMeta-learning
제목
A meta-learning enhanced deep reinforcement learning approach for generalizing across orienteering problem with time windows
저자
BARDE STEPHANE김현준
DOI
10.1016/j.trc.2025.105450
발행일
2026-02
유형
Article
저널명
Transportation Research Part C: Emerging Technologies
183
페이지
1 ~ 29