{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T05:46:27Z","timestamp":1773380787590,"version":"3.50.1"},"reference-count":0,"publisher":"Slovenian Association Informatika","issue":"9","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJCAI"],"abstract":"<jats:p>This study addresses the Time-Dependent Vehicle Routing Problem with Time Windows (TD-VRPTW) for a single-vehicle urban distribution system in Jakarta. Time-dependent travel times are constructed from one week of hourly historical congestion profiles obtained from TomTom Traffic and preprocessed into time-varying speed factors that are mapped to 40- and 50-customer delivery instances with a common service window of 08:00\u201319:00. A Deep Q-Network (DQN) enhanced with Double DQN and Prioritized Experience Replay (PER) is trained end-to-end using a multilayer perceptron with two hidden layers (128 and 64 units, ReLU activations) to approximate the state\u2013action value function. The reward function penalizes time-dependent travel time, lateness with respect to customer time windows, long inter-customer jumps, and inter-cluster moves, thereby shaping the policy toward both schedule adherence and congestion-aware routing. For each scenario, the agent is trained for 1,000 episodes under three random seeds and evaluated on three representative weekdays (Monday, Wednesday, and Friday). Across all settings, the learned policy achieves a 100% on-time delivery rate with zero late customers, with best time-dependent route costs of approximately 526\u2013539 minutes for 40 customers and 595\u2013617 minutes for 50 customers. Comparative experiments with Genetic Algorithm (GA) and Ant Colony Optimization (ACO) show that ACO attains the shortest travel times, while the proposed DQN+PER model yields routes that are only about 5\u20138% longer than ACO but reduce time-dependent travel cost by roughly 35\u201345% compared with GA in the same TD-VRPTW instances. Reward and loss trajectories exhibit smooth convergence, and a sensitivity analysis on the lateness penalty confirms that the main conclusions are robust to hyperparameter variations. These findings demonstrate that leveraging historical congestion to build time-dependent travel times enables DQN-based control to produce competitive, congestion-aware solutions for TD-VRPTW in realistic urban distribution networks.<\/jats:p>","DOI":"10.31449\/inf.v50i9.12122","type":"journal-article","created":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T22:21:12Z","timestamp":1773354072000},"source":"Crossref","is-referenced-by-count":0,"title":["Double Deep Q-Network with Experience Replay for Time Dependent Vehicle Routing Problem with Time Windows Under Historical Congestion Constraints Rina"],"prefix":"10.31449","volume":"50","author":[{"given":"Rina","family":"Refianti","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alifurrohman","family":"Alifurrohman","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eri Prasetyo","family":"Wibowo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ina Siti","family":"Hasanah","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Achmad Benny","family":"Mutiara","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"16141","published-online":{"date-parts":[[2026,3,12]]},"container-title":["Informatica"],"original-title":[],"link":[{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/12122\/6546","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/12122\/6546","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T22:21:13Z","timestamp":1773354073000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/view\/12122"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,12]]},"references-count":0,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2026,3,12]]}},"URL":"https:\/\/doi.org\/10.31449\/inf.v50i9.12122","relation":{},"ISSN":["1854-3871","0350-5596"],"issn-type":[{"value":"1854-3871","type":"electronic"},{"value":"0350-5596","type":"print"}],"subject":[],"published":{"date-parts":[[2026,3,12]]}}}