{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T14:29:02Z","timestamp":1780496942835,"version":"3.54.1"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,1,18]],"date-time":"2022-01-18T00:00:00Z","timestamp":1642464000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["61902134, 62011530437"],"award-info":[{"award-number":["61902134, 62011530437"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003819","name":"Hubei Natural Science Foundation","doi-asserted-by":"crossref","award":["2020CFB871"],"award-info":[{"award-number":["2020CFB871"]}],"id":[{"id":"10.13039\/501100003819","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["2019kfyXKJC021, 2019kfyXJJS091"],"award-info":[{"award-number":["2019kfyXKJC021, 2019kfyXJJS091"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2022,6,30]]},"abstract":"<jats:p>Online ride-hailing platforms have reduced significantly the amounts of the time that taxis are idle and that passengers spend on waiting. As a key component of these platforms, the fleet management problem can be naturally modeled as a Markov Decision Process, which enables us to use the deep reinforcement learning. However, existing studies are proposed based on simplified problem settings that fail to model the complicated supply-dynamics and restrict the performance in the real traffic environment. In this article, we propose a supply-demand-aware deep reinforcement learning algorithm for taxi dispatching, where we use a deep Q-network with action sampling policy, called AS-DQN, to learn an optimal dispatching policy. Furthermore, we utilize a dueling network architecture, called AS-DDQN, to improve the performance of AS-DQN. Extensive experiments on real-world datasets offer insight into the performance of our model and show that it is capable of outperforming the baseline approaches.<\/jats:p>","DOI":"10.1145\/3467979","type":"journal-article","created":{"date-parts":[[2022,1,18]],"date-time":"2022-01-18T15:36:16Z","timestamp":1642520176000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":21,"title":["Supply-Demand-aware Deep Reinforcement Learning for Dynamic Fleet Management"],"prefix":"10.1145","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8639-4570","authenticated-orcid":false,"given":"Bolong","family":"Zheng","sequence":"first","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lingfeng","family":"Ming","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qi","family":"Hu","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhipeng","family":"L\u00fc","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guanfeng","family":"Liu","sequence":"additional","affiliation":[{"name":"Macquarie University, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaofang","family":"Zhou","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Kowloon, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,1,18]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"2021. Didi. Retrieved from https:\/\/www.xiaojukeji.com."},{"key":"e_1_3_2_3_2","unstructured":"2021. Uber. Retrieved from https:\/\/www.uber.com."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.5555\/3367243.3367314"},{"key":"e_1_3_2_5_2","volume-title":"NeurIPS","author":"Bai Lei","year":"2020","unstructured":"Lei Bai, Lina Yao, Can Li, Xianzhi Wang, and Can Wang. 2020. Adaptive graph convolutional recurrent network for traffic forecasting. In NeurIPS."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2018.1763"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2018.1822"},{"key":"e_1_3_2_8_2","first-page":"3054","volume-title":"KDD","author":"Duan Lu","year":"2020","unstructured":"Lu Duan, Yang Zhan, Haoyuan Hu, Yu Gong, Jiangwen Wei, Xiaodong Zhang, and Yinghui Xu. 2020. Efficiently solving the practical vehicle routing problem: A novel joint learning approach. In KDD. 3054\u20133063."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1080\/13658816.2018.1458984"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313401"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3347146.3363349"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357978"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.5555\/1622737.1622748"},{"key":"e_1_3_2_14_2","volume-title":"ICLR","author":"Kaiser Lukasz","year":"2020","unstructured":"Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H. Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, Afroz Mohiuddin, Ryan Sepassi, George Tucker, and Henryk Michalewski. 2020. Model based reinforcement learning for atari. In ICLR."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313433"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397536.3427186"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41019-020-00142-0"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411871"},{"key":"e_1_3_2_20_2","volume-title":"ICLR (Poster)","author":"Lillicrap Timothy P.","year":"2016","unstructured":"Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016. Continuous control with deep reinforcement learning. In ICLR (Poster)."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219993"},{"key":"e_1_3_2_22_2","first-page":"1","article-title":"Context-aware taxi dispatching at city-scale using deep reinforcement learning","author":"Liu Zhidan","year":"2020","unstructured":"Zhidan Liu, Jiangzhou Li, and Kaishun Wu. 2020. Context-aware taxi dispatching at city-scale using deep reinforcement learning. IEEE Trans. Intell. Transport. Syst. (2020), 1\u201314.","journal-title":"IEEE Trans. Intell. Transport. Syst."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASE.2016.2529580"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397536.3427187"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.5555\/3045390.3045594"},{"key":"e_1_3_2_26_2","article-title":"Playing atari with deep reinforcement learning","volume":"1312","author":"Mnih Volodymyr","year":"2013","unstructured":"Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller. 2013. Playing atari with deep reinforcement learning. CoRR abs\/1312.5602.","journal-title":"CoRR"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_28_2","first-page":"2708","volume-title":"INFOCOM","author":"Oda Takuma","year":"2018","unstructured":"Takuma Oda and Carlee Joe-Wong. 2018. MOVI: A model-free approach to dynamic fleet management. In INFOCOM. 2708\u20132716."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623668"},{"key":"e_1_3_2_30_2","volume-title":"ICLR (Poster)","author":"Schaul Tom","year":"2016","unstructured":"Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. 2016. Prioritized experience replay. In ICLR (Poster)."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature16961"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.5555\/551283"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.5555\/3009657.3009806"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330724"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41019-020-00117-1"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.5555\/3016100.3016191"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-018-0095-1"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.5555\/3367722.3367882"},{"key":"e_1_3_2_39_2","first-page":"617","volume-title":"ICDM","author":"Wang Zhaodong","year":"2018","unstructured":"Zhaodong Wang, Zhiwei Qin, Xiaocheng Tang, Jieping Ye, and Hongtu Zhu. 2018. Deep reinforcement learning with knowledge transfer for online rides order dispatching. In ICDM. 617\u2013626."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.5555\/3045390.3045601"},{"key":"e_1_3_2_41_2","first-page":"220","volume-title":"ITSC","author":"Wen Jian","year":"2017","unstructured":"Jian Wen, Jinhua Zhao, and Patrick Jaillet. 2017. Rebalancing shared mobility-on-demand systems: A reinforcement learning approach. In ITSC. 220\u2013225."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219824"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380287"},{"key":"e_1_3_2_44_2","first-page":"6672","volume-title":"AAAI","author":"Ye Deheng","year":"2020","unstructured":"Deheng Ye, Zhao Liu, Mingfei Sun, Bei Shi, Peilin Zhao, Hao Wu, Hongsheng Yu, Shaojie Yang, Xipeng Wu, Qingwei Guo, Qiaobo Chen, Yinyuting Yin, Hao Zhang, Tengfei Shi, Liang Wang, Qiang Fu, Wei Yang, and Lanxiao Huang. 2020. Mastering complex control in MOBA games with deep reinforcement learning. In AAAI. 6672\u20136679."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.5555\/3304222.3304273"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","unstructured":"Bolong Zheng Qi Hu Lingfeng Ming Jilin Hu Lu Chen Kai Zheng and Christian S. Jensen. 2020. Spatial-temporal demand forecasting and competitive supply. DOI:10.1109\/TKDE.2021.3110778","DOI":"10.1109\/TKDE.2021.3110778"},{"key":"e_1_3_2_47_2","first-page":"2645","volume-title":"CIKM","author":"Zhou Ming","year":"2019","unstructured":"Ming Zhou, Jiarui Jin, Weinan Zhang, Zhiwei Qin, Yan Jiao, Chenxi Wang, Guobin Wu, Yong Yu, and Jieping Ye. 2019. Multi-agent reinforcement learning for order-dispatching via order-vehicle distribution matching. In CIKM. 2645\u20132653."}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3467979","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3467979","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:19:05Z","timestamp":1750191545000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3467979"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,18]]},"references-count":46,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,6,30]]}},"alternative-id":["10.1145\/3467979"],"URL":"https:\/\/doi.org\/10.1145\/3467979","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,18]]},"assertion":[{"value":"2021-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}