{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T07:56:43Z","timestamp":1772265403643,"version":"3.50.1"},"reference-count":32,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2021,4,21]],"date-time":"2021-04-21T00:00:00Z","timestamp":1618963200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61972161"],"award-info":[{"award-number":["61972161"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Hong Kong RGC General Research Fund","award":["152199\/17E"],"award-info":[{"award-number":["152199\/17E"]}]},{"DOI":"10.13039\/501100021171","name":"Guangdong Basic and Applied Basic Research Foundation","doi-asserted-by":"crossref","award":["2020A1515011496"],"award-info":[{"award-number":["2020A1515011496"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2021,6,30]]},"abstract":"<jats:p>Autonomous on-demand services, such as GOGOX (formerly GoGoVan) in Hong Kong, provide a platform for users to request services and for suppliers to meet such demands. In such a platform, the suppliers have autonomy to accept or reject the demands to be dispatched to him\/her, so it is challenging to make an online matching between demands and suppliers. Existing methods use round-based approaches to dispatch demands. In these works, the dispatching decision is based on the predicted response patterns of suppliers to demands in the current round, but they all fail to consider the impact of future demands and suppliers on the current dispatching decision. This could lead to taking a suboptimal dispatching decision from the future perspective. To solve this problem, we propose a novel demand dispatching model using deep reinforcement learning. In this model, we make each demand as an agent. The action of each agent, i.e., the dispatching decision of each demand, is determined by a centralized algorithm in a coordinated way. The model works in the following two steps. (1) It learns the demand\u2019s expected value in each spatiotemporal state using historical transition data. (2) Based on the learned values, it conducts a Many-To-Many dispatching using a combinatorial optimization algorithm by considering both immediate rewards and expected values of demands in the next round. In order to get a higher total reward, the demands with a high expected value (short response time) in the future may be delayed to the next round. On the contrary, the demands with a low expected value (long response time) in the future would be dispatched immediately. Through extensive experiments using real-world datasets, we show that the proposed model outperforms the existing models in terms of Cancellation Rate and Average Response Time.<\/jats:p>","DOI":"10.1145\/3442343","type":"journal-article","created":{"date-parts":[[2021,4,21]],"date-time":"2021-04-21T15:42:54Z","timestamp":1619019774000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Exploring Deep Reinforcement Learning for Task Dispatching in Autonomous On-Demand Services"],"prefix":"10.1145","volume":"15","author":[{"given":"Lei","family":"Yang","sequence":"first","affiliation":[{"name":"School of Software Engineering, South China University of Technology, Guangzhou, Guangdong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xi","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Software Engineering, South China University of Technology, Guangzhou, Guangdong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiannong","family":"Cao","sequence":"additional","affiliation":[{"name":"Department of Computing, The Hong Kong Polytechnic University, Kowloon, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuxun","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, South China University of Technology, Guangzhou, Guangdong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pan","family":"Zhou","sequence":"additional","affiliation":[{"name":"Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, and Huazhong University of Science and Technology, Wuhan, Hubei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,4,21]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the 8th International Conference on Autonomous Agents and Multiagent Systems. 21\u201328","author":"Alshamsi Aamena","year":"2009","unstructured":"Aamena Alshamsi , Sherief Abdallah , and Iyad Rahwan . 2009 . Multiagent self-organization for a taxi dispatch system . In Proceedings of the 8th International Conference on Autonomous Agents and Multiagent Systems. 21\u201328 . Aamena Alshamsi, Sherief Abdallah, and Iyad Rahwan. 2009. Multiagent self-organization for a taxi dispatch system. In Proceedings of the 8th International Conference on Autonomous Agents and Multiagent Systems. 21\u201328."},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the 10th IEEE International Conference on Collaborative Computing: Networking, Applications and Worksharing. 378\u2013387","author":"Arunapuram P.","year":"2014","unstructured":"P. Arunapuram , J. W. Bartel , and P. Dewan . 2014. Distribution, correlation and prediction of response times in Stack Overflow . In Proceedings of the 10th IEEE International Conference on Collaborative Computing: Networking, Applications and Worksharing. 378\u2013387 . DOI:https:\/\/doi.org\/10.4108\/icst.collaboratecom. 2014 .257265 10.4108\/icst.collaboratecom.2014.257265 P. Arunapuram, J. W. Bartel, and P. Dewan. 2014. Distribution, correlation and prediction of response times in Stack Overflow. In Proceedings of the 10th IEEE International Conference on Collaborative Computing: Networking, Applications and Worksharing. 378\u2013387. DOI:https:\/\/doi.org\/10.4108\/icst.collaboratecom.2014.257265"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the 2015 IEEE 14th International Conference on Machine Learning and Applications. 618\u2013623","author":"Burlutskiy N.","year":"2015","unstructured":"N. Burlutskiy , A. Fish , N. Ali , and M. Petridis . 2015. Prediction of users\u2019 response time in Q&A communities . In Proceedings of the 2015 IEEE 14th International Conference on Machine Learning and Applications. 618\u2013623 . DOI:https:\/\/doi.org\/10.1109\/ICMLA. 2015 .190 10.1109\/ICMLA.2015.190 N. Burlutskiy, A. Fish, N. Ali, and M. Petridis. 2015. Prediction of users\u2019 response time in Q&A communities. In Proceedings of the 2015 IEEE 14th International Conference on Machine Learning and Applications. 618\u2013623. DOI:https:\/\/doi.org\/10.1109\/ICMLA.2015.190"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the 1st International Conference on Distributed Artificial Intelligence. ACM, Article No. 7. https:\/\/doi.org\/10","author":"Chen Yong","year":"2018","unstructured":"Yong Chen , Ming Zhou , Ying Wen , Yaodong Yang , Yufeng Su , Weinan Zhang , Dell Zhang , Jun Wang , and Han Liu . 2018 . Factorized Q-learning for large-scale multi-agent systems . In Proceedings of the 1st International Conference on Distributed Artificial Intelligence. ACM, Article No. 7. https:\/\/doi.org\/10 .1145\/3356464.3357707 10.1145\/3356464.3357707 Yong Chen, Ming Zhou, Ying Wen, Yaodong Yang, Yufeng Su, Weinan Zhang, Dell Zhang, Jun Wang, and Han Liu. 2018. Factorized Q-learning for large-scale multi-agent systems. In Proceedings of the 1st International Conference on Distributed Artificial Intelligence. ACM, Article No. 7. https:\/\/doi.org\/10.1145\/3356464.3357707"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the 2017 IEEE 33rd International Conference on Data Engineering. IEEE, 997\u20131008","author":"Cheng P.","year":"2017","unstructured":"P. Cheng , X. Lian , L. Chen , and C. Shahabi . 2017. Prediction-based task assignment in spatial crowdsourcing . In Proceedings of the 2017 IEEE 33rd International Conference on Data Engineering. IEEE, 997\u20131008 . DOI:https:\/\/doi.org\/10.1109\/ICDE. 2017 .146 10.1109\/ICDE.2017.146 P. Cheng, X. Lian, L. Chen, and C. Shahabi. 2017. Prediction-based task assignment in spatial crowdsourcing. In Proceedings of the 2017 IEEE 33rd International Conference on Data Engineering. IEEE, 997\u20131008. DOI:https:\/\/doi.org\/10.1109\/ICDE.2017.146"},{"key":"e_1_2_1_6_1","volume-title":"C","author":"Geiger David","year":"2014","unstructured":"David Geiger and Martin Schader . 2014. Personalized task recommendation in crowdsourcing information systems\u2014current state of the art. Decision Support Systems 65 , C ( 2014 ), 3\u201316. David Geiger and Martin Schader. 2014. Personalized task recommendation in crowdsourcing information systems\u2014current state of the art. Decision Support Systems 65, C (2014), 3\u201316."},{"key":"e_1_2_1_7_1","unstructured":"GOGOX. 2020. GOGOX Hong Kong. Retrieved from https:\/\/www.gogox.com.hk.  GOGOX. 2020. GOGOX Hong Kong. Retrieved from https:\/\/www.gogox.com.hk."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357978"},{"key":"e_1_2_1_9_1","article-title":"Learning to delay in ride-sourcing systems: a multi-agent deep reinforcement learning framework","volume":"10","author":"Ke Jintao","year":"2020","unstructured":"Jintao Ke , Feng Xiao , Hai Yang , and Jieping Ye . 2020 . Learning to delay in ride-sourcing systems: a multi-agent deep reinforcement learning framework . IEEE Transactions on Knowledge and Data Engineering. DOI : 10 .1109\/TKDE.2020.3006084 10.1109\/TKDE.2020.3006084 Jintao Ke, Feng Xiao, Hai Yang, and Jieping Ye. 2020. Learning to delay in ride-sourcing systems: a multi-agent deep reinforcement learning framework. IEEE Transactions on Knowledge and Data Engineering. DOI:10.1109\/TKDE.2020.3006084","journal-title":"IEEE Transactions on Knowledge and Data Engineering. DOI"},{"key":"e_1_2_1_10_1","first-page":"3","article-title":"Context-aware hierarchical online learning for performance maximization in mobile crowdsourcing","volume":"26","author":"\u00e9e M\u00fcller S. Klos","year":"2018","unstructured":"S. Klos n \u00e9e M\u00fcller , C. Tekin , M. van der Schaar , and A. Klein . 2018 . Context-aware hierarchical online learning for performance maximization in mobile crowdsourcing . IEEE\/ACM Transactions on Networking 26 , 3 (Jun. 2018), 1334\u20131347. DOI:https:\/\/doi.org\/10.1109\/TNET.2018.2828415 10.1109\/TNET.2018.2828415 S. Klos n\u00e9e M\u00fcller, C. Tekin, M. van der Schaar, and A. Klein. 2018. Context-aware hierarchical online learning for performance maximization in mobile crowdsourcing. IEEE\/ACM Transactions on Networking 26, 3 (Jun. 2018), 1334\u20131347. DOI:https:\/\/doi.org\/10.1109\/TNET.2018.2828415","journal-title":"IEEE\/ACM Transactions on Networking"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.3141\/1882-23"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/MITS.2019.2921082"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219993"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 7th International AAAI Conference on Weblogs and Social Media.","author":"Mahmud Jalal","year":"2013","unstructured":"Jalal Mahmud , Jilin Chen , and Jeffrey Nichols . 2013 . When will you answer this? Estimating response time in Twitter . In Proceedings of the 7th International AAAI Conference on Weblogs and Social Media. Jalal Mahmud, Jilin Chen, and Jeffrey Nichols. 2013. When will you answer this? Estimating response time in Twitter. In Proceedings of the 7th International AAAI Conference on Weblogs and Social Media."},{"key":"e_1_2_1_15_1","volume-title":"Human-level control through deep reinforcement learning. Nature 518, 7540","author":"Mnih Volodymyr","year":"2015","unstructured":"Volodymyr Mnih , Koray Kavukcuoglu , David Silver , Andrei A. Rusu , Joel Veness , Marc G. Bellemare , Alex Graves , Martin Riedmiller , Andreas K. Fidjeland , Georg Ostrovski , Stig Petersen , Charles Beattie , Amir Sadik , Ioannis Antonoglou , Helen King , Dharshan Kumaran , Daan Wierstra , Shane Legg , and Demis Hassabis . 2015. Human-level control through deep reinforcement learning. Nature 518, 7540 ( 2015 ), 529\u2013533. DOI:https:\/\/doi.org\/10.1038\/nature14236 10.1038\/nature14236 Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015. Human-level control through deep reinforcement learning. Nature 518, 7540 (2015), 529\u2013533. DOI:https:\/\/doi.org\/10.1038\/nature14236"},{"key":"e_1_2_1_16_1","first-page":"196","article-title":"Algorithms for the assignment and transportation problems","volume":"10","author":"Munkres James","year":"1957","unstructured":"James Munkres . 1957 . Algorithms for the assignment and transportation problems . Journal of the Society for Industrial and Applied Mathematics 10 , 1 (1957), 196 \u2013 210 . James Munkres. 1957. Algorithms for the assignment and transportation problems. Journal of the Society for Industrial and Applied Mathematics 10, 1 (1957), 196\u2013210.","journal-title":"Journal of the Society for Industrial and Applied Mathematics"},{"key":"e_1_2_1_17_1","first-page":"3","article-title":"A collaborative multiagent taxi-dispatch system","volume":"7","author":"Seow K. T.","year":"2010","unstructured":"K. T. Seow , N. H. Dang , and D. Lee . 2010 . A collaborative multiagent taxi-dispatch system . IEEE Transactions on Automation Science and Engineering 7 , 3 (Jul. 2010), 607\u2013616. DOI:https:\/\/doi.org\/10.1109\/TASE.2009.2028577 10.1109\/TASE.2009.2028577 K. T. Seow, N. H. Dang, and D. Lee. 2010. A collaborative multiagent taxi-dispatch system. IEEE Transactions on Automation Science and Engineering 7, 3 (Jul. 2010), 607\u2013616. DOI:https:\/\/doi.org\/10.1109\/TASE.2009.2028577","journal-title":"IEEE Transactions on Automation Science and Engineering"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330724"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2729713"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994523"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2016.7498228"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137643"},{"key":"e_1_2_1_23_1","first-page":"2948863","article-title":"Two-sided online micro-task assignment in spatial crowdsourcing","volume":"2019","author":"Tong Y.","year":"2019","unstructured":"Y. Tong , Y. Zeng , B. Ding , L. Wang , and L. Chen . 2019 . Two-sided online micro-task assignment in spatial crowdsourcing . IEEE Transactions on Knowledge and Data Engineering. DOI:https:\/\/doi.org\/10.1109\/TKDE. 2019 . 2948863 10.1109\/TKDE.2019.2948863 Y. Tong, Y. Zeng, B. Ding, L. Wang, and L. Chen. 2019. Two-sided online micro-task assignment in spatial crowdsourcing. IEEE Transactions on Knowledge and Data Engineering. DOI:https:\/\/doi.org\/10.1109\/TKDE.2019.2948863","journal-title":"IEEE Transactions on Knowledge and Data Engineering. DOI:https:\/\/doi.org\/10.1109\/TKDE."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-019-00568-7"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the 2017 ACM Conference on Information and Knowledge Management. ACM","author":"Wang Yuqi","unstructured":"Yuqi Wang , Jiannong Cao , Lifang He , Wengen Li , Lichao Sun , and Philip S. Yu . 2017. Coupled sparse matrix factorization for response time prediction in logistics services . In Proceedings of the 2017 ACM Conference on Information and Knowledge Management. ACM , New York, NY, 939\u2013947. DOI:https:\/\/doi.org\/10.1145\/3132847.3132948 10.1145\/3132847.3132948 Yuqi Wang, Jiannong Cao, Lifang He, Wengen Li, Lichao Sun, and Philip S. Yu. 2017. Coupled sparse matrix factorization for response time prediction in logistics services. In Proceedings of the 2017 ACM Conference on Information and Knowledge Management. ACM, New York, NY, 939\u2013947. DOI:https:\/\/doi.org\/10.1145\/3132847.3132948"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 2019 IEEE 35th International Conference on Data Engineering. IEEE, 1478\u20131489","author":"Wang Y.","year":"2019","unstructured":"Y. Wang , Y. Tong , C. Long , P. Xu , K. Xu , and W. Lv . 2019. Adaptive dynamic bipartite graph matching: A reinforcement learning approach . In Proceedings of the 2019 IEEE 35th International Conference on Data Engineering. IEEE, 1478\u20131489 . DOI:https:\/\/doi.org\/10.1109\/ICDE. 2019 .00133 10.1109\/ICDE.2019.00133 Y. Wang, Y. Tong, C. Long, P. Xu, K. Xu, and W. Lv. 2019. Adaptive dynamic bipartite graph matching: A reinforcement learning approach. In Proceedings of the 2019 IEEE 35th International Conference on Data Engineering. IEEE, 1478\u20131489. DOI:https:\/\/doi.org\/10.1109\/ICDE.2019.00133"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00077"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219824"},{"key":"e_1_2_1_29_1","first-page":"2941680","article-title":"A novel demand dispatching model for autonomous on-demand services","volume":"2019","author":"Yang L.","year":"2019","unstructured":"L. Yang , X. Yu , J. Cao , W. Li , Y. Wang , and M. Szczecinski . 2019 . A novel demand dispatching model for autonomous on-demand services . IEEE Transactions on Services Computing. DOI:https:\/\/doi.org\/10.1109\/TSC. 2019 . 2941680 10.1109\/TSC.2019.2941680 L. Yang, X. Yu, J. Cao, W. Li, Y. Wang, and M. Szczecinski. 2019. A novel demand dispatching model for autonomous on-demand services. IEEE Transactions on Services Computing. DOI:https:\/\/doi.org\/10.1109\/TSC.2019.2941680","journal-title":"IEEE Transactions on Services Computing. DOI:https:\/\/doi.org\/10.1109\/TSC."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098138"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2017.2703848"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.14778\/3204028.3204030"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3442343","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3442343","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:15Z","timestamp":1750193295000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3442343"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,21]]},"references-count":32,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,6,30]]}},"alternative-id":["10.1145\/3442343"],"URL":"https:\/\/doi.org\/10.1145\/3442343","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,4,21]]},"assertion":[{"value":"2020-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-04-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}