{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T00:50:03Z","timestamp":1780534203374,"version":"3.54.1"},"reference-count":34,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,3,3]],"date-time":"2022-03-03T00:00:00Z","timestamp":1646265600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Science Foundation of China","doi-asserted-by":"crossref","award":["61876138"],"award-info":[{"award-number":["61876138"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2022,6,30]]},"abstract":"<jats:p>\n            Dynamic pricing plays an important role in solving the problems such as traffic load reduction, congestion control, and revenue improvement. Efficient dynamic pricing strategies can increase capacity utilization, total revenue of service providers, and the satisfaction of both passengers and drivers. Many proposed dynamic pricing technologies focus on short-term optimization and face poor scalability in modeling long-term goals for the limitations of solution optimality and prohibitive computation. In this article, a deep reinforcement learning framework is proposed to tackle the dynamic pricing problem for ride-hailing platforms. A soft actor-critic (SAC) algorithm is adopted in the reinforcement learning framework. First, the dynamic pricing problem is translated into a\n            <jats:bold>Markov Decision Process (MDP)<\/jats:bold>\n            and is set up in continuous action spaces, which is no need for the discretization of action space. Then, a new reward function is obtained by the order response rate and the KL-divergence between supply distribution and demand distribution. Experiments and case studies demonstrate that the proposed method outperforms the baselines in terms of order response rate and total revenue.\n          <\/jats:p>","DOI":"10.1145\/3474841","type":"journal-article","created":{"date-parts":[[2022,3,3]],"date-time":"2022-03-03T09:07:01Z","timestamp":1646298421000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":41,"title":["Deep Reinforcement Learning-based Trajectory Pricing on Ride-hailing Platforms"],"prefix":"10.1145","volume":"13","author":[{"given":"Jianbin","family":"Huang","sequence":"first","affiliation":[{"name":"Xidian University, XiAn, Shanxi, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Longji","family":"Huang","sequence":"additional","affiliation":[{"name":"Xidian University, XiAn, Shanxi, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meijuan","family":"Liu","sequence":"additional","affiliation":[{"name":"Xidian University, XiAn, Shanxi, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"He","family":"Li","sequence":"additional","affiliation":[{"name":"Xidian University, XiAn, Shanxi, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qinglin","family":"Tan","sequence":"additional","affiliation":[{"name":"Xidian University, XiAn, Shanxi, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaoke","family":"Ma","sequence":"additional","affiliation":[{"name":"Xidian University, XiAn, Shanxi, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiangtao","family":"Cui","sequence":"additional","affiliation":[{"name":"Xidian University, XiAn, Shanxi, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"De-Shuang","family":"Huang","sequence":"additional","affiliation":[{"name":"Guangxi Academy of Science, China and Tongji University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,3,3]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Didi Chuxing. Retrieved from http:\/\/www.didichuxing.com\/en\/."},{"key":"e_1_3_1_3_2","unstructured":"New York City Taxi and Limousine Commission Dataset. Retrieved from https:\/\/www1.nyc.gov\/site\/."},{"key":"e_1_3_1_4_2","unstructured":"New York City Trip Data. Retrieved from https:\/\/www1.nyc.gov\/site\/tlc\/about\/tlc-trip-record-data.page."},{"key":"e_1_3_1_5_2","unstructured":"Uber. Retrieved from https:\/\/www.uber.com."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3274895.3274928"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/2764468.2764527"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.trc.2019.08.019"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2018.1800"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3033274.3085098"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/2940716.2940798"},{"issue":"2859864","key":"e_1_3_1_12_2","article-title":"Pricing and matching with forward-looking buyers and sellers","author":"Chen Yiwei","year":"2018","unstructured":"Yiwei Chen and Ming Hu. 2018. Pricing and matching with forward-looking buyers and sellers. Rotman School Manag. Work. Pap.2859864 (2018).","journal-title":"Rotman School Manag. Work. Pap."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2018.3050"},{"key":"e_1_3_1_14_2","article-title":"Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor","author":"Haarnoja Tuomas","year":"2018","unstructured":"Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. arXiv preprint arXiv:1801.01290 (2018).","journal-title":"arXiv preprint arXiv:1801.01290"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSG.2015.2495145"},{"key":"e_1_3_1_16_2","article-title":"Continuous control with deep reinforcement learning","author":"Lillicrap Timothy P.","year":"2015","unstructured":"Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 (2015).","journal-title":"arXiv preprint arXiv:1509.02971"},{"key":"e_1_3_1_17_2","unstructured":"Jiaxi Liu Yidong Zhang Xiaoqing Wang Yuming Deng and Xingyu Wu. 2019. Dynamic pricing on e-commerce platform with deep reinforcement learning. arxiv:1912.02572 [cs.LG]"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.apenergy.2018.03.072"},{"key":"e_1_3_1_19_2","first-page":"120","volume-title":"Proceedings of SAI Intelligent Systems Conference","author":"Maestre Roberto","year":"2018","unstructured":"Roberto Maestre, Juan Duque, Alberto Rubio, and Juan Ar\u00e9valo. 2018. Reinforcement learning for fair dynamic pricing. In Proceedings of SAI Intelligent Systems Conference. Springer, 120\u2013135."},{"key":"e_1_3_1_20_2","article-title":"Playing Atari with deep reinforcement learning","author":"Mnih Volodymyr","year":"2013","unstructured":"Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing Atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 (2013).","journal-title":"arXiv preprint arXiv:1312.5602"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-013-5340-0"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2016.2614621"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.omega.2013.10.004"},{"issue":"6","key":"e_1_3_1_24_2","article-title":"A review of Uber, the growing alternative to traditional taxi service","volume":"51","author":"Rempel J.","year":"2014","unstructured":"J. Rempel. 2014. A review of Uber, the growing alternative to traditional taxi service. AFB AccessWorld\u00ae Mag. 51, 6 (2014).","journal-title":"AFB AccessWorld\u00ae Mag."},{"key":"e_1_3_1_25_2","volume-title":"Problem Solving with Reinforcement Learning","author":"Rummery Gavin Adrian","year":"1995","unstructured":"Gavin Adrian Rummery. 1995. Problem Solving with Reinforcement Learning. Ph.D. Dissertation. University of Cambridge."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2014.09.003"},{"key":"e_1_3_1_27_2","first-page":"1889","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Schulman John","year":"2015","unstructured":"John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015. Trust region policy optimization. In Proceedings of the International Conference on Machine Learning. 1889\u20131897."},{"key":"e_1_3_1_28_2","article-title":"Proximal policy optimization algorithms","author":"Schulman John","year":"2017","unstructured":"John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017).","journal-title":"arXiv preprint arXiv:1707.06347"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i02.5600"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.1998.712192"},{"key":"e_1_3_1_31_2","article-title":"Dynamic pricing and management for electric autonomous mobility on demand systems using reinforcement learning","author":"Turan Berkay","year":"2019","unstructured":"Berkay Turan, Ramtin Pedarsani, and Mahnoosh Alizadeh. 2019. Dynamic pricing and management for electric autonomous mobility on demand systems using reinforcement learning. arXiv preprint arXiv:1909.06962 (2019).","journal-title":"arXiv preprint arXiv:1909.06962"},{"key":"e_1_3_1_32_2","article-title":"Reinforcement learning for real-time pricing and scheduling control in EV charging stations","author":"Wang Shuoyao","year":"2019","unstructured":"Shuoyao Wang, Suzhi Bi, and Ying Jun Angela Zhang. 2019. Reinforcement learning for real-time pricing and scheduling control in EV charging stations. IEEE Trans. Industr. Inform. 30, 4 (2019), 2149\u20132159.","journal-title":"IEEE Trans. Industr. Inform."},{"key":"e_1_3_1_33_2","unstructured":"Christopher John Cornish Hellaby Watkins. 1989. Learning from delayed rewards. PhD Thesis. University of Cambridge England."},{"key":"e_1_3_1_34_2","article-title":"Dynamic pricing and matching in ride-hailing platforms","author":"Yan Chiwei","year":"2019","unstructured":"Chiwei Yan, Helin Zhu, Nikita Korolko, and Dawn Woodard. 2019. Dynamic pricing and matching in ride-hailing platforms. Nav. Res. Logist. 67, 8 (2019), 705\u2013724.","journal-title":"Nav. Res. Logist."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357799"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474841","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474841","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:28:44Z","timestamp":1750195724000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474841"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,3]]},"references-count":34,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,6,30]]}},"alternative-id":["10.1145\/3474841"],"URL":"https:\/\/doi.org\/10.1145\/3474841","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3,3]]},"assertion":[{"value":"2021-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-03-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}