{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,22]],"date-time":"2025-12-22T21:40:59Z","timestamp":1766439659650,"version":"3.48.0"},"reference-count":42,"publisher":"Springer Science and Business Media LLC","issue":"12","license":[{"start":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T00:00:00Z","timestamp":1761177600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T00:00:00Z","timestamp":1761177600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2025,12]]},"DOI":"10.1007\/s10994-025-06846-6","type":"journal-article","created":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T20:05:29Z","timestamp":1761249929000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Learning de-biased environment models for delivery incentive policy optimization on food delivery platforms"],"prefix":"10.1007","volume":"114","author":[{"given":"Yu-Ren","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiong-Hui","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Siyuan","family":"Xiao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyu","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xintong","family":"Qi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Linjun","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Yu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fangsheng","family":"Huang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,10,23]]},"reference":[{"key":"6846_CR1","doi-asserted-by":"crossref","unstructured":"Ai, M., Li, B., Gong, H., Yu, Q., Xue, S., Zhang, Y., Zhang, Y., & Jiang, P. (2022). LBCF: A large-scale budget-constrained causal forest algorithm. Proceedings of the 2022 ACM web conference (pp. 2310\u20132319). Lyon.","DOI":"10.1145\/3485447.3512103"},{"key":"6846_CR2","unstructured":"Alaa, A. M., & van der Schaar, M. (2018). Limits of estimating heterogeneous treatment effects: Guidelines for practical algorithm design. Proceedings of the 35th international conference on machine learning (Vol. 80, pp. 129\u2013138). Stockholm."},{"key":"6846_CR3","doi-asserted-by":"crossref","unstructured":"Anand, A. S., Kveen, J. E., Abu-Dakka, F. J., Gr\u00f8tli, E. I., & Gravdahl, J. T. (2022). Addressing sample efficiency and model-bias in model-based reinforcement learning. Proceedings of the 21st IEEE international conference on machine learning and applications (pp. 1\u20136). Nassau.","DOI":"10.1109\/ICMLA55696.2022.00009"},{"key":"6846_CR4","unstructured":"Assaad, S., Zeng, S., Tao, C., Datta, S., Mehta, N., Henao, R., Li, F., & Carin, L. (2021). Counterfactual representation learning with balancing weights. Proceedings of the 24th international conference on artificial intelligence and statistics (Vol. 130, pp. 1972\u20131980). Virtual Event."},{"key":"6846_CR5","unstructured":"Bica, I., Jordon, J., & van der Schaar, M. (2020). Estimating the effects of continuous-valued interventions using generative adversarial networks. Advances in neural information processing systems 33. Virtual Event."},{"key":"6846_CR6","doi-asserted-by":"crossref","unstructured":"Chen, X., He, B., Yu, Y., Li, Q., Qin, Z.T., Shang, W., Ye, J., & Ma, C. (2023). Sim2rec: A simulator-based decision-making approach to optimize real-world long-term user engagement in sequential recommender systems. Proceedings of the 39th IEEE international conference on data engineering (pp. 3389\u20133402). Anaheim.","DOI":"10.1109\/ICDE55515.2023.00260"},{"key":"6846_CR7","unstructured":"Chen, X., Li, S., Li, H., Jiang, S., Qi, Y., & Song, L. (2019). Generative adversarial user model for reinforcement learning based recommendation system. Proceedings of the 36th international conference on machine learning (pp. 1052\u20131061). Long Beach."},{"issue":"12","key":"6846_CR8","doi-asserted-by":"publisher","first-page":"15260","DOI":"10.1109\/TPAMI.2023.3317131","volume":"45","author":"X Chen","year":"2023","unstructured":"Chen, X., Luo, F., Yu, Y., Li, Q., Qin, Z., Shang, W., & Ye, J. (2023). Offline model-based adaptable policy learning for decision-making in out-of-support regions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(12), 15260\u201315274.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"6846_CR9","unstructured":"Chen, X., Yu, Y., Li, Q., Luo, F., Qin, Z.T., Shang, W., & Ye, J. (2021). Offline model- based adaptable policy learning. Advances in neural information processing systems 34 (pp. 8432\u20138443). Virtual Event."},{"key":"6846_CR10","unstructured":"Chen, X., Yu, Y., Zhu, Z., Yu, Z., Chen, Z., Wang, C., Wu, Y., Qin, R., Wu, H., Ding, R., & Huang, F. (2023). Adversarial counterfactual environment model learning. Advances in neural information processing systems 36. New Orleans."},{"key":"6846_CR11","doi-asserted-by":"crossref","unstructured":"Ding, X., Zhang, R., Mao, Z., Xing, K., Du, F., Liu, X., Wei, G., Yin, F., He, R., & Sun, Z. (2020). Delivery scope: A new way of restaurant retrieval for on-demand food delivery service. Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining (pp. 3026\u20133034). Virtual Event.","DOI":"10.1145\/3394486.3403353"},{"key":"6846_CR12","unstructured":"Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., & Bengio, Y. (2014). Generative adversarial nets. Advances in neural information processing systems 27 (pp. 2672\u20132680). Montr\u00e9al."},{"key":"6846_CR13","unstructured":"Ho, J., & Ermon, S. (2016). Generative adversarial imitation learning. Advances in neural information processing systems 29 (pp. 4565\u20134573). Barcelona."},{"key":"6846_CR14","unstructured":"Ie, E., Hsu, C., Mladenov, M., Jain, V., Narvekar, S., Wang, J., Wu, R., & Boutilier, C. (2019). RecSim: A configurable simulation platform for recommender systems. CoRR, abs\/1909.04847."},{"key":"6846_CR15","unstructured":"Jung, Y., Tian, J., & Bareinboim, E. (2020). Learning causal effects via weighted empirical risk minimization. Advances in neural information processing systems 33. Virtual Event."},{"key":"6846_CR16","unstructured":"Kidambi, R., Rajeswaran, A., Netrapalli, P., & Joachims, T. (2020). MOReL: Model- based offline reinforcement learning. Advances in neural information processing systems 33 (pp. 21810\u201321823). Virtual Event."},{"issue":"4","key":"6846_CR17","doi-asserted-by":"publisher","first-page":"969","DOI":"10.1007\/s00291-017-0501-3","volume":"40","author":"R Klein","year":"2018","unstructured":"Klein, R., Mackert, J., Neugebauer, M., & Steinhardt, C. (2018). A model-based approximation of opportunity cost for dynamic pricing in attended home delivery. OR Spectrum, 40(4), 969\u2013996.","journal-title":"OR Spectrum"},{"issue":"2","key":"6846_CR18","doi-asserted-by":"publisher","first-page":"633","DOI":"10.1016\/j.ejor.2020.04.002","volume":"287","author":"S Koch","year":"2020","unstructured":"Koch, S., & Klein, R. (2020). Route-based approximate dynamic programming for dynamic pricing in attended home delivery. European Journal of Operational Research, 287(2), 633\u2013652.","journal-title":"European Journal of Operational Research"},{"key":"6846_CR19","volume":"abs\/2005","author":"S Levine","year":"2020","unstructured":"Levine, S., Kumar, A., Tucker, G., & Fu, J. (2020). Offline reinforcement learning: Tutorial, review, and perspectives on open problems. CoRR, abs\/2005, Article 01643.","journal-title":"CoRR"},{"key":"6846_CR20","unstructured":"Liu, F., Tang, R., Li, X., Ye, Y., Chen, H., Guo, H., & Zhang, Y. (2018). Deep reinforcement learning based recommendation with explicit user-item interactions modeling. CoRR, abs\/1810.12027."},{"key":"6846_CR21","unstructured":"Luo, W., Li, H., Zhang, Z., Han, C., Lv, J., & Guo, T. (2024). SAMBO-RL: shifts-aware model-based offline reinforcement learning. CoRR, abs\/2408.12830."},{"issue":"7540","key":"6846_CR22","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529\u2013533.","journal-title":"Nature"},{"key":"6846_CR23","unstructured":"OpenAI, Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., & Zhang, L. (2019). Solving rubik\u2019s cube with a robot hand. CoRR, abs\/1910.07113."},{"key":"6846_CR24","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511803161","volume-title":"Causality","author":"J Pearl","year":"2009","unstructured":"Pearl, J. (2009). Causality. Cambridge University Press."},{"key":"6846_CR25","doi-asserted-by":"crossref","unstructured":"Peng, X. B., Andrychowicz, M., Zaremba, W., & Abbeel, P. (2018). Sim-to-real transfer of robotic control with dynamics randomization. Proceedings of the 2018 IEEE international conference on robotics and automation (pp. 1\u20138). Brisbane.","DOI":"10.1109\/ICRA.2018.8460528"},{"key":"6846_CR26","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1093\/biomet\/70.1.41","volume":"70","author":"PR Rosenbaum","year":"1983","unstructured":"Rosenbaum, P. R., & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika, 70, 41\u201355.","journal-title":"Biometrika"},{"key":"6846_CR27","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., & Moritz, P. (2015). Trust region policy optimization. Proceedings of the 32nd international conference on machine learning (Vol. 37, pp. 1889\u20131897). Lille."},{"key":"6846_CR28","doi-asserted-by":"crossref","unstructured":"Shang, W., Yu, Y., Li, Q., Qin, Z. T., Meng, Y., & Ye, J. (2019). Environment reconstruction with hidden confounders for reinforcement learning based recommendation. Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining (pp. 566\u2013576). Anchorage.","DOI":"10.1145\/3292500.3330933"},{"key":"6846_CR29","doi-asserted-by":"crossref","unstructured":"Shi, J., Yu, Y., Da, Q., Chen, S., & Zeng, A. (2019). Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning. Proceedings of the 33rd AAAI conference on artificial intelligence (pp. 4902\u20134909). Honolulu.","DOI":"10.1609\/aaai.v33i01.33014902"},{"issue":"2","key":"6846_CR30","doi-asserted-by":"publisher","first-page":"227","DOI":"10.1016\/S0378-3758(00)00115-4","volume":"90","author":"H Shimodaira","year":"2000","unstructured":"Shimodaira, H. (2000). Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal Of Statistical Planning And Inference, 90(2), 227\u2013244.","journal-title":"Journal Of Statistical Planning And Inference"},{"issue":"7587","key":"6846_CR31","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016). Mastering the game of go with deep neural networks and tree search. Nature, 529(7587), 484\u2013489.","journal-title":"Nature"},{"key":"6846_CR32","first-page":"1643","volume":"11","author":"P Spirtes","year":"2010","unstructured":"Spirtes, P. (2010). Introduction to causal inference. Journal Of Machine Learning Research, 11, 1643\u20131662.","journal-title":"Journal Of Machine Learning Research"},{"key":"6846_CR33","doi-asserted-by":"crossref","unstructured":"Swazinna, P., Udluft, S., & Runkler, T. A. (2020). Overcoming model bias for robust offline deep reinforcement learning. CoRR, abs\/2008.05533.","DOI":"10.1016\/j.engappai.2021.104366"},{"issue":"1","key":"6846_CR34","doi-asserted-by":"publisher","first-page":"75","DOI":"10.1287\/trsc.2020.1000","volume":"55","author":"MW Ulmer","year":"2021","unstructured":"Ulmer, M. W., Thomas, B. W., Campbell, A. M., & Woyak, N. (2021). The restaurant meal delivery problem: Dynamic pickup and delivery with deadlines and random ready times. Transportation Science, 55(1), 75\u2013100.","journal-title":"Transportation Science"},{"key":"6846_CR35","unstructured":"Wang, T., Qin, T., & Zhou, Z. (2023). Estimating possible causal effects with latent variables via adjustment. Proceedings of the 40th international conference on machine learning (Vol. 202, pp. 36308\u201336335). Honolulu."},{"key":"6846_CR36","unstructured":"Wang, T., Wu, X., Huang, S., & Zhou, Z. (2020). Cost-effectively identifying causal effects when only response variable is observable. Proceedings of the 37th international conference on machine learning (Vol. 119, pp. 10060\u201310069). Virtual Event."},{"key":"6846_CR37","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1023\/A:1022672621406","volume":"8","author":"RJ Williams","year":"1992","unstructured":"Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8, 229\u2013256.","journal-title":"Machine Learning"},{"key":"6846_CR38","doi-asserted-by":"crossref","unstructured":"Wu, Z., Wang, L., Huang, F., Zhou, L., Song, Y., Ye, C., Nie, P., Ren, H., Hao, J., He, R., et al. (2022). A framework for multi-stage bonus allocation in meal delivery platform. Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining (pp. 4195\u20134203). Washington, DC.","DOI":"10.1145\/3534678.3539202"},{"issue":"2","key":"6846_CR39","doi-asserted-by":"publisher","first-page":"473","DOI":"10.1287\/trsc.2014.0549","volume":"50","author":"X Yang","year":"2016","unstructured":"Yang, X., Strauss, A. K., Currie, C. S., & Eglese, R. (2016). Choice-based demand management and vehicle routing in e-fulfillment. Transportation Science, 50(2), 473\u2013488.","journal-title":"Transportation Science"},{"key":"6846_CR40","unstructured":"Yoon, J., Jordon, J., & van der Schaar, M. (2018). GANITE: Estimation of individualized treatment effects using generative adversarial nets. Proceedings of the 6th international conference on learning representations. Vancouver."},{"key":"6846_CR41","unstructured":"Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., & Ma, T. (2020). MOPO: model-based offline policy optimization. Advances in neural information processing systems 33. Virtual Event."},{"key":"6846_CR42","unstructured":"Zou, W. Y., Du, S., Lee, J., & Pedersen, J. O. (2020). Heterogeneous causal learning for effectiveness optimization in user marketing. CoRR, abs\/2004.09702."}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06846-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-025-06846-6","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06846-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,22]],"date-time":"2025-12-22T21:29:14Z","timestamp":1766438954000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-025-06846-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,23]]},"references-count":42,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["6846"],"URL":"https:\/\/doi.org\/10.1007\/s10994-025-06846-6","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"type":"print","value":"0885-6125"},{"type":"electronic","value":"1573-0565"}],"subject":[],"published":{"date-parts":[[2025,10,23]]},"assertion":[{"value":"5 September 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 May 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 July 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 October 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Our research, which focuses on the pricing and time limit setting mechanism for each delivery service provided by delivery drivers, involves human participants. It is important to note that this study is conducted within the framework of commercial operations and is fully aligned with the regulatory policies set forth by the government. To ensure compliance with ethical standards, we have adhered to the following principles:","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"All participants, i.e., the delivery driver, the platform, were informed about the nature of the research and its potential impacts on their work.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Informed consent"}},{"value":"The data collected from the participants was handled with strict confidentiality. Personal identifiers were removed to protect their privacy, and data was only used for the purposes outlined in the study.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Confidentiality and privacy"}},{"value":"Our research strictly followed all relevant government policies and regulations concerning commercial operations and human participant research.","order":6,"name":"Ethics","group":{"name":"EthicsHeading","label":"Regulatory compliance"}},{"value":"The research aimed to benefit the delivery drivers, the food delivery platform and the broader community by proposing a fair, reasonable, and efficient bonus allocation and time limit setting mechanism. It is crucial to highlight that the delivery incentive policy targets only about 5% of the overall orders, and primarily affects the tail-end orders with extremely long estimated delivery times (> 45\u00a0min). To protect delivery drivers\u2019 welfare, we implement several measures. Firstly, we establish a minimum time limit for incentives based on expected delivery times, with an added margin to allow for more flexible deadlines, which prevents undue pressure on delivery drivers. Secondly, the platform\u2019s ethics committee oversees the program, making adjustments to unrealistic time limits or inadequate bonuses through post-evaluation. Lastly, our experiments indicate that tight deadlines lead to lower order acceptance rates and increased customer complaints, thereby indirectly protecting delivery drivers from a policy optimization standpoint. Overall, these measures align the interests of the platform with those of the drivers, fostering a mutually beneficial relationship.","order":7,"name":"Ethics","group":{"name":"EthicsHeading","label":"Beneficence and non-maleficence"}}],"article-number":"262"}}