{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T10:08:07Z","timestamp":1775815687837,"version":"3.50.1"},"reference-count":55,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2023,11,29]],"date-time":"2023-11-29T00:00:00Z","timestamp":1701216000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,11,29]],"date-time":"2023-11-29T00:00:00Z","timestamp":1701216000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2024,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Many Internet applications adopt real-time bidding mechanisms to ensure different services (types of content) are shown to the users through fair competitions. The service offering the highest bid price gets the content slot to present a list of items in its candidate pool. Through user interactions with the recommended items, the service obtains the desired engagement activities. We propose a contextual-bandit framework to\u00a0jointly optimize\u00a0the\u00a0price to bid for the slot and the order to rank its candidates\u00a0for a given service in this type of recommendation systems. Our method can take as input any feature that describes the user and the candidates, including the outputs of other machine learning models. We train \u00a0reinforcement learning\u00a0policies using deep neural networks, and compute top-<jats:italic>K<\/jats:italic> Gaussian propensity scores to exclude the variance in the gradients\u00a0caused by randomness unrelated to the reward. This setup further facilitates us to automatically find accurate reward functions that trade off between budget spending and user engagements. In online A\/B experiments on two major services of Facebook Home Feed, Groups You Should Join and Friend Requests, our method statistically significantly\u00a0boosted the number of groups joined by 14.7%, the number of friend requests accepted by 7.0%, and the number of daily active Facebook\u00a0users by\u00a0about 1 million, against strong hand-tuned baselines that have been iterated in production over years.<\/jats:p>","DOI":"10.1007\/s10994-023-06444-4","type":"journal-article","created":{"date-parts":[[2023,11,29]],"date-time":"2023-11-29T18:02:13Z","timestamp":1701280933000},"page":"2559-2573","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Learning to bid and rank together in recommendation systems"],"prefix":"10.1007","volume":"113","author":[{"given":"Geng","family":"Ji","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wentao","family":"Jiang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiang","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fahmid Morshed","family":"Fahid","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengxing","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yinghua","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Xiao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chongxi","family":"Bao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheqing","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,11,29]]},"reference":[{"key":"6444_CR1","doi-asserted-by":"crossref","unstructured":"Abbeel, P., & Ng, A. Y. (2004). Apprenticeship learning via inverse reinforcement learning. In International conference on machine learning","DOI":"10.1145\/1015330.1015430"},{"key":"6444_CR2","unstructured":"Agarwal, A., Dud\u00edk, M., Kale, S., Langford, J., & Schapire, R. (2012). Contextual bandit learning with predictable rewards. In Artificial intelligence and statistics"},{"key":"6444_CR3","unstructured":"Agarwal, A., Hsu, D., Kale, S., Langford, J., Li, L., & Schapire, R. (2014). Taming the monster: A fast and simple algorithm for contextual bandits. In International conference on machine learning"},{"key":"6444_CR4","doi-asserted-by":"crossref","unstructured":"Agarwal, A., Takatsu, K., Zaitsev, I., & Joachims, T. (2019). A general framework for counterfactual learning-to-rank. In Conference on research and development in information retrieval","DOI":"10.1145\/3331184.3331202"},{"key":"6444_CR5","first-page":"1","volume":"18","author":"AG Baydin","year":"2018","unstructured":"Baydin, A. G., Pearlmutter, B. A., Radul, A. A., & Siskind, J. M. (2018). Automatic differentiation in machine learning: a survey. Journal of Machine Learning Research, 18, 1\u201343.","journal-title":"Journal of Machine Learning Research"},{"key":"6444_CR6","unstructured":"Bello, I., Kulkarni, S., Jain, S., Boutilier, C., Chi, E., Eban, E., Luo, X., Mackey, A., & Meshi, O. (2018). Seq2Slate: Re-ranking and slate optimization with RNNs. arXiv preprint arXiv:1810.02019"},{"key":"6444_CR7","unstructured":"Bottou, L., Peters, J., Qui\u00f1onero-Candela, J., Charles, D. X., Chickering, D. M., Portugaly, E., Ray, D., Simard, P., & Snelson, E. (2013). Counterfactual reasoning and learning systems: The example of computational advertising. Journal of Machine Learning Research, 14(101), 3207\u22123260."},{"key":"6444_CR8","doi-asserted-by":"crossref","unstructured":"Cai, H., Ren, K., Zhang, W., Malialis, K., Wang, J., Yu, Y., & Guo, D. (2017). Real-time bidding by reinforcement learning in display advertising. In International conference on web search and data mining","DOI":"10.1145\/3018661.3018702"},{"issue":"1","key":"6444_CR9","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1093\/biomet\/83.1.81","volume":"83","author":"G Casella","year":"1996","unstructured":"Casella, G., & Robert, C. P. (1996). Rao-blackwellisation of sampling schemes. Biometrika, 83(1), 81\u201394.","journal-title":"Biometrika"},{"key":"6444_CR10","doi-asserted-by":"crossref","unstructured":"Chen, M., Beutel, A., Covington, P., Jain, S., Belletti, F., & Chi, E. H. (2019). Top-$$K$$ off-policy correction for a REINFORCE recommender system. In International conference on web search and data mining","DOI":"10.1145\/3289600.3290999"},{"key":"6444_CR11","doi-asserted-by":"crossref","unstructured":"Cheng, H.-T., Koc, L., Harmsen, J., Shaked, T., Chandra, T., Aradhye, H., Anderson, G., Corrado, G., Chai, W., & Ispir, M. (2016). Wide & deep learning for recommender systems. In Workshop on deep learning for recommender systems","DOI":"10.1145\/2988450.2988454"},{"key":"6444_CR12","doi-asserted-by":"crossref","unstructured":"Clarke, C. L., Kolla, M., Cormack, G. V., Vechtomova, O., Ashkan, A., B\u00fcttcher, S., & MacKinnon, I. (2008). Novelty and diversity in information retrieval evaluation. In Conference on research and development in information retrieval","DOI":"10.1145\/1390334.1390446"},{"key":"6444_CR13","doi-asserted-by":"crossref","unstructured":"Craswell, N., Zoeter, O., Taylor, M., & Ramsey, B. (2008). An experimental comparison of click position-bias models. In International conference on web search and data mining","DOI":"10.1145\/1341531.1341545"},{"key":"6444_CR14","doi-asserted-by":"crossref","unstructured":"Ding, W., Govindaraj, D., & Vishwanathan, S. (2019). Whole page optimization with global constraints. In International conference on knowledge discovery & data mining","DOI":"10.1145\/3292500.3330675"},{"key":"6444_CR15","unstructured":"Dudik, M., Hsu, D., Kale, S., Karampatziakis, N., Langford, J., Reyzin, L., & Zhang, T. (2011). Efficient optimal learning for contextual bandits. arXiv preprint arXiv:1106.2369"},{"key":"6444_CR16","unstructured":"Facebook: Facebook News Feed: An Introduction for Content Creators. https:\/\/www.facebook.com\/business\/learn\/lessons\/facebook-news-feed-creators. Online; accessed January, 2023 (2022)"},{"key":"6444_CR17","unstructured":"Gauci, J., Conti, E., Liang, Y., Virochsiri, K., He, Y., Kaden, Z., Narayanan, V., Ye, X., Chen, Z., & Fujimoto, S. (2018). Horizon: Facebook\u2019s open source applied reinforcement learning platform. arXiv preprint arXiv:1811.00260"},{"key":"6444_CR18","doi-asserted-by":"crossref","unstructured":"Guo, H., Tang, R., Ye, Y., Li, Z., & He, X. (2017). Deepfm: A factorization-machine based neural network for ctr prediction. arXiv preprint arXiv:1703.04247","DOI":"10.24963\/ijcai.2017\/239"},{"key":"6444_CR19","doi-asserted-by":"crossref","unstructured":"Hill, D. N., Nassif, H., Liu, Y., Iyer, A., & Vishwanathan, S. (2017). An efficient bandit algorithm for realtime multivariate optimization. In International conference on knowledge discovery and data mining","DOI":"10.1145\/3097983.3098184"},{"key":"6444_CR20","unstructured":"Ho, J., & Ermon, S. (2016). Generative adversarial imitation learning. Advances in neural information processing systems"},{"issue":"2","key":"6444_CR21","doi-asserted-by":"publisher","first-page":"114","DOI":"10.1504\/IJEB.2008.018068","volume":"6","author":"BJ Jansen","year":"2008","unstructured":"Jansen, B. J., & Mullen, T. (2008). Sponsored search: An overview of the concept, history, and technology. International Journal of Electronic Business, 6(2), 114\u2013131.","journal-title":"International Journal of Electronic Business"},{"key":"6444_CR22","unstructured":"Ji, G., Sujono, D., & Sudderth, E. B. (2021). Marginalized stochastic natural gradients for black-box variational inference. In International conference on machine learning"},{"key":"6444_CR23","unstructured":"Jiang, R., Gowal, S., Mann, T. A., & Rezende, D. J. (2018). Beyond greedy ranking: Slate optimization via list-CVAE. arXiv preprint arXiv:1803.01682"},{"key":"6444_CR24","unstructured":"Kingma, D. P., & Welling, M. (2014). Auto-encoding variational Bayes. In International conference on learning representations"},{"key":"6444_CR25","unstructured":"Kingma, D. P., Ba, J. (2015). Adam: A method for stochastic optimization. In International conference on learning representations"},{"key":"6444_CR26","doi-asserted-by":"crossref","unstructured":"Lee, K.-C., Jalali, A., & Dasdan, A. (2013). Real time bid optimization with smooth budget delivery in online advertising. In International workshop on data mining for online advertising","DOI":"10.1145\/2501040.2501979"},{"key":"6444_CR27","doi-asserted-by":"crossref","unstructured":"Li, X., & Guan, D. (2014). Programmatic buying bidding strategies with win rate and winning price estimation in real time mobile advertising. In Pacific-Asia conference on knowledge discovery and data mining","DOI":"10.1007\/978-3-319-06608-0_37"},{"key":"6444_CR28","doi-asserted-by":"crossref","unstructured":"Li, L., Chen, S., Kleban, J., & Gupta, A. (2015). Counterfactual estimation and optimization of click metrics in search engines: A case study. In Proceedings of the 24th international conference on world wide web","DOI":"10.1145\/2740908.2742562"},{"key":"6444_CR29","doi-asserted-by":"crossref","unstructured":"Li, L., Chu, W., Langford, J., & Schapire, R. E. (2010). A contextual-bandit approach to personalized news article recommendation. In International conference on world wide web","DOI":"10.1145\/1772690.1772758"},{"key":"6444_CR30","doi-asserted-by":"crossref","unstructured":"Liu, Y., Chen, Z., Virochsiri, K., Wang, J., Wu, J., & Liang, F. (2021). Reinforcement learning-based product delivery frequency control. In AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v35i17.17803"},{"issue":"3","key":"6444_CR31","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1561\/1500000016","volume":"3","author":"T-Y Liu","year":"2009","unstructured":"Liu, T.-Y. (2009). Learning to rank for information retrieval. Foundations and Trends in Information Retrieval, 3(3), 225\u2013331.","journal-title":"Foundations and Trends in Information Retrieval"},{"key":"6444_CR32","doi-asserted-by":"crossref","unstructured":"McMahan, H. B., Holt, G., Sculley, D., Young, M., Ebner, D., Grady, J., Nie, L., Phillips, T., Davydov, E., & Golovin, D. (2013). Ad click prediction: A view from the trenches. In International conference on knowledge discovery and data mining","DOI":"10.1145\/2487575.2488200"},{"key":"6444_CR33","unstructured":"Mirhoseini, A., Pham, H., Le, Q. V., Steiner, B., Larsen, R., Zhou, Y., Kumar, N., Norouzi, M., Bengio, S., & Dean, J. (2017). Device placement optimization with reinforcement learning. In International conference on machine learning"},{"key":"6444_CR34","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing Atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602"},{"key":"6444_CR35","unstructured":"Naumov, M., Mudigere, D., Shi, H.-J. M., Huang, J., Sundaraman, N., Park, J., Wang, X., Gupta, U., Wu, C.-J., Azzolini, A. G., et al. (2019). Deep learning recommendation model for personalization and recommendation systems. arXiv preprint arXiv:1906.00091"},{"key":"6444_CR36","unstructured":"Ng, A. Y., Harada, D., & Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. In International conference on machine learning"},{"key":"6444_CR37","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., & Lerer, A. (2017). Automatic differentiation in PyTorch. In NIPS autodiff workshop"},{"key":"6444_CR38","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., & Antiga, L. (2019). PyTorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems"},{"key":"6444_CR39","unstructured":"Pinterest: New ways to control the ideas you see in your home feed. https:\/\/newsroom.pinterest.com\/en\/post\/new-ways-to-control-the-ideas-you-see-in-your-home-feed. Online; accessed January, 2023 (2019)"},{"issue":"4","key":"6444_CR40","doi-asserted-by":"publisher","first-page":"645","DOI":"10.1109\/TKDE.2017.2775228","volume":"30","author":"K Ren","year":"2017","unstructured":"Ren, K., Zhang, W., Chang, K., Rong, Y., Yu, Y., & Wang, J. (2017). Bidding machine: Learning to bid for directly optimizing profits in display advertising. Knowledge and Data Engineering, 30(4), 645\u2013659.","journal-title":"Knowledge and Data Engineering"},{"key":"6444_CR41","unstructured":"Rezende, D. J., Mohamed, S., & Wierstra, D. (2014). Stochastic backpropagation and approximate inference in deep generative models. In International conference on machine learning"},{"issue":"7676","key":"6444_CR42","doi-asserted-by":"publisher","first-page":"354","DOI":"10.1038\/nature24270","volume":"550","author":"D Silver","year":"2017","unstructured":"Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., & Bolton, A. (2017). Mastering the game of Go without human knowledge. Nature, 550(7676), 354\u2013359.","journal-title":"Nature"},{"key":"6444_CR43","unstructured":"Strehl, A., Langford, J., Li, L., & Kakade, S. M. (2010). Learning from logged implicit exploration data. Advances in neural information processing systems"},{"key":"6444_CR44","unstructured":"Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction. MIT Press"},{"key":"6444_CR45","unstructured":"Swaminathan, A., & Joachims, T. (2015). The self-normalized estimator for counterfactual learning. Advances in neural information processing systems"},{"issue":"1","key":"6444_CR46","first-page":"1731","volume":"16","author":"A Swaminathan","year":"2015","unstructured":"Swaminathan, A., & Joachims, T. (2015). Batch learning from logged bandit feedback through counterfactual risk minimization. Journal of Machine Learning Research, 16(1), 1731\u20131755.","journal-title":"Journal of Machine Learning Research"},{"key":"6444_CR47","doi-asserted-by":"crossref","unstructured":"Taylor, M., Guiver, J., Robertson, S., & Minka, T. (2008). Softrank: Optimizing non-smooth rank metrics. In International conference on web search and data mining","DOI":"10.1145\/1341531.1341544"},{"key":"6444_CR48","doi-asserted-by":"crossref","unstructured":"Wang, X., Golbandi, N., Bendersky, M., Metzler, D., & Najork, M. (2018). Position bias estimation for unbiased learning to rank in personal search. In International conference on web search and data mining","DOI":"10.1145\/3159652.3159732"},{"key":"6444_CR49","doi-asserted-by":"crossref","unstructured":"Wang, Y., Yin, D., Jie, L., Wang, P., Yamada, M., Chang, Y., & Mei, Q. (2016). Beyond ranking: Optimizing whole-page presentation. In International conference on web search and data mining","DOI":"10.1145\/2835776.2835824"},{"key":"6444_CR50","doi-asserted-by":"crossref","unstructured":"Wilhelm, M., Ramanathan, A., Bonomo, A., Jain, S., Chi, E. H., & Gillenwater, J. (2018). Practical diversified recommendations on youtube with determinantal point processes. In International conference on information and knowledge management","DOI":"10.1145\/3269206.3272018"},{"key":"6444_CR51","doi-asserted-by":"crossref","unstructured":"Xu, J., Lee, K.-c., Li, W., Qi, H., & Lu, Q. (2015). Smart pacing for effective online ad campaign optimization. In International conference on knowledge discovery and data mining","DOI":"10.1145\/2783258.2788615"},{"key":"6444_CR52","doi-asserted-by":"crossref","unstructured":"Xu, Z., Li, Z., Guan, Q., Zhang, D., Li, Q., Nan, J., Liu, C., Bian, W., & Ye, J. (2018). Large-scale order dispatch in on-demand ride-hailing platforms: A learning and planning approach. In International conference on knowledge discovery and data mining","DOI":"10.1145\/3219819.3219824"},{"key":"6444_CR53","doi-asserted-by":"crossref","unstructured":"Yuan, S., Wang, J., & Zhao, X. (2013). Real-time bidding for online advertising: Measurement and analysis. In International workshop on data mining for online advertising","DOI":"10.1145\/2501040.2501980"},{"key":"6444_CR54","doi-asserted-by":"crossref","unstructured":"Zhang, W., Yuan, S., & Wang, J. (2014). Optimal real-time bidding for display advertising. In International conference on knowledge discovery and data mining","DOI":"10.1145\/2623330.2623633"},{"key":"6444_CR55","unstructured":"Ziebart, B. D., Maas, A. L., Bagnell, J. A., & Dey, A. K. (2008). Maximum entropy inverse reinforcement learning. In AAAI conference on artificial intelligence"}],"updated-by":[{"DOI":"10.1007\/s10994-023-06496-6","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2024,1,4]],"date-time":"2024-01-04T00:00:00Z","timestamp":1704326400000}}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-023-06444-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-023-06444-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-023-06444-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,2]],"date-time":"2024-05-02T18:14:39Z","timestamp":1714673679000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-023-06444-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,29]]},"references-count":55,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,5]]}},"alternative-id":["6444"],"URL":"https:\/\/doi.org\/10.1007\/s10994-023-06444-4","relation":{"correction":[{"id-type":"doi","id":"10.1007\/s10994-023-06496-6","asserted-by":"object"}]},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,29]]},"assertion":[{"value":"20 March 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 August 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 October 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 November 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 January 2024","order":5,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Correction","order":6,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"A Correction to this paper has been published:","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"https:\/\/doi.org\/10.1007\/s10994-023-06496-6","URL":"https:\/\/doi.org\/10.1007\/s10994-023-06496-6","order":8,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}