{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,27]],"date-time":"2026-08-27T10:25:34Z","timestamp":1787826334543,"version":"build-2784847793"},"publisher-location":"New York, NY, USA","reference-count":47,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,7,6]],"date-time":"2022-07-06T00:00:00Z","timestamp":1657065600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key R\\&D Program of China","award":["2020YFB1406704"],"award-info":[{"award-number":["2020YFB1406704"]}]},{"name":"Tencent WeChat Rhino-Bird Focused Research Program","award":["JR-WXG-2021411"],"award-info":[{"award-number":["JR-WXG-2021411"]}]},{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"publisher","award":["ZR2021QF129"],"award-info":[{"award-number":["ZR2021QF129"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Fundamental Research Funds of Shandong University"},{"name":"Key Scientific and Technological Innovation Program of Shandong Province","award":["2019JZZY010129"],"award-info":[{"award-number":["2019JZZY010129"]}]},{"name":"Natural Science Foundation of China","award":["61902219, 61972234, 62072279, 62102234"],"award-info":[{"award-number":["61902219, 61972234, 62072279, 62102234"]}]},{"name":"Shandong University multidisciplinary research and innovation team of young scholars","award":["2020QNQT017"],"award-info":[{"award-number":["2020QNQT017"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,7,6]]},"DOI":"10.1145\/3477495.3531714","type":"proceedings-article","created":{"date-parts":[[2022,7,7]],"date-time":"2022-07-07T11:12:13Z","timestamp":1657192333000},"page":"1347-1357","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":31,"title":["Rethinking Reinforcement Learning for Recommendation"],"prefix":"10.1145","author":[{"given":"Xin","family":"Xin","sequence":"first","affiliation":[{"name":"Shandong University, Qingdao City, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tiago","family":"Pimentel","sequence":"additional","affiliation":[{"name":"University of Cambridge, Cambridge, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexandros","family":"Karatzoglou","sequence":"additional","affiliation":[{"name":"Google Research, London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pengjie","family":"Ren","sequence":"additional","affiliation":[{"name":"Shandong University, Qingdao City, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Konstantina","family":"Christakopoulou","sequence":"additional","affiliation":[{"name":"Google, Mountain View, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhaochun","family":"Ren","sequence":"additional","affiliation":[{"name":"Shandong University, Qingdao City, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,7,7]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"M Mehdi Afsar Trafford Crump and Behrouz Far. 2021. Reinforcement learning based recommender systems: A survey. arXiv preprint arXiv:2101.06286.  M Mehdi Afsar Trafford Crump and Behrouz Far. 2021. Reinforcement learning based recommender systems: A survey. arXiv preprint arXiv:2101.06286."},{"key":"e_1_3_2_2_2_1","volume-title":"International Conference on Machine Learning. PMLR, 224--232","author":"Arora Sanjeev","year":"2017","unstructured":"Sanjeev Arora , Rong Ge , Yingyu Liang , Tengyu Ma , and Yi Zhang . 2017 . Generalization and equilibrium in generative adversarial nets (gans) . In International Conference on Machine Learning. PMLR, 224--232 . Sanjeev Arora, Rong Ge, Yingyu Liang, Tengyu Ma, and Yi Zhang. 2017. Generalization and equilibrium in generative adversarial nets (gans). In International Conference on Machine Learning. PMLR, 224--232."},{"key":"e_1_3_2_2_3_1","volume-title":"Jamie Ryan Kiros, and Geoffrey E Hinton","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba , Jamie Ryan Kiros, and Geoffrey E Hinton . 2016 . Layer normalization. arXiv preprint arXiv:1607.06450. Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450."},{"key":"e_1_3_2_2_4_1","volume-title":"Science","volume":"153","author":"Bellman Richard","year":"1966","unstructured":"Richard Bellman . 1966 . Dynamic programming . Science , Vol. 153 , 3731, 34--37. Richard Bellman. 1966. Dynamic programming. Science, Vol. 153, 3731, 34--37."},{"key":"e_1_3_2_2_5_1","unstructured":"Jiawei Chen Hande Dong Xiang Wang Fuli Feng Meng Wang and Xiangnan He. 2020. Bias and debias in recommender system: A survey and future directions. arXiv preprint arXiv:2010.03240.  Jiawei Chen Hande Dong Xiang Wang Fuli Feng Meng Wang and Xiangnan He. 2020. Bias and debias in recommender system: A survey and future directions. arXiv preprint arXiv:2010.03240."},{"key":"e_1_3_2_2_6_1","unstructured":"Lili Chen Kevin Lu Aravind Rajeswaran Kimin Lee Aditya Grover Michael Laskin Pieter Abbeel Aravind Srinivas and Igor Mordatch. 2021 a. Decision transformer: Reinforcement learning via sequence modeling. In Advances in Neural Information Processing Systems.  Lili Chen Kevin Lu Aravind Rajeswaran Kimin Lee Aditya Grover Michael Laskin Pieter Abbeel Aravind Srinivas and Igor Mordatch. 2021 a. Decision transformer: Reinforcement learning via sequence modeling. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290999"},{"key":"e_1_3_2_2_8_1","volume-title":"Generative Adversarial User Model for Reinforcement Learning Based Recommendation System. In International Conference on Machine Learning. 1052--1061","author":"Chen Xinshi","year":"2019","unstructured":"Xinshi Chen , Shuang Li , Hui Li , Shaohua Jiang , Yuan Qi , and Le Song . 2019 b . Generative Adversarial User Model for Reinforcement Learning Based Recommendation System. In International Conference on Machine Learning. 1052--1061 . Xinshi Chen, Shuang Li, Hui Li, Shaohua Jiang, Yuan Qi, and Le Song. 2019 b. Generative Adversarial User Model for Reinforcement Learning Based Recommendation System. In International Conference on Machine Learning. 1052--1061."},{"key":"e_1_3_2_2_9_1","unstructured":"Xiaocong Chen Lina Yao Julian McAuley Guanglin Zhou and Xianzhi Wang. 2021 b. A survey of deep reinforcement learning in recommender systems: A systematic review and future directions. arXiv preprint arXiv:2109.03540.  Xiaocong Chen Lina Yao Julian McAuley Guanglin Zhou and Xianzhi Wang. 2021 b. A survey of deep reinforcement learning in recommender systems: A systematic review and future directions. arXiv preprint arXiv:2109.03540."},{"key":"e_1_3_2_2_10_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"2062","author":"Fujimoto Scott","year":"2019","unstructured":"Scott Fujimoto , David Meger , and Doina Precup . 2019 . Off-policy deep reinforcement learning without exploration . In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research , Vol. 97). PMLR, 2052-- 2062 . Scott Fujimoto, David Meger, and Doina Precup. 2019. Off-policy deep reinforcement learning without exploration. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97). PMLR, 2052--2062."},{"key":"e_1_3_2_2_11_1","unstructured":"Hado V Hasselt. 2010. Double Q-learning. In Advances in Neural Information Processing Systems. 2613--2621.  Hado V Hasselt. 2010. Double Q-learning. In Advances in Neural Information Processing Systems. 2613--2621."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2016.0030"},{"key":"e_1_3_2_2_14_1","volume-title":"4th International Conference on Learning Representations.","author":"Hidasi Bal\u00e1zs","year":"2016","unstructured":"Bal\u00e1zs Hidasi , Alexandros Karatzoglou , Linas Baltrunas , and Domonkos Tikk . 2016 . Session-based recommendations with recurrent neural networks . In 4th International Conference on Learning Representations. Bal\u00e1zs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2016. Session-based recommendations with recurrent neural networks. In 4th International Conference on Learning Representations."},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-015-0417-y"},{"key":"e_1_3_2_2_16_1","unstructured":"Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. In Advances in neural information processing systems. 4565--4573.  Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. In Advances in neural information processing systems. 4565--4573."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219846"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3383313.3412252"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/360"},{"key":"e_1_3_2_2_20_1","unstructured":"Michael Janner Qiyang Li and Sergey Levine. 2021. Offline Reinforcement Learning as One Big Sequence Modeling Problem. In Advances in Neural Information Processing Systems.  Michael Janner Qiyang Li and Sergey Levine. 2021. Offline Reinforcement Learning as One Big Sequence Modeling Problem. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/582415.582418"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00035"},{"key":"e_1_3_2_2_23_1","volume-title":"3rd International Conference on Learning Representations","author":"Kingma Diederik P","year":"2015","unstructured":"Diederik P Kingma and Jimmy Ba . 2015 . Adam: A method for stochastic optimization . In 3rd International Conference on Learning Representations , 2015. Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, 2015."},{"key":"e_1_3_2_2_24_1","unstructured":"Vijay R Konda and John N Tsitsiklis. 2000. Actor-critic algorithms. In Advances in neural information processing systems. 1008--1014.  Vijay R Konda and John N Tsitsiklis. 2000. Actor-critic algorithms. In Advances in neural information processing systems. 1008--1014."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557072"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2009.263"},{"key":"e_1_3_2_2_27_1","volume-title":"Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems","author":"Kumar Aviral","year":"2020","unstructured":"Aviral Kumar , Aurick Zhou , George Tucker , and Sergey Levine . 2020 . Conservative q-learning for offline reinforcement learning . In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020. Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020. Conservative q-learning for offline reinforcement learning. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020."},{"key":"e_1_3_2_2_28_1","unstructured":"Pengfei Liu Weizhe Yuan Jinlan Fu Zhengbao Jiang Hiroaki Hayashi and Graham Neubig. 2021. Pre-train prompt and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.  Pengfei Liu Weizhe Yuan Jinlan Fu Zhengbao Jiang Hiroaki Hayashi and Graham Neubig. 2021. Pre-train prompt and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_2_30_1","unstructured":"R\u00e9mi Munos Tom Stepleton Anna Harutyunyan and Marc Bellemare. 2016. Safe and efficient off-policy reinforcement learning. In Advances in Neural Information Processing Systems. 1054--1062.  R\u00e9mi Munos Tom Stepleton Anna Harutyunyan and Marc Bellemare. 2016. Safe and efficient off-policy reinforcement learning. In Advances in Neural Information Processing Systems. 1054--1062."},{"key":"e_1_3_2_2_31_1","volume-title":"Proceedings of the 37th International Conference on Machine Learning,2020 (Proceedings of Machine Learning Research","volume":"7498","author":"Parisotto Emilio","year":"2020","unstructured":"Emilio Parisotto , H Francis Song , Jack W Rae , Razvan Pascanu , Caglar Gulcehre , Siddhant M Jayakumar , Max Jaderberg , Raphael Lopez Kaufman , Aidan Clark , Seb Noury , 2020 . Stabilizing Transformers for Reinforcement Learning . In Proceedings of the 37th International Conference on Machine Learning,2020 (Proceedings of Machine Learning Research , Vol. 119). PMLR, 7487-- 7498 . Emilio Parisotto, H Francis Song, Jack W Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant M Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, et al. 2020. Stabilizing Transformers for Reinforcement Learning. In Proceedings of the 37th International Conference on Machine Learning,2020 (Proceedings of Machine Learning Research, Vol. 119). PMLR, 7487--7498."},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2010.127"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772773"},{"key":"e_1_3_2_2_34_1","unstructured":"Juergen Schmidhuber. 2019. Reinforcement Learning Upside Down: Don't Predict Rewards--Just Map Them to Actions. arXiv preprint arXiv:1912.02875.  Juergen Schmidhuber. 2019. Reinforcement Learning Upside Down: Don't Predict Rewards--Just Map Them to Actions. arXiv preprint arXiv:1912.02875."},{"key":"e_1_3_2_2_35_1","article-title":"An MDP-based recommender system","volume":"6","author":"Shani Guy","year":"2005","unstructured":"Guy Shani , David Heckerman , and Ronen I Brafman . 2005 . An MDP-based recommender system . Journal of Machine Learning Research , Vol. 6 , Sep, 1265--1295. Guy Shani, David Heckerman, and Ronen I Brafman. 2005. An MDP-based recommender system. Journal of Machine Learning Research, Vol. 6, Sep, 1265--1295.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature16961"},{"key":"e_1_3_2_2_37_1","volume-title":"NeurIPS Deep Reinforcement Learning Workshop.","author":"Srivastava Rupesh Kumar","year":"2019","unstructured":"Rupesh Kumar Srivastava , Pranav Shyam , Filipe Mutz , Wojciech Ja'skowski , and J\u00fcrgen Schmidhuber . 2019 . Training agents using upside-down reinforcement learning . In NeurIPS Deep Reinforcement Learning Workshop. Rupesh Kumar Srivastava, Pranav Shyam, Filipe Mutz, Wojciech Ja'skowski, and J\u00fcrgen Schmidhuber. 2019. Training agents using upside-down reinforcement learning. In NeurIPS Deep Reinforcement Learning Workshop."},{"key":"e_1_3_2_2_38_1","volume-title":"Proceedings of the 15th ACM International Conference on Web Search and Data Mining.","author":"Stamenkovic Dusan","year":"2021","unstructured":"Dusan Stamenkovic , Alexandros Karatzoglou , Ioannis Arapakis , Xin Xin , and Kleomenis Katevas . 2021 . Choosing the Best of Both Worlds: Diverse and Novel Recommendations through Multi-Objective Reinforcement Learning . In Proceedings of the 15th ACM International Conference on Web Search and Data Mining. Dusan Stamenkovic, Alexandros Karatzoglou, Ioannis Arapakis, Xin Xin, and Kleomenis Katevas. 2021. Choosing the Best of Both Worlds: Diverse and Novel Recommendations through Multi-Objective Reinforcement Learning. In Proceedings of the 15th ACM International Conference on Web Search and Data Mining."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357895"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3159652.3159656"},{"key":"e_1_3_2_2_41_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008."},{"key":"e_1_3_2_2_42_1","volume-title":"Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning","author":"Williams Ronald J","unstructured":"Ronald J Williams . 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning , Vol. 8 , 3--4, 229--256. Ronald J Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, Vol. 8, 3--4, 229--256."},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401147"},{"key":"e_1_3_2_2_44_1","volume-title":"Proceedings of the 15th ACM International Conference on Web Search and Data Mining.","author":"Xin Xin","year":"2021","unstructured":"Xin Xin , Alexandros Karatzoglou , Ioannis Arapakis , and Joemon M Jose . 2021 . Supervised Advantage Actor-Critic for Recommender Systems . In Proceedings of the 15th ACM International Conference on Web Search and Data Mining. Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M Jose. 2021. Supervised Advantage Actor-Critic for Recommender Systems. In Proceedings of the 15th ACM International Conference on Web Search and Data Mining."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290975"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Xiangyu Zhao Long Xia Jiliang Tang and Dawei Yin. 2019. \" Deep reinforcement learning for search recommendation and online advertising: a survey\" by Xiangyu Zhao Long Xia Jiliang Tang and Dawei Yin with Martin Vesely as coordinator. ACM SIGWEB Newsletter Spring 1--15.  Xiangyu Zhao Long Xia Jiliang Tang and Dawei Yin. 2019. \" Deep reinforcement learning for search recommendation and online advertising: a survey\" by Xiangyu Zhao Long Xia Jiliang Tang and Dawei Yin with Martin Vesely as coordinator. ACM SIGWEB Newsletter Spring 1--15.","DOI":"10.1145\/3320496.3320500"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330668"}],"event":{"name":"SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","location":"Madrid Spain","acronym":"SIGIR '22","sponsor":["SIGIR ACM Special Interest Group on Information Retrieval"]},"container-title":["Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3477495.3531714","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3477495.3531714","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T15:02:07Z","timestamp":1750172527000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3477495.3531714"}},"subtitle":["A Prompt Perspective"],"short-title":[],"issued":{"date-parts":[[2022,7,6]]},"references-count":47,"alternative-id":["10.1145\/3477495.3531714","10.1145\/3477495"],"URL":"https:\/\/doi.org\/10.1145\/3477495.3531714","relation":{},"subject":[],"published":{"date-parts":[[2022,7,6]]},"assertion":[{"value":"2022-07-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}