{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:46:25Z","timestamp":1787017585084,"version":"build-2736575974"},"reference-count":76,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2024,8,19]],"date-time":"2024-08-19T00:00:00Z","timestamp":1724025600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2024,11,30]]},"abstract":"<jats:p>Reinforcement learning serves as a potent tool for modeling dynamic user interests within recommender systems, garnering increasing research attention of late. However, a significant drawback persists: its poor data efficiency, stemming from its interactive nature. The training of reinforcement learning-based recommender systems demands expensive online interactions to amass adequate trajectories, essential for agents to learn user preferences. This inefficiency renders reinforcement learning-based recommender systems a formidable undertaking, necessitating the exploration of potential solutions. Recent strides in offline reinforcement learning present a new perspective. Offline reinforcement learning empowers agents to glean insights from offline datasets and deploy learned policies in online settings. Given that recommender systems possess extensive offline datasets, the framework of offline reinforcement learning aligns seamlessly. Despite being a burgeoning field, works centered on recommender systems utilizing offline reinforcement learning remain limited. This survey aims to introduce and delve into offline reinforcement learning within recommender systems, offering an inclusive review of existing literature in this domain. Furthermore, we strive to underscore prevalent challenges, opportunities, and future pathways, poised to propel research in this evolving field.<\/jats:p>","DOI":"10.1145\/3661996","type":"journal-article","created":{"date-parts":[[2024,4,29]],"date-time":"2024-04-29T12:47:21Z","timestamp":1714394841000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["On the Opportunities and Challenges of Offline Reinforcement Learning for Recommender Systems"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8849-4943","authenticated-orcid":false,"given":"Xiaocong","family":"Chen","sequence":"first","affiliation":[{"name":"Data61, CSIRO, Eveleigh, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-8726-5277","authenticated-orcid":false,"given":"Siyu","family":"Wang","sequence":"additional","affiliation":[{"name":"UNSW Sydney, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0955-7588","authenticated-orcid":false,"given":"Julian","family":"McAuley","sequence":"additional","affiliation":[{"name":"UCSD, La Jolla, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4698-8507","authenticated-orcid":false,"given":"Dietmar","family":"Jannach","sequence":"additional","affiliation":[{"name":"University of Klagenfurt, Klagenfurt, Austria"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4149-839X","authenticated-orcid":false,"given":"Lina","family":"Yao","sequence":"additional","affiliation":[{"name":"Data61, CSIRO, Eveleigh, Australia and UNSW Sydney, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,8,19]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3543846"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401196"},{"key":"e_1_3_2_4_2","first-page":"3676","volume-title":"Proceedings of the 40th International Conference on Machine Learning (ICML\u201923)","volume":"202","author":"Carta Thomas","year":"2023","unstructured":"Thomas Carta, Cl\u00e9ment Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer. 2023. Grounding large language models in interactive environments with online reinforcement learning. In Proceedings of the 40th International Conference on Machine Learning (ICML\u201923), Vol. 202, JMLR.org, 3676\u20133713."},{"issue":"3","key":"e_1_3_2_5_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3564284","article-title":"Bias and debias in recommender system: A survey and future directions","volume":"41","author":"Chen Jiawei","year":"2023","unstructured":"Jiawei Chen, Hande Dong, Xiang Wang, Fuli Feng, Meng Wang, and Xiangnan He. 2023. Bias and debias in recommender system: A survey and future directions. ACM Transactions on Information Systems 41, 3 (2023), 1\u201339.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290999"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3523227.3546758"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN48605.2020.9207010"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11280-023-01187-7"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM54844.2022.00102"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532015"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2023.110335"},{"key":"e_1_3_2_13_2","first-page":"1","volume-title":"Proceedings of the ACM SIGIR Forum","volume":"56","author":"Deffayet Romain","year":"2023","unstructured":"Romain Deffayet, Thibaut Thonet, Jean-Michel Renders, and Maarten de Rijke. 2023. Offline evaluation for reinforcement learning-based recommendation: A critical issue and some alternatives. In Proceedings of the ACM SIGIR Forum, Vol. 56, ACM, New York, NY, 1\u201314."},{"key":"e_1_3_2_14_2","first-page":"179","volume-title":"Proceedings of the 29th International Conference on Machine Learning (ICML\u201912)","author":"Degris Thomas","year":"2012","unstructured":"Thomas Degris, Martha White, and Richard S. Sutton. 2012. Off-policy actor-critic. In Proceedings of the 29th International Conference on Machine Learning (ICML\u201912). Omnipress, Madison, WI, 179\u2013186."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3439729"},{"key":"e_1_3_2_16_2","first-page":"8657","volume-title":"Proceedings of the 40th International Conference on Machine Learning (ICML\u201923)","volume":"202","author":"Du Yuqing","year":"2023","unstructured":"Yuqing Du, Olivia Watkins, Zihan Wang, C\u00e9dric Colas, Trevor Darrell, Pieter Abbeel, Abhishek Gupta, and Jacob Andreas. 2023. Guiding pretraining in reinforcement learning with large language models. In Proceedings of the 40th International Conference on Machine Learning (ICML\u201923), Vol. 202, JMLR.org, 8657\u20138677."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591636"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3594871"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383313.3412233"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485447.3511969"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106685"},{"key":"e_1_3_2_22_2","first-page":"1596","volume-title":"Proceedings of the 33rd International Conference on Machine Learning","volume":"48","author":"Hoiles William","year":"2016","unstructured":"William Hoiles and Mihaela Schaar. 2016. Bounded off-policy evaluation with missing data for course recommendation and curriculum design. In Proceedings of the 33rd International Conference on Machine Learning, Vol. 48, Maria Florina Balcan and Kilian Q. Weinberger (Eds.), PMLR, New York, NY, 1596\u20131604."},{"key":"e_1_3_2_23_2","first-page":"13157","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Hong Joey","year":"2023","unstructured":"Joey Hong, Branislav Kveton, Manzil Zaheer, Sumeet Katariya, and Mohammad Ghavamzadeh. 2023. Multi-task off-policy learning from bandit feedback. In Proceedings of the International Conference on Machine Learning. PMLR, New York, New York, 13157\u201313173."},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Chengkai Huang Tong Yu Kaige Xie Shuai Zhang Lina Yao and Julian McAuley. 2024. Foundation models for recommender systems: A survey and new perspectives. arXiv:2402.11143. Retrieved from https:\/\/arxiv.org\/abs\/2402.11143","DOI":"10.1200\/JCO.2024.42.16_suppl.11143"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290958"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460231.3474247"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3568029"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403175"},{"key":"e_1_3_2_29_2","unstructured":"Jiechuan Jiang Chen Dun Tiejun Huang and Zongqing Lu. 2018. Graph convolutional reinforcement learning. arXiv:1810.09202."},{"key":"e_1_3_2_30_2","first-page":"652","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Jiang Nan","year":"2016","unstructured":"Nan Jiang and Lihong Li. 2016. Doubly robust off-policy value evaluation for reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 652\u2013661."},{"key":"e_1_3_2_31_2","first-page":"1008","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Konda Vijay R.","year":"2000","unstructured":"Vijay R. Konda and John N. Tsitsiklis. 2000. Actor-critic algorithms. In Proceedings of the Advances in Neural Information Processing Systems. Citeseer, 1008\u20131014."},{"key":"e_1_3_2_32_2","unstructured":"Sergey Levine Aviral Kumar George Tucker and Justin Fu. 2020. Offline reinforcement learning: Tutorial review and perspectives on open problems. arXiv:2005.01643. Retrieved from https:\/\/arxiv.org\/abs\/2005.01643"},{"key":"e_1_3_2_33_2","unstructured":"Yuxi Li. 2017. Deep reinforcement learning: An overview. arXiv:1701.07274. Retrieved from https:\/\/arxiv.org\/abs\/1701.07274"},{"key":"e_1_3_2_34_2","unstructured":"Luofeng Liao Zuyue Fu Zhuoran Yang Yixin Wang Mladen Kolar and Zhaoran Wang. 2021. Instrumental variable value iteration for causal offline reinforcement learning. arXiv:2102.09907. Retrieved from https:\/\/arxiv.org\/abs\/2102.09907"},{"key":"e_1_3_2_35_2","unstructured":"Yen-Chen Lin Zhang-Wei Hong Yuan-Hong Liao Meng-Li Shih Ming-Yu Liu and Min Sun. 2017. Tactics of adversarial attack on deep reinforcement learning agents. arXiv:1703.06748. Retrieved from https:\/\/arxiv.org\/abs\/1703.06748"},{"key":"e_1_3_2_36_2","first-page":"13945","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201924)","author":"Liu Jinxin","year":"2023","unstructured":"Jinxin Liu, Ziqi Zhang, Zhenyu Wei, Zifeng Zhuang, Yachen Kang, Sibo Gai, and Donglin Wang. 2023. Beyond OOD state actions: Supported cross-domain offline reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201924). 13945\u201313953."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dss.2015.03.008"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380130"},{"key":"e_1_3_2_39_2","first-page":"3391","volume-title":"Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics","author":"Madjiheurem Sephora","year":"2019","unstructured":"Sephora Madjiheurem and Laura Toni. 2019. Representation learning on graphs: A reinforcement learning application. In Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 3391\u20133399."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1086\/354848"},{"key":"e_1_3_2_41_2","unstructured":"Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv:1301.3781."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460231.3474231"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313616"},{"key":"e_1_3_2_45_2","first-page":"32211","article-title":"Robust reinforcement learning using offline data","volume":"35","author":"Panaganti Kishan","year":"2022","unstructured":"Kishan Panaganti, Zaiyan Xu, Dileep Kalathil, and Mohammad Ghavamzadeh. 2022. Robust reinforcement learning using offline data. Advances in Neural Information Processing Systems 35 (2022), 32211\u201332224.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403121"},{"key":"e_1_3_2_47_2","first-page":"1889","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Schulman John","year":"2015","unstructured":"John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015. Trust region policy optimization. In Proceedings of the International Conference on Machine Learning. PMLR, 1889\u20131897."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33014902"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357895"},{"key":"e_1_3_2_50_2","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. MIT Press."},{"key":"e_1_3_2_51_2","first-page":"3632","article-title":"Off-policy evaluation for slate recommendation","volume":"30","author":"Swaminathan Adith","year":"2017","unstructured":"Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni. 2017. Off-policy evaluation for slate recommendation. Advances in Neural Information Processing Systems 30 (2017), 3632\u20133642.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-47426-3_2"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591648"},{"key":"e_1_3_2_54_2","unstructured":"Siyu Wang Xiaocong Chen and Lina Yao. 2024. Retentive decision transformer with adaptive masking for reinforcement learning based recommendation systems. arXiv:2403.17634. Retrieved from https:\/\/arxiv.org\/abs\/2403.17634"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485447.3512251"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2018.00074"},{"key":"e_1_3_2_57_2","unstructured":"Yanan Wang Yong Ge Li Li Rui Chen and Tong Xu. 2020. Offline meta-level model-based reinforcement learning approach for cold-start recommendation. arXiv:2012.02476. Retrieved from https:\/\/arxiv.org\/abs\/2012.02476"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992696"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3535101"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3604915.3610641"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i5.16579"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i8.20849"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1060"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557083"},{"key":"e_1_3_2_66_2","unstructured":"Junjie Zhang Ruobing Xie Yupeng Hou Wayne Xin Zhao Leyu Lin and Ji-Rong Wen. 2023. Recommendation as instruction following: A large language model empowered recommendation approach. arXiv:2305.07001."},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i10.21424"},{"key":"e_1_3_2_68_2","article-title":"Text-based interactive recommendation via constraint-augmented reinforcement learning","volume":"32","author":"Zhang Ruiyi","year":"2019","unstructured":"Ruiyi Zhang, Tong Yu, Yilin Shen, Hongxia Jin, and Changyou Chen. 2019. Text-based interactive recommendation via constraint-augmented reinforcement learning. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3158369"},{"key":"e_1_3_2_70_2","first-page":"5757","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics","author":"Zhang Xuezhou","year":"2022","unstructured":"Xuezhou Zhang, Yiding Chen, Xiaojin Zhu, and Wen Sun. 2022. Corruption-robust offline reinforcement learning. In Proceedings of the International Conference on Artificial Intelligence and Statistics. PMLR, 5757\u20135773."},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1561\/1500000066"},{"key":"e_1_3_2_72_2","unstructured":"Kesen Zhao Shuchang Liu Qingpeng Cai Xiangyu Zhao Ziru Liu Dong Zheng Peng Jiang and Kun Gai. 2023. KuaiSim: A comprehensive simulator for recommender systems. arXiv:2309.12645."},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401171"},{"key":"e_1_3_2_74_2","unstructured":"Wayne Xin Zhao Kun Zhou Junyi Li Tianyi Tang Xiaolei Wang Yupeng Hou Yingqian Min Beichen Zhang Junjie Zhang Zican Dong Yifan Du Chen Yang Yushuo Chen Zhipeng Chen Jinhao Jiang Ruiyang Ren Yifan Li Xinyu Tang Zikang Liu Peiyu Liu Jian-Yun Nie and Ji-Rong Wen. 2023. A survey of large language models. arXiv:2303.18223. Retrieved from https:\/\/arxiv.org\/abs\/2303.18223"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240323.3240374"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/3178876.3185994"},{"key":"e_1_3_2_77_2","unstructured":"Zheng-Mao Zhu Xiong-Hui Chen Hong-Long Tian Kun Zhang and Yang Yu. 2022. Offline reinforcement learning with causal structured world models. arXiv:2206.01474. Retrieved from https:\/\/arxiv.org\/abs\/2206.01474"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3661996","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3661996","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:06:22Z","timestamp":1750277182000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3661996"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,19]]},"references-count":76,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,11,30]]}},"alternative-id":["10.1145\/3661996"],"URL":"https:\/\/doi.org\/10.1145\/3661996","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,19]]},"assertion":[{"value":"2023-08-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}