{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T20:22:53Z","timestamp":1776889373311,"version":"3.51.2"},"publisher-location":"New York, NY, USA","reference-count":31,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,7,6]],"date-time":"2022-07-06T00:00:00Z","timestamp":1657065600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,7,6]]},"DOI":"10.1145\/3477495.3531796","type":"proceedings-article","created":{"date-parts":[[2022,7,7]],"date-time":"2022-07-07T15:12:13Z","timestamp":1657206733000},"page":"2008-2012","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["Value Penalized Q-Learning for Recommender Systems"],"prefix":"10.1145","author":[{"given":"Chengqian","family":"Gao","sequence":"first","affiliation":[{"name":"Tsinghua University, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ke","family":"Xu","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kuangqi","family":"Zhou","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lanqing","family":"Li","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xueqian","family":"Wang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bo","family":"Yuan","sequence":"additional","affiliation":[{"name":"Tsinghua University, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peilin","family":"Zhao","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Prefix, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,7,7]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"M Mehdi Afsar Trafford Crump and Behrouz Far. 2021. Reinforcement learning based recommender systems: A survey. arXiv:2101.06286  M Mehdi Afsar Trafford Crump and Behrouz Far. 2021. Reinforcement learning based recommender systems: A survey. arXiv:2101.06286"},{"key":"e_1_3_2_2_2_1","volume-title":"An Optimistic Perspective on Offline Reinforcement Learning. In ICML 2020","volume":"114","author":"Agarwal Rishabh","year":"2020","unstructured":"Rishabh Agarwal , Dale Schuurmans , and Mohammad Norouzi . 2020 . An Optimistic Perspective on Offline Reinforcement Learning. In ICML 2020 , 13--18 July 2020, Virtual Event (Proceedings of Machine Learning Research , Vol. 119). PMLR, 104-- 114 . http:\/\/proceedings.mlr.press\/v119\/agarwal20c.html Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. 2020. An Optimistic Perspective on Offline Reinforcement Learning. In ICML 2020, 13--18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119). PMLR, 104--114. http:\/\/proceedings.mlr.press\/v119\/agarwal20c.html"},{"key":"e_1_3_2_2_3_1","volume-title":"Statistical Estimates and Transformed Beta Variables. Almqvist & Wiksell","author":"Blom Gunnar","unstructured":"Gunnar Blom . 1958. Statistical Estimates and Transformed Beta Variables. Almqvist & Wiksell , John Wiley & Sons, Inc. , Sweden . Gunnar Blom. 1958. Statistical Estimates and Transformed Beta Variables. Almqvist & Wiksell, John Wiley & Sons, Inc., Sweden."},{"key":"e_1_3_2_2_4_1","volume-title":"Bellemare","author":"Buckman Jacob","year":"2020","unstructured":"Jacob Buckman , Carles Gelada , and Marc G . Bellemare . 2020 . The Importance of Pessimism in Fixed-Dataset Policy Optimization . arXiv:2009.06799 https: \/\/arxiv.org\/abs\/2009.06799 Jacob Buckman, Carles Gelada, and Marc G. Bellemare. 2020. The Importance of Pessimism in Fixed-Dataset Policy Optimization. arXiv:2009.06799 https: \/\/arxiv.org\/abs\/2009.06799"},{"key":"e_1_3_2_2_5_1","unstructured":"Justin Fu Aviral Kumar Ofir Nachum George Tucker and Sergey Levine. 2020. D4RL: Datasets for Deep Data-Driven Reinforcement Learning. arXiv:2004.07219 https:\/\/arxiv.org\/abs\/2004.07219  Justin Fu Aviral Kumar Ofir Nachum George Tucker and Sergey Levine. 2020. D4RL: Datasets for Deep Data-Driven Reinforcement Learning. arXiv:2004.07219 https:\/\/arxiv.org\/abs\/2004.07219"},{"key":"e_1_3_2_2_6_1","unstructured":"Scott Fujimoto and Shixiang Shane Gu. 2021. A Minimalist Approach to Offline Reinforcement Learning. arXiv:2106.06860 https:\/\/arxiv.org\/abs\/2106.06860  Scott Fujimoto and Shixiang Shane Gu. 2021. A Minimalist Approach to Offline Reinforcement Learning. arXiv:2106.06860 https:\/\/arxiv.org\/abs\/2106.06860"},{"key":"e_1_3_2_2_7_1","volume-title":"ICML 2019","volume":"2062","author":"Fujimoto Scott","year":"2019","unstructured":"Scott Fujimoto , David Meger , and Doina Precup . 2019 . Off-Policy Deep Reinforcement Learning without Exploration . In ICML 2019 , 9--15 June 2019, Long Beach, California, USA (Proceedings of Machine Learning Research , Vol. 97). PMLR, 2052-- 2062 . http:\/\/proceedings.mlr.press\/v97\/fujimoto19a.html Scott Fujimoto, David Meger, and Doina Precup. 2019. Off-Policy Deep Reinforcement Learning without Exploration. In ICML 2019, 9--15 June 2019, Long Beach, California, USA (Proceedings of Machine Learning Research, Vol. 97). PMLR, 2052--2062. http:\/\/proceedings.mlr.press\/v97\/fujimoto19a.html"},{"key":"e_1_3_2_2_8_1","volume-title":"Expected values of normal order statistics. Biometrika 48, 1 and 2","author":"Harter H. Leon","year":"1961","unstructured":"H. Leon Harter . 1961. Expected values of normal order statistics. Biometrika 48, 1 and 2 ( 1961 ), 151--165. H. Leon Harter. 1961. Expected values of normal order statistics. Biometrika 48, 1 and 2 (1961), 151--165."},{"key":"e_1_3_2_2_9_1","volume-title":"Craig Ferguson, \u00c0gata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard.","author":"Jaques Natasha","year":"2019","unstructured":"Natasha Jaques , Asma Ghandeharioun , Judy Hanwen Shen , Craig Ferguson, \u00c0gata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard. 2019 . Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog . arXiv:1907.00456 http:\/\/arxiv.org\/abs\/1907.00456 Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, \u00c0gata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard. 2019. Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog. arXiv:1907.00456 http:\/\/arxiv.org\/abs\/1907.00456"},{"key":"e_1_3_2_2_10_1","unstructured":"Ying Jin Zhuoran Yang and Zhaoran Wang. 2020. Is Pessimism Provably Efficient for Offline RL? arXiv:2012.15085 https:\/\/arxiv.org\/abs\/2012.15085  Ying Jin Zhuoran Yang and Zhaoran Wang. 2020. Is Pessimism Provably Efficient for Offline RL? arXiv:2012.15085 https:\/\/arxiv.org\/abs\/2012.15085"},{"key":"e_1_3_2_2_11_1","volume-title":"Self-Attentive Sequential Recommendation. In ICDM 2018","author":"Kang Wangcheng","year":"2018","unstructured":"Wangcheng Kang and Julian J . McAuley. 2018 . Self-Attentive Sequential Recommendation. In ICDM 2018 , Singapore, November 17--20 , 2018 . IEEE Computer Society, 197--206. https:\/\/doi.org\/10.1109\/ICDM.2018.00035 10.1109\/ICDM.2018.00035 Wangcheng Kang and Julian J. McAuley. 2018. Self-Attentive Sequential Recommendation. In ICDM 2018, Singapore, November 17--20, 2018. IEEE Computer Society, 197--206. https:\/\/doi.org\/10.1109\/ICDM.2018.00035"},{"key":"e_1_3_2_2_12_1","volume-title":"Actor-critic Algorithms. Ph. D. Dissertation","author":"Konda Vijaymohan","unstructured":"Vijaymohan Konda . 2002. Actor-critic Algorithms. Ph. D. Dissertation . Massachusetts Institute of Technology , Cambridge, MA, USA . http:\/\/hdl.handle.net\/ 1721.1\/8120 Vijaymohan Konda. 2002. Actor-critic Algorithms. Ph. D. Dissertation. Massachusetts Institute of Technology, Cambridge, MA, USA. http:\/\/hdl.handle.net\/ 1721.1\/8120"},{"key":"e_1_3_2_2_13_1","volume-title":"Offline Reinforcement Learning with Fisher Divergence Critic Regularization. In ICML 2021","volume":"5783","author":"Kostrikov Ilya","year":"2021","unstructured":"Ilya Kostrikov , Rob Fergus , Jonathan Tompson , and Ofir Nachum . 2021 . Offline Reinforcement Learning with Fisher Divergence Critic Regularization. In ICML 2021 , 18--24 July 2021, Virtual Event (Proceedings of Machine Learning Research , Vol. 139). PMLR, 5774-- 5783 . http:\/\/proceedings.mlr.press\/v139\/kostrikov21a.html Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum. 2021. Offline Reinforcement Learning with Fisher Divergence Critic Regularization. In ICML 2021, 18--24 July 2021, Virtual Event (Proceedings of Machine Learning Research, Vol. 139). PMLR, 5774--5783. http:\/\/proceedings.mlr.press\/v139\/kostrikov21a.html"},{"key":"e_1_3_2_2_14_1","volume-title":"NeurIPS","author":"Kumar Aviral","year":"2019","unstructured":"Aviral Kumar , Justin Fu , Matthew Soh , George Tucker , and Sergey Levine . 2019. Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction . In NeurIPS 2019 , December 8--14, 2019, Vancouver, BC, Canada . 11761--11771. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/ c2073ffa77b5357a498057413bb09d3a-Abstract.html Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. 2019. Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction. In NeurIPS 2019, December 8--14, 2019, Vancouver, BC, Canada. 11761--11771. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/ c2073ffa77b5357a498057413bb09d3a-Abstract.html"},{"key":"e_1_3_2_2_15_1","volume-title":"NeurIPS","author":"Kumar Aviral","year":"2020","unstructured":"Aviral Kumar , Aurick Zhou , George Tucker , and Sergey Levine . 2020. Conservative Q-Learning for Offline Reinforcement Learning . In NeurIPS 2020 , December 6--12, 2020, virtual. https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/ 0d2b2061826a5df3221116a5085a6052-Abstract.html Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020. Conservative Q-Learning for Offline Reinforcement Learning. In NeurIPS 2020, December 6--12, 2020, virtual. https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/ 0d2b2061826a5df3221116a5085a6052-Abstract.html"},{"key":"e_1_3_2_2_16_1","volume-title":"NeurIPS","author":"Lazic Nevena","year":"2018","unstructured":"Nevena Lazic , Craig Boutilier , Tyler Lu , Eehern Wong , Binz Roy , M. K. Ryu , and Greg Imwalle . 2018. Data center cooling using model-predictive control . In NeurIPS 2018 , December 3--8, 2018, Montr\u00e9al, Canada . 3818--3827. https: \/\/proceedings.neurips.cc\/paper\/2018\/hash\/059fdcd96baeb75112f09fa1dcc740ccAbstract.html Nevena Lazic, Craig Boutilier, Tyler Lu, Eehern Wong, Binz Roy, M. K. Ryu, and Greg Imwalle. 2018. Data center cooling using model-predictive control. In NeurIPS 2018, December 3--8, 2018, Montr\u00e9al, Canada. 3818--3827. https: \/\/proceedings.neurips.cc\/paper\/2018\/hash\/059fdcd96baeb75112f09fa1dcc740ccAbstract.html"},{"key":"e_1_3_2_2_17_1","unstructured":"Sergey Levine Aviral Kumar George Tucker and Justin Fu. 2020. Offline Reinforcement Learning: Tutorial Review and Perspectives on Open Problems. arXiv:2005.01643 https:\/\/arxiv.org\/abs\/2005.01643  Sergey Levine Aviral Kumar George Tucker and Justin Fu. 2020. Offline Reinforcement Learning: Tutorial Review and Perspectives on Open Problems. arXiv:2005.01643 https:\/\/arxiv.org\/abs\/2005.01643"},{"key":"e_1_3_2_2_18_1","volume-title":"Russell","author":"Ng Andrew Y.","year":"1999","unstructured":"Andrew Y. Ng , Daishi Harada , and Stuart J . Russell . 1999 . Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. In (ICML 1999), Bled, Slovenia, June 27 - 30, 1999. Morgan Kaufmann , 278--287. Andrew Y. Ng, Daishi Harada, and Stuart J. Russell. 1999. Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. In (ICML 1999), Bled, Slovenia, June 27 - 30, 1999. Morgan Kaufmann, 278--287."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.2307\/2347982"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3159652.3159656"},{"key":"e_1_3_2_2_21_1","volume-title":"OffPolicy Recommendation System Without Exploration. In PAKDD 2020, Singapore, May 11--14, 2020, Proceedings, Part I (Lecture Notes in Computer Science","volume":"27","author":"Wang Chengwei","year":"2020","unstructured":"Chengwei Wang , Tengfei Zhou , Chen Chen , Tianlei Hu , and Gang Chen . 2020 . OffPolicy Recommendation System Without Exploration. In PAKDD 2020, Singapore, May 11--14, 2020, Proceedings, Part I (Lecture Notes in Computer Science , Vol. 12084). Springer, 16-- 27 . https:\/\/doi.org\/10.1007\/978--3-030--47426--3_2 10.1007\/978--3-030--47426--3_2 Chengwei Wang, Tengfei Zhou, Chen Chen, Tianlei Hu, and Gang Chen. 2020. OffPolicy Recommendation System Without Exploration. In PAKDD 2020, Singapore, May 11--14, 2020, Proceedings, Part I (Lecture Notes in Computer Science, Vol. 12084). Springer, 16--27. https:\/\/doi.org\/10.1007\/978--3-030--47426--3_2"},{"key":"e_1_3_2_2_22_1","unstructured":"Yifan Wu George Tucker and Ofir Nachum. 2019. Behavior Regularized Offline Reinforcement Learning. arXiv:1911.11361 http:\/\/arxiv.org\/abs\/1911.11361  Yifan Wu George Tucker and Ofir Nachum. 2019. Behavior Regularized Offline Reinforcement Learning. arXiv:1911.11361 http:\/\/arxiv.org\/abs\/1911.11361"},{"key":"e_1_3_2_2_23_1","volume-title":"Uncertainty Weighted ActorCritic for Offline Reinforcement Learning. In ICML 2021","volume":"11328","author":"Wu Yue","year":"2021","unstructured":"Yue Wu , Shuangfei Zhai , Nitish Srivastava , Joshua M. Susskind , Jian Zhang , Ruslan Salakhutdinov , and Hanlin Goh . 2021 . Uncertainty Weighted ActorCritic for Offline Reinforcement Learning. In ICML 2021 , 18--24 July 2021, Virtual Event (Proceedings of Machine Learning Research , Vol. 139). PMLR, 11319-- 11328 . http:\/\/proceedings.mlr.press\/v139\/wu21i.html Yue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh. 2021. Uncertainty Weighted ActorCritic for Offline Reinforcement Learning. In ICML 2021, 18--24 July 2021, Virtual Event (Proceedings of Machine Learning Research, Vol. 139). PMLR, 11319--11328. http:\/\/proceedings.mlr.press\/v139\/wu21i.html"},{"key":"e_1_3_2_2_24_1","volume-title":"Self-Supervised Reinforcement Learning for Recommender Systems. In SIGIR 2020","author":"Xin Xin","year":"2020","unstructured":"Xin Xin , Alexandros Karatzoglou , Ioannis Arapakis , and Joemon M. Jose . 2020 . Self-Supervised Reinforcement Learning for Recommender Systems. In SIGIR 2020 , Virtual Event, China, July 25--30 , 2020 . ACM, 931--940. https:\/\/doi.org\/10. 1145\/3397271.3401147 Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M. Jose. 2020. Self-Supervised Reinforcement Learning for Recommender Systems. In SIGIR 2020, Virtual Event, China, July 25--30, 2020. ACM, 931--940. https:\/\/doi.org\/10. 1145\/3397271.3401147"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.6144"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290975"},{"key":"e_1_3_2_2_27_1","first-page":"236","article-title":"Cost-Sensitive Portfolio Selection via Deep Reinforcement Learning","volume":"34","author":"Zhang Yifan","year":"2022","unstructured":"Yifan Zhang , Peilin Zhao , Qingyao Wu , Bin Li , Junzhou Huang , and Mingkui Tan . 2022 . Cost-Sensitive Portfolio Selection via Deep Reinforcement Learning . IEEE TKDE. 34 , 1 (2022), 236 -- 248 . https:\/\/doi.org\/10.1109\/TKDE.2020.2979700 10.1109\/TKDE.2020.2979700 Yifan Zhang, Peilin Zhao, Qingyao Wu, Bin Li, Junzhou Huang, and Mingkui Tan. 2022. Cost-Sensitive Portfolio Selection via Deep Reinforcement Learning. IEEE TKDE. 34, 1 (2022), 236--248. https:\/\/doi.org\/10.1109\/TKDE.2020.2979700","journal-title":"IEEE TKDE."},{"key":"e_1_3_2_2_28_1","unstructured":"Xiangyu Zhao Changsheng Gu Haoshenglun Zhang Xiaobing Liu Xiwang Yang and Jiliang Tang. 2019. Deep Reinforcement Learning for Online Advertising in Recommender Systems. arXiv:1909.03602 http:\/\/arxiv.org\/abs\/1909.03602  Xiangyu Zhao Changsheng Gu Haoshenglun Zhang Xiaobing Liu Xiwang Yang and Jiliang Tang. 2019. Deep Reinforcement Learning for Online Advertising in Recommender Systems. arXiv:1909.03602 http:\/\/arxiv.org\/abs\/1909.03602"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240323.3240374"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219886"},{"key":"e_1_3_2_2_31_1","volume-title":"Xing Xie, and Zhenhui Li.","author":"Zheng Guanjie","year":"2018","unstructured":"Guanjie Zheng , Fuzheng Zhang , Zihan Zheng , Yang Xiang , Nicholas Jing Yuan , Xing Xie, and Zhenhui Li. 2018 . DRN : A Deep Reinforcement Learning Framework for News Recommendation. In WWW 2018, Lyon, France, April 23--27, 2018. ACM , 167--176. https:\/\/doi.org\/10.1145\/3178876.3185994 10.1145\/3178876.3185994 Guanjie Zheng, Fuzheng Zhang, Zihan Zheng, Yang Xiang, Nicholas Jing Yuan, Xing Xie, and Zhenhui Li. 2018. DRN: A Deep Reinforcement Learning Framework for News Recommendation. In WWW 2018, Lyon, France, April 23--27, 2018. ACM, 167--176. https:\/\/doi.org\/10.1145\/3178876.3185994"}],"event":{"name":"SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","location":"Madrid Spain","acronym":"SIGIR '22","sponsor":["SIGIR ACM Special Interest Group on Information Retrieval"]},"container-title":["Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3477495.3531796","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3477495.3531796","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:10:25Z","timestamp":1750183825000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3477495.3531796"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,6]]},"references-count":31,"alternative-id":["10.1145\/3477495.3531796","10.1145\/3477495"],"URL":"https:\/\/doi.org\/10.1145\/3477495.3531796","relation":{},"subject":[],"published":{"date-parts":[[2022,7,6]]},"assertion":[{"value":"2022-07-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}