{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:46:18Z","timestamp":1787017578116,"version":"build-2736575974"},"publisher-location":"New York, NY, USA","reference-count":72,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,9,13]],"date-time":"2021-09-13T00:00:00Z","timestamp":1631491200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-sa\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,9,13]]},"DOI":"10.1145\/3460231.3474236","type":"proceedings-article","created":{"date-parts":[[2021,9,13]],"date-time":"2021-09-13T17:45:02Z","timestamp":1631555102000},"page":"85-95","source":"Crossref","is-referenced-by-count":38,"title":["Values of User Exploration in Recommender Systems"],"prefix":"10.1145","author":[{"given":"Minmin","family":"Chen","sequence":"first","affiliation":[{"name":"Google, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuyan","family":"Wang","sequence":"additional","affiliation":[{"name":"Google Research, Brain Team, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Can","family":"Xu","sequence":"additional","affiliation":[{"name":"Google Inc, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ya","family":"Le","sequence":"additional","affiliation":[{"name":"Google AI, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohit","family":"Sharma","sequence":"additional","affiliation":[{"name":"Google, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lee","family":"Richardson","sequence":"additional","affiliation":[{"name":"Google, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Su-Lin","family":"Wu","sequence":"additional","affiliation":[{"name":"Google, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ed","family":"Chi","sequence":"additional","affiliation":[{"name":"Google, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,9,13]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015430"},{"key":"e_1_3_2_2_2_1","unstructured":"Joshua Achiam David Held Aviv Tamar and Pieter Abbeel. 2017. Constrained policy optimization. arXiv preprint arXiv:1705.10528(2017).  Joshua Achiam David Held Aviv Tamar and Pieter Abbeel. 2017. Constrained policy optimization. arXiv preprint arXiv:1705.10528(2017)."},{"key":"e_1_3_2_2_3_1","volume-title":"Proc. of the 1st International Workshop on Novelty and Diversity in Recommender Systems (DiveRS","author":"Adomavicius Gediminas","year":"2011","unstructured":"Gediminas Adomavicius and YoungOk Kwon . 2011 . Maximizing aggregate recommendation diversity: A graph-theoretic approach . In Proc. of the 1st International Workshop on Novelty and Diversity in Recommender Systems (DiveRS 2011). Citeseer, 3\u201310. Gediminas Adomavicius and YoungOk Kwon. 2011. Maximizing aggregate recommendation diversity: A graph-theoretic approach. In Proc. of the 1st International Workshop on Novelty and Diversity in Recommender Systems (DiveRS 2011). Citeseer, 3\u201310."},{"key":"e_1_3_2_2_4_1","volume-title":"Conference on learning theory. JMLR Workshop and Conference Proceedings, 39\u20131.","author":"Agrawal Shipra","year":"2012","unstructured":"Shipra Agrawal and Navin Goyal . 2012 . Analysis of thompson sampling for the multi-armed bandit problem . In Conference on learning theory. JMLR Workshop and Conference Proceedings, 39\u20131. Shipra Agrawal and Navin Goyal. 2012. Analysis of thompson sampling for the multi-armed bandit problem. In Conference on learning theory. JMLR Workshop and Conference Proceedings, 39\u20131."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380281"},{"key":"e_1_3_2_2_6_1","unstructured":"Marc\u00a0G Bellemare Sriram Srinivasan Georg Ostrovski Tom Schaul David Saxton and Remi Munos. 2016. Unifying count-based exploration and intrinsic motivation. arXiv preprint arXiv:1606.01868(2016).  Marc\u00a0G Bellemare Sriram Srinivasan Georg Ostrovski Tom Schaul David Saxton and Remi Munos. 2016. Unifying count-based exploration and intrinsic motivation. arXiv preprint arXiv:1606.01868(2016)."},{"key":"e_1_3_2_2_7_1","volume-title":"Latent dirichlet allocation. the Journal of machine Learning research 3","author":"Blei M","year":"2003","unstructured":"David\u00a0 M Blei , Andrew\u00a0 Y Ng , and Michael\u00a0 I Jordan . 2003. Latent dirichlet allocation. the Journal of machine Learning research 3 ( 2003 ), 993\u20131022. David\u00a0M Blei, Andrew\u00a0Y Ng, and Michael\u00a0I Jordan. 2003. Latent dirichlet allocation. the Journal of machine Learning research 3 (2003), 993\u20131022."},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Pablo Castells Sa\u00fal Vargas and Jun Wang. 2011. Novelty and diversity metrics for recommender systems: choice discovery and relevance. (2011).  Pablo Castells Sa\u00fal Vargas and Jun Wang. 2011. Novelty and diversity metrics for recommender systems: choice discovery and relevance. (2011).","DOI":"10.1145\/2043932.2044019"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454008.1454038"},{"key":"e_1_3_2_2_10_1","unstructured":"Olivier Chapelle and Lihong Li. 2011. An empirical evaluation of thompson sampling. In Advances in neural information processing systems. 2249\u20132257.  Olivier Chapelle and Lihong Li. 2011. An empirical evaluation of thompson sampling. In Advances in neural information processing systems. 2249\u20132257."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290999"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature04766"},{"key":"e_1_3_2_2_14_1","unstructured":"Gabriel Dulac-Arnold Richard Evans Hado van Hasselt Peter Sunehag Timothy Lillicrap Jonathan Hunt Timothy Mann Theophane Weber Thomas Degris and Ben Coppin. 2015. Deep reinforcement learning in large discrete action spaces. arXiv preprint arXiv:1512.07679(2015).  Gabriel Dulac-Arnold Richard Evans Hado van Hasselt Peter Sunehag Timothy Lillicrap Jonathan Hunt Timothy Mann Theophane Weber Thomas Degris and Ben Coppin. 2015. Deep reinforcement learning in large discrete action spaces. arXiv preprint arXiv:1512.07679(2015)."},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1864708.1864761"},{"key":"e_1_3_2_2_16_1","unstructured":"Dibya Ghosh Abhishek Gupta and Sergey Levine. 2018. Learning actionable representations with goal-conditioned policies. arXiv preprint arXiv:1811.07819(2018).  Dibya Ghosh Abhishek Gupta and Sergey Levine. 2018. Learning actionable representations with goal-conditioned policies. arXiv preprint arXiv:1811.07819(2018)."},{"key":"e_1_3_2_2_17_1","volume-title":"Multi-armed bandit allocation indices","author":"Gittins John","unstructured":"John Gittins , Kevin Glazebrook , and Richard Weber . 2011. Multi-armed bandit allocation indices . John Wiley & Sons . John Gittins, Kevin Glazebrook, and Richard Weber. 2011. Multi-armed bandit allocation indices. John Wiley & Sons."},{"key":"e_1_3_2_2_18_1","volume-title":"International Conference on Machine Learning. 2555\u20132565","author":"Hafner Danijar","year":"2019","unstructured":"Danijar Hafner , Timothy Lillicrap , Ian Fischer , Ruben Villegas , David Ha , Honglak Lee , and James Davidson . 2019 . Learning latent dynamics for planning from pixels . In International Conference on Machine Learning. 2555\u20132565 . Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. 2019. Learning latent dynamics for planning from pixels. In International Conference on Machine Learning. 2555\u20132565."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/963770.963772"},{"key":"e_1_3_2_2_20_1","volume-title":"Darla: Improving zero-shot transfer in reinforcement learning. arXiv preprint arXiv:1707.08475(2017).","author":"Higgins Irina","year":"2017","unstructured":"Irina Higgins , Arka Pal , Andrei\u00a0 A Rusu , Loic Matthey , Christopher\u00a0 P Burgess , Alexander Pritzel , Matthew Botvinick , Charles Blundell , and Alexander Lerchner . 2017 . Darla: Improving zero-shot transfer in reinforcement learning. arXiv preprint arXiv:1707.08475(2017). Irina Higgins, Arka Pal, Andrei\u00a0A Rusu, Loic Matthey, Christopher\u00a0P Burgess, Alexander Pritzel, Matthew Botvinick, Charles Blundell, and Alexander Lerchner. 2017. Darla: Improving zero-shot transfer in reinforcement learning. arXiv preprint arXiv:1707.08475(2017)."},{"key":"e_1_3_2_2_21_1","volume-title":"Vime: Variational information maximizing exploration. arXiv preprint arXiv:1605.09674(2016).","author":"Houthooft Rein","year":"2016","unstructured":"Rein Houthooft , Xi Chen , Yan Duan , John Schulman , Filip De\u00a0Turck , and Pieter Abbeel . 2016 . Vime: Variational information maximizing exploration. arXiv preprint arXiv:1605.09674(2016). Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De\u00a0Turck, and Pieter Abbeel. 2016. Vime: Variational information maximizing exploration. arXiv preprint arXiv:1605.09674(2016)."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219846"},{"key":"e_1_3_2_2_23_1","unstructured":"Eugene Ie Vihan Jain Jing Wang Sanmit Narvekar Ritesh Agarwal Rui Wu Heng-Tze Cheng Tushar Chandra and Craig Boutilier. 2019. SlateQ: A tractable decomposition for reinforcement learning with recommendation sets. (2019).  Eugene Ie Vihan Jain Jing Wang Sanmit Narvekar Ritesh Agarwal Rui Wu Heng-Tze Cheng Tushar Chandra and Craig Boutilier. 2019. SlateQ: A tractable decomposition for reinforcement learning with recommendation sets. (2019)."},{"key":"e_1_3_2_2_24_1","volume-title":"Bayesian surprise attracts human attention. Vision research 49, 10","author":"Itti Laurent","year":"2009","unstructured":"Laurent Itti and Pierre Baldi . 2009. Bayesian surprise attracts human attention. Vision research 49, 10 ( 2009 ), 1295\u20131306. Laurent Itti and Pierre Baldi. 2009. Bayesian surprise attracts human attention. Vision research 49, 10 (2009), 1295\u20131306."},{"key":"e_1_3_2_2_25_1","unstructured":"Max Jaderberg Volodymyr Mnih Wojciech\u00a0Marian Czarnecki Tom Schaul Joel\u00a0Z Leibo David Silver and Koray Kavukcuoglu. 2017. Reinforcement learning with unsupervised auxiliary tasks. In ICLR.  Max Jaderberg Volodymyr Mnih Wojciech\u00a0Marian Czarnecki Tom Schaul Joel\u00a0Z Leibo David Silver and Koray Kavukcuoglu. 2017. Reinforcement learning with unsupervised auxiliary tasks. In ICLR."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2926720"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913495721"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2009.263"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2016.08.014"},{"key":"e_1_3_2_2_30_1","volume-title":"Asymptotically efficient adaptive allocation rules. Advances in applied mathematics 6, 1","author":"Lai Tze\u00a0Leung","year":"1985","unstructured":"Tze\u00a0Leung Lai and Herbert Robbins . 1985. Asymptotically efficient adaptive allocation rules. Advances in applied mathematics 6, 1 ( 1985 ), 4\u201322. Tze\u00a0Leung Lai and Herbert Robbins. 1985. Asymptotically efficient adaptive allocation rules. Advances in applied mathematics 6, 1 (1985), 4\u201322."},{"key":"e_1_3_2_2_31_1","unstructured":"Thomas Laurent and James von Brecht. 2016. A recurrent neural network without chaos. arXiv preprint arXiv:1612.06212(2016).  Thomas Laurent and James von Brecht. 2016. A recurrent neural network without chaos. arXiv preprint arXiv:1612.06212(2016)."},{"key":"e_1_3_2_2_32_1","unstructured":"Sergey Levine Aviral Kumar George Tucker and Justin Fu. 2020. Offline reinforcement learning: Tutorial review and perspectives on open problems. arXiv preprint arXiv:2005.01643(2020).  Sergey Levine Aviral Kumar George Tucker and Justin Fu. 2020. Offline reinforcement learning: Tutorial review and perspectives on open problems. arXiv preprint arXiv:2005.01643(2020)."},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364917710318"},{"key":"e_1_3_2_2_34_1","unstructured":"Feng Liu Ruiming Tang Xutao Li Weinan Zhang Yunming Ye Haokun Chen Huifeng Guo and Yuzhou Zhang. 2018. Deep reinforcement learning based recommendation with explicit user-item interactions modeling. arXiv preprint arXiv:1810.12027(2018).  Feng Liu Ruiming Tang Xutao Li Weinan Zhang Yunming Ye Haokun Chen Huifeng Guo and Yuzhou Zhang. 2018. Deep reinforcement learning based recommendation with explicit user-item interactions modeling. arXiv preprint arXiv:1810.12027(2018)."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"crossref","unstructured":"Sean\u00a0M McNee John Riedl and Joseph\u00a0A Konstan. 2006. Being accurate is not enough: how accuracy metrics have hurt recommender systems. In CHI\u201906 extended abstracts on Human factors in computing systems. 1097\u20131101.  Sean\u00a0M McNee John Riedl and Joseph\u00a0A Konstan. 2006. Being accurate is not enough: how accuracy metrics have hurt recommender systems. In CHI\u201906 extended abstracts on Human factors in computing systems. 1097\u20131101.","DOI":"10.1145\/1125451.1125659"},{"key":"e_1_3_2_2_36_1","volume-title":"International Conference on Machine Learning. PMLR, 2430\u20132439","author":"Mirhoseini Azalia","year":"2017","unstructured":"Azalia Mirhoseini , Hieu Pham , Quoc\u00a0 V Le , Benoit Steiner , Rasmus Larsen , Yuefeng Zhou , Naveen Kumar , Mohammad Norouzi , Samy Bengio , and Jeff Dean . 2017 . Device placement optimization with reinforcement learning . In International Conference on Machine Learning. PMLR, 2430\u20132439 . Azalia Mirhoseini, Hieu Pham, Quoc\u00a0V Le, Benoit Steiner, Rasmus Larsen, Yuefeng Zhou, Naveen Kumar, Mohammad Norouzi, Samy Bengio, and Jeff Dean. 2017. Device placement optimization with reinforcement learning. In International Conference on Machine Learning. PMLR, 2430\u20132439."},{"key":"e_1_3_2_2_37_1","volume-title":"International Conference on Machine Learning. PMLR, 6987\u20136998","author":"Mladenov Martin","year":"2020","unstructured":"Martin Mladenov , Elliot Creager , Omer Ben-Porat , Kevin Swersky , Richard Zemel , and Craig Boutilier . 2020 . Optimizing long-term social welfare in recommender systems: A constrained matching approach . In International Conference on Machine Learning. PMLR, 6987\u20136998 . Martin Mladenov, Elliot Creager, Omer Ben-Porat, Kevin Swersky, Richard Zemel, and Craig Boutilier. 2020. Optimizing long-term social welfare in recommender systems: A constrained matching approach. In International Conference on Machine Learning. PMLR, 6987\u20136998."},{"key":"e_1_3_2_2_38_1","volume-title":"International conference on machine learning. 1928\u20131937","author":"Mnih Volodymyr","year":"2016","unstructured":"Volodymyr Mnih , Adria\u00a0Puigdomenech Badia , Mehdi Mirza , Alex Graves , Timothy Lillicrap , Tim Harley , David Silver , and Koray Kavukcuoglu . 2016 . Asynchronous methods for deep reinforcement learning . In International conference on machine learning. 1928\u20131937 . Volodymyr Mnih, Adria\u00a0Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016. Asynchronous methods for deep reinforcement learning. In International conference on machine learning. 1928\u20131937."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1111\/1468-0262.00321"},{"key":"e_1_3_2_2_40_1","unstructured":"Andrew\u00a0Y Ng Daishi Harada and Stuart Russell. 1999. Policy invariance under reward transformations: Theory and application to reward shaping. In Icml Vol.\u00a099. 278\u2013287.  Andrew\u00a0Y Ng Daishi Harada and Stuart Russell. 1999. Policy invariance under reward transformations: Theory and application to reward shaping. In Icml Vol.\u00a099. 278\u2013287."},{"key":"e_1_3_2_2_41_1","unstructured":"Kenta Oku and Fumio Hattori. 2011. Fusion-based Recommender System for Improving Serendipity.DiveRS@ RecSys 816(2011) 19\u201326.  Kenta Oku and Fumio Hattori. 2011. Fusion-based Recommender System for Improving Serendipity.DiveRS@ RecSys 816(2011) 19\u201326."},{"key":"e_1_3_2_2_42_1","unstructured":"Ian Osband Charles Blundell Alexander Pritzel and Benjamin Van\u00a0Roy. 2016. Deep exploration via bootstrapped DQN. arXiv preprint arXiv:1602.04621(2016).  Ian Osband Charles Blundell Alexander Pritzel and Benjamin Van\u00a0Roy. 2016. Deep exploration via bootstrapped DQN. arXiv preprint arXiv:1602.04621(2016)."},{"key":"e_1_3_2_2_43_1","volume-title":"International conference on machine learning. PMLR, 2721\u20132730","author":"Ostrovski Georg","year":"2017","unstructured":"Georg Ostrovski , Marc\u00a0 G Bellemare , A\u00e4ron Oord , and R\u00e9mi Munos . 2017 . Count-based exploration with neural density models . In International conference on machine learning. PMLR, 2721\u20132730 . Georg Ostrovski, Marc\u00a0G Bellemare, A\u00e4ron Oord, and R\u00e9mi Munos. 2017. Count-based exploration with neural density models. In International conference on machine learning. PMLR, 2721\u20132730."},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2017.70"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1982.10477845"},{"key":"e_1_3_2_2_46_1","unstructured":"Gabriel Pereyra George Tucker Jan Chorowski \u0141ukasz Kaiser and Geoffrey Hinton. 2017. Regularizing neural networks by penalizing confident output distributions. arXiv preprint arXiv:1701.06548(2017).  Gabriel Pereyra George Tucker Jan Chorowski \u0141ukasz Kaiser and Geoffrey Hinton. 2017. Regularizing neural networks by penalizing confident output distributions. arXiv preprint arXiv:1701.06548(2017)."},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/371920.372071"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.1991.170605"},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAMD.2010.2056368"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8462891"},{"key":"e_1_3_2_2_51_1","article-title":"An MDP-based recommender system","author":"Shani Guy","year":"2005","unstructured":"Guy Shani , David Heckerman , and Ronen\u00a0 I Brafman . 2005 . An MDP-based recommender system . Journal of Machine Learning Research 6 , Sep (2005). Guy Shani, David Heckerman, and Ronen\u00a0I Brafman. 2005. An MDP-based recommender system. Journal of Machine Learning Research 6, Sep (2005).","journal-title":"Journal of Machine Learning Research 6"},{"key":"e_1_3_2_2_52_1","volume-title":"Mastering the game of go without human knowledge. nature 550, 7676","author":"Silver David","year":"2017","unstructured":"David Silver , Julian Schrittwieser , Karen Simonyan , Ioannis Antonoglou , Aja Huang , Arthur Guez , Thomas Hubert , Lucas Baker , Matthew Lai , Adrian Bolton , 2017. Mastering the game of go without human knowledge. nature 550, 7676 ( 2017 ), 354\u2013359. David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, 2017. Mastering the game of go without human knowledge. nature 550, 7676 (2017), 354\u2013359."},{"key":"e_1_3_2_2_53_1","volume-title":"International conference on case-based reasoning. Springer, 347\u2013361","author":"Smyth Barry","year":"2001","unstructured":"Barry Smyth and Paul McClave . 2001 . Similarity vs. diversity . In International conference on case-based reasoning. Springer, 347\u2013361 . Barry Smyth and Paul McClave. 2001. Similarity vs. diversity. In International conference on case-based reasoning. Springer, 347\u2013361."},{"key":"e_1_3_2_2_54_1","volume-title":"Curl: Contrastive unsupervised representations for reinforcement learning. arXiv preprint arXiv:2004.04136(2020).","author":"Srinivas Aravind","year":"2020","unstructured":"Aravind Srinivas , Michael Laskin , and Pieter Abbeel . 2020 . Curl: Contrastive unsupervised representations for reinforcement learning. arXiv preprint arXiv:2004.04136(2020). Aravind Srinivas, Michael Laskin, and Pieter Abbeel. 2020. Curl: Contrastive unsupervised representations for reinforcement learning. arXiv preprint arXiv:2004.04136(2020)."},{"key":"e_1_3_2_2_55_1","unstructured":"Bradly\u00a0C Stadie Sergey Levine and Pieter Abbeel. 2015. Incentivizing exploration in reinforcement learning with deep predictive models. arXiv preprint arXiv:1507.00814(2015).  Bradly\u00a0C Stadie Sergey Levine and Pieter Abbeel. 2015. Incentivizing exploration in reinforcement learning with deep predictive models. arXiv preprint arXiv:1507.00814(2015)."},{"key":"e_1_3_2_2_56_1","volume-title":"Proceedings of the international conference on artificial neural networks","author":"Storck Jan","year":"1995","unstructured":"Jan Storck , Sepp Hochreiter , and J\u00fcrgen Schmidhuber . 1995 . Reinforcement driven information acquisition in non-deterministic environments . In Proceedings of the international conference on artificial neural networks , Paris, Vol.\u00a02. Citeseer, 159\u2013164. Jan Storck, Sepp Hochreiter, and J\u00fcrgen Schmidhuber. 1995. Reinforcement driven information acquisition in non-deterministic environments. In Proceedings of the international conference on artificial neural networks, Paris, Vol.\u00a02. Citeseer, 159\u2013164."},{"key":"e_1_3_2_2_57_1","volume-title":"Reinforcement learning: An introduction","author":"Sutton S","unstructured":"Richard\u00a0 S Sutton , Andrew\u00a0 G Barto , 1998. Reinforcement learning: An introduction . MIT press . Richard\u00a0S Sutton, Andrew\u00a0G Barto, 1998. Reinforcement learning: An introduction. MIT press."},{"key":"e_1_3_2_2_58_1","volume-title":"31st Conference on Neural Information Processing Systems (NIPS), Vol.\u00a030","author":"Tang Haoran","year":"2017","unstructured":"Haoran Tang , Rein Houthooft , Davis Foote , Adam Stooke , Xi Chen , Yan Duan , John Schulman , Filip De\u00a0Turck , and Pieter Abbeel . 2017 . # exploration: A study of count-based exploration for deep reinforcement learning . In 31st Conference on Neural Information Processing Systems (NIPS), Vol.\u00a030 . 1\u201318. Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De\u00a0Turck, and Pieter Abbeel. 2017. # exploration: A study of count-based exploration for deep reinforcement learning. In 31st Conference on Neural Information Processing Systems (NIPS), Vol.\u00a030. 1\u201318."},{"key":"e_1_3_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/25.3-4.285"},{"key":"e_1_3_2_2_60_1","volume-title":"Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8, 3-4","author":"Williams J","year":"1992","unstructured":"Ronald\u00a0 J Williams . 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8, 3-4 ( 1992 ), 229\u2013256. Ronald\u00a0J Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8, 3-4 (1992), 229\u2013256."},{"key":"e_1_3_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1080\/09540099108946587"},{"key":"e_1_3_2_2_62_1","unstructured":"Hongzhi Yin Bin Cui Jing Li Junjie Yao and Chen Chen. 2012. Challenging the long tail recommendation. arXiv preprint arXiv:1205.6700(2012).  Hongzhi Yin Bin Cui Jing Li Junjie Yao and Chen Chen. 2012. Challenging the long tail recommendation. arXiv preprint arXiv:1205.6700(2012)."},{"key":"e_1_3_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/1277741.1277790"},{"key":"e_1_3_2_2_64_1","volume-title":"Proceedings of the Web Conference","author":"Zhan Ruohan","year":"2021","unstructured":"Ruohan Zhan , Konstantina Christakopoulou , Ya Le , Jayden Ooi , Martin Mladenov , Alex Beutel , Craig Boutilier , Ed Chi , and Minmin Chen . 2021 . Towards Content Provider Aware Recommender Systems: A Simulation Study on the Interplay between User and Provider Utilities . In Proceedings of the Web Conference 2021. 3872\u20133883. Ruohan Zhan, Konstantina Christakopoulou, Ya Le, Jayden Ooi, Martin Mladenov, Alex Beutel, Craig Boutilier, Ed Chi, and Minmin Chen. 2021. Towards Content Provider Aware Recommender Systems: A Simulation Study on the Interplay between User and Provider Utilities. In Proceedings of the Web Conference 2021. 3872\u20133883."},{"key":"e_1_3_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3316481"},{"key":"e_1_3_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/2124295.2124300"},{"key":"e_1_3_2_2_67_1","volume-title":"Deep reinforcement learning for search, recommendation, and online advertising: a survey","author":"Zhao Xiangyu","year":"2019","unstructured":"Xiangyu Zhao , Long Xia , Jiliang Tang , and Dawei Yin . 2019. \u201d Deep reinforcement learning for search, recommendation, and online advertising: a survey \u201d by Xiangyu Zhao, Long Xia , Jiliang Tang, and Dawei Yin with Martin Vesely as coordinator. ACM SIGWEB NewsletterSpring ( 2019 ), 1\u201315. Xiangyu Zhao, Long Xia, Jiliang Tang, and Dawei Yin. 2019. \u201d Deep reinforcement learning for search, recommendation, and online advertising: a survey\u201d by Xiangyu Zhao, Long Xia, Jiliang Tang, and Dawei Yin with Martin Vesely as coordinator. ACM SIGWEB NewsletterSpring (2019), 1\u201315."},{"key":"e_1_3_2_2_68_1","unstructured":"Xiangyu Zhao Long Xia Dawei Yin and Jiliang Tang. 2019. Model-based reinforcement learning for whole-chain recommendations. arXiv preprint arXiv:1902.03987(2019).  Xiangyu Zhao Long Xia Dawei Yin and Jiliang Tang. 2019. Model-based reinforcement learning for whole-chain recommendations. arXiv preprint arXiv:1902.03987(2019)."},{"key":"e_1_3_2_2_69_1","unstructured":"Xiangyu Zhao Liang Zhang Long Xia Zhuoye Ding Dawei Yin and Jiliang Tang. 2017. Deep reinforcement learning for list-wise recommendations. arXiv preprint arXiv:1801.00209(2017).  Xiangyu Zhao Liang Zhang Long Xia Zhuoye Ding Dawei Yin and Jiliang Tang. 2017. Deep reinforcement learning for list-wise recommendations. arXiv preprint arXiv:1801.00209(2017)."},{"key":"e_1_3_2_2_70_1","volume-title":"DRN: A deep reinforcement learning framework for news recommendation.","author":"Zheng Guanjie","year":"2018","unstructured":"Guanjie Zheng , Fuzheng Zhang , Zihan Zheng , Yang Xiang , Nicholas\u00a0Jing Yuan , Xing Xie , and Zhenhui Li . 2018 . DRN: A deep reinforcement learning framework for news recommendation. (2018), 167\u2013176. Guanjie Zheng, Fuzheng Zhang, Zihan Zheng, Yang Xiang, Nicholas\u00a0Jing Yuan, Xing Xie, and Zhenhui Li. 2018. DRN: A deep reinforcement learning framework for news recommendation. (2018), 167\u2013176."},{"key":"e_1_3_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1000488107"},{"key":"e_1_3_2_2_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/1060745.1060754"}],"event":{"name":"RecSys '21: Fifteenth ACM Conference on Recommender Systems","location":"Amsterdam Netherlands","acronym":"RecSys '21","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGAI ACM Special Interest Group on Artificial Intelligence","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data","SIGIR ACM Special Interest Group on Information Retrieval","SIGCHI ACM Special Interest Group on Computer-Human Interaction","SIGecom Special Interest Group on Economics and Computation"]},"container-title":["Fifteenth ACM Conference on Recommender Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460231.3474236","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3460231.3474236","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:12:17Z","timestamp":1750176737000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460231.3474236"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,13]]},"references-count":72,"alternative-id":["10.1145\/3460231.3474236","10.1145\/3460231"],"URL":"https:\/\/doi.org\/10.1145\/3460231.3474236","relation":{},"subject":[],"published":{"date-parts":[[2021,9,13]]}}}