{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:42:32Z","timestamp":1787017352650,"version":"build-2736575974"},"publisher-location":"New York, NY, USA","reference-count":28,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,9,13]],"date-time":"2021-09-13T00:00:00Z","timestamp":1631491200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,9,13]]},"DOI":"10.1145\/3460231.3474601","type":"proceedings-article","created":{"date-parts":[[2021,9,13]],"date-time":"2021-09-13T17:45:04Z","timestamp":1631555104000},"page":"551-553","source":"Crossref","is-referenced-by-count":26,"title":["Exploration in Recommender Systems"],"prefix":"10.1145","author":[{"given":"Minmin","family":"Chen","sequence":"first","affiliation":[{"name":"Google, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,9,13]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Conference on learning theory. JMLR Workshop and Conference Proceedings, 39\u20131.","author":"Agrawal Shipra","year":"2012","unstructured":"Shipra Agrawal and Navin Goyal . 2012 . Analysis of thompson sampling for the multi-armed bandit problem . In Conference on learning theory. JMLR Workshop and Conference Proceedings, 39\u20131. Shipra Agrawal and Navin Goyal. 2012. Analysis of thompson sampling for the multi-armed bandit problem. In Conference on learning theory. JMLR Workshop and Conference Proceedings, 39\u20131."},{"key":"e_1_3_2_2_2_1","unstructured":"Marc\u00a0G Bellemare Sriram Srinivasan Georg Ostrovski Tom Schaul David Saxton and Remi Munos. 2016. Unifying count-based exploration and intrinsic motivation. arXiv preprint arXiv:1606.01868(2016).  Marc\u00a0G Bellemare Sriram Srinivasan Georg Ostrovski Tom Schaul David Saxton and Remi Munos. 2016. Unifying count-based exploration and intrinsic motivation. arXiv preprint arXiv:1606.01868(2016)."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454008.1454038"},{"key":"e_1_3_2_2_4_1","unstructured":"Olivier Chapelle and Lihong Li. 2011. An empirical evaluation of thompson sampling. In Advances in neural information processing systems. 2249\u20132257.  Olivier Chapelle and Lihong Li. 2011. An empirical evaluation of thompson sampling. In Advances in neural information processing systems. 2249\u20132257."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290999"},{"key":"e_1_3_2_2_6_1","volume-title":"Multi-armed bandit allocation indices","author":"Gittins John","unstructured":"John Gittins , Kevin Glazebrook , and Richard Weber . 2011. Multi-armed bandit allocation indices . John Wiley & Sons . John Gittins, Kevin Glazebrook, and Richard Weber. 2011. Multi-armed bandit allocation indices. John Wiley & Sons."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/963770.963772"},{"key":"e_1_3_2_2_8_1","volume-title":"Vime: Variational information maximizing exploration. arXiv preprint arXiv:1605.09674(2016).","author":"Houthooft Rein","year":"2016","unstructured":"Rein Houthooft , Xi Chen , Yan Duan , John Schulman , Filip De\u00a0Turck , and Pieter Abbeel . 2016 . Vime: Variational information maximizing exploration. arXiv preprint arXiv:1605.09674(2016). Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De\u00a0Turck, and Pieter Abbeel. 2016. Vime: Variational information maximizing exploration. arXiv preprint arXiv:1605.09674(2016)."},{"key":"e_1_3_2_2_9_1","unstructured":"Eugene Ie Vihan Jain Jing Wang Sanmit Narvekar Ritesh Agarwal Rui Wu Heng-Tze Cheng Tushar Chandra and Craig Boutilier. 2019. SlateQ: A tractable decomposition for reinforcement learning with recommendation sets. (2019).  Eugene Ie Vihan Jain Jing Wang Sanmit Narvekar Ritesh Agarwal Rui Wu Heng-Tze Cheng Tushar Chandra and Craig Boutilier. 2019. SlateQ: A tractable decomposition for reinforcement learning with recommendation sets. (2019)."},{"key":"e_1_3_2_2_10_1","volume-title":"Bayesian surprise attracts human attention. Vision research 49, 10","author":"Itti Laurent","year":"2009","unstructured":"Laurent Itti and Pierre Baldi . 2009. Bayesian surprise attracts human attention. Vision research 49, 10 ( 2009 ), 1295\u20131306. Laurent Itti and Pierre Baldi. 2009. Bayesian surprise attracts human attention. Vision research 49, 10 (2009), 1295\u20131306."},{"key":"e_1_3_2_2_11_1","volume-title":"Asymptotically efficient adaptive allocation rules. Advances in applied mathematics 6, 1","author":"Lai Tze\u00a0Leung","year":"1985","unstructured":"Tze\u00a0Leung Lai and Herbert Robbins . 1985. Asymptotically efficient adaptive allocation rules. Advances in applied mathematics 6, 1 ( 1985 ), 4\u201322. Tze\u00a0Leung Lai and Herbert Robbins. 1985. Asymptotically efficient adaptive allocation rules. Advances in applied mathematics 6, 1 (1985), 4\u201322."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772758"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"crossref","unstructured":"Sean\u00a0M McNee John Riedl and Joseph\u00a0A Konstan. 2006. Being accurate is not enough: how accuracy metrics have hurt recommender systems. In CHI\u201906 extended abstracts on Human factors in computing systems. 1097\u20131101.  Sean\u00a0M McNee John Riedl and Joseph\u00a0A Konstan. 2006. Being accurate is not enough: how accuracy metrics have hurt recommender systems. In CHI\u201906 extended abstracts on Human factors in computing systems. 1097\u20131101.","DOI":"10.1145\/1125451.1125659"},{"key":"e_1_3_2_2_14_1","unstructured":"Ian Osband Charles Blundell Alexander Pritzel and Benjamin Van\u00a0Roy. 2016. Deep exploration via bootstrapped DQN. arXiv preprint arXiv:1602.04621(2016).  Ian Osband Charles Blundell Alexander Pritzel and Benjamin Van\u00a0Roy. 2016. Deep exploration via bootstrapped DQN. arXiv preprint arXiv:1602.04621(2016)."},{"key":"e_1_3_2_2_15_1","volume-title":"International conference on machine learning. PMLR, 2721\u20132730","author":"Ostrovski Georg","year":"2017","unstructured":"Georg Ostrovski , Marc\u00a0 G Bellemare , A\u00e4ron Oord , and R\u00e9mi Munos . 2017 . Count-based exploration with neural density models . In International conference on machine learning. PMLR, 2721\u20132730 . Georg Ostrovski, Marc\u00a0G Bellemare, A\u00e4ron Oord, and R\u00e9mi Munos. 2017. Count-based exploration with neural density models. In International conference on machine learning. PMLR, 2721\u20132730."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2017.70"},{"key":"e_1_3_2_2_17_1","unstructured":"Carlos Riquelme George Tucker and Jasper Snoek. 2018. Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling. arXiv preprint arXiv:1802.09127(2018).  Carlos Riquelme George Tucker and Jasper Snoek. 2018. Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling. arXiv preprint arXiv:1802.09127(2018)."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.1991.170605"},{"key":"e_1_3_2_2_19_1","unstructured":"Bradly\u00a0C Stadie Sergey Levine and Pieter Abbeel. 2015. Incentivizing exploration in reinforcement learning with deep predictive models. arXiv preprint arXiv:1507.00814(2015).  Bradly\u00a0C Stadie Sergey Levine and Pieter Abbeel. 2015. Incentivizing exploration in reinforcement learning with deep predictive models. arXiv preprint arXiv:1507.00814(2015)."},{"key":"e_1_3_2_2_20_1","volume-title":"Proceedings of the international conference on artificial neural networks","author":"Storck Jan","year":"1995","unstructured":"Jan Storck , Sepp Hochreiter , and J\u00fcrgen Schmidhuber . 1995 . Reinforcement driven information acquisition in non-deterministic environments . In Proceedings of the international conference on artificial neural networks , Paris, Vol.\u00a02. Citeseer, 159\u2013164. Jan Storck, Sepp Hochreiter, and J\u00fcrgen Schmidhuber. 1995. Reinforcement driven information acquisition in non-deterministic environments. In Proceedings of the international conference on artificial neural networks, Paris, Vol.\u00a02. Citeseer, 159\u2013164."},{"key":"e_1_3_2_2_21_1","volume-title":"Reinforcement learning: An introduction","author":"Sutton S","unstructured":"Richard\u00a0 S Sutton , Andrew\u00a0 G Barto , 1998. Reinforcement learning: An introduction . MIT press . Richard\u00a0S Sutton, Andrew\u00a0G Barto, 1998. Reinforcement learning: An introduction. MIT press."},{"key":"e_1_3_2_2_22_1","volume-title":"31st Conference on Neural Information Processing Systems (NIPS), Vol.\u00a030","author":"Tang Haoran","year":"2017","unstructured":"Haoran Tang , Rein Houthooft , Davis Foote , Adam Stooke , Xi Chen , Yan Duan , John Schulman , Filip De\u00a0Turck , and Pieter Abbeel . 2017 . # exploration: A study of count-based exploration for deep reinforcement learning . In 31st Conference on Neural Information Processing Systems (NIPS), Vol.\u00a030 . 1\u201318. Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De\u00a0Turck, and Pieter Abbeel. 2017. # exploration: A study of count-based exploration for deep reinforcement learning. In 31st Conference on Neural Information Processing Systems (NIPS), Vol.\u00a030. 1\u201318."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/25.3-4.285"},{"key":"e_1_3_2_2_24_1","unstructured":"Tom Zahavy and Shie Mannor. 2019. Neural Linear Bandits: Overcoming Catastrophic Forgetting through Likelihood Matching. (2019).  Tom Zahavy and Shie Mannor. 2019. Neural Linear Bandits: Overcoming Catastrophic Forgetting through Likelihood Matching. (2019)."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2124295.2124300"},{"key":"e_1_3_2_2_26_1","volume-title":"Deep reinforcement learning for search, recommendation, and online advertising: a survey","author":"Zhao Xiangyu","year":"2019","unstructured":"Xiangyu Zhao , Long Xia , Jiliang Tang , and Dawei Yin . 2019. \u201d Deep reinforcement learning for search, recommendation, and online advertising: a survey \u201d by Xiangyu Zhao, Long Xia , Jiliang Tang, and Dawei Yin with Martin Vesely as coordinator. ACM SIGWEB NewsletterSpring ( 2019 ), 1\u201315. Xiangyu Zhao, Long Xia, Jiliang Tang, and Dawei Yin. 2019. \u201d Deep reinforcement learning for search, recommendation, and online advertising: a survey\u201d by Xiangyu Zhao, Long Xia, Jiliang Tang, and Dawei Yin with Martin Vesely as coordinator. ACM SIGWEB NewsletterSpring (2019), 1\u201315."},{"key":"e_1_3_2_2_27_1","unstructured":"Xiangyu Zhao Long Xia Dawei Yin and Jiliang Tang. 2019. Model-based reinforcement learning for whole-chain recommendations. arXiv preprint arXiv:1902.03987(2019).  Xiangyu Zhao Long Xia Dawei Yin and Jiliang Tang. 2019. Model-based reinforcement learning for whole-chain recommendations. arXiv preprint arXiv:1902.03987(2019)."},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1060745.1060754"}],"event":{"name":"RecSys '21: Fifteenth ACM Conference on Recommender Systems","location":"Amsterdam Netherlands","acronym":"RecSys '21","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGAI ACM Special Interest Group on Artificial Intelligence","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data","SIGIR ACM Special Interest Group on Information Retrieval","SIGCHI ACM Special Interest Group on Computer-Human Interaction","SIGecom Special Interest Group on Economics and Computation"]},"container-title":["Fifteenth ACM Conference on Recommender Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460231.3474601","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3460231.3474601","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:12:17Z","timestamp":1750176737000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460231.3474601"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,13]]},"references-count":28,"alternative-id":["10.1145\/3460231.3474601","10.1145\/3460231"],"URL":"https:\/\/doi.org\/10.1145\/3460231.3474601","relation":{},"subject":[],"published":{"date-parts":[[2021,9,13]]}}}