{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:42:33Z","timestamp":1787017353539,"version":"build-2736575974"},"publisher-location":"New York, NY, USA","reference-count":24,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,9,13]],"date-time":"2021-09-13T00:00:00Z","timestamp":1631491200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,9,13]]},"DOI":"10.1145\/3460231.3474231","type":"proceedings-article","created":{"date-parts":[[2021,9,13]],"date-time":"2021-09-13T17:45:02Z","timestamp":1631555102000},"page":"372-379","source":"Crossref","is-referenced-by-count":8,"title":["Debiased Off-Policy Evaluation for Recommendation Systems"],"prefix":"10.1145","author":[{"given":"Yusuke","family":"Narita","sequence":"first","affiliation":[{"name":"Yale University, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shota","family":"Yasui","sequence":"additional","affiliation":[{"name":"AILab CyberAgent, Inc., Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kohei","family":"Yata","sequence":"additional","affiliation":[{"name":"Department of Economics Yale University, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,9,13]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"Greg Brockman Vicki Cheung Ludwig Pettersson Jonas Schneider John Schulman Jie Tang and Wojciech Zaremba. 2016. OpenAI Gym. arXiv preprint arXiv:1606.01540(2016).  Greg Brockman Vicki Cheung Ludwig Pettersson Jonas Schneider John Schulman Jie Tang and Wojciech Zaremba. 2016. OpenAI Gym. arXiv preprint arXiv:1606.01540(2016)."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1111\/ectj.12097"},{"key":"e_1_3_2_2_3_1","volume-title":"Locally Robust Semiparametric Estimation. Arxiv","author":"Chernozhukov Victor","year":"2018","unstructured":"Victor Chernozhukov , Juan\u00a0Carlos Escanciano , Hidehiko Ichimura , Whitney\u00a0 K. Newey , and James\u00a0 M. Robins . 2018. Locally Robust Semiparametric Estimation. Arxiv ( 2018 ). Victor Chernozhukov, Juan\u00a0Carlos Escanciano, Hidehiko Ichimura, Whitney\u00a0K. Newey, and James\u00a0M. Robins. 2018. Locally Robust Semiparametric Estimation. Arxiv (2018)."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1214\/14-STS500"},{"key":"e_1_3_2_2_5_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning. 1447\u20131456","author":"Farajtabar Mehrdad","year":"2018","unstructured":"Mehrdad Farajtabar , Yinlam Chow , and Mohammad Ghavamzadeh . 2018 . More Robust Doubly Robust Off-policy Evaluation . In Proceedings of the 35th International Conference on Machine Learning. 1447\u20131456 . Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh. 2018. More Robust Doubly Robust Off-policy Evaluation. In Proceedings of the 35th International Conference on Machine Learning. 1447\u20131456."},{"key":"e_1_3_2_2_6_1","volume-title":"Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining. WSDM, 198\u2013\u2013206","author":"Gilotte Alexandre","year":"2018","unstructured":"Alexandre Gilotte , Cl\u00e9ment Calauz\u00e8nes , Thomas Nedelec , Alexandre Abraham , and Simon Doll\u00e9 . 2018 . Offline A\/B Testing for Recommender Systems , In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining. WSDM, 198\u2013\u2013206 . Alexandre Gilotte, Cl\u00e9ment Calauz\u00e8nes, Thomas Nedelec, Alexandre Abraham, and Simon Doll\u00e9. 2018. Offline A\/B Testing for Recommender Systems, In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining. WSDM, 198\u2013\u2013206."},{"key":"e_1_3_2_2_7_1","unstructured":"Alex Irpan Kanishka Rao Konstantinos Bousmalis Chris Harris Julian Ibarz and Sergey Levine. 2019. Off-Policy Evaluation via Off-Policy Classification. In Advances in Neural Information Processing Systems 32.  Alex Irpan Kanishka Rao Konstantinos Bousmalis Chris Harris Julian Ibarz and Sergey Levine. 2019. Off-Policy Evaluation via Off-Policy Classification. In Advances in Neural Information Processing Systems 32."},{"key":"e_1_3_2_2_8_1","volume-title":"Proceedings of the 33rd International Conference on Machine Learning. 652\u2013661","author":"Jiang Nan","year":"2016","unstructured":"Nan Jiang and Lihong Li . 2016 . Doubly Robust Off-policy Value Evaluation for Reinforcement Learning . In Proceedings of the 33rd International Conference on Machine Learning. 652\u2013661 . Nan Jiang and Lihong Li. 2016. Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. In Proceedings of the 33rd International Conference on Machine Learning. 652\u2013661."},{"key":"e_1_3_2_2_9_1","volume-title":"Proceedings of the 37th International Conference on Machine Learning. ICML.","author":"Kallus Nathan","year":"2020","unstructured":"Nathan Kallus and Masatoshi Uehara . 2020 . Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes , In Proceedings of the 37th International Conference on Machine Learning. ICML. Nathan Kallus and Masatoshi Uehara. 2020. Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes, In Proceedings of the 37th International Conference on Machine Learning. ICML."},{"key":"e_1_3_2_2_10_1","volume-title":"Lightgbm: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems 30. 3146\u20133154.","author":"Ke Guolin","year":"2017","unstructured":"Guolin Ke , Qi Meng , Thomas Finley , Taifeng Wang , Wei Chen , Weidong Ma , Qiwei Ye , and Tie-Yan Liu . 2017 . Lightgbm: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems 30. 3146\u20133154. Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017. Lightgbm: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems 30. 3146\u20133154."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772758"},{"key":"e_1_3_2_2_12_1","volume-title":"Proceedings of the Fourth ACM International Conference on Web Search and Data Mining. WSDM, 297\u2013306","author":"Li Lihong","year":"2011","unstructured":"Lihong Li , Wei Chu , John Langford , and Xuanhui Wang . 2011 . Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms , In Proceedings of the Fourth ACM International Conference on Web Search and Data Mining. WSDM, 297\u2013306 . Lihong Li, Wei Chu, John Langford, and Xuanhui Wang. 2011. Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms, In Proceedings of the Fourth ACM International Conference on Web Search and Data Mining. WSDM, 297\u2013306."},{"key":"e_1_3_2_2_13_1","unstructured":"Yao Liu Omer Gottesman Aniruddh Raghu Matthieu Komorowski Aldo\u00a0A Faisal Finale Doshi-Velez and Emma Brunskill. 2018. Representation balancing mdps for off-policy policy evaluation. In Advances in Neural Information Processing Systems 31. 2644\u20132653.  Yao Liu Omer Gottesman Aniruddh Raghu Matthieu Komorowski Aldo\u00a0A Faisal Finale Doshi-Velez and Emma Brunskill. 2018. Representation balancing mdps for off-policy policy evaluation. In Advances in Neural Information Processing Systems 31. 2644\u20132653."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33014634"},{"key":"e_1_3_2_2_15_1","volume-title":"Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation. Arxiv","author":"Newey K.","year":"2018","unstructured":"Whitney\u00a0 K. Newey and James\u00a0 M. Robins . 2018. Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation. Arxiv ( 2018 ). Whitney\u00a0K. Newey and James\u00a0M. Robins. 2018. Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation. Arxiv (2018)."},{"key":"e_1_3_2_2_16_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas Kopf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library , In Advances in Neural Information Processing Systems 32. NIPS, 8024\u20138035. Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library, In Advances in Neural Information Processing Systems 32. NIPS, 8024\u20138035."},{"key":"e_1_3_2_2_17_1","volume-title":"Proceedings of the 17th International Conference on Machine Learning. ICML, 759\u2013766","author":"Precup Doina","year":"2000","unstructured":"Doina Precup , Richard\u00a0 S. Sutton , and Satinder Singh . 2000 . Eligibility Traces for Off-Policy Policy Evaluation , In Proceedings of the 17th International Conference on Machine Learning. ICML, 759\u2013766 . Doina Precup, Richard\u00a0S. Sutton, and Satinder Singh. 2000. Eligibility Traces for Off-Policy Policy Evaluation, In Proceedings of the 17th International Conference on Machine Learning. ICML, 759\u2013766."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/82.4.805"},{"key":"e_1_3_2_2_19_1","unstructured":"Alex Strehl John Langford Lihong Li and Sham\u00a0M Kakade. 2010. Learning from Logged Implicit Exploration Data In Advances in Neural Information Processing Systems 23. NIPS 2217\u20132225.  Alex Strehl John Langford Lihong Li and Sham\u00a0M Kakade. 2010. Learning from Logged Implicit Exploration Data In Advances in Neural Information Processing Systems 23. NIPS 2217\u20132225."},{"key":"e_1_3_2_2_20_1","unstructured":"Adith Swaminathan Akshay Krishnamurthy Alekh Agarwal Miro Dudik John Langford Damien Jose and Imed Zitouni. 2017. Off-policy Evaluation for Slate Recommendation In Advances in Neural Information Processing Systems 30. NIPS 3635\u20133645.  Adith Swaminathan Akshay Krishnamurthy Alekh Agarwal Miro Dudik John Langford Damien Jose and Imed Zitouni. 2017. Off-policy Evaluation for Slate Recommendation In Advances in Neural Information Processing Systems 30. NIPS 3635\u20133645."},{"key":"e_1_3_2_2_21_1","volume-title":"Proceedings of the 33rd International Conference on Machine Learning. 2139\u20132148","author":"Thomas Philip","year":"2016","unstructured":"Philip Thomas and Emma Brunskill . 2016 . Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning . In Proceedings of the 33rd International Conference on Machine Learning. 2139\u20132148 . Philip Thomas and Emma Brunskill. 2016. Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning. In Proceedings of the 33rd International Conference on Machine Learning. 2139\u20132148."},{"key":"e_1_3_2_2_22_1","volume-title":"Proceedings of the 33rd International Conference on Machine Learning. ICML, 2139\u20132148","author":"Thomas Philip","year":"2016","unstructured":"Philip Thomas and Emma Brunskill . 2016 . Data-efficient Off-policy Policy Evaluation for Reinforcement Learning , In Proceedings of the 33rd International Conference on Machine Learning. ICML, 2139\u20132148 . Philip Thomas and Emma Brunskill. 2016. Data-efficient Off-policy Policy Evaluation for Reinforcement Learning, In Proceedings of the 33rd International Conference on Machine Learning. ICML, 2139\u20132148."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/3524938.3525833"},{"key":"e_1_3_2_2_24_1","volume-title":"Proceedings of the 34th International Conference on Machine Learning. ICML, 3589\u20133597","author":"Wang Yu-Xiang","year":"2017","unstructured":"Yu-Xiang Wang , Alekh Agarwal , and Miroslav Dudik . 2017 . Optimal and Adaptive Off-policy Evaluation in Contextual Bandits , In Proceedings of the 34th International Conference on Machine Learning. ICML, 3589\u20133597 . Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik. 2017. Optimal and Adaptive Off-policy Evaluation in Contextual Bandits, In Proceedings of the 34th International Conference on Machine Learning. ICML, 3589\u20133597."}],"event":{"name":"RecSys '21: Fifteenth ACM Conference on Recommender Systems","location":"Amsterdam Netherlands","acronym":"RecSys '21","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGAI ACM Special Interest Group on Artificial Intelligence","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data","SIGIR ACM Special Interest Group on Information Retrieval","SIGCHI ACM Special Interest Group on Computer-Human Interaction","SIGecom Special Interest Group on Economics and Computation"]},"container-title":["Fifteenth ACM Conference on Recommender Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460231.3474231","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3460231.3474231","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:12:17Z","timestamp":1750176737000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460231.3474231"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,13]]},"references-count":24,"alternative-id":["10.1145\/3460231.3474231","10.1145\/3460231"],"URL":"https:\/\/doi.org\/10.1145\/3460231.3474231","relation":{},"subject":[],"published":{"date-parts":[[2021,9,13]]}}}