{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T04:58:59Z","timestamp":1784177939318,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":27,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,9,13]],"date-time":"2021-09-13T00:00:00Z","timestamp":1631491200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,9,13]]},"DOI":"10.1145\/3460231.3474245","type":"proceedings-article","created":{"date-parts":[[2021,9,13]],"date-time":"2021-09-13T21:45:02Z","timestamp":1631569502000},"page":"114-123","source":"Crossref","is-referenced-by-count":21,"title":["Evaluating the Robustness of Off-Policy Evaluation"],"prefix":"10.1145","author":[{"given":"Yuta","family":"Saito","sequence":"first","affiliation":[{"name":"Hanjuku-kaso Co., Ltd., Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Takuma","family":"Udagawa","sequence":"additional","affiliation":[{"name":"Sony Group Corporation, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haruka","family":"Kiyohara","sequence":"additional","affiliation":[{"name":"Tokyo Institute of Technology, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kazuki","family":"Mogi","sequence":"additional","affiliation":[{"name":"Stanford University, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yusuke","family":"Narita","sequence":"additional","affiliation":[{"name":"Yale University, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kei","family":"Tateno","sequence":"additional","affiliation":[{"name":"Sony Group Corporation, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,9,13]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098155"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557040"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1214\/14-STS500"},{"key":"e_1_3_2_2_4_1","volume-title":"Implementation Matters in Deep RL: A Case Study on PPO and TRPO. In International Conference on Learning Representations.","author":"Engstrom Logan","year":"2020","unstructured":"Logan Engstrom , Andrew Ilyas , Shibani Santurkar , Dimitris Tsipras , Firdaus Janoos , Larry Rudolph , and Aleksander Madry . 2020 . Implementation Matters in Deep RL: A Case Study on PPO and TRPO. In International Conference on Learning Representations. Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry. 2020. Implementation Matters in Deep RL: A Case Study on PPO and TRPO. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_5_1","volume-title":"International Conference on Machine Learning, Vol.\u00a080","author":"Farajtabar Mehrdad","year":"2018","unstructured":"Mehrdad Farajtabar , Yinlam Chow , and Mohammad Ghavamzadeh . 2018 . More robust doubly robust off-policy evaluation . In International Conference on Machine Learning, Vol.\u00a080 . PMLR, 1447\u20131456. Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh. 2018. More robust doubly robust off-policy evaluation. In International Conference on Machine Learning, Vol.\u00a080. PMLR, 1447\u20131456."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3159652.3159687"},{"key":"e_1_3_2_2_7_1","volume-title":"International Conference on Machine Learning, Vol.\u00a048","author":"Jiang Nan","year":"2016","unstructured":"Nan Jiang and Lihong Li . 2016 . Doubly robust off-policy value evaluation for reinforcement learning . In International Conference on Machine Learning, Vol.\u00a048 . PMLR, 652\u2013661. Nan Jiang and Lihong Li. 2016. Doubly robust off-policy value evaluation for reinforcement learning. In International Conference on Machine Learning, Vol.\u00a048. PMLR, 652\u2013661."},{"key":"e_1_3_2_2_8_1","volume-title":"Proceedings of the 37th International Conference on Machine Learning. PMLR, 4962\u20134973","author":"Jordan Scott","year":"2020","unstructured":"Scott Jordan , Yash Chandak , Daniel Cohen , Mengxue Zhang , and Philip Thomas . 2020 . Evaluating the performance of reinforcement learning algorithms . In Proceedings of the 37th International Conference on Machine Learning. PMLR, 4962\u20134973 . Scott Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, and Philip Thomas. 2020. Evaluating the performance of reinforcement learning algorithms. In Proceedings of the 37th International Conference on Machine Learning. PMLR, 4962\u20134973."},{"key":"e_1_3_2_2_9_1","volume-title":"NeurIPS 2018 Workshop on Critiquing and Correcting Trends in Machine Learning.","author":"Jordan M","year":"2018","unstructured":"Scott\u00a0 M Jordan , Daniel Cohen , and Philip\u00a0 S Thomas . 2018 . Using cumulative distribution based performance analysis to benchmark models . In NeurIPS 2018 Workshop on Critiquing and Correcting Trends in Machine Learning. Scott\u00a0M Jordan, Daniel Cohen, and Philip\u00a0S Thomas. 2018. Using cumulative distribution based performance analysis to benchmark models. In NeurIPS 2018 Workshop on Critiquing and Correcting Trends in Machine Learning."},{"key":"e_1_3_2_2_10_1","volume-title":"Proceedings of the 38th International Conference on Machine Learning, Vol.\u00a0139","author":"Kallus Nathan","year":"2021","unstructured":"Nathan Kallus , Yuta Saito , and Masatoshi Uehara . 2021 . Optimal Off-Policy Evaluation from Multiple Logging Policies . In Proceedings of the 38th International Conference on Machine Learning, Vol.\u00a0139 . PMLR, 5247\u20135256. Nathan Kallus, Yuta Saito, and Masatoshi Uehara. 2021. Optimal Off-Policy Evaluation from Multiple Logging Policies. In Proceedings of the 38th International Conference on Machine Learning, Vol.\u00a0139. PMLR, 5247\u20135256."},{"key":"e_1_3_2_2_11_1","unstructured":"Nathan Kallus and Masatoshi Uehara. 2019. Intrinsically Efficient Stable and Bounded Off-Policy Evaluation for Reinforcement Learning. In Advances in Neural Information Processing Systems Vol.\u00a032. 3325\u20133334.  Nathan Kallus and Masatoshi Uehara. 2019. Intrinsically Efficient Stable and Bounded Off-Policy Evaluation for Reinforcement Learning. In Advances in Neural Information Processing Systems Vol.\u00a032. 3325\u20133334."},{"key":"e_1_3_2_2_12_1","unstructured":"Masahiro Kato Shota Yasui and Masatoshi Uehara. 2020. Off-Policy Evaluation and Learning for External Validity under a Covariate Shift. In Advances in Neural Information Processing Systems Vol.\u00a033. 49\u201361.  Masahiro Kato Shota Yasui and Masatoshi Uehara. 2020. Off-Policy Evaluation and Learning for External Validity under a Covariate Shift. In Advances in Neural Information Processing Systems Vol.\u00a033. 49\u201361."},{"key":"e_1_3_2_2_13_1","unstructured":"Anqi Liu Hao Liu Anima Anandkumar and Yisong Yue. 2019. Triply Robust Off-Policy Evaluation. arXiv preprint arXiv:1911.05811(2019).  Anqi Liu Hao Liu Anima Anandkumar and Yisong Yue. 2019. Triply Robust Off-Policy Evaluation. arXiv preprint arXiv:1911.05811(2019)."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33014634"},{"key":"e_1_3_2_2_15_1","unstructured":"Yusuke Narita Shota Yasui and Kohei Yata. 2020. Off-policy Bandit and Reinforcement Learning. arXiv preprint arXiv:2002.08536(2020).  Yusuke Narita Shota Yasui and Kohei Yata. 2020. Off-policy Bandit and Reinforcement Learning. arXiv preprint arXiv:2002.08536(2020)."},{"key":"e_1_3_2_2_16_1","volume-title":"Eligibility traces for off-policy policy evaluation","author":"Precup Doina","year":"2000","unstructured":"Doina Precup . 2000. Eligibility traces for off-policy policy evaluation . Computer Science Department Faculty Publication Series ( 2000 ), 80. Doina Precup. 2000. Eligibility traces for off-policy policy evaluation. Computer Science Department Faculty Publication Series (2000), 80."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3383313.3412262"},{"key":"e_1_3_2_2_18_1","unstructured":"Yuta Saito Shunsuke Aihara Megumi Matsutani and Yusuke Narita. 2020. Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation. arXiv preprint arXiv:2008.07146(2020).  Yuta Saito Shunsuke Aihara Megumi Matsutani and Yusuke Narita. 2020. Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation. arXiv preprint arXiv:2008.07146(2020)."},{"key":"e_1_3_2_2_19_1","first-page":"2217","article-title":"Learning from Logged Implicit Exploration Data, In Advances in Neural Information Processing Systems","volume":"23","author":"Strehl Alex","year":"2010","unstructured":"Alex Strehl , John Langford , Lihong Li , and Sham\u00a0 M Kakade . 2010 . Learning from Logged Implicit Exploration Data, In Advances in Neural Information Processing Systems . Advances in Neural Information Processing Systems 23 , 2217 \u2013 2225 . Alex Strehl, John Langford, Lihong Li, and Sham\u00a0M Kakade. 2010. Learning from Logged Implicit Exploration Data, In Advances in Neural Information Processing Systems. Advances in Neural Information Processing Systems 23, 2217\u20132225.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_20_1","volume-title":"International Conference on Machine Learning, Vol.\u00a0119","author":"Su Yi","year":"2020","unstructured":"Yi Su , Maria Dimakopoulou , Akshay Krishnamurthy , and Miroslav Dud\u00edk . 2020 . Doubly robust off-policy evaluation with shrinkage . In International Conference on Machine Learning, Vol.\u00a0119 . PMLR, 9167\u20139176. Yi Su, Maria Dimakopoulou, Akshay Krishnamurthy, and Miroslav Dud\u00edk. 2020. Doubly robust off-policy evaluation with shrinkage. In International Conference on Machine Learning, Vol.\u00a0119. PMLR, 9167\u20139176."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/3524938.3525791"},{"key":"e_1_3_2_2_22_1","volume-title":"International Conference on Machine Learning, Vol.\u00a097","author":"Su Yi","year":"2019","unstructured":"Yi Su , Lequn Wang , Michele Santacatterina , and Thorsten Joachims . 2019 . Cab: Continuous adaptive blending for policy evaluation and learning . In International Conference on Machine Learning, Vol.\u00a097 . PMLR, 6005\u20136014. Yi Su, Lequn Wang, Michele Santacatterina, and Thorsten Joachims. 2019. Cab: Continuous adaptive blending for policy evaluation and learning. In International Conference on Machine Learning, Vol.\u00a097. PMLR, 6005\u20136014."},{"key":"e_1_3_2_2_23_1","unstructured":"Adith Swaminathan and Thorsten Joachims. 2015. The self-normalized estimator for counterfactual learning. In Advances in Neural Information Processing Systems Vol.\u00a028. 3231\u20133239.  Adith Swaminathan and Thorsten Joachims. 2015. The self-normalized estimator for counterfactual learning. In Advances in Neural Information Processing Systems Vol.\u00a028. 3231\u20133239."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/25.3-4.285"},{"key":"e_1_3_2_2_25_1","volume-title":"International Conference on Machine Learning, Vol.\u00a097","author":"Vlassis Nikos","year":"2019","unstructured":"Nikos Vlassis , Aurelien Bibaut , Maria Dimakopoulou , and Tony Jebara . 2019 . On the design of estimators for bandit off-policy evaluation . In International Conference on Machine Learning, Vol.\u00a097 . PMLR, 6468\u20136476. Nikos Vlassis, Aurelien Bibaut, Maria Dimakopoulou, and Tony Jebara. 2019. On the design of estimators for bandit off-policy evaluation. In International Conference on Machine Learning, Vol.\u00a097. PMLR, 6468\u20136476."},{"key":"e_1_3_2_2_26_1","unstructured":"Cameron Voloshin Hoang\u00a0M Le Nan Jiang and Yisong Yue. 2019. Empirical study of off-policy policy evaluation for reinforcement learning. arXiv preprint arXiv:1911.06854(2019).  Cameron Voloshin Hoang\u00a0M Le Nan Jiang and Yisong Yue. 2019. Empirical study of off-policy policy evaluation for reinforcement learning. arXiv preprint arXiv:1911.06854(2019)."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305890.3306052"}],"event":{"name":"RecSys '21: Fifteenth ACM Conference on Recommender Systems","location":"Amsterdam Netherlands","acronym":"RecSys '21","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGAI ACM Special Interest Group on Artificial Intelligence","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data","SIGIR ACM Special Interest Group on Information Retrieval","SIGCHI ACM Special Interest Group on Computer-Human Interaction","SIGecom Special Interest Group on Economics and Computation"]},"container-title":["Fifteenth ACM Conference on Recommender Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460231.3474245","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3460231.3474245","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:12:17Z","timestamp":1750191137000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460231.3474245"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,13]]},"references-count":27,"alternative-id":["10.1145\/3460231.3474245","10.1145\/3460231"],"URL":"https:\/\/doi.org\/10.1145\/3460231.3474245","relation":{},"subject":[],"published":{"date-parts":[[2021,9,13]]}}}