{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T23:18:52Z","timestamp":1777591132359,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":45,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,4,19]],"date-time":"2021-04-19T00:00:00Z","timestamp":1618790400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,4,19]]},"DOI":"10.1145\/3442381.3449982","type":"proceedings-article","created":{"date-parts":[[2021,6,3]],"date-time":"2021-06-03T19:00:27Z","timestamp":1622746827000},"page":"2291-2303","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["Unifying Offline Causal Inference and Online Bandit Learning for Data Driven Decision"],"prefix":"10.1145","author":[{"given":"Ye","family":"Li","sequence":"first","affiliation":[{"name":"The Chinese University of Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hong","family":"Xie","sequence":"additional","affiliation":[{"name":"Chongqing University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yishi","family":"Lin","sequence":"additional","affiliation":[{"name":"Tencent Inc., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"John C.S.","family":"Lui","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,6,3]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Yasin Abbasi-Yadkori D\u00e1vid P\u00e1l and Csaba Szepesv\u00e1ri. 2011. Improved algorithms for linear stochastic bandits. In Advances in Neural Information Processing Systems. 2312\u20132320.  Yasin Abbasi-Yadkori D\u00e1vid P\u00e1l and Csaba Szepesv\u00e1ri. 2011. Improved algorithms for linear stochastic bandits. In Advances in Neural Information Processing Systems. 2312\u20132320."},{"key":"e_1_3_2_1_2_1","volume-title":"Conference on Learning Theory. 39\u20131.","author":"Agrawal Shipra","year":"2012","unstructured":"Shipra Agrawal and Navin Goyal . 2012 . Analysis of thompson sampling for the multi-armed bandit problem . In Conference on Learning Theory. 39\u20131. Shipra Agrawal and Navin Goyal. 2012. Analysis of thompson sampling for the multi-armed bandit problem. In Conference on Learning Theory. 39\u20131."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1214\/18-AOS1709"},{"key":"e_1_3_2_1_4_1","volume-title":"Finite-time analysis of the multiarmed bandit problem. Machine learning 47, 2-3","author":"Auer Peter","year":"2002","unstructured":"Peter Auer , Nicolo Cesa-Bianchi , and Paul Fischer . 2002. Finite-time analysis of the multiarmed bandit problem. Machine learning 47, 2-3 ( 2002 ), 235\u2013256. Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer. 2002. Finite-time analysis of the multiarmed bandit problem. Machine learning 47, 2-3 (2002), 235\u2013256."},{"key":"e_1_3_2_1_5_1","series-title":"SIAM journal on computing 32, 1","volume-title":"The nonstochastic multiarmed bandit problem","author":"Auer Peter","year":"2002","unstructured":"Peter Auer , Nicolo Cesa-Bianchi , Yoav Freund , and Robert\u00a0 E Schapire . 2002. The nonstochastic multiarmed bandit problem . SIAM journal on computing 32, 1 ( 2002 ), 48\u201377. Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert\u00a0E Schapire. 2002. The nonstochastic multiarmed bandit problem. SIAM journal on computing 32, 1 (2002), 48\u201377."},{"key":"e_1_3_2_1_6_1","volume-title":"An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate behavioral research 46, 3","author":"Austin C","year":"2011","unstructured":"Peter\u00a0 C Austin . 2011. An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate behavioral research 46, 3 ( 2011 ), 399\u2013424. Peter\u00a0C Austin. 2011. An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate behavioral research 46, 3 (2011), 399\u2013424."},{"key":"e_1_3_2_1_7_1","unstructured":"Elias Bareinboim Andrew Forney and Judea Pearl. 2015. Bandits with unobserved confounders: A causal approach. In Advances in Neural Information Processing Systems. 1342\u20131350.  Elias Bareinboim Andrew Forney and Judea Pearl. 2015. Bandits with unobserved confounders: A causal approach. In Advances in Neural Information Processing Systems. 1342\u20131350."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1056\/NEJM200006223422506"},{"key":"e_1_3_2_1_9_1","volume-title":"How much should we trust differences-in-differences estimates?The Quarterly journal of economics 119, 1","author":"Bertrand Marianne","year":"2004","unstructured":"Marianne Bertrand , Esther Duflo , and Sendhil Mullainathan . 2004. How much should we trust differences-in-differences estimates?The Quarterly journal of economics 119, 1 ( 2004 ), 249\u2013275. Marianne Bertrand, Esther Duflo, and Sendhil Mullainathan. 2004. How much should we trust differences-in-differences estimates?The Quarterly journal of economics 119, 1 (2004), 249\u2013275."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/2567709.2567766"},{"key":"e_1_3_2_1_11_1","unstructured":"Jinzhi Bu David Simchi-Levi and Yunzong Xu. 2019. Online pricing with offline data: Phase transition and inverse square law. arXiv preprint arXiv:1910.08693(2019).  Jinzhi Bu David Simchi-Levi and Yunzong Xu. 2019. Online pricing with offline data: Phase transition and inverse square law. arXiv preprint arXiv:1910.08693(2019)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_2_1_13_1","volume-title":"Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. 208\u2013214","author":"Chu Wei","year":"2011","unstructured":"Wei Chu , Lihong Li , Lev Reyzin , and Robert Schapire . 2011 . Contextual bandits with linear payoff functions . In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. 208\u2013214 . Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire. 2011. Contextual bandits with linear payoff functions. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. 208\u2013214."},{"key":"e_1_3_2_1_14_1","unstructured":"Varsha Dani Thomas\u00a0P Hayes and Sham\u00a0M Kakade. 2008. Stochastic linear optimization under bandit feedback. In COLT.  Varsha Dani Thomas\u00a0P Hayes and Sham\u00a0M Kakade. 2008. Stochastic linear optimization under bandit feedback. In COLT."},{"key":"e_1_3_2_1_15_1","unstructured":"Maria Dimakopoulou Susan Athey and Guido Imbens. 2017. Estimation considerations in contextual bandits. arXiv preprint arXiv:1711.07077(2017).  Maria Dimakopoulou Susan Athey and Guido Imbens. 2017. Estimation considerations in contextual bandits. arXiv preprint arXiv:1711.07077(2017)."},{"key":"e_1_3_2_1_16_1","unstructured":"Shi Dong and Benjamin Van\u00a0Roy. 2018. An information-theoretic analysis for Thompson sampling with many actions. In Advances in Neural Information Processing Systems. 4157\u20134165.  Shi Dong and Benjamin Van\u00a0Roy. 2018. An information-theoretic analysis for Thompson sampling with many actions. In Advances in Neural Information Processing Systems. 4157\u20134165."},{"key":"e_1_3_2_1_17_1","unstructured":"Miroslav Dud\u00edk John Langford and Lihong Li. 2011. Doubly robust policy evaluation and learning. arXiv preprint arXiv:1103.4601(2011).  Miroslav Dud\u00edk John Langford and Lihong Li. 2011. Doubly robust policy evaluation and learning. arXiv preprint arXiv:1103.4601(2011)."},{"key":"e_1_3_2_1_18_1","unstructured":"Rapha\u00ebl F\u00e9raud Robin Allesiardo Tanguy Urvoy and Fabrice Cl\u00e9rot. 2016. Random forest for the contextual bandit problem. In Artificial Intelligence and Statistics.  Rapha\u00ebl F\u00e9raud Robin Allesiardo Tanguy Urvoy and Fabrice Cl\u00e9rot. 2016. Random forest for the contextual bandit problem. In Artificial Intelligence and Statistics."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305501"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11749"},{"key":"e_1_3_2_1_21_1","volume-title":"Large sample properties of generalized method of moments estimators. Econometrica: Journal of the Econometric Society","author":"Hansen Lars\u00a0Peter","year":"1982","unstructured":"Lars\u00a0Peter Hansen . 1982. Large sample properties of generalized method of moments estimators. Econometrica: Journal of the Econometric Society ( 1982 ), 1029\u20131054. Lars\u00a0Peter Hansen. 1982. Large sample properties of generalized method of moments estimators. Econometrica: Journal of the Econometric Society (1982), 1029\u20131054."},{"key":"e_1_3_2_1_22_1","volume-title":"The Collected Works of Wassily Hoeffding","author":"Hoeffding Wassily","unstructured":"Wassily Hoeffding . 1994. Probability inequalities for sums of bounded random variables . In The Collected Works of Wassily Hoeffding . Springer , 409\u2013426. Wassily Hoeffding. 1994. Probability inequalities for sums of bounded random variables. In The Collected Works of Wassily Hoeffding. Springer, 409\u2013426."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305556"},{"key":"e_1_3_2_1_24_1","unstructured":"Nathan Kallus. 2018. Balanced policy evaluation and learning. In Advances in Neural Information Processing Systems. 8895\u20138906.  Nathan Kallus. 2018. Balanced policy evaluation and learning. In Advances in Neural Information Processing Systems. 8895\u20138906."},{"key":"e_1_3_2_1_25_1","unstructured":"Volodymyr Kuleshov and Doina Precup. 2014. Algorithms for multi-armed bandit problems. arXiv preprint arXiv:1402.6028(2014).  Volodymyr Kuleshov and Doina Precup. 2014. Algorithms for multi-armed bandit problems. arXiv preprint arXiv:1402.6028(2014)."},{"key":"e_1_3_2_1_26_1","unstructured":"Finnian Lattimore Tor Lattimore and Mark\u00a0D Reid. 2016. Causal bandits: Learning good interventions via causal inference. In Advances in Neural Information Processing Systems. 1181\u20131189.  Finnian Lattimore Tor Lattimore and Mark\u00a0D Reid. 2016. Causal bandits: Learning good interventions via causal inference. In Advances in Neural Information Processing Systems. 1181\u20131189."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/248"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Lihong Li. 2015. Offline evaluation and optimization for interactive systems. (2015).  Lihong Li. 2015. Offline evaluation and optimization for interactive systems. (2015).","DOI":"10.1145\/2684822.2697040"},{"key":"e_1_3_2_1_29_1","volume-title":"Proceedings of the Workshop on On-line Trading of Exploration and Exploitation 2. 19\u201336","author":"Li Lihong","year":"2012","unstructured":"Lihong Li , Wei Chu , John Langford , Taesup Moon , and Xuanhui Wang . 2012 . An unbiased offline evaluation of contextual bandit algorithms with generalized linear models . In Proceedings of the Workshop on On-line Trading of Exploration and Exploitation 2. 19\u201336 . Lihong Li, Wei Chu, John Langford, Taesup Moon, and Xuanhui Wang. 2012. An unbiased offline evaluation of contextual bandit algorithms with generalized linear models. In Proceedings of the Workshop on On-line Trading of Exploration and Exploitation 2. 19\u201336."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772758"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Travis Mandel Yun-En Liu Emma Brunskill and Zoran Popovic. 2015. The Queue Method: Handling Delay Heuristics Prior Data and Evaluation in Bandits.. In AAAI. 2849\u20132856.  Travis Mandel Yun-En Liu Emma Brunskill and Zoran Popovic. 2015. The Queue Method: Handling Delay Heuristics Prior Data and Evaluation in Bandits.. In AAAI. 2849\u20132856.","DOI":"10.1609\/aaai.v29i1.9604"},{"key":"e_1_3_2_1_32_1","volume-title":"Propensity score estimation with boosted regression for evaluating causal effects in observational studies.Psychological methods 9, 4","author":"McCaffrey F","year":"2004","unstructured":"Daniel\u00a0 F McCaffrey , Greg Ridgeway , and Andrew\u00a0 R Morral . 2004. Propensity score estimation with boosted regression for evaluating causal effects in observational studies.Psychological methods 9, 4 ( 2004 ), 403. Daniel\u00a0F McCaffrey, Greg Ridgeway, and Andrew\u00a0R Morral. 2004. Propensity score estimation with boosted regression for evaluating causal effects in observational studies.Psychological methods 9, 4 (2004), 403."},{"key":"e_1_3_2_1_33_1","volume-title":"Causality: models, reasoning and inference. Vol.\u00a029","author":"Pearl Judea","unstructured":"Judea Pearl . 2000. Causality: models, reasoning and inference. Vol.\u00a029 . Springer . Judea Pearl. 2000. Causality: models, reasoning and inference. Vol.\u00a029. Springer."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/70.1.41"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1198\/016214504000001880"},{"key":"e_1_3_2_1_36_1","unstructured":"Pannagadatta Shivaswamy and Thorsten Joachims. 2012. Multi-armed bandit problems with history. In Artificial Intelligence and Statistics. 1046\u20131054.  Pannagadatta Shivaswamy and Thorsten Joachims. 2012. Multi-armed bandit problems with history. In Artificial Intelligence and Statistics. 1046\u20131054."},{"key":"e_1_3_2_1_37_1","unstructured":"Brandon Stewart. 2016. Causality with Measured Confounding. https:\/\/scholar.princeton.edu\/sites\/default\/files\/bstewart\/files\/lecture10handout.pdf  Brandon Stewart. 2016. Causality with Measured Confounding. https:\/\/scholar.princeton.edu\/sites\/default\/files\/bstewart\/files\/lecture10handout.pdf"},{"key":"e_1_3_2_1_38_1","volume-title":"Matching methods for causal inference: A review and a look forward. Statistical science: a review journal of the Institute of Mathematical Statistics 25, 1","author":"Stuart A","year":"2010","unstructured":"Elizabeth\u00a0 A Stuart . 2010. Matching methods for causal inference: A review and a look forward. Statistical science: a review journal of the Institute of Mathematical Statistics 25, 1 ( 2010 ), 1. Elizabeth\u00a0A Stuart. 2010. Matching methods for causal inference: A review and a look forward. Statistical science: a review journal of the Institute of Mathematical Statistics 25, 1 (2010), 1."},{"key":"e_1_3_2_1_39_1","volume-title":"International Conference on Machine Learning. 814\u2013823","author":"Swaminathan Adith","year":"2015","unstructured":"Adith Swaminathan and Thorsten Joachims . 2015 . Counterfactual risk minimization: Learning from logged bandit feedback . In International Conference on Machine Learning. 814\u2013823 . Adith Swaminathan and Thorsten Joachims. 2015. Counterfactual risk minimization: Learning from logged bandit feedback. In International Conference on Machine Learning. 814\u2013823."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.2017.1319839"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2983323.2983847"},{"key":"e_1_3_2_1_42_1","unstructured":"Yixin Wang Dawen Liang Laurent Charlin and David\u00a0M Blei. 2018. The deconfounded recommender: A causal inference approach to recommendation. arXiv preprint arXiv:1808.06581(2018).  Yixin Wang Dawen Liang Laurent Charlin and David\u00a0M Blei. 2018. The deconfounded recommender: A causal inference approach to recommendation. arXiv preprint arXiv:1808.06581(2018)."},{"key":"e_1_3_2_1_43_1","unstructured":"Yahoo. 2020. Yahoo! Front Page Today Module User Click Log Dataset version 1.0 Link:webscope.sandbox.yahoo.com\/catalog.php?datatype=r&did=49.  Yahoo. 2020. Yahoo! Front Page Today Module User Click Log Dataset version 1.0 Link:webscope.sandbox.yahoo.com\/catalog.php?datatype=r&did=49."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"crossref","unstructured":"Li Ye Hong Xie Yishi Lin and John\u00a0C.S. Lui. 2021. Supplementary material Code and Data for \u201dUnifying Offline Causal Inference and Online Bandit Learning for Data Driven Decision Link:https:\/\/github.com\/lonyle\/causal_bandit.  Li Ye Hong Xie Yishi Lin and John\u00a0C.S. Lui. 2021. Supplementary material Code and Data for \u201dUnifying Offline Causal Inference and Online Bandit Learning for Data Driven Decision Link:https:\/\/github.com\/lonyle\/causal_bandit.","DOI":"10.1145\/3442381.3449982"},{"key":"e_1_3_2_1_45_1","volume-title":"Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback. In International Conference on Machine Learning. 7335\u20137344","author":"Zhang Chicheng","year":"2019","unstructured":"Chicheng Zhang , Alekh Agarwal , Hal\u00a0Daum\u00e9 Iii , John Langford , and Sahand Negahban . 2019 . Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback. In International Conference on Machine Learning. 7335\u20137344 . Chicheng Zhang, Alekh Agarwal, Hal\u00a0Daum\u00e9 Iii, John Langford, and Sahand Negahban. 2019. Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback. In International Conference on Machine Learning. 7335\u20137344."}],"event":{"name":"WWW '21: The Web Conference 2021","location":"Ljubljana Slovenia","acronym":"WWW '21","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"]},"container-title":["Proceedings of the Web Conference 2021"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3442381.3449982","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3442381.3449982","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:24:45Z","timestamp":1750195485000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3442381.3449982"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,19]]},"references-count":45,"alternative-id":["10.1145\/3442381.3449982","10.1145\/3442381"],"URL":"https:\/\/doi.org\/10.1145\/3442381.3449982","relation":{},"subject":[],"published":{"date-parts":[[2021,4,19]]},"assertion":[{"value":"2021-06-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}