{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T13:28:47Z","timestamp":1787318927340,"version":"build-2736575974"},"reference-count":31,"publisher":"National Academy of Sciences","issue":"27","license":[{"start":{"date-parts":[[2017,1,5]],"date-time":"2017-01-05T00:00:00Z","timestamp":1483574400000},"content-version":"vor","delay-in-days":184,"URL":"http:\/\/www.pnas.org\/preview_site\/misc\/userlicense.xhtml"}],"funder":[{"DOI":"10.13039\/100006112","name":"Microsoft Research","doi-asserted-by":"publisher","award":["0"],"award-info":[{"award-number":["0"]}],"id":[{"id":"10.13039\/100006112","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["www.pnas.org"],"crossmark-restriction":true},"short-container-title":["Proc. Natl. Acad. Sci. U.S.A."],"published-print":{"date-parts":[[2016,7,5]]},"abstract":"<jats:p>In this paper we propose methods for estimating heterogeneity in causal effects in experimental and observational studies and for conducting hypothesis tests about the magnitude of differences in treatment effects across subsets of the population. We provide a data-driven approach to partition the data into subpopulations that differ in the magnitude of their treatment effects. The approach enables the construction of valid confidence intervals for treatment effects, even with many covariates relative to the sample size, and without \u201csparsity\u201d assumptions. We propose an \u201chonest\u201d approach to estimation, whereby one sample is used to construct the partition and another to estimate treatment effects for each subpopulation. Our approach builds on regression tree methods, modified to optimize for goodness of fit in treatment effects and to account for honest estimation. Our model selection criterion anticipates that bias will be eliminated by honest estimation and also accounts for the effect of making additional splits on the variance of treatment effect estimates within each subpopulation. We address the challenge that the \u201cground truth\u201d for a causal effect is not observed for any individual unit, so that standard approaches to cross-validation must be modified. Through a simulation study, we show that for our preferred method honest estimation results in nominal coverage for 90% confidence intervals, whereas coverage ranges between 74% and 84% for nonhonest approaches. Honest estimation requires estimating the model with a smaller sample size; the cost in terms of mean squared error of treatment effects for our preferred method ranges between 7\u201322%.<\/jats:p>","DOI":"10.1073\/pnas.1510489113","type":"journal-article","created":{"date-parts":[[2016,7,5]],"date-time":"2016-07-05T14:11:32Z","timestamp":1467727892000},"page":"7353-7360","update-policy":"https:\/\/doi.org\/10.1073\/pnas.cm10313","source":"Crossref","is-referenced-by-count":1209,"title":["Recursive partitioning for heterogeneous causal effects"],"prefix":"10.1073","volume":"113","author":[{"given":"Susan","family":"Athey","sequence":"first","affiliation":[{"name":"Stanford Graduate School of Business, Stanford University, Stanford, CA 94305"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guido","family":"Imbens","sequence":"additional","affiliation":[{"name":"Stanford Graduate School of Business, Stanford University, Stanford, CA 94305"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"341","published-online":{"date-parts":[[2016,7,5]]},"reference":[{"key":"e_1_3_3_1_2","first-page":"688","article-title":"Estimating causal effects of treatments in randomized and non-randomized studies","volume":"66","author":"Rubin D","year":"1974","unstructured":"D Rubin, Estimating causal effects of treatments in randomized and non-randomized studies. Educ Psychol 66, 688\u2013701 (1974).","journal-title":"Educ Psychol"},{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1986.10478354"},{"key":"e_1_3_3_3_2","doi-asserted-by":"crossref","first-page":"159","DOI":"10.1017\/CBO9781139025751","volume-title":"Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction","author":"Imbens G","year":"2015","unstructured":"G Imbens, D Rubin Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction (Cambridge Univ Press, Cambridge, UK), pp. 159 (2015)."},{"key":"e_1_3_3_4_2","volume-title":"The Elements of Statistical Learning: Data Mining, Inference, and Prediction","author":"Hastie T","year":"2011","unstructured":"T Hastie, R Tibshirani, J Friedman The Elements of Statistical Learning: Data Mining, Inference, and Prediction (Springer, 2nd Ed, New York, 2011).","edition":"2"},{"key":"e_1_3_3_5_2","volume-title":"Classification and Regression Trees","author":"Breiman L","year":"1984","unstructured":"L Breiman, J Friedman, R Olshen, C Stone Classification and Regression Trees (Wadsworth, Belmont, CA, 1984)."},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1010933404324"},{"key":"e_1_3_3_7_2","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","article-title":"Regression shrinkage and selection via the lasso","volume":"58","author":"Tibshirani R","year":"1996","unstructured":"R Tibshirani, Regression shrinkage and selection via the lasso. J R Stat Soc B 58, 267\u2013288 (1996).","journal-title":"J R Stat Soc B"},{"key":"e_1_3_3_8_2","volume-title":"Statistical Learning Theory","author":"Vapnik V","year":"1998","unstructured":"V Vapnik Statistical Learning Theory (Wiley, New York, 1998)."},{"key":"e_1_3_3_9_2","unstructured":"S Wager S Athey Estimation and inference of heterogeneous treatment effects using random forests. Available at arxiv.org\/abs\/1510.04342. (2015)."},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1214\/aos\/1176344064"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/70.1.41"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1111\/j.1468-0262.2006.00655.x"},{"key":"e_1_3_3_13_2","volume-title":"Causality: Models, Reasoning and Inference","author":"Pearl J","year":"2000","unstructured":"J Pearl Causality: Models, Reasoning and Inference (Cambridge Univ Press, Cambridge, UK, 2000)."},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-3692-2"},{"key":"e_1_3_3_15_2","doi-asserted-by":"crossref","unstructured":"A Beygelzimer J Langford The offset tree for learning with partial labels. arxiv:0812.4044. (2009).","DOI":"10.1145\/1557019.1557040"},{"key":"e_1_3_3_16_2","unstructured":"M Dudik J Langford L Li Doubly robust policy evaluation and learning. Proceedings of the 28th International Conference on Machine Learning (International Machine Learning Society). (2011)."},{"key":"e_1_3_3_17_2","unstructured":"J Sigovitch Identifying informative biological markers in high-dimensional genomic data and clinical trials. PhD thesis (Harvard Univ Cambridge MA). (2007)."},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1177\/1740774515588096"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1198\/106186008X319331"},{"key":"e_1_3_3_20_2","first-page":"141","article-title":"Subgroup analysis via recursive partitioning","volume":"10","author":"Su X","year":"2009","unstructured":"X Su, C Tsai, H Wang, D Nickerson, B Li, Subgroup analysis via recursive partitioning. J Mach Learn Res 10, 141\u2013158 (2009).","journal-title":"J Mach Learn Res"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1111\/1468-0262.00442"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1162\/rest.90.3.389"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.2014.951443"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1002\/sim.4322"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1214\/12-AOAS593"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1093\/poq\/nfs036"},{"key":"e_1_3_3_27_2","unstructured":"M Taddy M Gardner L Chen D Draper Heterogeneous treatment effects in digital experimentation. arXiv:1412.8563. (2015)."},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4419-9782-1"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/asr055"},{"key":"e_1_3_3_30_2","unstructured":"S Wager G Walther Uniform convergence of random forests via adaptive concentration. arxiv:1503.06388. (2015)."},{"key":"e_1_3_3_31_2","doi-asserted-by":"crossref","unstructured":"J List A Shaikh Y Xu Multiple hypothesis testing in experimental economics. NBER Working Paper No. 21875 (National Bureau of Economic Research Cambridge MA). (2016).","DOI":"10.3386\/w21875"}],"container-title":["Proceedings of the National Academy of Sciences"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.pnas.org\/syndication\/doi\/10.1073\/pnas.1510489113","content-type":"unspecified","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/pnas.org\/doi\/pdf\/10.1073\/pnas.1510489113","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,17]],"date-time":"2024-06-17T23:37:47Z","timestamp":1718667467000},"score":1,"resource":{"primary":{"URL":"https:\/\/pnas.org\/doi\/full\/10.1073\/pnas.1510489113"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,7,5]]},"references-count":31,"journal-issue":{"issue":"27","published-print":{"date-parts":[[2016,7,5]]}},"alternative-id":["10.1073\/pnas.1510489113"],"URL":"https:\/\/doi.org\/10.1073\/pnas.1510489113","relation":{},"ISSN":["0027-8424","1091-6490"],"issn-type":[{"value":"0027-8424","type":"print"},{"value":"1091-6490","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,7,5]]},"assertion":[{"value":"2016-07-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}