{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,9,22]],"date-time":"2026-09-22T14:01:46Z","timestamp":1790085706098,"version":"4.0.1"},"reference-count":64,"publisher":"Cambridge University Press (CUP)","issue":"2","license":[{"start":{"date-parts":[[2017,1,4]],"date-time":"2017-01-04T00:00:00Z","timestamp":1483488000000},"content-version":"unspecified","delay-in-days":5847,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Polit. anal."],"published-print":{"date-parts":[[2001]]},"abstract":"<jats:p>We study rare events data, binary dependent variables with dozens to thousands of times fewer ones (events, such as wars, vetoes, cases of political activism, or epidemiological infections) than zeros (\u201cnonevents\u201d). In many literatures, these variables have proven difficult to explain and predict, a problem that seems to have at least two sources. First, popular statistical procedures, such as logistic regression, can sharply underestimate the probability of rare events. We recommend corrections that outperform existing methods and change the estimates of absolute and relative risks by as much as some estimated effects reported in the literature. Second, commonly used data collection strategies are grossly inefficient for rare events data. The fear of collecting data with too few events has led to data collections with huge numbers of observations but relatively few, and poorly measured, explanatory variables, such as in international conflict data with more than a quarter-million dyads, only a few of which are at war. As it turns out, more efficient sampling designs exist for making valid inferences, such as sampling all available events (e.g., wars) and a tiny fraction of nonevents (peace). This enables scholars to save as much as 99% of their (nonfixed) data collection costs or to collect much more meaningful explanatory variables. We provide methods that link these two results, enabling both types of corrections to work simultaneously, and software that implements the methods developed.<\/jats:p>","DOI":"10.1093\/oxfordjournals.pan.a004868","type":"journal-article","created":{"date-parts":[[2012,1,21]],"date-time":"2012-01-21T23:59:26Z","timestamp":1327190366000},"page":"137-163","source":"Crossref","is-referenced-by-count":3197,"title":["Logistic Regression in Rare Events Data"],"prefix":"10.1017","volume":"9","author":[{"given":"Gary","family":"King","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Langche","family":"Zeng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"56","published-online":{"date-parts":[[2017,1,4]]},"reference":[{"key":"S1047198700003740_ref51","doi-asserted-by":"crossref","DOI":"10.2307\/j.ctv1pnc1k7","volume-title":"Voice and Equality: Civic Voluntarism in American Politics","author":"Verba","year":"1995"},{"key":"S1047198700003740_ref49","unstructured":"Tucker Richard . 1999. \u201cBTSCS: A Binary Time-Series-Cross-Section Data Analysis Utility,\u201d Version 3.0.4. http:\/\/www.fas.harvard.edu\/\u223crtucker\/programs\/btscs\/btscs.html."},{"key":"S1047198700003740_ref48","unstructured":"Tucker Richard . 1998. \u201cThe Interstate Dyad-Year Dataset, 1816\u20131997,\u201d Version 3.0. http:\/\/www.fas.harvard.edu\/\u223crtucker\/data\/dyadyear\/."},{"key":"S1047198700003740_ref47","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-4024-2"},{"key":"S1047198700003740_ref46","volume-title":"Bayesian Statistics","author":"Smith","year":"1998"},{"key":"S1047198700003740_ref45","doi-asserted-by":"publisher","DOI":"10.1111\/0020-8833.00113"},{"key":"S1047198700003740_ref44","doi-asserted-by":"publisher","DOI":"10.2307\/2585396"},{"key":"S1047198700003740_ref43","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1111\/j.2517-6161.1986.tb01400.x","article-title":"Fitting Logistic Models Under Case-Control or Choice Based Sampling","volume":"48","author":"Scott","year":"1986","journal-title":"Journal of the Royal Statistical Society, B"},{"key":"S1047198700003740_ref39","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511812651"},{"key":"S1047198700003740_ref38","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/66.3.403"},{"key":"S1047198700003740_ref37","doi-asserted-by":"publisher","DOI":"10.1002\/sim.4780140806"},{"key":"S1047198700003740_ref33","doi-asserted-by":"publisher","DOI":"10.2307\/2938740"},{"key":"S1047198700003740_ref32","volume-title":"Structural Analysis of Discrete Data with Econometric Applications","author":"Manski","year":"1981"},{"key":"S1047198700003740_ref31","doi-asserted-by":"publisher","DOI":"10.2307\/1914121"},{"key":"S1047198700003740_ref30","volume-title":"Nonlinear Statistical Inference: Essays in Honor of Takeshi Amemiya","author":"Manski","year":"1999"},{"key":"S1047198700003740_ref27","doi-asserted-by":"publisher","DOI":"10.1016\/0304-4076(94)01698-4"},{"key":"S1047198700003740_ref35","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4899-3242-6"},{"key":"S1047198700003740_ref26","doi-asserted-by":"publisher","DOI":"10.2307\/2669316"},{"key":"S1047198700003740_ref24","unstructured":"King Gary , and Zeng Langche . 2000b. \u201cExplaining Rare Events in International Relations.\u201d International Organization (in press)."},{"key":"S1047198700003740_ref23","unstructured":"King Gary , and Zeng Langche . 2000a. \u201cInference in Case-Control Studies with Limited Auxilliary Information\u201d (in press). (Preprint at http:\/\/Gking.harvard.edu.)"},{"key":"S1047198700003740_ref22","doi-asserted-by":"publisher","DOI":"10.2307\/2951544"},{"key":"S1047198700003740_ref20","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1985.10478165"},{"key":"S1047198700003740_ref17","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511521713"},{"key":"S1047198700003740_ref16","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4899-4467-2"},{"key":"S1047198700003740_ref15","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/80.1.27"},{"key":"S1047198700003740_ref13","doi-asserted-by":"publisher","DOI":"10.2307\/1912755"},{"key":"S1047198700003740_ref10","doi-asserted-by":"crossref","DOI":"10.2307\/j.ctt1bh4dhm","volume-title":"War and Reason: Domestic and International Imperatives","author":"Bueno de Mesquita","year":"1992"},{"key":"S1047198700003740_ref9","volume-title":"The War Trap","author":"Bueno de Mesquita","year":"1981"},{"key":"S1047198700003740_ref7","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1996.10476660"},{"key":"S1047198700003740_ref5","unstructured":"Bennett D. Scott , and Stam Allan C. III . 1998a. EUGene: Expected Utility Generation and Data Management Program, Version 1.12. http:\/\/wizard.ucr.edu\/cps\/eugene\/eugene.html."},{"key":"S1047198700003740_ref3","doi-asserted-by":"publisher","DOI":"10.2307\/1913609"},{"key":"S1047198700003740_ref2","doi-asserted-by":"publisher","DOI":"10.1214\/ss\/1177011454"},{"key":"S1047198700003740_fn10","unstructured":"We translated the different format in which Bennett and Stam (1998b) report relative risk to our percentage figure. If r is their measure, ours is 100 \u00d7 (r \u2212 1)."},{"key":"S1047198700003740_fn9","unstructured":"Deriving as an approximately unbiased estimator involves some approximations not required for the optimal Bayesian version derived in Appendix E. The problem is that instead of expanding a random \u03c0i around a fixed \u03b2 as in the Bayesian version, we now must expand a random around a fixed \u03b2. Thus, to take the expectation and compute Ci , we need to imagine that in the correction term, is a reasonable estimate of \u03c0i in this context. This is obviously an undesirable approximation but it is better than setting it to zero or one (i.e., the equivalent of setting Ci = 0), and as our Monte Carlos show below, is indeed approximately unbiased."},{"key":"S1047198700003740_ref8","volume-title":"Statistical Methods in Cancer Research","author":"Breslow","year":"1980"},{"key":"S1047198700003740_fn1","unstructured":"Bennett and Stam (1998b) analyze a data set with 684,000 dyad-years and (1998a) have even developed sophisticated software for managing the larger, 1.2 million-dyad data set they distribute."},{"key":"S1047198700003740_ref42","doi-asserted-by":"publisher","DOI":"10.1002\/sim.4780020108"},{"key":"S1047198700003740_ref41","volume-title":"Modern Epidemiology","author":"Rothman","year":"1998"},{"key":"S1047198700003740_ref14","volume-title":"Structural Analysis of Discrete Data with Econometric Applications","author":"Cosslett","year":"1981b"},{"key":"S1047198700003740_ref40","volume-title":"In Search of Global Patterns","author":"Rosenau","year":"1976"},{"key":"S1047198700003740_ref50","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511583483"},{"key":"S1047198700003740_ref29","first-page":"2120","volume-title":"Behavior, Society, and Nuclear War","volume":"1","author":"Levy","year":"1989"},{"key":"S1047198700003740_ref6","unstructured":"Bennett D. Scott , and Stam Allan C. III . 1998b. \u201cTheories of Conflict Initiation and Escalation: Comparative Testing, 1816\u20131980,\u201d Presented at the annual meeting of the International Studies Association Minneapolis."},{"key":"S1047198700003740_fn5","unstructured":"We analyze the problem of absolute risk directly and then compute relative risk as the ratio of two absolute risks. Although we do not pursue other options here because our estimates of relative risk clearly outperform existing methods, it seems possible that even better methods could be developed that estimate relative risk directly."},{"key":"S1047198700003740_fn7","unstructured":"More formally, suppose P(X | Y = j) = Normal(X | \u00b5j , 1), for j = 0, 1. Then the logit model should classify an observation as 1 if the probability is greater than 0.5 or equivalently X \u00bb T (\u00b5 0, \u00b5 1) = [ln(1 \u2013 \u03c4) \u2013 ln(\u03c4)]\/(\u00b5 1 \u2013 \u00b5 0) + (\u00b5 0 + \u00b5 1)\/2. A logit of Y on a constant term and X is fully saturated and hence equivalent to estimating \u00b5j with (the mean of X i for all i in which Yi = j). However, the estimated classification boundary, , will be larger than T(\u00b5 0, \u00b5 1) when \u03c4 < 0.5 (and thus ln[(1 \u2013 \u03c4)\/\u03c4] > 0), since, by Jensen's inequality, . Hence, the threshold will be too far to the right in Fig. 1 and will underestimate the probability of a one in finite samples."},{"key":"S1047198700003740_ref34","volume-title":"Tensor Methods in Statistics","author":"McCullagh","year":"1987"},{"key":"S1047198700003740_ref36","volume-title":"Exact Inference for Categorical Data","author":"Mehta","year":"1997"},{"key":"S1047198700003740_ref53","doi-asserted-by":"publisher","DOI":"10.1177\/0049124189017003003"},{"key":"S1047198700003740_ref21","doi-asserted-by":"publisher","DOI":"10.2307\/1957394"},{"key":"S1047198700003740_ref4","doi-asserted-by":"publisher","DOI":"10.1017\/S0003055400220078"},{"key":"S1047198700003740_ref18","volume-title":"Econometric Analysis","author":"Greene","year":"1993"},{"key":"S1047198700003740_ref1","unstructured":"Achen Christopher A. 1999. \u201cRetrospective Sampling in International Relations,\u201d Presented at the annual meetings of the Midwest Political Science Association, Chicago."},{"key":"S1047198700003740_ref52","doi-asserted-by":"publisher","DOI":"10.1016\/0378-3758(93)E0091-T"},{"key":"S1047198700003740_ref19","doi-asserted-by":"publisher","DOI":"10.1177\/0193841X8801200301"},{"key":"S1047198700003740_ref28","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1016\/0304-4076(95)01756-9","article-title":"Efficient Estimation and Stratified Sampling","volume":"74","author":"Lancaster","year":"1996b","journal-title":"Journal of Econometrics"},{"key":"S1047198700003740_ref25","doi-asserted-by":"crossref","DOI":"10.1515\/9781400821211","volume-title":"Designing Social Inquiry: Scientific Inference in Qualitative Research","author":"King","year":"1994"},{"key":"S1047198700003740_ref11","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-0258(19970315)16:5<545::AID-SIM421>3.0.CO;2-3"},{"key":"S1047198700003740_fn3","unstructured":"We have found no discussion in political science of the effects of finite samples and rare events on logistic regression or of most of the methods we discuss that allow selection on Y. There is a brief discussion of one method of correcting selection on Y in asymptotic samples by Bueno de Mesquita and Lalman (1992, Appendix) and in an unpublished paper they cite that has recently become available (Achen 1999)."},{"key":"S1047198700003740_ref12","doi-asserted-by":"crossref","first-page":"629","DOI":"10.1111\/j.2517-6161.1991.tb01852.x","article-title":"Bias Correction in Generalized Linear Models","volume":"53","author":"Cordeiro","year":"1991","journal-title":"Journal of the Royal Statistical Society, B"},{"key":"#cr-split#-S1047198700003740_fn8.1","unstructured":"An elegant result due to Firth (1993) shows that bias can also be corrected during the maximization procedure by applying Jeffrey's invariant prior to the logistic likelihood and using the maximum posterior estimate. We have applied this work to weighting and prior correction and run experiments to compare the methods. Consistent with Firth's examples, we find that the methods give answers that are always numerically very close (almost always less than half a percent). An advantage of Firth's procedure is that it gives answers even when the MLE is undefined, as in cases of perfect discrimination"},{"key":"#cr-split#-S1047198700003740_fn8.2","unstructured":"a disadvantage is computational in that the analytical gradient and Hessian are much more complicated. Another approach to bias reduction is based on jackknife methods, which replace analytical derivations with easy computations, although systematic comparisons by Bull et al. (1997) show that they do not generally work as well as the analytical approaches."},{"key":"S1047198700003740_fn2","unstructured":"The fixed costs involved in gearing up to collect data would be borne with either data collection strategy, and so selecting on the dependent variable as we suggest saves something less in research dollars than the fraction of observations not collected."},{"key":"S1047198700003740_fn4","unstructured":"King and Zeng (2000a), building on results of Manski (1999), modify the methods in this paper for the situation when \u03c4 is unknown or partially known. King and Zeng use \u201crobust bayesian analysis\u201d to specify classes of prior distributions on \u03c4, representing full or partial ignorance. For example, the user can specify that \u03c4 is completely unknown or known to fall with some probability to lie only in a given interval. The result is classes of posterior distributions (instead of a single posterior) that, in many cases, provide informative estimates of quantities of interest."},{"key":"S1047198700003740_fn6","unstructured":"\u201cExact\u201d tests are a good solution to the problem when all variables are discrete and sufficient (often massive) computational power is available (see Agresti 1992; Mehta and Patel 1997). These tests compute exact finite sample distributions based on permutations of the data tables."}],"container-title":["Political Analysis"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1047198700003740","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,18]],"date-time":"2024-04-18T04:18:23Z","timestamp":1713413903000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1047198700003740\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2001]]},"references-count":64,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2001]]}},"alternative-id":["S1047198700003740"],"URL":"https:\/\/doi.org\/10.1093\/oxfordjournals.pan.a004868","relation":{},"ISSN":["1047-1987","1476-4989"],"issn-type":[{"value":"1047-1987","type":"print"},{"value":"1476-4989","type":"electronic"}],"subject":[],"published":{"date-parts":[[2001]]}}}