{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T23:27:30Z","timestamp":1783121250555,"version":"3.54.6"},"reference-count":58,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2022,1,29]],"date-time":"2022-01-29T00:00:00Z","timestamp":1643414400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000038","name":"Natural Sciences and Engineering Research Council","doi-asserted-by":"publisher","award":["123456 288332"],"award-info":[{"award-number":["123456 288332"]}],"id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Background: Many methods under the umbrella of inter-rater agreement (IRA) have been proposed to evaluate how well two or more medical experts agree on a set of outcomes. The objective of this work was to assess key IRA statistics in the context of multiple raters with binary outcomes. Methods: We simulated the responses of several raters (2\u20135) with 20, 50, 300, and 500 observations. For each combination of raters and observations, we estimated the expected value and variance of four commonly used inter-rater agreement statistics (Fleiss\u2019 Kappa, Light\u2019s Kappa, Conger\u2019s Kappa, and Gwet\u2019s AC1). Results: In the case of equal outcome prevalence (symmetric), the estimated expected values of all four statistics were equal. In the asymmetric case, only the estimated expected values of the three Kappa statistics were equal. In the symmetric case, Fleiss\u2019 Kappa yielded a higher estimated variance than the other three statistics. In the asymmetric case, Gwet\u2019s AC1 yielded a lower estimated variance than the three Kappa statistics for each scenario. Conclusion: Since the population-level prevalence of a set of outcomes may not be known a priori, Gwet\u2019s AC1 statistic should be favored over the three Kappa statistics. For meaningful direct comparisons between IRA measures, transformations between statistics should be conducted.<\/jats:p>","DOI":"10.3390\/sym14020262","type":"journal-article","created":{"date-parts":[[2022,1,29]],"date-time":"2022-01-29T01:43:27Z","timestamp":1643420607000},"page":"262","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":25,"title":["An Empirical Comparative Assessment of Inter-Rater Agreement of Binary Outcomes and Multiple Raters"],"prefix":"10.3390","volume":"14","author":[{"given":"Menelaos","family":"Konstantinidis","sequence":"first","affiliation":[{"name":"Division of Biostatistics, Dalla Lana School of Public Health, University of Toronto, Toronto, ON M5T 3M7, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lisa. W.","family":"Le","sequence":"additional","affiliation":[{"name":"Department of Biostatistics, Princess Margaret Cancer Centre, University Health Network, Toronto, ON M5G 2C1, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8057-5100","authenticated-orcid":false,"given":"Xin","family":"Gao","sequence":"additional","affiliation":[{"name":"Department of Mathematics and Statistics, York University, Toronto, ON M3J 1P3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,1,29]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1016\/S0140-6736(86)90837-8","article-title":"Statistical Methods for Assessing Agreement Between Two Methods of Clinical Measurement","volume":"327","author":"Altman","year":"1986","journal-title":"Lancet"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1177\/001316446002000104","article-title":"A Coefficient of Agreement for Nominal Scales","volume":"20","author":"Cohen","year":"1960","journal-title":"Educ. Psychol. Meas."},{"key":"ref_3","unstructured":"Gwet, K.L. (2014). Handbook of Inter-Rater Reliability, Advanced Analytics. [4th ed.]."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"921","DOI":"10.1177\/0013164488484007","article-title":"A Generalization of Cohen\u2019s Kappa Agreement Measure to Interval Measurement and Multiple Raters","volume":"48","author":"Berry","year":"1988","journal-title":"Educ. Psychol. Meas."},{"key":"ref_5","first-page":"1","article-title":"Disagreement on Agreement: Two Alternative Agreement Coefficients","volume":"186","author":"Blood","year":"2007","journal-title":"SAS Glob. Forum"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"330","DOI":"10.1016\/j.sapharm.2012.04.004","article-title":"Interrater agreement and interrater reliability: Key concepts, approaches, and applications","volume":"9","author":"Gisev","year":"2013","journal-title":"Res. Soc. Adm. Pharm."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zapf, A., Castell, S., Morawietz, L., and Karch, A. (2016). Measuring inter-rater reliability for nominal data\u2014Which coefficients and confidence intervals are appropriate?. BMC Med. Res. Methodol., 16.","DOI":"10.1186\/s12874-016-0200-9"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"108","DOI":"10.1016\/j.cllc.2012.06.003","article-title":"Capturing Acute Toxicity Data During Lung Radiotherapy by Using a Patient-Reported Assessment Tool","volume":"14","author":"Tang","year":"2013","journal-title":"Clin. Lung Cancer"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Wongpakaran, N., Wongpakaran, T., Wedding, D., and Gwet, K.L. (2013). A comparison of Cohen\u2019s Kappa and Gwet\u2019s AC1 when calculating inter-rater reliability coefficients: A study conducted with personality disorder samples. BMC Med. Res. Methodol., 13.","DOI":"10.1186\/1471-2288-13-61"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1027\/1614-2241\/a000119","article-title":"Misunderstanding Reliability","volume":"12","author":"Krippendorff","year":"2016","journal-title":"Methodology"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1037\/h0031643","article-title":"Measures of response agreement for qualitative data: Some generalizations and alternatives","volume":"76","author":"Light","year":"1971","journal-title":"Psychol. Bull."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"423","DOI":"10.1016\/0895-4356(93)90018-V","article-title":"Bias, prevalence and Kappa","volume":"46","author":"Byrt","year":"1993","journal-title":"J. Clin. Epidemiol."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"811","DOI":"10.1002\/bimj.4710370705","article-title":"Raking Kappa: Describing Potential Impact of Marginal Distributions on Measures of Agreement","volume":"37","author":"Agresti","year":"1995","journal-title":"Biom. J."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1037\/h0026256","article-title":"Weighted Kappa: Nominal scale agreement provision for scaled disagreement or partial credit","volume":"70","author":"Cohen","year":"1968","journal-title":"Psychol. Bull."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"378","DOI":"10.1037\/h0031619","article-title":"Measuring Nominal Scale agreement amongst many raters","volume":"76","author":"Fleiss","year":"1971","journal-title":"Psychol. Bull."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"322","DOI":"10.1037\/0033-2909.88.2.322","article-title":"Integration and generalization of kappas for multiple raters","volume":"88","author":"Conger","year":"1980","journal-title":"Psychol. Bull."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1177\/001316447003000105","article-title":"Estimating the Reliability, Systematic Error and Random Error of Interval Data","volume":"30","author":"Krippendorff","year":"1970","journal-title":"Educ. Psychol. Meas."},{"key":"ref_18","first-page":"7","article-title":"Agree or Disagree? A Demonstration of An Alternative Statistic to Cohen\u2019 s Kappa for Measuring the Extent and Reliability of Agreement between Observers","volume":"3","author":"Xie","year":"2013","journal-title":"FCSM Res. Conf."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Higgins, J.P., Thomas, J., Chandlerr, J., Cumpston, M., Li, T., Page, M.J., and Welch, V.A. (2019). Cochrane Handbook for Systematic Reviews of Interventions, John Wiley & Sons, Ltd.. [2nd ed.].","DOI":"10.1002\/9781119536604"},{"key":"ref_20","unstructured":"Garritty, C., Gartlehner, G., Kamel, C., King, V.J., Nussbaumer-Streit, B., Stevens, A., Hamel, C., and Affengruber, L. (2020). Cochrane Rapid Reviews, Cochrane Community. Interim Guidence Cochrane Rapid Reviews Methods Group."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Munn, Z., Peters, M.D.J., Stern, C., Tufanaru, C., McArthur, A., and Aromataris, E. (2018). Systematic review or scoping review? Guidance for authors when choosing between a systematic or scoping review approach. BMC Med. Res. Methodol., 18.","DOI":"10.1186\/s12874-018-0611-x"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Kastner, M., Tricco, A.C., Soobiah, C., Lillie, E., Perrier, L., Horsley, T., Welch, V., Cogo, E., Antony, J., and Straus, S.E. (2012). What is the most appropriate knowledge synthesis method to conduct a review? Protocol for a scoping review. BMC Med. Res. Methodol., 12.","DOI":"10.1186\/1471-2288-12-114"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"l4898","DOI":"10.1136\/bmj.l4898","article-title":"RoB 2: A revised tool for assessing risk of bias in randomised trials","volume":"366","author":"Sterne","year":"2019","journal-title":"BMJ"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"i4919","DOI":"10.1136\/bmj.i4919","article-title":"ROBINS-I: A tool for assessing risk of bias in non-randomised studies of interventions","volume":"355","author":"Sterne","year":"2016","journal-title":"BMJ"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Pieper, D., Jacobs, A., Weikert, B., Fishta, A., and Wegewitz, U. (2017). Inter-rater reliability of AMSTAR is dependent on the pair of reviewers. BMC Med. Res. Methodol., 17.","DOI":"10.1186\/s12874-017-0380-y"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1186\/s13643-020-1271-6","article-title":"Inter-rater reliability and concurrent validity of ROBINS-I: Protocol for a cross-sectional study","volume":"9","author":"Jeyaraman","year":"2020","journal-title":"Syst. Rev."},{"key":"ref_27","unstructured":"Hartling, L., Hamm, M., Milne, A., Vandermeer, B., Santaguida, P.L., Ansari, M., Tsertsvadze, A., Hempel, S., Shekelle, P., and Drydem, D.M. (2012). Validity and Inter-Rater Reliability Testing of Quality Assessment Instruments, Agency for Healthcare Research and Quality."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"837","DOI":"10.1177\/0049124118799372","article-title":"Interrater Reliability in Systematic Review Methodology","volume":"50","author":"Belur","year":"2021","journal-title":"Sociol. Methods Res."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Woo, S.A., Cragg, A., Wickham, M.E., Peddie, D., Balka, E., Scheuermeyer, F., Villanyi, D., and Hohl, C.M. (2018). Methods for evaluating adverse drug event preventability in emergency department patients. BMC Med. Res. Methodol., 18.","DOI":"10.1186\/s12874-018-0617-4"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"640","DOI":"10.1111\/j.1553-2712.2012.01379.x","article-title":"Clinical decision rules to improve the detection of adverse drug events in emergency department patients","volume":"19","author":"Hohl","year":"2012","journal-title":"Acad. Emerg. Med."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1015","DOI":"10.1111\/acem.13407","article-title":"Prospective validation of clinical criteria to identify emergency department patients at high risk for adverse drug events","volume":"25","author":"Hohl","year":"2018","journal-title":"Acad. Emerg. Med."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"2002","DOI":"10.1056\/NEJMsa1103053","article-title":"Emergency hospitalizations for adverse drug events in older Americans","volume":"365","author":"Budnitz","year":"2011","journal-title":"N. Engl. J. Med."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1563","DOI":"10.1503\/cmaj.071594","article-title":"Incidence, severity and preventability of medication-related visits to the emergency department: A prospective study","volume":"178","author":"Zed","year":"2008","journal-title":"CMAJ"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Hamilton, H.J., Gallagher, P.F., and O\u2019Mahony, D. (2009). Inappropriate prescribing and adverse drug events in older people. BMC Geriatr., 9.","DOI":"10.1186\/1471-2318-9-5"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1016\/j.jclinepi.2008.04.007","article-title":"Diagnostic test accuracy may vary with prevalence: Implications for evidence-based diagnosis","volume":"62","author":"Leeflang","year":"2009","journal-title":"J. Clin. Epidemiol."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"499","DOI":"10.1016\/S0895-4356(99)00174-2","article-title":"Bias and prevalence effects on Kappa viewed in terms of sensitivity and specificity","volume":"53","author":"Hoehler","year":"2000","journal-title":"J. Clin. Epidemiol."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"677","DOI":"10.1016\/j.annepidem.2017.09.001","article-title":"Summary measures of agreement and association between many raters\u2019 ordinal classifications","volume":"27","author":"Mitani","year":"2017","journal-title":"Ann. Epidemiol."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"512","DOI":"10.1016\/0047-259X(88)90145-5","article-title":"Estimating multiple rater agreement for a rare diagnosis","volume":"27","author":"Verducci","year":"1988","journal-title":"J. Multivar. Anal."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"277","DOI":"10.22237\/jmasm\/1509495300","article-title":"Modeling Agreement between Binary Classifications of Multiple Raters in R and SAS","volume":"16","author":"Mitani","year":"2017","journal-title":"J. Mod. Appl. Stat. Methods"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"639","DOI":"10.1002\/bimj.201700078","article-title":"Evaluating the effects of rater and subject factors on measures of association","volume":"60","author":"Nelson","year":"2018","journal-title":"Biom. J."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1135","DOI":"10.1007\/s00228-019-02670-9","article-title":"Adverse drug reaction causality assessment tools for drug-induced Stevens-Johnson syndrome and toxic epidermal necrolysis: Room for improvement","volume":"75","author":"Goldman","year":"2019","journal-title":"Eur. J. Clin. Pharmacol."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1016\/j.jclinepi.2010.03.002","article-title":"Guidelines for Reporting Reliability and Agreement Studies (GRRAS) were proposed","volume":"64","author":"Kottner","year":"2011","journal-title":"J. Clin. Epidemiol."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"462","DOI":"10.1161\/STROKEAHA.112.678615","article-title":"Reliability (Inter-rater Agreement) of the Barthel Index for Assessment of Stroke Survivors","volume":"44","author":"Duffy","year":"2013","journal-title":"Stroke"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"200","DOI":"10.1111\/j.1747-4949.2009.00271.x","article-title":"Functional outcome measures in contemporary stroke trials","volume":"4","author":"Quinn","year":"2009","journal-title":"Int. J. Stroke"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"1146","DOI":"10.1161\/STROKEAHA.110.598540","article-title":"Barthel index for stroke trials: Development, properties, and application","volume":"42","author":"Quinn","year":"2011","journal-title":"Stroke"},{"key":"ref_46","first-page":"61","article-title":"Functional evaluation: The Barthel Index: A simple index of independence useful in scoring improvement in the rehabilitation of the chronically ill","volume":"14","author":"Mahoney","year":"1965","journal-title":"Md. State Med. J."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"3638","DOI":"10.1007\/s00330-015-3759-3","article-title":"Diagnostic performance of the automated breast volume scanner: A systematic review of inter-rater reliability\/agreement and meta-analysis of diagnostic accuracy for differentiating benign and malignant breast lesions","volume":"25","author":"Meng","year":"2015","journal-title":"Eur. Radiol."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"5","DOI":"10.5070\/D397D8T291","article-title":"Treatment of severe drug reactions: Stevens-Johnson syndrome, toxic epidermal necrolysis and hypersensitivity syndrome","volume":"8","author":"Ghislain","year":"2002","journal-title":"Dermatol. Online J."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Gallagher, R.M., Kirkham, J.J., Mason, J.R., Bird, K.A., Williamson, P.R., Nunn, A.J., Turner, M.A., Smyth, R.L., and Pirmohamed, M. (2011). Development and Inter-Rater Reliability of the Liverpool Adverse Drug Reaction Causality Assessment Tool. PLoS ONE, 6.","DOI":"10.1371\/journal.pone.0028096"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1016\/0197-2456(86)90046-2","article-title":"Meta-analysis in clinical trials","volume":"7","author":"DerSimonian","year":"1986","journal-title":"Control. Clin. Trials"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1007\/s10742-011-0077-3","article-title":"Meta-analysis of Cohen\u2019s Kappa","volume":"11","author":"Sun","year":"2011","journal-title":"Health Serv. Outcomes Res. Methodol."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Bornmann, L., Mutz, R., and Daniel, H.-D. (2010). A Reliability-Generalization Study of Journal Peer Reviews: A Multilevel Meta-Analysis of Inter-Rater Reliability and Its Determinants. PLoS ONE, 5.","DOI":"10.1371\/journal.pone.0014331"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Honda, C., and Ohyama, T. (2020). Homogeneity score test of AC1 statistics and estimation of common AC1 in multiple or stratified inter-rater agreement studies. BMC Med. Res. Methodol., 20.","DOI":"10.1186\/s12874-019-0887-5"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"876","DOI":"10.1002\/sim.4780130809","article-title":"A goodness-of-fit approach to inference procedures for the kappa statistic: Confidence interval construction, significance-testing and sample size estimation","volume":"13","author":"Kraemer","year":"1994","journal-title":"Stat. Med."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"271","DOI":"10.1007\/s11634-010-0073-4","article-title":"Inequalities between multi-rater kappas","volume":"4","author":"Warrens","year":"2010","journal-title":"Adv. Data Anal. Classif."},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"3","DOI":"10.2307\/3315487","article-title":"Beyond kappa: A rev interrater agreemen","volume":"27","author":"Banerjee","year":"2019","journal-title":"Can. J. Stat."},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1002\/bimj.201700016","article-title":"Asymptotic distributions of kappa statistics and their differences with many raters, many rating categories and two conditions","volume":"60","author":"Grassano","year":"2018","journal-title":"Biom. J."},{"key":"ref_58","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1348\/000711006X126600","article-title":"Computing inter-rater reliability and its variance in the presence of high agreement","volume":"61","author":"Gwet","year":"2008","journal-title":"Br. J. Math. Stat. Psychol."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/2\/262\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:10:35Z","timestamp":1760134235000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/2\/262"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,29]]},"references-count":58,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["sym14020262"],"URL":"https:\/\/doi.org\/10.3390\/sym14020262","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,29]]}}}