{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T22:52:14Z","timestamp":1783810334999,"version":"3.55.0"},"reference-count":71,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,9,11]],"date-time":"2020-09-11T00:00:00Z","timestamp":1599782400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,9,11]],"date-time":"2020-09-11T00:00:00Z","timestamp":1599782400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Med Inform Decis Mak"],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Background<\/jats:title><jats:p>We focus on the importance of interpreting the quality of the labeling used as the input of predictive models to understand the reliability of their output in support of human decision-making, especially in critical domains, such as medicine.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>Accordingly, we propose a framework distinguishing the reference labeling (or Gold Standard) from the set of annotations from which it is usually derived (the Diamond Standard). We define a set of quality dimensions and related metrics: representativeness (are the available data representative of its reference population?); reliability (do the raters agree with each other in their ratings?); and accuracy (are the raters\u2019 annotations a true representation?). The metrics for these dimensions are, respectively, the<jats:italic>degree of correspondence<\/jats:italic>,<jats:italic>\u03a8<\/jats:italic>, the<jats:italic>degree of weighted concordance<\/jats:italic><jats:italic>\u03f1<\/jats:italic>, and the<jats:italic>degree of fineness<\/jats:italic>,<jats:italic>\u03a6<\/jats:italic>. We apply and evaluate these metrics in a diagnostic user study involving 13 radiologists.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>We evaluate<jats:italic>\u03a8<\/jats:italic>against hypothesis-testing techniques, highlighting that our metrics can better evaluate distribution similarity in high-dimensional spaces. We discuss how<jats:italic>\u03a8<\/jats:italic>could be used to assess the reliability of new predictions or for train-test selection. We report the value of<jats:italic>\u03f1<\/jats:italic>for our case study and compare it with traditional reliability metrics, highlighting both their theoretical properties and the reasons that they differ. Then, we report the<jats:italic>degree of fineness<\/jats:italic>as an estimate of the accuracy of the collected annotations and discuss the relationship between this latter degree and the<jats:italic>degree of weighted concordance<\/jats:italic>, which we find to be moderately but significantly correlated. Finally, we discuss the implications of the proposed dimensions and metrics with respect to the context of Explainable Artificial Intelligence (XAI).<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusion<\/jats:title><jats:p>We propose different dimensions and related metrics to assess the quality of the datasets used to build predictive models and Medical Artificial Intelligence (MAI). We argue that the proposed metrics are feasible for application in real-world settings for the continuous development of trustable and interpretable MAI systems.<\/jats:p><\/jats:sec>","DOI":"10.1186\/s12911-020-01224-9","type":"journal-article","created":{"date-parts":[[2020,9,11]],"date-time":"2020-09-11T12:02:41Z","timestamp":1599825761000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":42,"title":["As if sand were stone. New concepts and metrics to probe the ground on which to build trustable AI"],"prefix":"10.1186","volume":"20","author":[{"given":"Federico","family":"Cabitza","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrea","family":"Campagner","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Luca Maria","family":"Sconfienza","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,9,11]]},"reference":[{"key":"1224_CR1","doi-asserted-by":"crossref","unstructured":"Cabitza F, Zeitoun J-D. The proof of the pudding: in praise of a culture of real-world validation for medical artificial intelligence. Ann Transl Med. 2019; 7(8). http:\/\/atm.amegroups.com\/article\/view\/25300.","DOI":"10.21037\/atm.2019.04.07"},{"issue":"2","key":"1224_CR2","first-page":"47","volume":"31","author":"K Siau","year":"2018","unstructured":"Siau K, Wang W. Building trust in artificial intelligence, machine learning, and robotics. Cutter Bus Technol J. 2018; 31(2):47\u201353.","journal-title":"Cutter Bus Technol J"},{"issue":"11","key":"1224_CR3","doi-asserted-by":"crossref","first-page":"1134","DOI":"10.1145\/1968.1972","volume":"27","author":"LG Valiant","year":"1984","unstructured":"Valiant LG. A theory of the learnable. Commun. ACM. 1984; 27(11):1134\u201342. https:\/\/doi-org.proxy.unimib.it\/10.1145\/1968.1972.","journal-title":"Commun. ACM"},{"key":"1224_CR4","volume-title":"Machine learning","author":"TM Mitchell","year":"1997","unstructured":"Mitchell TM. Machine learning. The address: McGraw-Hill Education; 1997."},{"issue":"11","key":"1224_CR5","doi-asserted-by":"crossref","first-page":"4014","DOI":"10.3390\/app10114014","volume":"10","author":"F Cabitza","year":"2020","unstructured":"Cabitza F, Campagner A, Albano D, Aliprandi A, Bruno A, Chianca V, Corazza A, Di Pietto F, Gambino A, Gitto S, et al.The elephant in the machine: Proposing a new metric of data reliability and its application to a medical case to assess classification reliability. Appl Sci. 2020; 10(11):4014.","journal-title":"Appl Sci"},{"issue":"2","key":"1224_CR6","doi-asserted-by":"crossref","first-page":"709","DOI":"10.5465\/amr.2007.24348410","volume":"32","author":"D Schoorman","year":"2007","unstructured":"Schoorman D, Mayer R, Davis J. An integrative model of organizational trust: Past, present, and future. Acad Manag Rev. 2007; 32(2):709\u201334.","journal-title":"Acad Manag Rev"},{"key":"1224_CR7","doi-asserted-by":"crossref","unstructured":"Campagner A, Cabitza F, Ciucci D. Exploring medical data classification with three-way decision tree. In: Proceedings of the 12th BIOSTEC International Joint Conference, vol. 5: 2019. p. 147\u201358.","DOI":"10.5220\/0007571001470158"},{"key":"1224_CR8","doi-asserted-by":"crossref","unstructured":"Holzinger A. From machine learning to explainable ai. In: 2018 World Symposium on Digital Intelligence for Systems and Machines (DISA). IEEE: 2018. p. 55\u201366.","DOI":"10.1109\/DISA.2018.8490530"},{"issue":"2","key":"1224_CR9","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1007\/s00769-012-0885-3","volume":"17","author":"P De Bi\u00e8vre","year":"2012","unstructured":"De Bi\u00e8vre P. The 2012 international vocabulary of metrology:vim. Accred Qual Assur. 2012; 17(2):231\u20132.","journal-title":"Accred Qual Assur"},{"issue":"2","key":"1224_CR10","doi-asserted-by":"crossref","first-page":"413","DOI":"10.1037\/0033-2909.88.2.413","volume":"88","author":"FE Saal","year":"1980","unstructured":"Saal FE, Downey RG, Lahey MA. Rating the ratings: Assessing the psychometric quality of rating data. Psychol Bull. 1980; 88(2):413.","journal-title":"Psychol Bull"},{"issue":"3","key":"1224_CR11","doi-asserted-by":"crossref","first-page":"475","DOI":"10.1177\/1460458218824705","volume":"25","author":"F Cabitza","year":"2019","unstructured":"Cabitza F, Locoro A, Alderighi C, Rasoini R, Compagnone D, Berjano P. The elephant in the record: on the multiplicity of data recording work. Health Inform J. 2019; 25(3):475\u201390.","journal-title":"Health Inform J"},{"issue":"1","key":"1224_CR12","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1080\/19312450709336664","volume":"1","author":"AF Hayes","year":"2007","unstructured":"Hayes AF, Krippendorff K. Answering the call for a standard reliability measure for coding data. Commun Methods Measures. 2007; 1(1):77\u201389.","journal-title":"Commun Methods Measures"},{"issue":"4","key":"1224_CR13","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1080\/00031305.2016.1141708","volume":"70","author":"D Quarfoot","year":"2016","unstructured":"Quarfoot D, Levine RA. How robust are multirater interrater reliability indices to changes in frequency distribution?Am Stat. 2016; 70(4):373\u201384.","journal-title":"Am Stat"},{"issue":"2","key":"1224_CR14","doi-asserted-by":"crossref","first-page":"325","DOI":"10.1214\/aoms\/1177698950","volume":"38","author":"A Dempster","year":"1967","unstructured":"Dempster A. Upper and lower probabilities induced by a multivalued mapping. Ann Math Stat. 1967; 38(2):325\u201339.","journal-title":"Ann Math Stat"},{"key":"1224_CR15","doi-asserted-by":"crossref","DOI":"10.1515\/9780691214696","volume-title":"A Mathematical Theory of Evidence vol. 42","author":"G Shafer","year":"1976","unstructured":"Shafer G. A Mathematical Theory of Evidence vol. 42. Princeton, New Jersey: Princeton university press; 1976."},{"issue":"4","key":"1224_CR16","doi-asserted-by":"crossref","first-page":"361","DOI":"10.1016\/j.inffus.2005.06.005","volume":"7","author":"R Haenni","year":"2006","unstructured":"Haenni R, Hartmann S. Modeling partially reliable information sources: a general approach based on dempster\u2013shafer theory. Inf Fusion. 2006; 7(4):361\u201379.","journal-title":"Inf Fusion"},{"issue":"3","key":"1224_CR17","doi-asserted-by":"crossref","first-page":"449","DOI":"10.1016\/j.ijar.2010.10.004","volume":"52","author":"J Schubert","year":"2011","unstructured":"Schubert J. Conflict management in dempster\u2013shafer theory using the degree of falsity. Int J Approx Reason. 2011; 52(3):449\u201360.","journal-title":"Int J Approx Reason"},{"issue":"3-4","key":"1224_CR18","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1016\/S0020-0255(03)00172-5","volume":"155","author":"B Scotney","year":"2003","unstructured":"Scotney B, McClean S. Database aggregation of imprecise and uncertain evidence. Inf Sci. 2003; 155(3-4):245\u201363.","journal-title":"Inf Sci"},{"issue":"5","key":"1224_CR19","doi-asserted-by":"crossref","first-page":"1487","DOI":"10.3390\/s18051487","volume":"18","author":"F Xiao","year":"2018","unstructured":"Xiao F, Qin B. A weighted combination method for conflicting evidence in multi-sensor data fusion. Sensors. 2018; 18(5):1487.","journal-title":"Sensors"},{"key":"1224_CR20","volume-title":"Handbook of inter-rater reliability: the definitive guide to measuring the extent of agreement among raters","author":"KL Gwet","year":"2014","unstructured":"Gwet KL. Handbook of inter-rater reliability: the definitive guide to measuring the extent of agreement among raters. Gaithersburg, MD: Advanced Analytics, LLC; 2014."},{"key":"1224_CR21","doi-asserted-by":"crossref","DOI":"10.2172\/800792","volume-title":"Combination of evidence in dempster-shafer theory. Technical report","author":"K Sentz","year":"2002","unstructured":"Sentz K, Ferson S. Combination of evidence in dempster-shafer theory. Technical report. Albuquerque: Sandia National Laboratories; 2002."},{"key":"1224_CR22","volume-title":"Probabilistic models for some intelligence and attainment tests, 1960","author":"G Rasch","year":"1980","unstructured":"Rasch G. Probabilistic models for some intelligence and attainment tests, 1960. Copenhagen: Danish Institute for Educational Research; 1980."},{"key":"1224_CR23","volume-title":"Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, vol 7","author":"S Heinecke","year":"2019","unstructured":"Heinecke S, Reyzin L. Crowdsourced pac learning under classification. In: Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, vol 7. Austin: AAAI Press: 2019. p. 41\u201349."},{"issue":"3","key":"1224_CR24","doi-asserted-by":"crossref","first-page":"175","DOI":"10.1016\/S0004-9514(14)60377-9","volume":"44","author":"K Bennell","year":"1998","unstructured":"Bennell K, Talbot R, Wajswelner H, Techovanich W, Kelly D, Hall AJ. Intra-rater and inter-rater reliability of a weight-bearing lunge measure of ankle dorsiflexion. Aust J Physiother. 1998; 44(3):175\u201380.","journal-title":"Aust J Physiother"},{"issue":"5","key":"1224_CR25","doi-asserted-by":"crossref","first-page":"0124290","DOI":"10.1371\/journal.pone.0124290","volume":"10","author":"ME Gianinazzi","year":"2015","unstructured":"Gianinazzi ME, et al.Intra-rater and inter-rater reliability of a medical record abstraction study on transition of care after childhood cancer. PloS ONE. 2015; 10(5):0124290.","journal-title":"PloS ONE"},{"key":"1224_CR26","volume-title":"International Cross-Domain Conference for Machine Learning and Knowledge Extraction","author":"F Cabitza","year":"2019","unstructured":"Cabitza F, Campagner A, Ciucci D. New frontiers in explainable ai: Understanding the gi to interpret the go. In: International Cross-Domain Conference for Machine Learning and Knowledge Extraction. Cham: Springer: 2019. p. 27\u201347."},{"key":"1224_CR27","volume-title":"Advances in Neural Information Processing Systems","author":"R Jin","year":"2003","unstructured":"Jin R, Ghahramani Z. Learning with multiple labels. In: Advances in Neural Information Processing Systems. Cambridge: MIT Press: 2003. p. 921\u20138."},{"key":"1224_CR28","unstructured":"Sriperumbudur BK, Fukumizu K, Gretton A, Sch\u00f6lkopf B, Lanckriet GR. On integral probability metrics, \u03d5-divergences and binary classification. arXiv preprint arXiv:0901.2698. 2009."},{"key":"1224_CR29","volume-title":"Nonparametric statistics: a step-by-step approach","author":"GW Corder","year":"2014","unstructured":"Corder GW, Foreman DI. Nonparametric statistics: a step-by-step approach. Hoboken, New Jersey: John Wiley & Sons; 2014."},{"key":"1224_CR30","doi-asserted-by":"crossref","unstructured":"Gretton A, Borgwardt K, Rasch M, Sch\u00f6lkopf B, Smola AJ. A kernel method for the two-sample-problem. In: Advances in Neural Information Processing Systems. Curran Associates, Inc.: 2007. p. 513\u201320.","DOI":"10.7551\/mitpress\/7503.003.0069"},{"issue":"7","key":"1224_CR31","doi-asserted-by":"crossref","first-page":"3797","DOI":"10.1109\/TIT.2014.2320500","volume":"60","author":"T Van Erven","year":"2014","unstructured":"Van Erven T, Harremos P. R\u00e9nyi divergence and kullback-leibler divergence. IEEE Trans Inf Theory. 2014; 60(7):3797\u2013820.","journal-title":"IEEE Trans Inf Theory"},{"issue":"7","key":"1224_CR32","doi-asserted-by":"crossref","first-page":"1858","DOI":"10.1109\/TIT.2003.813506","volume":"49","author":"DM Endres","year":"2003","unstructured":"Endres DM, Schindelin JE. A new metric for probability distributions. IEEE Trans Inf Theory. 2003; 49(7):1858\u201360.","journal-title":"IEEE Trans Inf Theory"},{"key":"1224_CR33","unstructured":"P\u00e9rez-Cruz F. Estimation of information theoretic measures for continuous random variables. In: Advances in Neural Information Processing Systems. Curran Associates, Inc.: 2009. p. 1257\u201364."},{"issue":"11","key":"1224_CR34","doi-asserted-by":"crossref","first-page":"5847","DOI":"10.1109\/TIT.2010.2068870","volume":"56","author":"X Nguyen","year":"2010","unstructured":"Nguyen X, Wainwright MJ, Jordan MI. Estimating divergence functionals and the likelihood ratio by convex risk minimization. IEEE Trans Inf Theory. 2010; 56(11):5847\u201361.","journal-title":"IEEE Trans Inf Theory"},{"key":"1224_CR35","volume-title":"Statistical inference based on divergence measures","author":"L Pardo","year":"2005","unstructured":"Pardo L. Statistical inference based on divergence measures. Boca Raton, FL: CRC press; 2005."},{"key":"1224_CR36","volume-title":"Handbook of Biological Statistics vol. 2","author":"JH McDonald","year":"2009","unstructured":"McDonald JH. Handbook of Biological Statistics vol. 2. Baltimore: sparky house publishing Baltimore, MD; 2009."},{"key":"1224_CR37","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9781139644150","volume-title":"Aggregation Functions vol. 127","author":"M Grabisch","year":"2009","unstructured":"Grabisch M, Marichal J-L, Mesiar R, Pap E. Aggregation Functions vol. 127. Cambridge, United Kingdom: Cambridge University Press; 2009."},{"issue":"3","key":"1224_CR38","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1016\/S0167-7152(97)00020-5","volume":"35","author":"A Justel","year":"1997","unstructured":"Justel A, Pe\u00f1a D, Zamar R. A multivariate kolmogorov-smirnov test of goodness of fit. Stat Probab Lett. 1997; 35(3):251\u20139.","journal-title":"Stat Probab Lett"},{"issue":"4","key":"1224_CR39","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1111\/j.1467-9868.2005.00513.x","volume":"67","author":"PR Rosenbaum","year":"2005","unstructured":"Rosenbaum PR. An exact distribution-free test comparing two multivariate distributions based on adjacency. J R Stat Soc Ser B Stat Methodol. 2005; 67(4):515\u201330.","journal-title":"J R Stat Soc Ser B Stat Methodol"},{"key":"1224_CR40","volume-title":"Twenty-Ninth AAAI Conference on Artificial Intelligence","author":"A Ramdas","year":"2015","unstructured":"Ramdas A, Reddi SJ, P\u00f3czos B, Singh A, Wasserman L. On the decreasing power of kernel and distance based nonparametric hypothesis tests in high dimensions. In: Twenty-Ninth AAAI Conference on Artificial Intelligence. Austin: AAAI Press: 2015."},{"key":"1224_CR41","doi-asserted-by":"crossref","unstructured":"Ahuja K. Estimating kullback-leibler divergence using kernel machines. In: 2019 53rd Asilomar Conference on Signals, Systems, and Computers. IEEE: 2019. p. 690\u20136.","DOI":"10.1109\/IEEECONF44664.2019.9049082"},{"key":"1224_CR42","doi-asserted-by":"crossref","unstructured":"Boltz S, Debreuve E, Barlaud M. knn-based high-dimensional kullback-leibler distance for tracking. In: Eighth International Workshop on Image Analysis for Multimedia Interactive Services (WIAMIS\u201907). IEEE: 2007. p. 16\u201316.","DOI":"10.1109\/WIAMIS.2007.53"},{"issue":"11","key":"1224_CR43","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1371\/journal.pmed.1002699","volume":"15","author":"N Bien","year":"2018","unstructured":"Bien N, Rajpurkar P, Ball RL, Irvin J, Park A, Jones E, Bereket M, Patel BN, Yeom KW, Shpanskaya K, Halabi S, Zucker E, Fanton G, Amanatullah DF, Beaulieu CF, Riley GM, Stewart RJ, Blankenberg FG, Larson DB, Jones RH, Langlotz CP, Ng AY, Lungren MP. Deep-learning-assisted diagnosis for knee magnetic resonance imaging: Development and retrospective validation of mrnet. PLOS Medicine. 2018; 15(11):1\u201319. https:\/\/doi.org\/10.1371\/journal.pmed.1002699.","journal-title":"PLOS Medicine"},{"key":"1224_CR44","unstructured":"Paxton C, Niculescu-Mizil A, Saria S. Developing predictive models using electronic medical records: challenges and pitfalls. In: AMIA Annual Symposium Proceedings, vol. 2013. AMIA: 2013. p. 1109."},{"issue":"1","key":"1224_CR45","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1038\/s41746-020-0232-8","volume":"3","author":"A Kiani","year":"2020","unstructured":"Kiani A, Uyumazturk B, Rajpurkar P, Wang A, Gao R, Jones E, Yu Y, Langlotz CP, Ball RL, Montine TJ, et al.Impact of a deep learning assistant on the histopathologic classification of liver cancer. NPJ Digit Med. 2020; 3(1):1\u20138.","journal-title":"NPJ Digit Med"},{"key":"1224_CR46","doi-asserted-by":"crossref","unstructured":"Cabitza F, Campagner A, Balsano C. Bridging the last mile gap between ai implementation and operation: data awareness that matters. Ann Transl Med. 2020; 8(7). http:\/\/atm.amegroups.com\/article\/view\/39228.","DOI":"10.21037\/atm.2020.03.63"},{"issue":"1","key":"1224_CR47","doi-asserted-by":"crossref","first-page":"159","DOI":"10.2307\/2529310","volume":"33","author":"JR Landis","year":"1977","unstructured":"Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977; 33(1):159\u201374.","journal-title":"Biometrics"},{"key":"1224_CR48","volume-title":"Content analysis: an introduction to its methodology","author":"K Krippendorff","year":"2018","unstructured":"Krippendorff K. Content analysis: an introduction to its methodology. Sage UK: London, England: Sage publications; 2018."},{"key":"1224_CR49","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-90503-7","volume-title":"Organizing for the Digital World","author":"F Cabitza","year":"2019","unstructured":"Cabitza F, Ciucci D, Rasoini R. A giant with feet of clay: on the validity of the data that feed machine learning in medicine. In: Organizing for the Digital World. Cham: Springer: 2019. p. 121\u2013136."},{"key":"1224_CR50","unstructured":"Jiang H, Nachum O. Identifying and correcting label bias in machine learning. arXiv preprint arXiv:1901.04966. 2019."},{"issue":"4","key":"1224_CR51","doi-asserted-by":"crossref","first-page":"363","DOI":"10.5271\/sjweh.555","volume":"26","author":"J Stand","year":"2000","unstructured":"Stand J. The hawthorne effect what did the original hawthorne studies actually show. Scand J Work Environ Health. 2000; 26(4):363\u20137.","journal-title":"Scand J Work Environ Health"},{"issue":"1","key":"1224_CR52","doi-asserted-by":"crossref","first-page":"47","DOI":"10.1148\/radiol.2491072025","volume":"249","author":"D Gur","year":"2008","unstructured":"Gur D, Bandos AI, Cohen CS, Hakim CM, Hardesty LA, Ganott MA, Perrin RL, Poller WR, Shah R, Sumkin JH, et al.The \u201claboratory\u201d effect: comparing radiologists\u2019 performance and variability during prospective clinical and laboratory mammography interpretations. Radiology. 2008; 249(1):47\u201353.","journal-title":"Radiology"},{"issue":"Suppl 2","key":"1224_CR53","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1136\/bmjqs-2012-001615","volume":"22","author":"ML Graber","year":"2013","unstructured":"Graber ML. The incidence of diagnostic error in medicine. BMJ Qual Saf. 2013; 22(Suppl 2):21\u201327.","journal-title":"BMJ Qual Saf"},{"key":"1224_CR54","doi-asserted-by":"crossref","unstructured":"Oakden-Rayner L, Dunnmon J, Carneiro G, R\u00e9 C. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. arXiv preprint arXiv:1909.12475. 2019.","DOI":"10.1145\/3368555.3384468"},{"key":"1224_CR55","unstructured":"Campagner A, Ciucci D, Svensson C-M, Figge MT, Cabitza F. Ground truthing from multi-rater labelling with three-way decisions and possibility theory. In: Cambiare rivista in Information Sciences. Elsevier: 2019."},{"key":"1224_CR56","doi-asserted-by":"crossref","unstructured":"Svensson C-M, Figge MT, H\u00fcbler R. Automated classification of circulating tumor cells and the impact of interobsever variability on classifier training and performance. J Immunol Res. 2015; 2015. https:\/\/pubmed-ncbi-nlm-nih-gov.proxy.unimib.it\/26504857\/.","DOI":"10.1155\/2015\/573165"},{"key":"1224_CR57","unstructured":"Chatterjee S. Learning and memorization. In: International Conference on Machine Learning. PMLR: 2018. p. 755\u201363."},{"key":"1224_CR58","volume-title":"Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing","author":"V Feldman","year":"2020","unstructured":"Feldman V. Does learning require memorization? a short tale about a long tail. In: Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing. New York: Association for Computing Machinery: 2020. p. 954\u20139."},{"issue":"6","key":"1224_CR59","doi-asserted-by":"crossref","first-page":"1435","DOI":"10.1002\/hbm.24886","volume":"41","author":"B Heinrichs","year":"2019","unstructured":"Heinrichs B, Eickhoff SB. Your evidence? machine learning algorithms for medical diagnosis and prediction. Hum Brain Mapp. 2019; 41(6):1435\u201344.","journal-title":"Hum Brain Mapp"},{"issue":"12","key":"1224_CR60","doi-asserted-by":"crossref","first-page":"3799","DOI":"10.1007\/s11229-014-0616-x","volume":"192","author":"C Kelp","year":"2015","unstructured":"Kelp C. Understanding phenomena. Synthese. 2015; 192(12):3799\u2013816.","journal-title":"Synthese"},{"key":"1224_CR61","unstructured":"Lipton ZC, Steinhardt J. Troubling trends in machine learning scholarship. arXiv preprint arXiv:1807.03341. 2018."},{"key":"1224_CR62","unstructured":"Sculley D, Snoek J, Wiltschko A, Rahimi A. Winner\u2019s curse? on pace, progress, and empirical rigor. OpenReview. 2018."},{"key":"1224_CR63","volume-title":"Machines We Trust - Getting Along with Artificial Intelligence","author":"F Cabitza","year":"2020","unstructured":"Cabitza F. Cobra ai: exploring some unintended consequences of artificial intelligence In: Pelillo M, Scantamburlo T, editors. Machines We Trust - Getting Along with Artificial Intelligence. Boston, MA, USA: MIT Press: 2020. p. 93\u2013112. Chap. 6."},{"key":"1224_CR64","volume-title":"Reliability Data Collection and Analysis","author":"D Dubois","year":"1992","unstructured":"Dubois D, Prade H. On the combination of evidence in various mathematical frameworks. In: Reliability Data Collection and Analysis. Dordrecht: Springer: 1992. p. 213\u201341."},{"key":"1224_CR65","volume-title":"Representation, propagation, and aggregation of uncertainty. Technical report","author":"S Ferson","year":"2002","unstructured":"Ferson S, Kreinovich V. Representation, propagation, and aggregation of uncertainty. Technical report. Albuquerque: Sandia National Laboratories; 2002."},{"key":"1224_CR66","doi-asserted-by":"crossref","first-page":"1134","DOI":"10.1145\/1968.1972","volume":"27","author":"LG Valiant","year":"1984","unstructured":"Valiant LG. A theory of the learnable. Commun. ACM. 1984; 27:1134\u20131142. https:\/\/doi-org.proxy.unimib.it\/10.1145\/1968.1972.","journal-title":"Commun. ACM"},{"issue":"4","key":"1224_CR67","first-page":"343","volume":"2","author":"D Angluin","year":"1988","unstructured":"Angluin D, Laird P. Learning from noisy examples. Mach Learn. 1988; 2(4):343\u201370.","journal-title":"Mach Learn"},{"key":"1224_CR68","doi-asserted-by":"crossref","first-page":"264","DOI":"10.1137\/1116025","volume":"17","author":"V N. Vapnik","year":"1971","unstructured":"N. Vapnik V, Ya. Chervonenkis A. On the uniform convergence of relative frequencies of events to their probabilities. Theor Probab Applicactions. 1971; 17:264\u201380.","journal-title":"Theor Probab Applicactions"},{"key":"1224_CR69","volume-title":"Probabilistic graphical models: principles and techniques","author":"D Koller","year":"2009","unstructured":"Koller D, Friedman N. Probabilistic graphical models: principles and techniques. Cambridge: MIT press; 2009."},{"key":"1224_CR70","unstructured":"Ramshaw L, Tarjan RE. On minimum-cost assignments in unbalanced bipartite graphs. Labs, HP, Palo Alto, CA, USA, Tech. Rep. HPL-2012-40R1. 2012."},{"key":"1224_CR71","unstructured":"Mastin A, Jaillet P. Greedy online bipartite matching on random graphs. arXiv preprint arXiv:1307.2536. 2013."}],"container-title":["BMC Medical Informatics and Decision Making"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-020-01224-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12911-020-01224-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-020-01224-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,7]],"date-time":"2023-10-07T06:33:28Z","timestamp":1696660408000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcmedinformdecismak.biomedcentral.com\/articles\/10.1186\/s12911-020-01224-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,11]]},"references-count":71,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["1224"],"URL":"https:\/\/doi.org\/10.1186\/s12911-020-01224-9","relation":{},"ISSN":["1472-6947"],"issn-type":[{"value":"1472-6947","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,9,11]]},"assertion":[{"value":"16 January 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 August 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 September 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"No permissions were required to use the MRNet dataset (all rights reserved by the Stanford University School of Medicine) as this was used according to the Research Use Agreement. We obtained the consent to publish the original annotations of this dataset by the involved clinicians. These latter ones are mentioned in the acknowledgements section below.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"219"}}