{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T06:16:16Z","timestamp":1783664176189,"version":"3.55.0"},"reference-count":171,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T00:00:00Z","timestamp":1734652800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T00:00:00Z","timestamp":1734652800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003065","name":"University of Vienna","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100003065","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Machine learning (ML) and artificial intelligence (AI) approaches are often criticized for their inherent bias and for their lack of control, accountability, and transparency. Consequently, regulatory bodies struggle with containing this technology\u2019s potential negative side effects. High-level requirements such as fairness and robustness need to be formalized into concrete specification metrics, imperfect proxies that capture isolated aspects of the underlying requirements. Given possible trade-offs between different metrics and their vulnerability to over-optimization, integrating specification metrics in system development processes is not trivial. This paper defines\n                    <jats:italic>specification overfitting<\/jats:italic>\n                    , a scenario where systems focus excessively on specified metrics to the detriment of high-level requirements and task performance. We present an extensive literature survey to categorize how researchers propose, measure, and optimize specification metrics in several AI fields (e.g., natural language processing, computer vision, reinforcement learning). Using a keyword-based search on papers from major AI conferences and journals between 2018 and mid-2023, we identify and analyze 74 papers that propose or optimize specification metrics. We find that although most papers implicitly address specification overfitting (e.g., by reporting more than one specification metric), they rarely discuss which role specification metrics should play in system development or explicitly define the scope and assumptions behind metric formulations.\n                  <\/jats:p>","DOI":"10.1007\/s10462-024-11040-6","type":"journal-article","created":{"date-parts":[[2024,12,19]],"date-time":"2024-12-19T22:57:19Z","timestamp":1734649039000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Specification overfitting in artificial intelligence"],"prefix":"10.1007","volume":"58","author":[{"given":"Benjamin","family":"Roth","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pedro Henrique","family":"Luz de Araujo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuxi","family":"Xia","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Saskia","family":"Kaltenbrunner","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christoph","family":"Korab","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,12,20]]},"reference":[{"key":"11040_CR1","doi-asserted-by":"publisher","first-page":"109669","DOI":"10.1016\/j.knosys.2022.109669","volume":"254","author":"V Ace\u00f1a","year":"2022","unstructured":"Ace\u00f1a V, Mart\u00edn de Diego I, et al (2022) Minimally overfitted learners: a general framework for ensemble learning. Knowl-Based Syst 254:109669. https:\/\/doi.org\/10.1016\/j.knosys.2022.109669","journal-title":"Knowl-Based Syst"},{"key":"11040_CR2","unstructured":"Angwin J, Jeff Larson SM, Kirchner L (2016) Machine bias. Pro Publica https:\/\/www.propublica.org\/article\/machine-bias-risk-assessments-in-criminal-sentencing"},{"issue":"6","key":"11040_CR3","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1109\/MSP.2017.2743240","volume":"34","author":"K Arulkumaran","year":"2017","unstructured":"Arulkumaran K, Deisenroth MP, Brundage M et al (2017) Deep reinforcement learning: a brief survey. IEEE Signal Process Mag 34(6):26\u201338. https:\/\/doi.org\/10.1109\/MSP.2017.2743240","journal-title":"IEEE Signal Process Mag"},{"key":"11040_CR4","unstructured":"Bahdanau D, Cho K, Bengio Y (2015) Neural machine translation by jointly learning to align and translate. In: Bengio Y, LeCun Y (eds) 3rd International conference on learning representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, http:\/\/arxiv.org\/abs\/1409.0473"},{"key":"11040_CR5","first-page":"671","volume":"104","author":"S Barocas","year":"2016","unstructured":"Barocas S, Selbst AD (2016) Big data\u2019s disparate impact. Calif L Rev 104:671","journal-title":"Calif L Rev"},{"key":"11040_CR6","unstructured":"Barocas S, Hardt M, Narayanan A (2019) Fairness and machine learning: limitations and opportunities. fairmlbook.org, http:\/\/www.fairmlbook.org"},{"key":"11040_CR7","doi-asserted-by":"publisher","unstructured":"Bartolo M, Thrush T, Jia R, et al (2021) Improving question answering model robustness with synthetic adversarial data generation. In: Moens MF, Huang X, Specia L, et al (eds) Proceedings of the 2021 conference on empirical methods in natural language processing. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, pp. 8830\u20138848, https:\/\/doi.org\/10.18653\/v1\/2021.emnlp-main.696","DOI":"10.18653\/v1\/2021.emnlp-main.696"},{"issue":"1","key":"11040_CR8","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1007\/s10994-009-5152-4","volume":"79","author":"S Ben-David","year":"2010","unstructured":"Ben-David S, Blitzer J, Crammer K et al (2010) A theory of learning from different domains. Mach Learn 79(1):151\u2013175. https:\/\/doi.org\/10.1007\/s10994-009-5152-4","journal-title":"Mach Learn"},{"issue":"1","key":"11040_CR9","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1007\/s42786-020-00020-3","volume":"4","author":"S Bhatore","year":"2020","unstructured":"Bhatore S, Mohan L, Reddy YR (2020) Machine learning techniques for credit risk evaluation: a systematic literature review. J Bank Financ Technol 4(1):111\u2013138. https:\/\/doi.org\/10.1007\/s42786-020-00020-3","journal-title":"J Bank Financ Technol"},{"key":"11040_CR10","doi-asserted-by":"publisher","unstructured":"Black E, Yeom S, Fredrikson M (2020) FlipTest: fairness testing via optimal transport. In: Proceedings of the 2020 conference on fairness, accountability, and transparency. ACM, Barcelona Spain, pp 111\u2013121, https:\/\/doi.org\/10.1145\/3351095.3372845","DOI":"10.1145\/3351095.3372845"},{"issue":"3","key":"11040_CR11","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1007\/s11948-023-00443-3","volume":"29","author":"H Bleher","year":"2023","unstructured":"Bleher H, Braun M (2023) Reflections on putting AI ethics into practice: how three AI ethics approaches conceptualize theory and practice. Sci Eng Ethics 29(3):21. https:\/\/doi.org\/10.1007\/s11948-023-00443-3","journal-title":"Sci Eng Ethics"},{"key":"11040_CR12","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2022.3229161","author":"V Borisov","year":"2022","unstructured":"Borisov V, Leemann T, Se\u00dfler K et al (2022) Deep neural networks and tabular data: a survey. IEEE Trans Neural Networks Learning Syst. https:\/\/doi.org\/10.1109\/TNNLS.2022.3229161","journal-title":"IEEE Trans Neural Networks Learning Syst"},{"issue":"4","key":"11040_CR13","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1109\/MSP.2017.2693418","volume":"34","author":"MM Bronstein","year":"2017","unstructured":"Bronstein MM, Bruna J, LeCun Y et al (2017) Geometric deep learning: going beyond euclidean data. IEEE Signal Process Mag 34(4):18\u201342. https:\/\/doi.org\/10.1109\/MSP.2017.2693418","journal-title":"IEEE Signal Process Mag"},{"key":"11040_CR14","first-page":"31871","volume":"35","author":"D Buffelli","year":"2022","unstructured":"Buffelli D, Li\u00f3 P, Vandin F (2022) SizeShiftReg: a regularization method for improving size-generalization in graph neural networks. Adv Neural Inf Process Syst 35:31871\u201331885","journal-title":"Adv Neural Inf Process Syst"},{"key":"11040_CR15","unstructured":"Buolamwini J, Gebru T (2018) Gender shades: intersectional accuracy disparities in commercial gender classification. In: Friedler SA, Wilson C (eds) Proceedings of the 1st conference on fairness, accountability and transparency, Proceedings of Machine Learning Research, vol\u00a081. PMLR, pp 77\u201391, https:\/\/proceedings.mlr.press\/v81\/buolamwini18a.html"},{"key":"11040_CR16","doi-asserted-by":"publisher","unstructured":"Chen J, Kallus N, Mao X, et al (2019) Fairness under unawareness: assessing disparity when protected class is unobserved. In: Proceedings of the conference on fairness, accountability, and transparency. Association for Computing Machinery, New York, NY, USA, FAT* \u201919, pp. 339\u2013348, https:\/\/doi.org\/10.1145\/3287560.3287594","DOI":"10.1145\/3287560.3287594"},{"key":"11040_CR17","unstructured":"Chen Y, Zhou K, Bian Y, et al (2022) Pareto invariant risk minimization: towards mitigating the optimization dilemma in out-of-distribution generalization. In: The eleventh international conference on learning representations"},{"key":"11040_CR18","doi-asserted-by":"publisher","unstructured":"Cheng M, Wei W, Hsieh CJ (2019) Evaluating and enhancing the robustness of dialogue systems: a case study on a negotiation agent. In: Burstein J, Doran C, Solorio T (eds) Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, pp. 3325\u20133335, https:\/\/doi.org\/10.18653\/v1\/N19-1336","DOI":"10.18653\/v1\/N19-1336"},{"issue":"04","key":"11040_CR19","doi-asserted-by":"publisher","first-page":"3601","DOI":"10.1609\/aaai.v34i04.5767","volume":"34","author":"M Cheng","year":"2020","unstructured":"Cheng M, Yi J, Chen PY et al (2020) Seq2Sick: Evaluating the robustness of sequence-to-sequence models with adversarial examples. Proc AAAI Conf Artif Intell 34(04):3601\u20133608. https:\/\/doi.org\/10.1609\/aaai.v34i04.5767","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"11040_CR20","doi-asserted-by":"publisher","unstructured":"Cheng M, Lei Q, Chen PY, et al (2022) CAT: customized adversarial training for improved robustness. In: Raedt LD (ed) Proceedings of the thirty-first international joint conference on artificial intelligence, IJCAI 2022, Vienna, Austria, 23-29 July 2022. ijcai.org, pp 673\u2013679, https:\/\/doi.org\/10.24963\/IJCAI.2022\/95","DOI":"10.24963\/IJCAI.2022\/95"},{"key":"11040_CR21","unstructured":"Clark K, Luong MT, Le QV, et al (2020) ELECTRA: pre-training text encoders as discriminators rather than generators. In: International conference on learning representations, https:\/\/openreview.net\/forum?id=r1xMH1BtvB"},{"key":"11040_CR22","unstructured":"Clarysse J, H\u00f6rrmann J, Yang F (2022) Why adversarial training can hurt robust accuracy. In: The eleventh international conference on learning representations"},{"key":"11040_CR23","doi-asserted-by":"publisher","unstructured":"Coston A, Mishler A, Kennedy EH, et al (2020) Counterfactual risk assessments, evaluation, and fairness. In: Proceedings of the 2020 conference on fairness, accountability, and transparency. Association for Computing Machinery, New York, NY, USA, FAT* \u201920, pp 582\u2013593, https:\/\/doi.org\/10.1145\/3351095.3372851, arXiv:1909.00066","DOI":"10.1145\/3351095.3372851"},{"key":"11040_CR24","unstructured":"Cotter A, Gupta M, Jiang H, et al (2019) Training well-generalizing classifiers for fairness metrics and other data-dependent constraints. In: Proceedings of the 36th international conference on machine learning. PMLR, pp 1397\u20131405"},{"key":"11040_CR25","unstructured":"Croce F, Andriushchenko M, Sehwag V, et al (2021) RobustBench: a standardized adversarial robustness benchmark. In: Thirty-fifth conference on neural information processing systems datasets and benchmarks track (Round 2), arXiv:2010.09670"},{"issue":"1","key":"11040_CR26","first-page":"10237","volume":"23","author":"A D\u2019Amour","year":"2022","unstructured":"D\u2019Amour A, Heller K, Moldovan D et al (2022) Underspecification presents challenges for credibility in modern machine learning. J Mach Learn Res 23(1):10237\u201310297","journal-title":"J Mach Learn Res"},{"key":"11040_CR27","doi-asserted-by":"crossref","unstructured":"Dapello J, Kar K, Schrimpf M, et al (2022) Aligning model and macaque inferior temporal cortex representations improves model-to-human behavioral alignment and adversarial robustness. In: The eleventh international conference on learning representations. Cold Spring Harbor Laboratory, pp 2022\u201307","DOI":"10.1101\/2022.07.01.498495"},{"issue":"1","key":"11040_CR28","doi-asserted-by":"publisher","first-page":"512","DOI":"10.1609\/icwsm.v11i1.14955","volume":"11","author":"T Davidson","year":"2017","unstructured":"Davidson T, Warmsley D, Macy M et al (2017) Automated hate speech detection and the problem of offensive language. Proc Int AAAI Conf Web Soc Media 11(1):512\u2013515","journal-title":"Proc Int AAAI Conf Web Soc Media"},{"key":"11040_CR29","doi-asserted-by":"publisher","first-page":"1066","DOI":"10.1162\/tacl_a_00590","volume":"11","author":"PH Luz de Araujo","year":"2023","unstructured":"Luz de Araujo PH, Roth B (2023) Cross-functional analysis of generalization in behavioral learning. Trans Assoc Comput Linguist 11:1066\u20131081. https:\/\/doi.org\/10.1162\/tacl_a_00590","journal-title":"Trans Assoc Comput Linguist"},{"key":"11040_CR30","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, et al (2009) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition, IEEE, pp 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"11040_CR31","unstructured":"Deng Z, Zhang J, Zhang L, et al (2023) FIFA: making fairness more generalizable in classifiers trained on imbalanced data. In: The eleventh international conference on learning representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, https:\/\/openreview.net\/pdf?id=zVrw4OH1Lch"},{"key":"11040_CR32","unstructured":"Devlin J, Chang MW, Lee K, et al (2019) BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 conference of the north american chapter of the association for computational linguistics: human language technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, pp 4171\u20134186, https:\/\/www.aclweb.org\/anthology\/N19-1423"},{"key":"11040_CR33","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, et al (2021) An image is worth 16 x 16 words: transformers for image recognition at scale. In: International conference on learning representations, https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"key":"11040_CR34","doi-asserted-by":"publisher","unstructured":"Elkahky A, Webster K, Andor D, et al (2018) A challenge set and methods for noun-verb ambiguity. In: Proceedings of the 2018 conference on empirical methods in natural language processing. Association for Computational Linguistics, Brussels, Belgium, pp 2562\u20132572, https:\/\/doi.org\/10.18653\/v1\/D18-1277","DOI":"10.18653\/v1\/D18-1277"},{"issue":"1","key":"11040_CR35","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1038\/s41746-020-00376-2","volume":"4","author":"A Esteva","year":"2021","unstructured":"Esteva A, Chou K, Yeung S et al (2021) Deep learning-enabled medical computer vision. npj Digit Med 4(1):5. https:\/\/doi.org\/10.1038\/s41746-020-00376-2","journal-title":"npj Digit Med"},{"key":"#cr-split#-11040_CR36.1","unstructured":"European Parliament and Council of the European Union (2012) Regulation"},{"key":"#cr-split#-11040_CR36.2","unstructured":"(EU) No 1025\/2012 of the European Parliament and of the Council of 25 October 2012 on European standardisation. https:\/\/eur-lex.europa.eu\/eli\/reg\/2012\/1025\/oj"},{"key":"11040_CR37","unstructured":"European Parliament and Council of the European Union (2022) Proposal for a Directive of the European Parliament and of the Council on adapting non-contractual civil liability rules to artificial intelligence (AI Liability Directive). https:\/\/eur-lex.europa.eu\/legal-content\/EN\/TXT\/?uri=CELEX"},{"key":"11040_CR38","unstructured":"European Parliament and Council of the European Union (2024) Regulation of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial intelligence Act)"},{"key":"11040_CR39","doi-asserted-by":"publisher","unstructured":"Fan W, Ma Y, Li Q, et al (2019) Graph neural networks for social recommendation. In: The world wide web conference. Association for Computing Machinery, New York, NY, USA, WWW \u201919, p 417-426, https:\/\/doi.org\/10.1145\/3308558.3313488","DOI":"10.1145\/3308558.3313488"},{"key":"11040_CR40","doi-asserted-by":"publisher","unstructured":"Fatemi Z, Xing C, Liu W, et al (2023) Improving gender fairness of pre-trained language models without catastrophic forgetting. In: Rogers A, Boyd-Graber J, Okazaki N (eds) Proceedings of the 61st annual meeting of the association for computational linguistics (Vol 2: Short Papers). Association for Computational Linguistics, Toronto, Canada, pp 1249\u20131262, https:\/\/doi.org\/10.18653\/v1\/2023.acl-short.108","DOI":"10.18653\/v1\/2023.acl-short.108"},{"key":"11040_CR41","volume-title":"Principled artificial intelligence: mapping consensus in ethical and rights-based approaches to principles for AI","author":"J Fjeld","year":"2020","unstructured":"Fjeld J, Achten N, Hilligoss H et al (2020) Principled artificial intelligence: mapping consensus in ethical and rights-based approaches to principles for AI. Berkman Klein Center Research Publication, Cambridge"},{"key":"11040_CR42","doi-asserted-by":"publisher","DOI":"10.1609\/icwsm.v12i1.14991","author":"A Founta","year":"2018","unstructured":"Founta A, Djouvas C, Chatzakou D et al (2018) Large scale crowdsourcing and characterization of twitter abusive behavior. Proc Int AAAI Conf Web Soc Media. https:\/\/doi.org\/10.1609\/icwsm.v12i1.14991","journal-title":"Proc Int AAAI Conf Web Soc Media"},{"issue":"4","key":"11040_CR43","doi-asserted-by":"publisher","first-page":"136","DOI":"10.1145\/3433949","volume":"64","author":"SA Friedler","year":"2021","unstructured":"Friedler SA, Scheidegger C, Venkatasubramanian S (2021) The (im)possibility of fairness: different value systems require different mechanisms for fair decision making. Commun ACM 64(4):136\u2013143. https:\/\/doi.org\/10.1145\/3433949","journal-title":"Commun ACM"},{"issue":"4","key":"11040_CR44","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1007\/BF00344251","volume":"36","author":"K Fukushima","year":"1980","unstructured":"Fukushima K (1980) Neocognitron: a self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biol Cybern 36(4):193\u2013202. https:\/\/doi.org\/10.1007\/BF00344251","journal-title":"Biol Cybern"},{"key":"11040_CR45","doi-asserted-by":"publisher","unstructured":"Gan WC, Ng HT (2019) Improving the robustness of question answering systems to question paraphrasing. In: Proceedings of the 57th annual meeting of the association for computational linguistics. Association for Computational Linguistics, Florence, Italy, pp 6065\u20136075, https:\/\/doi.org\/10.18653\/v1\/P19-1610","DOI":"10.18653\/v1\/P19-1610"},{"key":"11040_CR46","doi-asserted-by":"publisher","unstructured":"Gardner M, Artzi Y, Basmov V, et al (2020) Evaluating models\u2019 local decision boundaries via contrast sets. In: Findings of the association for computational linguistics: EMNLP 2020. Association for Computational Linguistics, Online, pp 1307\u20131323, https:\/\/doi.org\/10.18653\/v1\/2020.findings-emnlp.117","DOI":"10.18653\/v1\/2020.findings-emnlp.117"},{"key":"11040_CR47","unstructured":"Geirhos R, Rubisch P, Michaelis C, et al (2018) ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In: International conference on learning representations"},{"key":"11040_CR48","unstructured":"Goodfellow I, Shlens J, Szegedy C (2015) Explaining and harnessing adversarial examples. In: International confserence on learning representations, arXiv:1412.6572"},{"key":"11040_CR49","first-page":"4218","volume":"34","author":"S Gowal","year":"2021","unstructured":"Gowal S, Rebuffi SA, Wiles O et al (2021) Improving robustness using generated data. Adv Neural Inform Proc Syst 34:4218\u20134233","journal-title":"Adv Neural Inform Proc Syst"},{"key":"11040_CR50","unstructured":"Guo Y, Zhang C, Zhang C, et al (2018) Sparse DNNs with improved adversarial robustness. In: Advances in neural information processing systems, vol\u00a031. Curran Associates, Inc."},{"key":"11040_CR51","doi-asserted-by":"publisher","unstructured":"Guo J, Chen Y, Hao Y, et al (2022) Towards comprehensive testing on the robustness of cooperative multi-agent reinforcement learning. In: IEEE\/CVF conference on computer vision and pattern recognition workshops, CVPR Workshops 2022, New Orleans, LA, USA, June 19\u201320, 2022. IEEE, pp 114\u2013121, https:\/\/doi.org\/10.1109\/CVPRW56347.2022.00022","DOI":"10.1109\/CVPRW56347.2022.00022"},{"issue":"1","key":"11040_CR52","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1007\/s11023-020-09517-8","volume":"30","author":"T Hagendorff","year":"2020","unstructured":"Hagendorff T (2020) The ethics of AI ethics: an evaluation of guidelines. Mind Mach 30(1):99\u2013120. https:\/\/doi.org\/10.1007\/s11023-020-09517-8","journal-title":"Mind Mach"},{"issue":"1","key":"11040_CR53","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","volume":"45","author":"K Han","year":"2023","unstructured":"Han K, Wang Y, Chen H et al (2023) A survey on vision transformer. IEEE Trans Pattern Anal Mach Intell 45(1):87\u2013110. https:\/\/doi.org\/10.1109\/TPAMI.2022.3152247","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11040_CR54","unstructured":"Hardt M, Price E, Srebro N (2016) Equality of opportunity in supervised learning. Adv Neural Inform Proc Syst 29"},{"key":"11040_CR55","unstructured":"Havasi M, Jenatton R, Fort S, et al (2020) Training independent subnetworks for robust prediction. In: International conference on learning representations"},{"issue":"3","key":"11040_CR56","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1109\/TPAMI.2005.55","volume":"27","author":"X He","year":"2005","unstructured":"He X, Yan S, Hu Y et al (2005) Face recognition using Laplacianfaces. IEEE Trans Pattern Anal Mach Intell 27(3):328\u2013340. https:\/\/doi.org\/10.1109\/TPAMI.2005.55","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11040_CR57","doi-asserted-by":"publisher","unstructured":"He K, Zhang X, Ren S, et al (2016) Deep residual learning for image recognition. In: 2016 IEEE conference on computer vision and pattern recognition (CVPR), pp 770\u2013778, https:\/\/doi.org\/10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"11040_CR58","unstructured":"Hendrycks D, Dietterich T (2018) Benchmarking neural network robustness to common corruptions and perturbations. In: International conference on learning representations"},{"key":"11040_CR59","unstructured":"Hendrycks D, Mazeika M, Kadavath S, et al (2019a) Using self-supervised learning can improve model robustness and uncertainty. In: Advances in neural information processing systems, vol\u00a032. Curran Associates, Inc"},{"key":"11040_CR60","unstructured":"Hendrycks D, Mu N, Cubuk ED, et al (2019b) AugMix: a simple data processing method to improve robustness and uncertainty. In: International conference on learning representations"},{"key":"11040_CR61","unstructured":"High-Level Expert Group on AI (2019) Ethics guidelines for trustworthy AI, High-Level Expert Group on AI. https:\/\/digital-strategy.ec.europa.eu\/en\/library\/ethics-guidelines-trustworthy-ai"},{"key":"11040_CR62","doi-asserted-by":"crossref","unstructured":"Hu Y, Yang J, Chen L, et al (2023) Planning-oriented autonomous driving. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR52729.2023.01712"},{"key":"11040_CR63","doi-asserted-by":"publisher","DOI":"10.1007\/s43681-023-00353-x","author":"R Iniesta","year":"2023","unstructured":"Iniesta R (2023) The human role to guarantee an ethical AI in healthcare: a five-facts approach. AI Ethics. https:\/\/doi.org\/10.1007\/s43681-023-00353-x","journal-title":"AI Ethics"},{"key":"11040_CR64","doi-asserted-by":"publisher","unstructured":"Jackson M (1995) The world and the machine. In: Proceedings of the 17th international conference on software engineering. Association for Computing Machinery, New York, NY, USA, ICSE \u201995, pp 283-292, https:\/\/doi.org\/10.1145\/225014.225041","DOI":"10.1145\/225014.225041"},{"key":"11040_CR65","doi-asserted-by":"publisher","unstructured":"Jacobs AZ, Wallach H (2021) Measurement and fairness. In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. Association for Computing Machinery, New York, NY, USA, FAccT \u201921, p 375-385, https:\/\/doi.org\/10.1145\/3442188.3445901","DOI":"10.1145\/3442188.3445901"},{"key":"11040_CR66","doi-asserted-by":"publisher","unstructured":"Jakubovitz D, Giryes R (2018) Improving DNN robustness to adversarial attacks using jacobian regularization. In: Ferrari V, Hebert M, Sminchisescu C, et al (eds) Computer vision\u2014ECCV 2018. Springer, Cham, Lecture Notes in Computer Science, pp 525\u2013541, https:\/\/doi.org\/10.1007\/978-3-030-01258-8_32","DOI":"10.1007\/978-3-030-01258-8_32"},{"key":"11040_CR67","doi-asserted-by":"publisher","DOI":"10.1145\/3571730","author":"Z Ji","year":"2023","unstructured":"Ji Z, Lee N, Frieske R et al (2023) Survey of hallucination in natural language generation. ACM Comput Surv. https:\/\/doi.org\/10.1145\/3571730","journal-title":"ACM Comput Surv"},{"issue":"9","key":"11040_CR68","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1038\/s42256-019-0088-2","volume":"1","author":"A Jobin","year":"2019","unstructured":"Jobin A, Ienca M, Vayena E (2019) The global landscape of AI ethics guidelines. Nature Mach Intell 1(9):389\u2013399","journal-title":"Nature Mach Intell"},{"key":"11040_CR69","unstructured":"Jung S, Park T, Chun S, et al (2022) Re-weighting based group fairness regularization via classwise robust optimization. In: The eleventh international conference on learning representations"},{"key":"11040_CR70","doi-asserted-by":"publisher","unstructured":"Karpukhin V, Levy O, Eisenstein J, et al (2019) Training on synthetic noise improves robustness to natural noise in machine translation. In: Proceedings of the 5th workshop on noisy user-generated text (W-NUT 2019). Association for Computational Linguistics, Hong Kong, China, pp 42\u201347, https:\/\/doi.org\/10.18653\/v1\/D19-5506","DOI":"10.18653\/v1\/D19-5506"},{"key":"11040_CR71","doi-asserted-by":"publisher","first-page":"102274","DOI":"10.1016\/j.lindif.2023.102274","volume":"103","author":"E Kasneci","year":"2023","unstructured":"Kasneci E, Sessler K, K\u00fcchemann S et al (2023) ChatGPT for good? On opportunities and challenges of large language models for education. Learn Individ Differ 103:102274. https:\/\/doi.org\/10.1016\/j.lindif.2023.102274","journal-title":"Learn Individ Differ"},{"key":"11040_CR72","doi-asserted-by":"publisher","unstructured":"Kiden S, Stahl B, Townsend B, et al (2024) Responsible AI governance: a response to UN interim report on governing AI for humanity. https:\/\/eprints.soton.ac.uk\/488908\/, https:\/\/doi.org\/10.5258\/SOTON\/PP0057","DOI":"10.5258\/SOTON\/PP0057"},{"key":"11040_CR73","unstructured":"Kirichenko P, Izmailov P, Wilson AG (2022) Last layer re-training is sufficient for robustness to spurious correlations. In: The eleventh international conference on learning representations"},{"key":"11040_CR74","doi-asserted-by":"publisher","unstructured":"Kirk H, Vidgen B, Rottger P, et al (2022) Hatemoji: a test suite and adversarially-generated dataset for benchmarking and detecting emoji-based hate. In: Proceedings of the 2022 conference of the north american chapter of the association for computational linguistics: human language technologies. Association for Computational Linguistics, Seattle, United States, pp 1352\u20131368, https:\/\/doi.org\/10.18653\/v1\/2022.naacl-main.97","DOI":"10.18653\/v1\/2022.naacl-main.97"},{"key":"11040_CR75","doi-asserted-by":"publisher","unstructured":"Kleinberg J, Mullainathan S, Raghavan M (2017) Inherent trade-offs in the fair determination of risk scores. In: Papadimitriou CH (ed) 8th innovations in theoretical computer science conference (ITCS 2017), Leibniz International Proceedings in Informatics (LIPIcs), vol\u00a067. Schloss Dagstuhl-Leibniz-Zentrum fuer Informatik, pp.43, https:\/\/doi.org\/10.4230\/LIPIcs.ITCS.2017.43","DOI":"10.4230\/LIPIcs.ITCS.2017.43"},{"key":"11040_CR76","unstructured":"Komiyama J, Takeda A, Honda J, et al (2018) Nonconvex optimization for regression with fairness constraints. In: Proceedings of the 35th international conference on machine learning. PMLR, pp 2737\u20132746"},{"issue":"1","key":"11040_CR77","doi-asserted-by":"publisher","first-page":"89","DOI":"10.1016\/S0933-3657(01)00077-X","volume":"23","author":"I Kononenko","year":"2001","unstructured":"Kononenko I (2001) Machine learning for medical diagnosis: history, state of the art and perspective. Artif Intell Med 23(1):89\u2013109. https:\/\/doi.org\/10.1016\/S0933-3657(01)00077-X","journal-title":"Artif Intell Med"},{"key":"11040_CR78","unstructured":"Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: Pereira F, Burges C, Bottou L, et al (eds) Advances in neural information processing systems, vol\u00a025. Curran Associates, Inc., https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2012\/file\/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf"},{"key":"11040_CR79","unstructured":"Lake B, Baroni M (2018) Generalization without systematicity: on the compositional skills of sequence-to-sequence recurrent networks. In: Dy J, Krause A (eds) Proceedings of the 35th international conference on machine learning, Proceedings of machine learning research, vol\u00a080. PMLR, pp 2873\u20132882, https:\/\/proceedings.mlr.press\/v80\/lake18a.html"},{"key":"11040_CR80","doi-asserted-by":"publisher","unstructured":"Lample G, Ballesteros M, Subramanian S, et al (2016) Neural architectures for named entity recognition. In: Knight K, Nenkova A, Rambow O (eds) Proceedings of the 2016 conference of the North American Chapter of the association for computational linguistics: human language technologies. Association for Computational Linguistics, San Diego, California, pp 260\u2013270, https:\/\/doi.org\/10.18653\/v1\/N16-1030,","DOI":"10.18653\/v1\/N16-1030"},{"issue":"4","key":"11040_CR81","doi-asserted-by":"publisher","first-page":"541","DOI":"10.1162\/neco.1989.1.4.541","volume":"1","author":"Y LeCun","year":"1989","unstructured":"LeCun Y, Boser B, Denker JS et al (1989) Backpropagation applied to handwritten zip code recognition. Neural Comput 1(4):541\u2013551. https:\/\/doi.org\/10.1162\/neco.1989.1.4.541","journal-title":"Neural Comput"},{"key":"11040_CR82","doi-asserted-by":"publisher","unstructured":"Lee J, Kim G, Olfat M, et al (2022) Fast and efficient MMD-based fair PCA via optimization over stiefel manifold. In: Proceedings of the AAAI conference on artificial intelligence, pp 7363\u20137371, https:\/\/doi.org\/10.1609\/aaai.v36i7.20699,","DOI":"10.1609\/aaai.v36i7.20699"},{"issue":"39","key":"11040_CR83","first-page":"1","volume":"17","author":"S Levine","year":"2016","unstructured":"Levine S, Finn C, Darrell T et al (2016) End-to-end training of deep visuomotor policies. J Mach Learn Res 17(39):1\u201340","journal-title":"J Mach Learn Res"},{"key":"11040_CR84","doi-asserted-by":"publisher","unstructured":"Levy I, Bogin B, Berant J (2023) Diverse demonstrations improve in-context compositional generalization. In: Rogers A, Boyd-Graber JL, Okazaki N (eds) Proceedings of the 61st annual meeting of the association for computational linguistics (Vol 1: Long Papers). Association for Computational Linguistics, Toronto, Canada, pp 1401\u20131422, https:\/\/doi.org\/10.18653\/v1\/2023.acl-long.78","DOI":"10.18653\/v1\/2023.acl-long.78"},{"key":"11040_CR85","doi-asserted-by":"crossref","unstructured":"Ley M (2002) The DBLP computer science bibliography: evolution, research issues, perspectives. In: International symposium on string processing and information retrieval, Springer, pp 1\u201310","DOI":"10.1007\/3-540-45735-6_1"},{"key":"11040_CR86","unstructured":"Li Q, Guo Y, Zuo W, et al (2023a) Squeeze training for adversarial robustness. In: The eleventh international conference on learning representations, https:\/\/openreview.net\/forum?id=Z_tmYu060Kr"},{"key":"11040_CR87","doi-asserted-by":"publisher","unstructured":"Li X, Wu P, Su J (2023b) Accurate fairness: improving individual fairness without trading accuracy. In: Proceedings of the thirty-seventh AAAI conference on artificial intelligence and thirty-fifth conference on innovative applications of artificial intelligence and thirteenth symposium on educational advances in artificial intelligence, AAAI\u201923\/IAAI\u201923\/EAAI\u201923, vol\u00a037. AAAI Press, pp 14312\u201314320, https:\/\/doi.org\/10.1609\/aaai.v37i12.26674","DOI":"10.1609\/aaai.v37i12.26674"},{"key":"11040_CR88","unstructured":"Liang Y, Sun Y, Zheng R, et al (2022) Efficient adversarial training without attacking: worst-case-aware robust reinforcement learning. In: Advances in neural information processing systems"},{"key":"11040_CR89","doi-asserted-by":"publisher","unstructured":"Lin S, Hilton J, Evans O (2022) TruthfulQA: measuring how models mimic human falsehoods. In: Muresan S, Nakov P, Villavicencio A (eds) Proceedings of the 60th annual meeting of the association for computational linguistics (Vol 1: Long Papers). Association for Computational Linguistics, Dublin, Ireland, pp 3214\u20133252, https:\/\/doi.org\/10.18653\/v1\/2022.acl-long.229,","DOI":"10.18653\/v1\/2022.acl-long.229"},{"key":"11040_CR90","doi-asserted-by":"publisher","unstructured":"Liu NF, Schwartz R, Smith NA (2019a) Inoculation by fine-tuning: a method for analyzing challenge datasets. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, Vol 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, pp 2171\u20132179, https:\/\/doi.org\/10.18653\/v1\/N19-1225","DOI":"10.18653\/v1\/N19-1225"},{"key":"11040_CR91","unstructured":"Liu Y, Ott M, Goyal N, et al (2019b) RoBERTa: a robustly optimized BERT pretraining approach. CoRR abs\/1907.11692. arXiv:1907.11692"},{"key":"11040_CR92","unstructured":"Liu EZ, Haghgoo B, Chen AS, et al (2021) Just train twice: improving group robustness without training group information. In: Proceedings of the 38th international conference on machine learning. PMLR, pp 6781\u20136792"},{"key":"11040_CR93","unstructured":"L\u00fctjens B, Everett M, How JP (2020) Certified adversarial robustness for deep reinforcement learning. In: Kaelbling LP, Kragic D, Sugiura K (eds) Proceedings of the conference on robot learning, Proceedings of machine learning research, vol 100. PMLR, pp 1328\u20131337, https:\/\/proceedings.mlr.press\/v100\/lutjens20a.html"},{"key":"11040_CR94","doi-asserted-by":"crossref","unstructured":"Ma M, Ren J, Zhao L, et al (2022) Are multimodal transformers robust to missing modality? In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 18177\u201318186","DOI":"10.1109\/CVPR52688.2022.01764"},{"key":"11040_CR95","doi-asserted-by":"publisher","unstructured":"Madaio MA, Stark L, Wortman Vaughan J, et al (2020) Co-designing checklists to understand organizational challenges and opportunities around fairness in AI. In: Proceedings of the 2020 CHI conference on human factors in computing systems. Association for Computing Machinery, New York, NY, USA, CHI \u201920, p 1\u201314https:\/\/doi.org\/10.1145\/3313831.3376445","DOI":"10.1145\/3313831.3376445"},{"key":"11040_CR96","unstructured":"Madras D, Pitassi T, Zemel R (2018) Predict responsibly: improving fairness and accuracy by learning to defer. In: Advances in neural information processing systems, vol\u00a031. Curran Associates, Inc"},{"key":"11040_CR97","unstructured":"Malik MM (2020) A hierarchy of limitations in machine learning. CoRR arXiv:2002.05193"},{"key":"11040_CR98","doi-asserted-by":"publisher","DOI":"10.1145\/3457607","author":"N Mehrabi","year":"2021","unstructured":"Mehrabi N, Morstatter F, Saxena N et al (2021) A survey on bias and fairness in machine learning. ACM Comput Surv. https:\/\/doi.org\/10.1145\/3457607","journal-title":"ACM Comput Surv"},{"key":"11040_CR99","doi-asserted-by":"publisher","unstructured":"Min J, McCoy RT, Das D, et al (2020) Syntactic data augmentation increases robustness to inference heuristics. In: Proceedings of the 58th annual meeting of the association for computational linguistics. Association for Computational Linguistics, Online, pp 2339\u20132352, https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.212","DOI":"10.18653\/v1\/2020.acl-main.212"},{"issue":"7","key":"11040_CR100","doi-asserted-by":"publisher","first-page":"3523","DOI":"10.1109\/TPAMI.2021.3059968","volume":"44","author":"S Minaee","year":"2022","unstructured":"Minaee S, Boykov Y, Porikli F et al (2022) Image segmentation using deep learning: a survey. IEEE Trans Pattern Anal Mach Intell 44(7):3523\u20133542. https:\/\/doi.org\/10.1109\/TPAMI.2021.3059968","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11040_CR101","doi-asserted-by":"publisher","unstructured":"Mishler A, Kennedy EH, Chouldechova A (2021) Fairness in risk assessment instruments: post-processing to achieve counterfactual equalized odds. In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. Association for Computing Machinery, New York, NY, USA, FAccT \u201921, pp 386\u2013400, https:\/\/doi.org\/10.1145\/3442188.3445902","DOI":"10.1145\/3442188.3445902"},{"issue":"7540","key":"11040_CR102","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D et al (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529\u2013533. https:\/\/doi.org\/10.1038\/nature14236","journal-title":"Nature"},{"key":"11040_CR103","unstructured":"Naik A, Ravichander A, Sadeh N, et al (2018) Stress test evaluation for natural language inference. In: Proceedings of the 27th international conference on computational linguistics. Association for Computational Linguistics, Santa Fe, New Mexico, USA, pp 2340\u20132353"},{"key":"11040_CR104","doi-asserted-by":"publisher","unstructured":"Nangia N, Vania C, Bhalerao R, et al (2020) CrowS-Pairs: a challenge dataset for measuring social biases in masked language models. In: Webber B, Cohn T, He Y, et al (eds) Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP). Association for Computational Linguistics, Online, pp 1953\u20131967, https:\/\/doi.org\/10.18653\/v1\/2020.emnlp-main.154, https:\/\/aclanthology.org\/2020.emnlp-main.154","DOI":"10.18653\/v1\/2020.emnlp-main.154"},{"key":"11040_CR105","unstructured":"Narasimhan H, Cotter A, Gupta M (2019) Optimizing generalized rate metrics with three players. In: Advances in neural information processing systems, vol\u00a032. Curran Associates, Inc"},{"key":"11040_CR106","unstructured":"OECD (2019) OECD AI Principles Overview. OECD AI Policy Observatory https:\/\/oecd.ai\/en\/ai-principles"},{"key":"11040_CR107","first-page":"27730","volume":"35","author":"L Ouyang","year":"2022","unstructured":"Ouyang L, Wu J, Jiang X et al (2022) Training language models to follow instructions with human feedback. Adv Neural Inform Proc Syst 35:27730\u201327744","journal-title":"Adv Neural Inform Proc Syst"},{"issue":"3","key":"11040_CR108","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3494672","volume":"55","author":"D Pessach","year":"2023","unstructured":"Pessach D, Shmueli E (2023) A review on fairness in machine learning. ACM Comput Surv 55(3):1\u201344. https:\/\/doi.org\/10.1145\/3494672","journal-title":"ACM Comput Surv"},{"key":"11040_CR109","doi-asserted-by":"publisher","unstructured":"Petersen E, Ganz M, Holm S, et al (2023) On (assessing) the fairness of risk score models. In: Proceedings of the 2023 ACM conference on fairness, accountability, and transparency. Association for Computing Machinery, New York, NY, USA, FAccT \u201923, pp 817\u2013829, https:\/\/doi.org\/10.1145\/3593013.3594045","DOI":"10.1145\/3593013.3594045"},{"key":"11040_CR110","doi-asserted-by":"publisher","unstructured":"Pfohl S, Xu Y, Foryciarz A, et al (2022a) Net benefit, calibration, threshold selection, and training objectives for algorithmic fairness in healthcare. In: Proceedings of the 2022 ACM conference on fairness, accountability, and transparency. Association for Computing Machinery, New York, NY, USA, FAccT \u201922, pp 1039\u20131052, https:\/\/doi.org\/10.1145\/3531146.3533166","DOI":"10.1145\/3531146.3533166"},{"issue":"1","key":"11040_CR111","doi-asserted-by":"publisher","first-page":"3254","DOI":"10.1038\/s41598-022-07167-7","volume":"12","author":"SR Pfohl","year":"2022","unstructured":"Pfohl SR, Zhang H, Xu Y et al (2022) A comparison of approaches to improve worst-case predictive model performance over patient subpopulations. Sci Rep 12(1):3254","journal-title":"Sci Rep"},{"key":"11040_CR112","doi-asserted-by":"publisher","unstructured":"Qiu L, Shaw P, Pasupat P, et al (2022) Improving compositional generalization with latent structure and data augmentation. In: Proceedings of the 2022 conference of the North American chapter of the association for computational linguistics: human language technologies. Association for Computational Linguistics, Seattle, United States, pp 4341\u20134362, https:\/\/doi.org\/10.18653\/v1\/2022.naacl-main.323","DOI":"10.18653\/v1\/2022.naacl-main.323"},{"key":"11040_CR113","unstructured":"Radford A, Kim JW, Hallacy C, et al (2021) Learning transferable visual models from natural language supervision. In: Meila M, Zhang T (eds) Proceedings of the 38th international conference on machine learning, proceedings of machine learning research, vol 139. PMLR, pp 8748\u20138763, https:\/\/proceedings.mlr.press\/v139\/radford21a.html"},{"issue":"140","key":"11040_CR114","first-page":"1","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel C, Shazeer N, Roberts A et al (2020) Exploring the limits of transfer learning with a unified text-to-text transformer. J Mach Learn Res 21(140):1\u201367","journal-title":"J Mach Learn Res"},{"key":"11040_CR115","doi-asserted-by":"publisher","unstructured":"Rahmattalabi A, Jabbari S, Lakkaraju H, et al (2021) Fair influence maximization: a welfare optimization approach. In: Proceedings of the AAAI conference on artificial intelligence, pp 11630\u201311638, https:\/\/doi.org\/10.1609\/aaai.v35i13.17383,","DOI":"10.1609\/aaai.v35i13.17383"},{"key":"11040_CR116","unstructured":"Ramesh A, Pavlov M, Goh G, et al (2021) Zero-shot text-to-image generation. In: Meila M, Zhang T (eds) Proceedings of the 38th international conference on machine learning, proceedings of machine learning research, vol 139. PMLR, pp 8821\u20138831, https:\/\/proceedings.mlr.press\/v139\/ramesh21a.html"},{"key":"11040_CR117","unstructured":"Rebuffi SA, Gowal S, Calian DA, et al (2021) Data augmentation can improve robustness. In: Advances in neural information processing systems, arXiv:2111.05328"},{"key":"11040_CR118","doi-asserted-by":"crossref","unstructured":"Ribeiro MT, Singh S, Guestrin C (2016) Why should I trust you?: Explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, ACM, pp 1135\u20131144","DOI":"10.1145\/2939672.2939778"},{"key":"11040_CR119","doi-asserted-by":"publisher","unstructured":"Ribeiro MT, Wu T, Guestrin C, et al (2020) Beyond accuracy: behavioral testing of NLP models with checkList. In: Proceedings of the 58th annual meeting of the association for computational linguistics. Association for Computational Linguistics, Online, pp 4902\u20134912, https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.442","DOI":"10.18653\/v1\/2020.acl-main.442"},{"key":"11040_CR120","unstructured":"Roelofs R, Shankar V, Recht B, et al (2019) A meta-analysis of overfitting in machine learning. Adv Neural Inform Proc Syst 32"},{"key":"11040_CR121","unstructured":"Roh Y, Lee K, Whang SE, et al (2021) Sample selection for fair and robust training. Adv Neural Inform Proc Syst"},{"key":"11040_CR122","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11504","author":"A Ross","year":"2018","unstructured":"Ross A, Doshi-Velez F (2018) Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients. Proceedings AAAI Conf Artif Intell. https:\/\/doi.org\/10.1609\/aaai.v32i1.11504","journal-title":"Proceedings AAAI Conf Artif Intell"},{"key":"11040_CR123","doi-asserted-by":"publisher","unstructured":"R\u00f6ttger P, Vidgen B, Nguyen D, et al (2021) HateCheck: functional tests for hate speech detection models. In: Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (Volume 1: Long Papers). Association for Computational Linguistics, Online, pp 41\u201358, https:\/\/doi.org\/10.18653\/v1\/2021.acl-long.4","DOI":"10.18653\/v1\/2021.acl-long.4"},{"key":"11040_CR124","first-page":"19861","volume":"33","author":"L Ruis","year":"2020","unstructured":"Ruis L, Andreas J, Baroni M et al (2020) A benchmark for systematic generalization in grounded language understanding. Adv Neural Inform Proc Syst 33:19861\u201319872","journal-title":"Adv Neural Inform Proc Syst"},{"issue":"3","key":"11040_CR125","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky O, Deng J, Su H et al (2015) ImageNet large scale visual recognition challenge. Int J Comput Vision 115(3):211\u2013252. https:\/\/doi.org\/10.1007\/s11263-015-0816-y","journal-title":"Int J Comput Vision"},{"key":"11040_CR126","doi-asserted-by":"publisher","first-page":"1408","DOI":"10.1162\/tacl_a_00434","volume":"9","author":"T Schick","year":"2021","unstructured":"Schick T, Udupa S, Sch\u00fctze H (2021) Self-diagnosis and self-debiasing: a proposal for reducing corpus-based bias in NLP. Trans Assoc Comput Linguist 9:1408\u20131424. https:\/\/doi.org\/10.1162\/tacl_a_00434","journal-title":"Trans Assoc Comput Linguist"},{"key":"11040_CR127","first-page":"11539","volume":"33","author":"S Schneider","year":"2020","unstructured":"Schneider S, Rusak E, Eck L et al (2020) Improving robustness against common corruptions by covariate shift adaptation. Adv Neural Inform Proc Syst 33:11539\u201311551","journal-title":"Adv Neural Inform Proc Syst"},{"key":"11040_CR128","unstructured":"Sehwag V, Mahloujifar S, Handina T, et al (2022) Robust learning meets generative models: can proxy distributions improve adversarial robustness? In: The tenth international conference on learning representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net"},{"key":"11040_CR129","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9781107298019","volume-title":"Understanding machine learning: from theory to algorithms","author":"S Shalev-Shwartz","year":"2014","unstructured":"Shalev-Shwartz S, Ben-David S (2014) Understanding machine learning: from theory to algorithms. Cambridge University Press, Cambridge"},{"key":"11040_CR130","unstructured":"Sinha A, Namkoong H, Duchi J (2019) Certifying some distributional robustness with principled adversarial training. In: International conference on learning representations"},{"key":"11040_CR131","unstructured":"Skalse JMV, Howe NHR, Krasheninnikov D, et al (2022) Defining and characterizing reward gaming. In: Oh AH, Agarwal A, Belgrave D, et al (eds) Advances in neural information processing systems, https:\/\/openreview.net\/forum?id=yb3HOXO3lX2"},{"key":"11040_CR132","doi-asserted-by":"crossref","unstructured":"Socher R, Perelygin A, Wu J, et al (2013) Recursive deep models for semantic compositionality over a sentiment treebank. In: Yarowsky D, Baldwin T, Korhonen A, et al (eds) Proceedings of the 2013 conference on empirical methods in natural language processing. Association for Computational Linguistics, Seattle, Washington, USA, pp 1631\u20131642, https:\/\/aclanthology.org\/D13-1170","DOI":"10.18653\/v1\/D13-1170"},{"issue":"1","key":"11040_CR133","doi-asserted-by":"publisher","first-page":"48","DOI":"10.1186\/s40537-019-0212-5","volume":"6","author":"G Sreenu","year":"2019","unstructured":"Sreenu G, Saleem Durai MA (2019) Intelligent video surveillance: a review through deep learning techniques for crowd analysis. Journal of Big Data 6(1):48. https:\/\/doi.org\/10.1186\/s40537-019-0212-5","journal-title":"Journal of Big Data"},{"key":"11040_CR134","unstructured":"Sun Y, Wang X, Liu Z, et al (2020) Test-time training with self-supervision for generalization under distribution shifts. In: Proceedings of the 37th international conference on machine learning. PMLR, pp 9229\u20139248"},{"issue":"8","key":"11040_CR135","doi-asserted-by":"publisher","first-page":"7693","DOI":"10.1109\/TKDE.2022.3201243","volume":"35","author":"L Sun","year":"2023","unstructured":"Sun L, Dou Y, Yang C et al (2023) Adversarial attack and defense on graph data: a survey. IEEE Trans Knowl Data Eng 35(8):7693\u20137711. https:\/\/doi.org\/10.1109\/TKDE.2022.3201243","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"11040_CR136","volume-title":"Reinforcement learning: an introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton RS, Barto AG (2018) Reinforcement learning: an introduction. A Bradford Book, Cambridge"},{"key":"11040_CR137","doi-asserted-by":"publisher","unstructured":"Szegedy C, Vanhoucke V, Ioffe S, et al (2016) Rethinking the inception architecture for computer vision. In: 2016 IEEE conference on computer vision and pattern recognition (CVPR), pp 2818\u20132826, https:\/\/doi.org\/10.1109\/CVPR.2016.308","DOI":"10.1109\/CVPR.2016.308"},{"key":"11040_CR138","doi-asserted-by":"publisher","unstructured":"Taskesen B, Blanchet J, Kuhn D, et al (2021) A statistical test for probabilistic fairness. In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. Association for Computing Machinery, New York, NY, USA, FAccT \u201921, pp 648\u2013665, https:\/\/doi.org\/10.1145\/3442188.3445927","DOI":"10.1145\/3442188.3445927"},{"key":"11040_CR139","unstructured":"The European Comission (2003) General guidelines for the cooperation between CEN, Cenelec and ETSI and the European Commission and the European Free Trade Association. https:\/\/eur-lex.europa.eu\/legal-content\/EN\/ALL\/?uri=CELEX:52003XC0416(03)"},{"key":"11040_CR140","unstructured":"The European Comission (2008) New legislative framework. https:\/\/single-market-economy.ec.europa.eu\/single-market\/goods\/new-legislative-framework_en"},{"key":"11040_CR141","unstructured":"The European Comission (2018) Factsheet: artificial intelligence for Europe. https:\/\/digital-strategy.ec.europa.eu\/en\/library\/ethics-guidelines-trustworthy-ai"},{"issue":"8","key":"11040_CR142","doi-asserted-by":"publisher","first-page":"1930","DOI":"10.1038\/s41591-023-02448-8","volume":"29","author":"AJ Thirunavukarasu","year":"2023","unstructured":"Thirunavukarasu AJ, Ting DSJ, Elangovan K et al (2023) Large language models in medicine. Nat Med 29(8):1930\u20131940. https:\/\/doi.org\/10.1038\/s41591-023-02448-8","journal-title":"Nat Med"},{"key":"11040_CR143","unstructured":"Tjeng V, Xiao KY, Tedrake R (2019) Evaluating robustness of neural networks with mixed integer programming. In: 7th International conference on learning representations, ICLR 2019, New Orleans, LA, USA, May 6\u20139, 2019. OpenReview.net"},{"key":"11040_CR144","doi-asserted-by":"publisher","first-page":"621","DOI":"10.1162\/tacl_a_00335","volume":"8","author":"L Tu","year":"2020","unstructured":"Tu L, Lalwani G, Gella S et al (2020) An empirical study on robustness to spurious correlations using pre-trained language models. Trans Assoc Comput Linguist 8:621\u2013633. https:\/\/doi.org\/10.1162\/tacl_a_00335","journal-title":"Trans Assoc Comput Linguist"},{"key":"11040_CR145","unstructured":"Vaswani A, Shazeer N, Parmar N, et al (2017) Attention is all you need. In: Advances in neural information processing systems, pp 5998\u20136008"},{"issue":"4","key":"11040_CR146","doi-asserted-by":"publisher","first-page":"97","DOI":"10.9785\/cri-2021-220402","volume":"22","author":"M Veale","year":"2021","unstructured":"Veale M, Borgesius FZ (2021) Demystifying the draft EU artificial intelligence act-analysing the good, the bad, and the unclear elements of the proposed approach. Comput Law Rev Int 22(4):97\u2013112","journal-title":"Comput Law Rev Int"},{"key":"11040_CR147","first-page":"841","volume":"31","author":"S Wachter","year":"2017","unstructured":"Wachter S, Mittelstadt B, Russell C (2017) Counterfactual explanations without opening the black box: automated decisions and the GDPR. Harv JL & Tech 31:841","journal-title":"Harv JL & Tech"},{"key":"11040_CR148","doi-asserted-by":"publisher","unstructured":"Wang Y, Bansal M (2018) Robust machine comprehension models via adversarial training. In: Proceedings of the 2018 conference of the North American chapter of the association for computational linguistics: human language technologies, Volume 2 (Short Papers). Association for Computational Linguistics, New Orleans, Louisiana, pp 575\u2013581, https:\/\/doi.org\/10.18653\/v1\/N18-2091","DOI":"10.18653\/v1\/N18-2091"},{"key":"11040_CR149","unstructured":"Wang Y, Zou D, Yi J, et al (2019) Improving Adversarial robustness requires revisiting misclassified examples. In: International conference on learning representations"},{"key":"11040_CR150","unstructured":"Wang B, Wang S, Cheng Y, et al (2020a) InfoBERT: improving robustness of language models from an information theoretic perspective. In: International conference on learning representations"},{"key":"11040_CR151","first-page":"5190","volume":"33","author":"S Wang","year":"2020","unstructured":"Wang S, Guo W, Narasimhan H et al (2020b) Robust optimization for fairness with noisy protected groups. Adv Neural Inform Proc Syst 33:5190\u20135203","journal-title":"Adv Neural Inform Proc Syst"},{"key":"11040_CR152","doi-asserted-by":"publisher","unstructured":"Wang T, Sridhar R, Yang D, et al (2022a) Identifying and mitigating spurious correlations for improving robustness in NLP models. In: Findings of the association for computational linguistics: NAACL 2022. Association for Computational Linguistics, Seattle, United States, pp 1719\u20131729, https:\/\/doi.org\/10.18653\/v1\/2022.findings-naacl.130","DOI":"10.18653\/v1\/2022.findings-naacl.130"},{"key":"11040_CR153","doi-asserted-by":"publisher","unstructured":"Wang X, Wang H, Yang D (2022b) Measure and improve robustness in NLP models: a survey. In: Proceedings of the 2022 conference of the North American chapter of the association for computational linguistics: human language technologies. Association for Computational Linguistics, Seattle, United States, pp 4569\u20134586, https:\/\/doi.org\/10.18653\/v1\/2022.naacl-main.339,","DOI":"10.18653\/v1\/2022.naacl-main.339"},{"key":"11040_CR154","doi-asserted-by":"publisher","first-page":"744","DOI":"10.1007\/978-0-387-30164-8_623","volume-title":"Overfitting","author":"GI Webb","year":"2010","unstructured":"Webb GI (2010) Overfitting. Springer, Boston, pp 744\u2013744. https:\/\/doi.org\/10.1007\/978-0-387-30164-8_623"},{"key":"11040_CR155","unstructured":"Wei J, Bosma M, Zhao V, et al (2022a) Finetuned language models are zero-shot learners. In: International conference on learning representations, https:\/\/openreview.net\/forum?id=gEZrGCozdqR"},{"key":"11040_CR156","unstructured":"Wei J, Tay Y, Bommasani R, et al (2022b) Emergent abilities of large language models. Transactions on machine learning research https:\/\/openreview.net\/forum?id=yzkSU5zdwD"},{"key":"11040_CR157","unstructured":"Weng TW, Zhang H, Chen PY, et al (2018) Evaluating the robustness of neural networks: an extreme value theory approach. In: International conference on learning representations (ICLR)"},{"issue":"3410","key":"11040_CR158","doi-asserted-by":"publisher","first-page":"1355","DOI":"10.1126\/science.131.3410.1355","volume":"131","author":"N Wiener","year":"1960","unstructured":"Wiener N (1960) Some moral and technical consequences of automation: as machines learn they may develop unforeseen strategies at rates that baffle their programmers. Science 131(3410):1355\u20131358","journal-title":"Science"},{"key":"11040_CR159","unstructured":"Wu Y, Jiang A, Ba J, et al (2020) INT: an inequality benchmark for evaluating generalization in theorem proving. In: International conference on learning representations"},{"issue":"1","key":"11040_CR160","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1109\/TNNLS.2020.2978386","volume":"32","author":"Z Wu","year":"2021","unstructured":"Wu Z, Pan S, Chen F et al (2021) A comprehensive survey on graph neural networks. IEEE Trans Neural Networks Learn Syst 32(1):4\u201324. https:\/\/doi.org\/10.1109\/TNNLS.2020.2978386","journal-title":"IEEE Trans Neural Networks Learn Syst"},{"key":"11040_CR161","doi-asserted-by":"crossref","unstructured":"Xie C, Wu Y, van der Maaten L, et al (2019) Feature denoising for improving adversarial robustness. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 501\u2013509","DOI":"10.1109\/CVPR.2019.00059"},{"issue":"16","key":"11040_CR162","doi-asserted-by":"publisher","first-page":"8749","DOI":"10.1021\/acs.jmedchem.9b00959","volume":"63","author":"Z Xiong","year":"2020","unstructured":"Xiong Z, Wang D, Liu X et al (2020) Pushing the boundaries of molecular representation for drug discovery with the graph attention mechanism. J Med Chem 63(16):8749\u20138760. https:\/\/doi.org\/10.1021\/acs.jmedchem.9b00959","journal-title":"J Med Chem"},{"key":"11040_CR163","doi-asserted-by":"publisher","unstructured":"Yuan A, Coenen A, Reif E, et al (2022) Wordcraft: story writing with large language models. In: 27th International conference on intelligent user interfaces. Association for Computing Machinery, New York, NY, USA, IUI \u201922, p 841-852, https:\/\/doi.org\/10.1145\/3490099.3511105","DOI":"10.1145\/3490099.3511105"},{"key":"11040_CR164","doi-asserted-by":"publisher","unstructured":"Zan D, Chen B, Zhang F, et al (2023) Large language models meet NL2Code: a survey. In: Rogers A, Boyd-Graber J, Okazaki N (eds) Proceedings of the 61st annual meeting of the association for computational linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Toronto, Canada, pp 7443\u20137464, https:\/\/doi.org\/10.18653\/v1\/2023.acl-long.411,","DOI":"10.18653\/v1\/2023.acl-long.411"},{"key":"11040_CR165","unstructured":"Zhang H, Chen H, Xiao C, et al (2019) Towards stable and efficient training of verifiably robust neural networks. In: International conference on learning representations"},{"key":"11040_CR166","doi-asserted-by":"publisher","DOI":"10.1145\/3374217","author":"WE Zhang","year":"2020","unstructured":"Zhang WE, Sheng QZ, Alhazmi A et al (2020) Adversarial attacks on deep-learning models in natural language processing: a survey. ACM Trans Intell Syst Technol. https:\/\/doi.org\/10.1145\/3374217","journal-title":"ACM Trans Intell Syst Technol"},{"key":"11040_CR167","unstructured":"Zhang C, Tian Y, Ju M, et al (2022a) Chasing all-round graph representation robustness: model, training, and optimization. In: The eleventh international conference on learning representations"},{"key":"11040_CR168","first-page":"38629","volume":"35","author":"M Zhang","year":"2022","unstructured":"Zhang M, Levine S, Finn C (2022) MEMO: test time robustness via adaptation and augmentation. Adv Neural Inf Process Syst 35:38629\u201338642","journal-title":"Adv Neural Inf Process Syst"},{"key":"11040_CR169","unstructured":"Zhang M, Sohoni NS, Zhang HR, et al (2022c) Correct-N-contrast: a contrastive approach for improving robustness to spurious correlations. In: Proceedings of the 39th international conference on machine learning. PMLR, pp 26484\u201326516, arXiv:2203.01517"},{"key":"11040_CR170","doi-asserted-by":"publisher","unstructured":"Zhuo TY, Li Z, Huang Y, et al (2023) On robustness of prompt-based semantic parsing with large pre-trained language model: an empirical study on codex. In: Vlachos A, Augenstein I (eds) Proceedings of the 17th conference of the european chapter of the association for computational linguistics. Association for Computational Linguistics, Dubrovnik, Croatia, pp 1090\u20131102, https:\/\/doi.org\/10.18653\/v1\/2023.eacl-main.77","DOI":"10.18653\/v1\/2023.eacl-main.77"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-024-11040-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-024-11040-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-024-11040-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,29]],"date-time":"2025-01-29T22:46:44Z","timestamp":1738190804000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-024-11040-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,20]]},"references-count":171,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["11040"],"URL":"https:\/\/doi.org\/10.1007\/s10462-024-11040-6","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-4106577\/v1","asserted-by":"object"}]},"ISSN":["1573-7462"],"issn-type":[{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,20]]},"assertion":[{"value":"18 November 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 December 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no competing interests to declare that are relevant to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"35"}}