{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T02:37:28Z","timestamp":1784255848189,"version":"3.55.0"},"reference-count":0,"publisher":"Association for the Advancement of Artificial Intelligence (AAAI)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AIES"],"abstract":"<jats:p>Quantitative Artificial Intelligence (AI) Benchmarks have\nemerged as fundamental tools for evaluating the\nperformance, capability, and safety of AI models and\nsystems. Currently, they shape the direction of AI\ndevelopment and are playing an increasingly prominent role\nin regulatory frameworks. As their influence grows,\nhowever, so too does concerns about how and with what\neffects they evaluate highly sensitive topics such as\ncapabilities, including high-impact capabilities, safety\nand systemic risks. This paper presents an\ninterdisciplinary meta-review of about 110 studies that\ndiscuss shortcomings in quantitative benchmarking\npractices, published in the last 10 years. It brings\ntogether many fine-grained issues in the design and\napplication of benchmarks (such as biases in dataset\ncreation, inadequate documentation, data contamination, and\nfailures to distinguish signal from noise) with broader\nsociotechnical issues (such as an over-focus on evaluating\ntext-based AI models according to one-time testing logic\nthat fails to account for how AI models are increasingly\nmultimodal and interact with humans and other technical\nsystems). Our review also highlights a series of systemic\nflaws in current benchmarking practices, such as misaligned\nincentives, construct validity issues, unknown unknowns,\nand problems with the gaming of benchmark results.\nFurthermore, it underscores how benchmark practices are\nfundamentally shaped by cultural, commercial and\ncompetitive dynamics that often prioritise state-of-the-art\nperformance at the expense of broader societal concerns. By\nproviding an overview of risks associated with existing\nbenchmarking procedures, we problematise disproportionate\ntrust placed in benchmarks and contribute to ongoing\nefforts to improve the accountability and relevance of\nquantitative AI benchmarks within the complexities of\nreal-world scenarios.<\/jats:p>","DOI":"10.1609\/aies.v8i1.36595","type":"journal-article","created":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T13:19:47Z","timestamp":1760534387000},"page":"850-864","source":"Crossref","is-referenced-by-count":18,"title":["Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation"],"prefix":"10.1609","volume":"8","author":[{"given":"Maria","family":"Eriksson","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Erasmo","family":"Purificato","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Arman","family":"Noroozian","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jo\u00e3o","family":"Vinagre","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guillaume","family":"Chaslot","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Emilia","family":"Gomez","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David","family":"Fernandez-Llorca","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"9382","published-online":{"date-parts":[[2025,10,15]]},"container-title":["Proceedings of the AAAI\/ACM Conference on AI, Ethics, and Society"],"original-title":[],"link":[{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/download\/36595\/38733","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/download\/36595\/38733","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T13:19:47Z","timestamp":1760534387000},"score":1,"resource":{"primary":{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/view\/36595"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,15]]},"references-count":0,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,10,15]]}},"URL":"https:\/\/doi.org\/10.1609\/aies.v8i1.36595","relation":{},"ISSN":["3065-8365"],"issn-type":[{"value":"3065-8365","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,15]]}}}