{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T09:07:59Z","timestamp":1784020079867,"version":"3.55.0"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,9,30]],"date-time":"2024-09-30T00:00:00Z","timestamp":1727654400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Data and Information Quality"],"published-print":{"date-parts":[[2024,9,30]]},"abstract":"<jats:p>Machine learning (ML) and artificial intelligence (AI) systems rely heavily on human-annotated data for training and evaluation. A major challenge in this context is the occurrence of annotation errors, as their effects can degrade model performance. This paper presents a predictive error model trained to detect potential errors in search relevance annotation tasks for three industry-scale ML applications (music streaming, video streaming, and mobile apps). Drawing on data from an extensive search relevance annotation program, we demonstrate that errors can be predicted with moderate model performance (AUC=0.65-0.75) and that model performance generalizes well across applications (i.e., a global, task-agnostic model performs on par with task-specific models). In contrast to past research, which has often focused on predicting annotation labels from task-specific features, our model is trained to predict errors directly from a combination of task features and behavioral features derived from the annotation process, in order to achieve a high degree of generalizability. We demonstrate the usefulness of the model in the context of auditing, where prioritizing tasks with high predicted error probabilities considerably increases the amount of corrected annotation errors (e.g., 40% efficiency gains for the music streaming application). These results highlight that behavioral error detection models can yield considerable improvements in the efficiency and quality of data annotation processes. Our findings reveal critical insights into effective error management in the data annotation process, thereby contributing to the broader field of human-in-the-loop ML.<\/jats:p>","DOI":"10.1145\/3688394","type":"journal-article","created":{"date-parts":[[2024,9,26]],"date-time":"2024-09-26T11:14:21Z","timestamp":1727349261000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Generalizable Error Modeling for Human Data Annotation: Evidence From an Industry-Scale Search Data Annotation Program"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0571-6388","authenticated-orcid":false,"given":"Heinrich","family":"Peters","sequence":"first","affiliation":[{"name":"Apple Inc, New York, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0561-6130","authenticated-orcid":false,"given":"Alireza","family":"Hashemi","sequence":"additional","affiliation":[{"name":"Apple Inc, Seattle, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-3828-4972","authenticated-orcid":false,"given":"James","family":"Rae","sequence":"additional","affiliation":[{"name":"Apple Inc, Seattle, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,10,9]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.142"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/2702123.2702509"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1182"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1609\/aimag.v36i1.2564"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-4802"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-022-28818-3"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-28645-5_29"},{"key":"e_1_3_3_11_2","first-page":"3986","volume-title":"Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916)","author":"Hollenstein Nora","year":"2016","unstructured":"Nora Hollenstein, Nathan Schneider, and Bonnie Webber. 2016. Inconsistency detection in semantic annotation. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916). European Language Resources Association (ELRA), Portoro\u017e, Slovenia, 3986\u20133990. https:\/\/aclanthology.org\/L16-1629"},{"key":"e_1_3_3_12_2","unstructured":"Jan-Christoph Klie Bonnie Webber and Iryna Gurevych. 2022. Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future. http:\/\/arxiv.org\/abs\/2206.02280arXiv:2206.02280 [cs]."},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.3115\/1072228.1072249"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.442"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295230"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1609\/hcomp.v4i1.13287"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","unstructured":"Christopher Meek. 2016. A Characterization of Prediction Errors. 10.48550\/arXiv.1611.05955arXiv:1611.05955 [cs].","DOI":"10.48550\/arXiv.1611.05955"},{"key":"e_1_3_3_18_2","unstructured":"Curtis G. Northcutt Anish Athalye and Jonas Mueller. 2021. Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. https:\/\/arxiv.org\/abs\/2103.14749v4"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","unstructured":"Long Ouyang Jeff Wu Xu Jiang Diogo Almeida Carroll L. Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray John Schulman Jacob Hilton Fraser Kelton Luke Miller Maddie Simens Amanda Askell Peter Welinder Paul Christiano Jan Leike and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. 10.48550\/arXiv.2203.02155arXiv:2203.02155 [cs].","DOI":"10.48550\/arXiv.2203.02155"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijhcs.2022.102772"},{"key":"e_1_3_3_21_2","unstructured":"Verified Market Research. 2023. Data Annotation Service Market Size Share Trends & Forecast. https:\/\/www.verifiedmarketresearch.com\/product\/data-annotation-service-market\/"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.13"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445518"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","unstructured":"Ravid Shwartz-Ziv and Amitai Armon. 2021. Tabular Data: Deep Learning is Not All You Need. 10.48550\/arXiv.2106.03253arXiv:2106.03253 [cs].","DOI":"10.48550\/arXiv.2106.03253"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","unstructured":"Tina Tseng Amanda Stent and Domenic Maida. 2020. Best practices for managing data annotation projects. (2020). 10.13140\/RG.2.2.34497.58727arXiv:2009.11654 [cs].","DOI":"10.13140\/RG.2.2.34497.58727"},{"key":"e_1_3_3_26_2","first-page":"48","volume-title":"Proceedings of the COLING-2000 Workshop on Linguistically Interpreted Corpora","author":"Halteren Hans van","year":"2000","unstructured":"Hans van Halteren. 2000. The detection of inconsistency in manually tagged text. In Proceedings of the COLING-2000 Workshop on Linguistically Interpreted Corpora. International Committee on Computational Linguistics, Centre Universitaire, Luxembourg, 48\u201355. https:\/\/aclanthology.org\/W00-1907"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1150"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.14778\/3055540.3055547"}],"container-title":["Journal of Data and Information Quality"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3688394","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3688394","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:05:56Z","timestamp":1750291556000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3688394"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,30]]},"references-count":27,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,9,30]]}},"alternative-id":["10.1145\/3688394"],"URL":"https:\/\/doi.org\/10.1145\/3688394","relation":{},"ISSN":["1936-1955","1936-1963"],"issn-type":[{"value":"1936-1955","type":"print"},{"value":"1936-1963","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,30]]},"assertion":[{"value":"2023-12-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-21","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}