{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T06:24:52Z","timestamp":1783405492823,"version":"3.54.6"},"reference-count":54,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,5,23]],"date-time":"2024-05-23T00:00:00Z","timestamp":1716422400000},"content-version":"vor","delay-in-days":143,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,5,16]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>The annotation of ambiguous or subjective NLP tasks is usually addressed by various annotators. In most datasets, these annotations are aggregated into a single ground truth. However, this omits divergent opinions of annotators, hence missing individual perspectives. We propose FLEAD (Federated Learning for Exploiting Annotators\u2019 Disagreements), a methodology built upon federated learning to independently learn from the opinions of all the annotators, thereby leveraging all their underlying information without relying on a single ground truth. We conduct an extensive experimental study and analysis in diverse text classification tasks to show the contribution of our approach with respect to mainstream approaches based on majority voting and other recent methodologies that also learn from annotator disagreements.<\/jats:p>","DOI":"10.1162\/tacl_a_00664","type":"journal-article","created":{"date-parts":[[2024,5,23]],"date-time":"2024-05-23T17:46:17Z","timestamp":1716486377000},"page":"630-648","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":8,"title":["Federated Learning for Exploiting Annotators\u2019 Disagreements in Natural Language Processing"],"prefix":"10.1162","volume":"12","author":[{"given":"Nuria","family":"Rodr\u00edguez-Barroso","sequence":"first","affiliation":[{"name":"Department of Computer Science and Artificial Intelligence, Andalusian Research Institute in, Data Science and Computational Intelligence (DaSCI), University of Granada, Spain. rbnuria@ugr.es"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eugenio Mart\u00ednez","family":"C\u00e1mara","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Ja\u00e9n, Spain. emcamara@ujaen.es"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jose Camacho","family":"Collados","sequence":"additional","affiliation":[{"name":"Cardiff University, Cardiff, United Kingdom. CamachoColladosJ@cardiff.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"M. Victoria","family":"Luz\u00f3n","sequence":"additional","affiliation":[{"name":"Department of Software Engineering, Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI), University of Granada, Spain. luzon@ugr.es"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Francisco","family":"Herrera","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Artificial Intelligence, Andalusian Research Institute in, Data Science and Computational Intelligence (DaSCI), University of Granada, Spain. fherrera@decsai.ugr.es"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,5,16]]},"reference":[{"issue":"5","key":"2024052819541836100_bib1","doi-asserted-by":"publisher","first-page":"1313","DOI":"10.1109\/TMI.2016.2528120","article-title":"Aggnet: Deep learning from crowds for mitosis detection in breast cancer histology images","volume":"35","author":"Albarqouni","year":"2016","journal-title":"IEEE Transactions on Medical Imaging"},{"issue":"8","key":"2024052819541836100_bib2","doi-asserted-by":"publisher","first-page":"331","DOI":"10.3390\/info12080331","article-title":"A survey on sentiment analysis and opinion mining in Greek social media","volume":"12","author":"Alexandridis","year":"2021","journal-title":"Information"},{"key":"2024052819541836100_bib3","doi-asserted-by":"publisher","DOI":"10.2139\/ssrn.4166108","article-title":"Politics and virality in the time of twitter: A large-scale cross-party sentiment analysis in Greece, Spain and United Kingdom","author":"Antypas","year":"2022","journal-title":"CoRR"},{"key":"2024052819541836100_bib4","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.124","article-title":"Stop measuring calibration when humans disagree","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Baan","year":"2022"},{"key":"2024052819541836100_bib5","doi-asserted-by":"publisher","first-page":"1644","DOI":"10.18653\/v1\/2020.findings-emnlp.148","article-title":"TweetEval: Unified benchmark and comparative evaluation for tweet classification","volume-title":"Proceedings of Findings of EMNLP","author":"Barbieri","year":"2020"},{"key":"2024052819541836100_bib6","first-page":"258","article-title":"XLM-T: Multilingual language models in Twitter for sentiment analysis and beyond","volume-title":"Proceedings of the Thirteenth Language Resources and Evaluation Conference","author":"Barbieri","year":"2022"},{"key":"2024052819541836100_bib7","doi-asserted-by":"publisher","first-page":"441","DOI":"10.1007\/978-3-030-77091-4_26","article-title":"It\u2019s the end of the gold standard as we know it","volume-title":"International Conference of the Italian Association for Artificial Intelligence","author":"Basile","year":"2021"},{"key":"2024052819541836100_bib8","article-title":"Toward a perspectivist turn in ground truthing for predictive computing","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Basile","year":"2023"},{"key":"2024052819541836100_bib9","doi-asserted-by":"publisher","first-page":"15","DOI":"10.18653\/v1\/2021.bppf-1.3","article-title":"We need to consider disagreement in evaluation","volume-title":"Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future","author":"Basile","year":"2021"},{"key":"2024052819541836100_bib10","article-title":"Are we done with imagenet?","author":"Beyer","year":"2020","journal-title":"CoRR"},{"issue":"1","key":"2024052819541836100_bib11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12911-020-01224-9","article-title":"As if sand were stone. New concepts and metrics to probe the ground on which to build trustable AI","volume":"20","author":"Cabitza","year":"2020","journal-title":"BMC Medical Informatics and Decision Making"},{"issue":"3","key":"2024052819541836100_bib12","doi-asserted-by":"publisher","first-page":"475","DOI":"10.1177\/1460458218824705","article-title":"The elephant in the record: On the multiplicity of data recording work","volume":"25","author":"Cabitza","year":"2019","journal-title":"Health Informatics Journal"},{"key":"2024052819541836100_bib13","article-title":"Spanish pre-trained BERT model and evaluation data","volume-title":"Practical Machine Learning for Developing Countries at ICLR 2020","author":"Ca\u00f1ete","year":"2020"},{"key":"2024052819541836100_bib14","doi-asserted-by":"publisher","first-page":"7388","DOI":"10.18653\/v1\/2021.emnlp-main.587","article-title":"ConvAbuse: Data, analysis, and benchmarks for nuanced abuse detection in conversational AI","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Curry","year":"2021"},{"key":"2024052819541836100_bib15","doi-asserted-by":"publisher","first-page":"8772","DOI":"10.18653\/v1\/2020.acl-main.774","article-title":"Uncertain natural language inference","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Chen","year":"2020"},{"key":"2024052819541836100_bib16","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.747","article-title":"Unsupervised cross-lingual representation learning at scale","author":"Conneau","year":"2019","journal-title":"CoRR"},{"key":"2024052819541836100_bib17","doi-asserted-by":"publisher","first-page":"92","DOI":"10.1162\/tacl_a_00449","article-title":"Dealing with disagreements: Looking beyond the majority vote in subjective annotations","volume":"10","author":"Davani","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024052819541836100_bib18","doi-asserted-by":"publisher","first-page":"291","DOI":"10.18653\/v1\/D15-1035","article-title":"Noise or additional information? Leveraging crowdsource annotation item agreement for natural language tasks","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing","author":"Jamison","year":"2015"},{"key":"2024052819541836100_bib19","doi-asserted-by":"publisher","first-page":"1357","DOI":"10.1162\/tacl_a_00523","article-title":"Investigating reasons for disagreement in natural language inference","volume":"10","author":"Jiang","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"1\u20132","key":"2024052819541836100_bib20","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1561\/9781680837896","article-title":"Advances and open problems in federated learning","volume":"14","author":"Kairouz","year":"2021","journal-title":"Foundations and Trends\u00aein Machine Learning"},{"issue":"1","key":"2024052819541836100_bib21","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1007\/s10579-021-09569-x","article-title":"Introducing the Gab Hate Corpus: Defining and applying hate-based rhetoric to social media posts at scale","volume":"56","author":"Kennedy","year":"2022","journal-title":"Language Resources and Evaluation"},{"key":"2024052819541836100_bib22","doi-asserted-by":"publisher","first-page":"1886","DOI":"10.18653\/v1\/N18-1171","article-title":"Sentiment analysis: It\u2019s complicated!","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies)","author":"Kenyon-Dean","year":"2018"},{"issue":"5","key":"2024052819541836100_bib23","doi-asserted-by":"publisher","first-page":"102643","DOI":"10.1016\/j.ipm.2021.102643","article-title":"Offensive, aggressive, and hate speech analysis: From data-centric to human-centered approach","volume":"58","author":"Koco\u0144","year":"2021","journal-title":"Information Processing & Management"},{"key":"2024052819541836100_bib24","doi-asserted-by":"publisher","first-page":"159","DOI":"10.2307\/2529310","article-title":"The measurement of observer agreement for categorical data","author":"Richard Landis","year":"1977","journal-title":"Biometrics"},{"key":"2024052819541836100_bib25","doi-asserted-by":"publisher","first-page":"2304","DOI":"10.18653\/v1\/2023.semeval-1.314","article-title":"SemEval-2023 task 11: Learning with disagreements (LeWiDi)","volume-title":"Proceedings of the The 17th International Workshop on Semantic Evaluation (SemEval-2023)","author":"Leonardelli","year":"2023"},{"key":"2024052819541836100_bib26","doi-asserted-by":"publisher","first-page":"10528","DOI":"10.18653\/v1\/2021.emnlp-main.822","article-title":"Agreeing to disagree: Annotating offensive language datasets with annotators\u2019 disagreement","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Leonardelli","year":"2021"},{"key":"2024052819541836100_bib27","article-title":"Roberta: A robustly optimized BERT pretraining approach","author":"Liu","year":"2019","journal-title":"CoRR"},{"key":"2024052819541836100_bib28","first-page":"13","article-title":"Overview of tass 2018: Opinions, health and emotions","volume-title":"Proceedings of TASS 2018: Workshop on Semantic Analysis at SEPLN (TASS 2018)","author":"Mart\u00ednez-C\u00e1mara","year":"2018"},{"key":"2024052819541836100_bib29","unstructured":"Philip M.\n              McCarthy\n            \n          . 2005. An Assessment of the Range and Usefulness of Lexical Diversity Measures and the Potential of the Measure of Textual, Lexical Diversity (MTLD). Ph.D. thesis, The University of Memphis."},{"key":"2024052819541836100_bib30","first-page":"1273","article-title":"Communication-efficient learning of deep networks from decentralized data","volume-title":"Proceedings of the 20th International Conference on Artificial Intelligence and Statistics","author":"McMahan","year":"2017"},{"issue":"1\u20132","key":"2024052819541836100_bib31","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1561\/9781601981516","article-title":"Opinion mining and sentiment analysis","volume":"2","author":"Bo","year":"2008","journal-title":"Foundations and Trends\u00ae in Information Retrieval"},{"key":"2024052819541836100_bib32","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.eacl-main.130","article-title":"Don\u2019t blame the annotator: Bias already starts in the annotation instructions","author":"Parmar","year":"2022","journal-title":"CoRR"},{"key":"2024052819541836100_bib33","doi-asserted-by":"publisher","first-page":"571","DOI":"10.1162\/tacl_a_00040","article-title":"Comparing Bayesian models of annotation","volume":"6","author":"Paun","year":"2018","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024052819541836100_bib34","doi-asserted-by":"publisher","first-page":"677","DOI":"10.1162\/tacl_a_00293","article-title":"Inherent disagreements in human textual inferences","volume":"7","author":"Pavlick","year":"2019","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024052819541836100_bib35","doi-asserted-by":"publisher","first-page":"9616","DOI":"10.1109\/ICCV.2019.00971","article-title":"Human uncertainty makes classification more robust","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Peterson","year":"2019"},{"key":"2024052819541836100_bib36","doi-asserted-by":"publisher","first-page":"742","DOI":"10.3115\/v1\/E14-1078","article-title":"Learning part-of-speech taggers with inter-annotator agreement loss","volume-title":"Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics","author":"Plank","year":"2014"},{"key":"2024052819541836100_bib37","doi-asserted-by":"publisher","first-page":"507","DOI":"10.3115\/v1\/P14-2083","article-title":"Linguistically debatable or just plain wrong?","volume-title":"Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Plank","year":"2014"},{"key":"2024052819541836100_bib38","first-page":"1492","article-title":"EmoEvent: A multilingual emotion corpus based on different events","volume-title":"Proceedings of the Twelfth Language Resources and Evaluation Conference","author":"del Arco","year":"2020"},{"key":"2024052819541836100_bib39","doi-asserted-by":"publisher","first-page":"8","DOI":"10.3115\/1611628.1611631","article-title":"Exploiting \u2018subjective\u2019 annotations","volume-title":"Proceedings of the Workshop on Human Judgements in Computational Linguistics","author":"Reidsma","year":"2008"},{"issue":"1","key":"2024052819541836100_bib40","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.1609\/aaai.v32i1.11506","article-title":"Deep learning from crowds","volume":"32","author":"Rodrigues","year":"2018","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2024052819541836100_bib41","doi-asserted-by":"publisher","first-page":"957","DOI":"10.1007\/0-387-25465-X_45","article-title":"Ensemble methods for classifiers","volume-title":"Data Mining and Knowledge Discovery Handbook","author":"Rokach","year":"2005"},{"key":"2024052819541836100_bib42","doi-asserted-by":"publisher","first-page":"208","DOI":"10.18653\/v1\/P18-1020","article-title":"Efficient online scalar annotation with bounded support","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Sakaguchi","year":"2018"},{"key":"2024052819541836100_bib43","doi-asserted-by":"publisher","first-page":"2428","DOI":"10.18653\/v1\/2023.eacl-main.178","article-title":"Why don\u2019t you do it right? Analysing annotators\u2019 disagreement in subjective tasks","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics","author":"Sandri","year":"2023"},{"key":"2024052819541836100_bib44","doi-asserted-by":"publisher","first-page":"2420","DOI":"10.18653\/v1\/2023.eacl-main.178","article-title":"Why don\u2019t you do it right? Analysing annotators\u2019 disagreement in subjective tasks","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics","author":"Sandri","year":"2023"},{"key":"2024052819541836100_bib45","article-title":"Distilbert, a distilled version of BERT: Smaller, faster, cheaper and lighter","author":"Sanh","year":"2019","journal-title":"CoRR"},{"key":"2024052819541836100_bib46","doi-asserted-by":"publisher","first-page":"94","DOI":"10.18653\/v1\/2023.semeval-1.12","article-title":"SafeWebUH at SemEval-2023 task 11: Learning annotator disagreement in derogatory text: Comparison of direct training vs aggregation","volume-title":"Proceedings of the The 17th International Workshop on Semantic Evaluation (SemEval-2023)","author":"Shahriar","year":"2023"},{"key":"2024052819541836100_bib47","doi-asserted-by":"publisher","first-page":"614","DOI":"10.1145\/1401890.1401965","article-title":"Get another label? Improving data quality and data mining using multiple, noisy labelers","volume-title":"Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Sheng","year":"2008"},{"key":"2024052819541836100_bib48","doi-asserted-by":"publisher","first-page":"978","DOI":"10.18653\/v1\/2023.semeval-1.135","article-title":"University at Buffalo at SemEval-2023 task 11: MASDA\u2013modelling annotator sensibilities through DisAggregation","volume-title":"Proceedings of the The 17th International Workshop on Semantic Evaluation (SemEval-2023)","author":"Sullivan","year":"2023"},{"key":"2024052819541836100_bib49","doi-asserted-by":"publisher","first-page":"1385","DOI":"10.1613\/jair.1.12752","article-title":"Learning from disagreement: A survey","volume":"72","author":"Uma","year":"2022","journal-title":"Journal Artificial Intelligence Research"},{"key":"2024052819541836100_bib50","article-title":"GSI-UPM at IberLEF2021: Emotion analysis of Spanish tweets by fine-tuning the XLM-RoBERTa language model","volume-title":"Proceedings of the Iberian Languages Evaluation Forum","author":"Vera","year":"2021"},{"key":"2024052819541836100_bib51","doi-asserted-by":"publisher","DOI":"10.3115\/997939.998008","article-title":"Identifying subjective characters in narrative","volume-title":"COLING 1990 Volume 2: Papers presented to the 13th International Conference on Computational Linguistics","author":"Wiebe","year":"1990"},{"key":"2024052819541836100_bib52","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"Transformers: State-of-the-art natural language processing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Wolf","year":"2020"},{"key":"2024052819541836100_bib53","doi-asserted-by":"publisher","first-page":"902","DOI":"10.1609\/icwsm.v17i1.22198","article-title":"Annobert: Effectively representing multiple annotators\u2019 label choices to improve hate speech detection","volume-title":"Proceedings of the International AAAI Conference on Web and Social Media","author":"Yin","year":"2023"},{"key":"2024052819541836100_bib54","article-title":"Federated learning with non-iid data","author":"Zhao","year":"2022","journal-title":"CoRR"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00664\/2374794\/tacl_a_00664.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00664\/2374794\/tacl_a_00664.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,28]],"date-time":"2024-05-28T19:55:11Z","timestamp":1716926111000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00664\/121195\/Federated-Learning-for-Exploiting-Annotators"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":54,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00664","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}