{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,25]],"date-time":"2026-07-25T18:46:48Z","timestamp":1785005208549,"version":"3.55.0"},"reference-count":91,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2022,2,8]],"date-time":"2022-02-08T00:00:00Z","timestamp":1644278400000},"content-version":"vor","delay-in-days":38,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,1,31]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Majority voting and averaging are common approaches used to resolve annotator disagreements and derive single ground truth labels from multiple annotations. However, annotators may systematically disagree with one another, often reflecting their individual biases and values, especially in the case of subjective tasks such as detecting affect, aggression, and hate speech. Annotator disagreements may capture important nuances in such tasks that are often ignored while aggregating annotations to a single ground truth. In order to address this, we investigate the efficacy of multi-annotator models. In particular, our multi-task based approach treats predicting each annotators\u2019 judgements as separate subtasks, while sharing a common learned representation of the task. We show that this approach yields same or better performance than aggregating labels in the data prior to training across seven different binary classification tasks. Our approach also provides a way to estimate uncertainty in predictions, which we demonstrate better correlate with annotation disagreements than traditional methods. Being able to model uncertainty is especially useful in deployment scenarios where knowing when not to make a prediction is important.<\/jats:p>","DOI":"10.1162\/tacl_a_00449","type":"journal-article","created":{"date-parts":[[2022,2,8]],"date-time":"2022-02-08T14:56:44Z","timestamp":1644332204000},"page":"92-110","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":136,"title":["Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations"],"prefix":"10.1162","volume":"10","author":[{"given":"Aida Mostafazadeh","family":"Davani","sequence":"first","affiliation":[{"name":"University of Southern California, USA. mostafaz@usc.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mark","family":"D\u00edaz","sequence":"additional","affiliation":[{"name":"Google Research, USA. markdiaz@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vinodkumar","family":"Prabhakaran","sequence":"additional","affiliation":[{"name":"Google Research, USA. vinodkpg@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2022,1,31]]},"reference":[{"key":"2022033118510369200_bib1","first-page":"107","article-title":"Subjective natural language problems: Motivations, applications, characterizations, and implications","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies","author":"Alm","year":"2011"},{"key":"2022033118510369200_bib2","unstructured":"Ebba Cecilia Ovesdotter Alm . 2008. Affect in* Text and Speech. Ph.D. thesis, University of Illinois at Urbana-Champaign."},{"key":"2022033118510369200_bib3","doi-asserted-by":"publisher","first-page":"89","DOI":"10.18653\/v1\/W15-2711","article-title":"Predicting word sense annotation agreement","volume-title":"Proceedings of the First Workshop on Linking Computational Models of Lexical, Sentential and Discourse-level Semantics","author":"Alonso","year":"2015"},{"key":"2022033118510369200_bib4","doi-asserted-by":"publisher","first-page":"196","DOI":"10.1007\/978-3-540-74628-7_27","article-title":"Identifying expressions of emotion in text","volume-title":"International Conference on Text, Speech andDialogue","author":"Aman","year":"2007"},{"key":"2022033118510369200_bib5","doi-asserted-by":"publisher","first-page":"4964","DOI":"10.1109\/ICASSP.2018.8461299","article-title":"Soft-target training with ambiguous emotional utterances for dnn-based speech emotion classification","volume-title":"2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Ando","year":"2018"},{"key":"2022033118510369200_bib6","article-title":"Crowd truth: Harnessing disagreement in crowdsourcing a relation extraction gold standard","volume":"2013","author":"Aroyo","year":"2013","journal-title":"WebSci2013. ACM"},{"key":"2022033118510369200_bib7","doi-asserted-by":"publisher","first-page":"1664","DOI":"10.18653\/v1\/D19-1176","article-title":"Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Breitfeller","year":"2019"},{"key":"2022033118510369200_bib8","doi-asserted-by":"publisher","first-page":"578","DOI":"10.18653\/v1\/E17-2092","article-title":"Emobank: Studying the impact of annotation perspective and representation format on dimensional emotion analysis","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers","author":"Buechel","year":"2017"},{"issue":"CSCW","key":"2022033118510369200_bib9","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3359276","article-title":"Crossmod: A cross-community learning-based system to assist reddit moderators","volume":"3","author":"Chandrasekharan","year":"2019","journal-title":"Proceedings of the ACM on human-computer interaction"},{"key":"2022033118510369200_bib10","doi-asserted-by":"publisher","first-page":"105","DOI":"10.1007\/978-3-030-01364-6_12","article-title":"Crowd disagreement about medical images is informative","volume-title":"Intravascular Imaging and Computer Assisted Stenting and Large- scale Annotation of Biomedical Data and ExpertLabel Synthesis","author":"Cheplygina","year":"2018"},{"key":"2022033118510369200_bib11","doi-asserted-by":"crossref","first-page":"5886","DOI":"10.1109\/ICASSP.2019.8682170","article-title":"Every rating matters: Joint learning of subjective labels and individual annotators for speech emotion classification","volume-title":"ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Chou","year":"2019"},{"key":"2022033118510369200_bib12","first-page":"32","article-title":"Modelling annotator bias with multi-task Gaussian processes: An application to machine translation quality estimation","volume-title":"Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Cohn","year":"2013"},{"issue":"2","key":"2022033118510369200_bib13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3377323","article-title":"A multilingual evaluation for online hate speech detection","volume":"20","author":"Corazza","year":"2020","journal-title":"ACM Transactions on Internet Technology (TOIT)"},{"issue":"4","key":"2022033118510369200_bib14","doi-asserted-by":"publisher","first-page":"300","DOI":"10.1111\/1471-6402.00110","article-title":"Empathy, ways of knowing, and interdependence as mediators of gender differences in attitudes toward hate speech and freedom of speech","volume":"27","author":"Cowan","year":"2003","journal-title":"Psychology of Women Quarterly"},{"issue":"1","key":"2022033118510369200_bib15","doi-asserted-by":"publisher","first-page":"69","DOI":"10.1177\/1529100619850176","article-title":"Mapping the passions: Toward a high-dimensional taxonomy of emotional experience and expression","volume":"20","author":"Cowen","year":"2019","journal-title":"Psychological Science in the Public Interest"},{"key":"2022033118510369200_bib16","author":"Crowdflower","year":"2016"},{"key":"2022033118510369200_bib17","doi-asserted-by":"publisher","first-page":"25","DOI":"10.18653\/v1\/W19-3504","article-title":"Racial bias in hate speech and abusive language detection datasets","volume-title":"Proceedings of the Third Workshop on Abusive Language Online","author":"Davidson","year":"2019"},{"key":"2022033118510369200_bib18","doi-asserted-by":"crossref","DOI":"10.1609\/icwsm.v11i1.14955","article-title":"Automated hate speech detection and the problem of offensive language","volume-title":"Proceedings of the International AAAI Conference on Web and Social Media","author":"Davidson","year":"2017"},{"issue":"1","key":"2022033118510369200_bib19","doi-asserted-by":"publisher","first-page":"20","DOI":"10.2307\/2346806","article-title":"Maximum likelihood estimation of observer error-rates using the em algorithm","volume":"28","author":"Dawid","year":"1979","journal-title":"Journal of the Royal Statistical Society: Series C(Applied Statistics)"},{"issue":"2","key":"2022033118510369200_bib20","doi-asserted-by":"publisher","first-page":"301","DOI":"10.1162\/COLI_a_00097","article-title":"Did it happen? The pragmatic complexity of veridicality assessment","volume":"38","author":"Marneffe","year":"2012","journal-title":"Computational Linguistics"},{"key":"2022033118510369200_bib21","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.372","article-title":"GoEmotions: A dataset of fine-grained emotions","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Demszky","year":"2020"},{"issue":"16","key":"2022033118510369200_bib22","doi-asserted-by":"publisher","first-page":"6351","DOI":"10.1016\/j.eswa.2013.05.050","article-title":"Emotion detection in suicide notes","volume":"40","author":"Desmet","year":"2013","journal-title":"Expert Systems with Applications"},{"key":"2022033118510369200_bib23","article-title":"Bert: Pre-training of deep bidirectional transformers for language understanding","volume-title":"NAACL-HLT","author":"Devlin","year":"2019"},{"key":"2022033118510369200_bib24","unstructured":"Mark D\u00edaz . 2020. Biases as Values: Evaluating Algorithms in Context. Ph.D. thesis, Northwestern University."},{"key":"2022033118510369200_bib25","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3173574.3173986","article-title":"Addressing age-related bias in sentiment analysis","volume-title":"Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems","author":"D\u00edaz","year":"2018"},{"key":"2022033118510369200_bib26","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1145\/3278721.3278729","article-title":"Measuring and mitigating unintended bias in text classification","volume-title":"Proceedings of the 2018 AAAI\/ACM Conference on AI, Ethics, and Society","author":"Dixon","year":"2018"},{"key":"2022033118510369200_bib27","doi-asserted-by":"publisher","first-page":"701","DOI":"10.1007\/978-3-319-18818-8_43","article-title":"Crowdsourcing disagreement for collecting semantic annotation","volume-title":"European Semantic Web Conference","author":"Dumitrache","year":"2015"},{"issue":"3\u20134","key":"2022033118510369200_bib28","doi-asserted-by":"publisher","first-page":"169","DOI":"10.1080\/02699939208411068","article-title":"An argument for basic emotions","volume":"6","author":"Ekman","year":"1992","journal-title":"Cognition & Emotion"},{"key":"2022033118510369200_bib29","doi-asserted-by":"publisher","first-page":"566","DOI":"10.1109\/IJCNN.2016.7727250","article-title":"Modeling subjectiveness in emotion recognition with deep neural networks: Ensembles vs soft labels","volume-title":"2016 International Joint Conference on Neural Networks (IJCNN)","author":"Fayek","year":"2016"},{"key":"2022033118510369200_bib30","doi-asserted-by":"publisher","first-page":"2591","DOI":"10.18653\/v1\/2021.naacl-main.204","article-title":"Beyond black & white: Leveraging annotator disagreement via soft-label multi-task learning","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Fornaciari","year":"2021"},{"key":"2022033118510369200_bib31","unstructured":"Gavin Gaffney . 2018. Pushshift gab corpus. https:\/\/files.pushshift.io\/gab\/. Accessed: 2019-5-23."},{"key":"2022033118510369200_bib32","first-page":"1050","article-title":"Dropout as a bayesian approximation: Representing model uncertainty in deep learning","volume-title":"International Conference on Machine Learning","author":"Gal","year":"2016"},{"key":"2022033118510369200_bib33","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.18653\/v1\/D19-1107","article-title":"Are we modeling the task or the annotator? An investigation of annotator bias in natural language understanding datasets","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Geva","year":"2019"},{"key":"2022033118510369200_bib34","doi-asserted-by":"publisher","first-page":"4202","DOI":"10.1109\/ICCVW.2019.00517","article-title":"Characterizing sources of uncertainty to proxy calibration and disambiguate annotator and data bias","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision Workshop (ICCVW)","author":"Ghandeharioun","year":"2019"},{"key":"2022033118510369200_bib35","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445423","article-title":"The disagreement deconvolution: Bringing machine learning performance metrics in line with reality","volume-title":"Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems","author":"Gordon","year":"2021"},{"key":"2022033118510369200_bib36","doi-asserted-by":"publisher","DOI":"10.4324\/9781315648156","volume-title":"Social cognition: How individuals construct social reality","author":"Greifeneder","year":"2017"},{"key":"2022033118510369200_bib37","article-title":"A baseline for detecting misclassified and out-of-distribution examples in neural networks","author":"Hendrycks","year":"2017","journal-title":"Proceedings of International Conference on Learning Representations"},{"key":"2022033118510369200_bib38","article-title":"Experiments in emotional speech","volume-title":"ISCA & IEEE Workshop on Spontaneous Speech Processing and Recognition","author":"Hirschberg","year":"2003"},{"issue":"6245","key":"2022033118510369200_bib39","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1126\/science.aaa8685","article-title":"Advances in natural language processing","volume":"349","author":"Hirschberg","year":"2015","journal-title":"Science"},{"key":"2022033118510369200_bib40","first-page":"1120","article-title":"Learning whom to trust with mace","volume-title":"Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Hovy","year":"2013"},{"key":"2022033118510369200_bib41","doi-asserted-by":"publisher","first-page":"pages 5491\u2013pages 5501","DOI":"10.18653\/v1\/2020.acl-main.487","article-title":"Social biases in NLP models as barriers for persons with disabilities","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Hutchinson","year":"2020"},{"key":"2022033118510369200_bib42","article-title":"Toxic comment classification challenge","author":"Jigsaw","year":"2018"},{"key":"2022033118510369200_bib43","article-title":"Unintended bias in toxicity classification","author":"Jigsaw","year":"2019"},{"key":"2022033118510369200_bib44","doi-asserted-by":"publisher","first-page":"3658","DOI":"10.18653\/v1\/P19-1357","article-title":"A just and comprehensive strategy for using NLP to address online abuse","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Jurgens","year":"2019"},{"key":"2022033118510369200_bib45","doi-asserted-by":"publisher","first-page":"1637","DOI":"10.1145\/2818048.2820016","article-title":"Parting crowds: Characterizing divergent interpretations in crowdsourced annotation tasks","volume-title":"Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work & Social Computing","author":"Kairam","year":"2016"},{"key":"2022033118510369200_bib46","unstructured":"Brendan Kennedy , MohammadAtari, Aida MostafazadehDavani, LeighYeh, AliOmrani, YehsongKim, KrisCoombsJr., ShreyaHavaldar, GwenythPortillo-Wightman, ElaineGonzalez, JoeHoover, AidaAzatian, GabrielCardenas, AlyzehHussain, AustinLara, AdamOmary, ChristinaPark, XinWang, ClarisaWijaya, YongZhang, BethMeyerowitz, and MortezaDehghani. 2020. The gab hate corpus: A collection of 27k posts annotated for hate speech. 10.31234\/osf.io\/hqjxn"},{"key":"2022033118510369200_bib47","article-title":"Adam: A method for stochastic optimization","volume-title":"3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7\u20139, 2015, Conference Track Proceedings","author":"Kingma","year":"2015"},{"key":"2022033118510369200_bib48","doi-asserted-by":"publisher","first-page":"431","DOI":"10.1007\/978-3-319-99229-7_36","article-title":"Uncertainty in machine learning applications: A practice-driven classification of uncertainty","volume-title":"International Conference on Computer Safety, Reliability, and Security","author":"Kl\u00e4s","year":"2018"},{"issue":"2","key":"2022033118510369200_bib49","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1080\/19312458.2011.568376","article-title":"Agreement and information in the reliability of coding","volume":"5","author":"Krippendorff","year":"2011","journal-title":"Communication Methods and Measures"},{"key":"2022033118510369200_bib50","doi-asserted-by":"crossref","DOI":"10.21437\/Eurospeech.2003-306","article-title":"Classifying subject ratings of emotional speech using acoustic features","volume-title":"Eighth European Conference on Speech Communication and Technology","author":"Liscombe","year":"2003"},{"issue":"2010","key":"2022033118510369200_bib51","first-page":"627","article-title":"Sentiment analysis and subjectivity.","volume":"2","author":"Liu","year":"2010","journal-title":"Handbook of Natural Language Processing"},{"key":"2022033118510369200_bib52","doi-asserted-by":"publisher","first-page":"125","DOI":"10.1145\/604045.604067","article-title":"A model of textual affect sensing using real-world knowledge","volume-title":"Proceedings of the 8th International Conference on Intelligent User Interfaces","author":"Liu","year":"2003"},{"key":"2022033118510369200_bib53","article-title":"Human-in-the-loop learning from crowdsourcing and social media","author":"Liu","year":"2020"},{"key":"2022033118510369200_bib54","doi-asserted-by":"crossref","first-page":"4487","DOI":"10.18653\/v1\/P19-1441","article-title":"Multi-task deep neural networks for natural language understanding","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Liu","year":"2019"},{"key":"2022033118510369200_bib55","doi-asserted-by":"publisher","first-page":"3296","DOI":"10.18653\/v1\/2020.findings-emnlp.296","article-title":"Detecting stance in media on global warming","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Luo","year":"2020"},{"key":"2022033118510369200_bib56","first-page":"139","article-title":"A corpus-based approach to finding happiness.","volume-title":"AAAI Spring Symposium: Computational Approaches to Analyzing Weblogs","author":"Mihalcea","year":"2006"},{"key":"2022033118510369200_bib57","article-title":"Tackling online abuse: A survey of automated abuse detection methods","author":"Mishra","year":"2019","journal-title":"arXiv preprint arXiv:1908.06024"},{"key":"2022033118510369200_bib58","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/ACII.2009.5349500","article-title":"Interpreting ambiguous emotional expressions","volume-title":"2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops","author":"Mower","year":"2009"},{"key":"2022033118510369200_bib59","doi-asserted-by":"publisher","first-page":"928","DOI":"10.1007\/978-3-030-36687-2_77","article-title":"A bert-based transfer learning approach for hate speech detection in online social media","volume-title":"International Conference on Complex Networks and Their Applications","author":"Mozafari","year":"2019"},{"key":"2022033118510369200_bib60","doi-asserted-by":"publisher","first-page":"557","DOI":"10.1145\/1743384.1743478","article-title":"How reliable are annotations via crowdsourcing: a study about inter-annotator agreement for multi-label image annotation","volume-title":"Proceedings of the International Conference on Multimedia Information Retrieval","author":"Nowak","year":"2010"},{"key":"2022033118510369200_bib61","doi-asserted-by":"publisher","first-page":"271","DOI":"10.3115\/1218955.1218990","article-title":"A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts","volume-title":"Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04)","author":"Bo","year":"2004"},{"key":"2022033118510369200_bib62","doi-asserted-by":"publisher","first-page":"311","DOI":"10.1162\/tacl_a_00185","article-title":"The benefits of a model of annotation","volume":"2","author":"Passonneau","year":"2014","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022033118510369200_bib63","doi-asserted-by":"publisher","DOI":"10.24251\/HICSS.2019.260","article-title":"Annotating social media data from vulnerable populations: Evaluating disagreement between domain experts and graduate student annotators","volume-title":"Proceedings of the 52nd Hawaii International Conference on System Sciences","author":"Patton","year":"2019"},{"key":"2022033118510369200_bib64","doi-asserted-by":"publisher","first-page":"571","DOI":"10.1162\/tacl_a_00040","article-title":"Comparing bayesian models of annotation","volume":"6","author":"Paun","year":"2018","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022033118510369200_bib65","doi-asserted-by":"publisher","first-page":"742","DOI":"10.3115\/v1\/E14-1078","article-title":"Learning part-of-speech taggers with inter-annotator agreement loss","volume-title":"Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics","author":"Plank","year":"2014"},{"key":"2022033118510369200_bib66","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1016\/B978-0-12-558701-3.50007-7","article-title":"A general psychoevolutionary theory of emotion","volume-title":"Theories of Emotion","author":"Plutchik","year":"1980"},{"key":"2022033118510369200_bib67","doi-asserted-by":"publisher","first-page":"100943","DOI":"10.1109\/ACCESS.2019.2929050","article-title":"Emotion recognition in conversation: Research challenges, datasets, and recent advances","volume":"7","author":"Poria","year":"2019","journal-title":"IEEE Access"},{"key":"2022033118510369200_bib68","first-page":"57","article-title":"Statistical modality tagging from rule-based annotations and crowdsourcing","volume-title":"Proceedings of the Workshop on Extra-Propositional Aspects of Meaning in Computational Linguistics","author":"Prabhakaran","year":"2012"},{"key":"2022033118510369200_bib69","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/2021.law-1.14","article-title":"On releasing annotator-level labels and information in datasets","volume-title":"Proceedings of the 15th Linguistic Annotation Workshop","author":"Prabhakaran","year":"2021"},{"key":"2022033118510369200_bib70","doi-asserted-by":"publisher","first-page":"5740","DOI":"10.18653\/v1\/D19-1578","article-title":"Perturbation sensitivity analysis to detect unintended model biases","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Prabhakaran","year":"2019"},{"key":"2022033118510369200_bib71","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/2020.alw-1.1","article-title":"Online abuse and human rights: WOAH satellite session at RightsCon 2020","volume-title":"Proceedings of the Fourth Workshop on Online Abuse and Harms","author":"Prabhakaran","year":"2020"},{"key":"2022033118510369200_bib72","doi-asserted-by":"publisher","first-page":"114","DOI":"10.18653\/v1\/2020.alw-1.15","article-title":"Six attributes of unhealthy conversations","volume-title":"Proceedings of the Fourth Workshop on Online Abuse and Harms","author":"Price","year":"2020"},{"key":"2022033118510369200_bib73","doi-asserted-by":"publisher","first-page":"842","DOI":"10.21437\/Interspeech.2013-239","article-title":"\u201csure, i did the right thing\u201d: A system for sarcasm detection in speech.","volume-title":"Interspeech","author":"Rakov","year":"2013"},{"key":"2022033118510369200_bib74","doi-asserted-by":"publisher","first-page":"2863","DOI":"10.1145\/1753846.1753873","article-title":"Who are the crowdworkers? Shifting demographics in mechanical turk","volume-title":"CHI\u201910 Extended Abstracts on Human Factors in Computing Systems","author":"Ross","year":"2010"},{"issue":"1","key":"2022033118510369200_bib75","doi-asserted-by":"publisher","first-page":"145","DOI":"10.1037\/0033-295X.110.1.145","article-title":"Core affect and the psychological construction of emotion.","volume":"110","author":"Russell","year":"2003","journal-title":"Psychological Review"},{"key":"2022033118510369200_bib76","first-page":"859","article-title":"Corpus annotation through crowdsourcing: Towards best practice guidelines.","volume-title":"LREC","author":"Sabou","year":"2014"},{"key":"2022033118510369200_bib77","doi-asserted-by":"crossref","first-page":"1668","DOI":"10.18653\/v1\/P19-1163","article-title":"The risk of racial bias in hate speech detection","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Sap","year":"2019"},{"key":"2022033118510369200_bib78","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/W17-1101","article-title":"A survey on hate speech detection using natural language processing","volume-title":"Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media","author":"Schmidt","year":"2017"},{"key":"2022033118510369200_bib79","article-title":"CXPlain: Causal Explanations for Model Interpretation under Uncertainty","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Schwab","year":"2019"},{"key":"2022033118510369200_bib80","doi-asserted-by":"publisher","first-page":"254","DOI":"10.3115\/1613715.1613751","article-title":"Cheap and fast \u2013 but is it good? Evaluating non-expert annotations for natural language tasks","volume-title":"Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing","author":"Snow","year":"2008"},{"key":"2022033118510369200_bib81","doi-asserted-by":"publisher","first-page":"70","DOI":"10.3115\/1621474.1621487","article-title":"Semeval-2007 task 14: Affective text","volume-title":"Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007)","author":"Strapparava","year":"2007"},{"key":"2022033118510369200_bib82","doi-asserted-by":"publisher","first-page":"1667","DOI":"10.18653\/v1\/2021.acl-long.132","article-title":"Learning from the worst: Dynamically generated datasets to improve online hate detection","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Vidgen","year":"2021"},{"key":"2022033118510369200_bib83","first-page":"19","article-title":"Detecting hate speech on the world wide web","volume-title":"Proceedings of the Second Workshop on Language in Social Media","author":"Warner","year":"2012"},{"key":"2022033118510369200_bib84","doi-asserted-by":"publisher","first-page":"138","DOI":"10.18653\/v1\/W16-5618","article-title":"Are you a racist or am i seeing things? Annotator influence on hate speech detection on Twitter","volume-title":"Proceedings of the First Workshop on NLP and Computational Social Science","author":"Waseem","year":"2016"},{"key":"2022033118510369200_bib85","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-3012","article-title":"Understanding abuse: A typology of abusive language detection subtasks","author":"Waseem","year":"2017","journal-title":"arXiv preprint arXiv:1705.09899"},{"key":"2022033118510369200_bib86","doi-asserted-by":"publisher","first-page":"88","DOI":"10.18653\/v1\/N16-2013","article-title":"Hateful symbols or hateful people? Predictive features for hate speech detection on twitter","volume-title":"Proceedings of the NAACL Student Research Workshop","author":"Waseem","year":"2016"},{"key":"2022033118510369200_bib87","doi-asserted-by":"publisher","first-page":"623","DOI":"10.1145\/2441776.2441846","article-title":"Pay by the bit: an information-theoretic metric for collective human judgment","volume-title":"Proceedings of the 2013 Conference on Computer Supported Cooperative Work","author":"Waterhouse","year":"2013"},{"issue":"3","key":"2022033118510369200_bib88","doi-asserted-by":"publisher","first-page":"277","DOI":"10.1162\/0891201041850885","article-title":"Learning subjective language","volume":"30","author":"Wiebe","year":"2004","journal-title":"Computational Linguistics"},{"key":"2022033118510369200_bib89","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"Transformers: State-of-the-art natural language processing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Wolf","year":"2020"},{"key":"2022033118510369200_bib90","doi-asserted-by":"publisher","first-page":"3143","DOI":"10.18653\/v1\/2021.eacl-main.274","article-title":"Challenges in automated debiasing for toxic language detection","author":"Zhou","year":"2021"},{"key":"2022033118510369200_bib91","doi-asserted-by":"crossref","first-page":"127","DOI":"10.18653\/v1\/2020.louhi-1.14","article-title":"Identifying personal experience tweets of medication effects using pre-trained RoBERTa language model and its updating","volume-title":"Proceedings of the 11th International Workshop on Health Text Mining and Information Analysis","author":"Zhu","year":"2020"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00449\/1986597\/tacl_a_00449.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00449\/1986597\/tacl_a_00449.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,26]],"date-time":"2023-01-26T08:15:29Z","timestamp":1674720929000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00449\/109286\/Dealing-with-Disagreements-Looking-Beyond-the"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"references-count":91,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00449","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022]]},"published":{"date-parts":[[2022]]}}}