{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T04:27:30Z","timestamp":1781152050315,"version":"3.54.1"},"reference-count":46,"publisher":"Cambridge University Press (CUP)","issue":"6","license":[{"start":{"date-parts":[[2023,8,3]],"date-time":"2023-08-03T00:00:00Z","timestamp":1691020800000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2024,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Anti-Asian speech during the COVID-19 pandemic has been a serious problem with severe consequences. A hate speech wave swept social media platforms. The timely detection of Anti-Asian COVID-19-related hate speech is of utmost importance, not only to allow the application of preventive mechanisms but also to anticipate and possibly prevent other similar discriminatory situations. In this paper, we address the problem of detecting Anti-Asian COVID-19-related hate speech from social media data. Previous approaches that tackled this problem used a transformer-based model, BERT\/RoBERTa, trained on the homologous annotated dataset and achieved good performance on this task. However, this requires extensive and annotated datasets with a strong connection to the topic. Both goals are difficult to meet without employing reliable, vast, and costly resources. In this paper, we propose a robust semi-supervised model, SSL-GAN-RoBERTa, that learns from a limited heterogeneous dataset and whose performance is further enhanced by using vast amounts of unlabeled data from another related domain. Compared with the RoBERTa baseline model, the experimental results show that the model has substantial performance gains in terms of Accuracy and Macro-F1 score in different scenarios that use data from different domains. Our proposed model achieves state-of-the-art performance results while efficiently using unlabeled data, showing promising applicability to other complex classification tasks where large amounts of labeled examples are difficult to obtain.<\/jats:p>","DOI":"10.1017\/s1351324923000396","type":"journal-article","created":{"date-parts":[[2023,8,3]],"date-time":"2023-08-03T08:04:56Z","timestamp":1691049896000},"page":"1161-1180","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":13,"title":["SSL-GAN-RoBERTa: A robust semi-supervised model for detecting Anti-Asian COVID-19 hate speech on social media"],"prefix":"10.1017","volume":"30","author":[{"given":"Xuanyu","family":"Su","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yansong","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9917-3694","authenticated-orcid":false,"given":"Paula","family":"Branco","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0202-2444","authenticated-orcid":false,"given":"Diana","family":"Inkpen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"56","published-online":{"date-parts":[[2023,8,3]]},"reference":[{"key":"S1351324923000396_ref3","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2022.101365"},{"key":"S1351324923000396_ref13","first-page":"10","volume-title":"The International Conference on Learning Representations (ICLR)","author":"Denton","year":"2017"},{"key":"S1351324923000396_ref23","first-page":"607","volume-title":"Proceedings of the International AAAI Conference on Web and Social Media","volume":"16","author":"Li","year":"2022"},{"key":"S1351324923000396_ref11","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-3504"},{"key":"S1351324923000396_ref33","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"key":"S1351324923000396_ref17","doi-asserted-by":"publisher","DOI":"10.1145\/3487351.3488324"},{"key":"S1351324923000396_ref12","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5102"},{"key":"S1351324923000396_ref20","author":"Kennedy","year":"2021"},{"key":"S1351324923000396_ref32","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-demos.2"},{"key":"S1351324923000396_ref9","first-page":"6513","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems. NIPS\u201917","author":"Dai","year":"2017"},{"key":"S1351324923000396_ref39","unstructured":"Thain, N. , Dixon, L. and Wulczyn, E. (2017). Wikipedia Talk Labels: Toxicity. https:\/\/figshare.com\/articles\/dataset\/ Wikipedia_Talk_Labels_Toxicity\/4563973"},{"key":"S1351324923000396_ref40","volume-title":"Thirty-Second AAAI Conference on Artificial Intelligence. Online: Thirty-Second AAAI Conference on Artificial Intelligence","author":"Ulyanov","year":"2018"},{"key":"S1351324923000396_ref45","unstructured":"Wei, B. , Li, J. , Gupta, A. , Umair, H. , Vovor, A. and Durzynski, N. (2021). Offensive language and hate speech detection with deep learning and transfer learning. arXiv preprint arXiv: 2108.03305."},{"key":"S1351324923000396_ref43","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5446"},{"key":"S1351324923000396_ref37","article-title":"Unsupervised and Semi-supervised Learning with Categorical Generative Adversarial Networks","author":"Springenberg","year":"2016","journal-title":"CoRR"},{"key":"S1351324923000396_ref46","unstructured":"Yuan, L. , Wang, T. , Ferraro, G. , Suominen, H. and Rizoiu, M.-A. (2019). Transfer learning for hate speech detection in social media. arXiv preprint arXiv: 1906.03829."},{"key":"S1351324923000396_ref38","volume-title":"Workshop Proceedings of the 16th International AAAI Conference on Web and Social Media","author":"Tekumalla","year":"2022"},{"key":"S1351324923000396_ref1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3109091"},{"key":"S1351324923000396_ref19","article-title":"Atlanta spa attacks shine a light on anti-Asian hate crimes around the world","author":"Johnson","year":"2021","journal-title":"CNN"},{"key":"S1351324923000396_ref26","unstructured":"Lora, J. , Palumbo, D. and Brown, D. (2021). Coronavirus: How the pandemic has changed the world economy (accessed 17 August 2021). https:\/\/www.bbc.co.uk\/news\/business-51706225"},{"key":"S1351324923000396_ref27","first-page":"142","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language","author":"Maas","year":"2011"},{"key":"S1351324923000396_ref10","first-page":"512","volume-title":"Proceedings of the International AAAI Conference on Web and Social Media","volume":"11","author":"Davidson","year":"2017"},{"key":"S1351324923000396_ref36","first-page":"1631","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Socher","year":"2013"},{"key":"S1351324923000396_ref22","doi-asserted-by":"crossref","first-page":"577","DOI":"10.1016\/j.procs.2019.11.159","article-title":"Semi-supervised learning for sentiment classification using small number of labeled data","volume":"161","author":"Lee","year":"2019","journal-title":"Procedia Computer Science"},{"key":"S1351324923000396_ref7","doi-asserted-by":"publisher","DOI":"10.33972\/jhs.198"},{"key":"S1351324923000396_ref24","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219956"},{"key":"S1351324923000396_ref18","unstructured":"He, P. , Liu, X. , Gao, J. and Chen, W. (2020). DeBERTa: Decoding-enhanced BERT with Disentangled Attention. arXiv preprint arXiv:2006.03654"},{"key":"S1351324923000396_ref6","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"S1351324923000396_ref34","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324923000396_ref15","doi-asserted-by":"publisher","DOI":"10.1002\/pra2.313"},{"key":"S1351324923000396_ref14","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"S1351324923000396_ref28","first-page":"14867","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Mathew","year":"2021"},{"key":"S1351324923000396_ref41","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0243300"},{"key":"S1351324923000396_ref4","volume-title":"Advances in Neural Information Processing Systems","volume":"32","author":"Berthelot","year":"2019"},{"key":"S1351324923000396_ref8","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.191"},{"key":"S1351324923000396_ref2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2021.115632"},{"key":"S1351324923000396_ref31","doi-asserted-by":"publisher","DOI":"10.3389\/frai.2023.1023281"},{"key":"S1351324923000396_ref35","first-page":"2234","volume-title":"Advances in Neural Information Processing Systems","volume":"29","author":"Salimans","year":"2016"},{"key":"S1351324923000396_ref44","volume-title":"Advances in Neural Information Processing Systems","volume":"32","author":"Wang","year":"2019"},{"key":"S1351324923000396_ref42","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.alw-1.19"},{"key":"S1351324923000396_ref29","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2858821"},{"key":"S1351324923000396_ref21","article-title":"ALBERT: A Lite BERT for Self-supervised Learning of Language Representations","author":"Lan","year":"2019","journal-title":"CoRR"},{"key":"S1351324923000396_ref5","first-page":"15","volume-title":"Proceedings of the First Workshop on Language Technology for Equality, Diversity and Inclusion","author":"Bigoulaeva","year":"2021"},{"key":"S1351324923000396_ref16","volume-title":"Proceedings of the International AAAI Conference on Web and Social Media","volume":"12","author":"Founta","year":"2018"},{"key":"S1351324923000396_ref25","article-title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","author":"Liu","year":"2019","journal-title":"CoRR"},{"key":"S1351324923000396_ref30","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-36687-2_77"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324923000396","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,12]],"date-time":"2024-12-12T11:12:25Z","timestamp":1734001945000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324923000396\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,3]]},"references-count":46,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,11]]}},"alternative-id":["S1351324923000396"],"URL":"https:\/\/doi.org\/10.1017\/s1351324923000396","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,3]]},"assertion":[{"value":"\u00a9 The Author(s), 2023. Published by Cambridge University Press","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (http:\/\/creativecommons.org\/licenses\/by\/4.0\/), which permits unrestricted re-use, distribution and reproduction, provided the original article is properly cited.","name":"license","label":"License","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This content has been made available to all.","name":"free","label":"Free to read"}]}}