{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T13:28:08Z","timestamp":1783690088078,"version":"3.55.0"},"reference-count":63,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2023,10,21]],"date-time":"2023-10-21T00:00:00Z","timestamp":1697846400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2023,10,21]],"date-time":"2023-10-21T00:00:00Z","timestamp":1697846400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"funder":[{"name":"MobileTeleSystems","award":["MTS-Skoltech laboratory on AI"],"award-info":[{"award-number":["MTS-Skoltech laboratory on AI"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Lang Resources &amp; Evaluation"],"published-print":{"date-parts":[[2024,6]]},"DOI":"10.1007\/s10579-023-09682-z","type":"journal-article","created":{"date-parts":[[2023,10,21]],"date-time":"2023-10-21T12:02:10Z","timestamp":1697889730000},"page":"459-504","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements"],"prefix":"10.1007","volume":"58","author":[{"given":"Nikolay","family":"Babakov","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Varvara","family":"Logacheva","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6097-6118","authenticated-orcid":false,"given":"Alexander","family":"Panchenko","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,10,21]]},"reference":[{"key":"9682_CR1","unstructured":"Babakov, N., Logacheva, V., Kozlova, O., Semenov, N., & Panchenko, A. (2021). Detecting inappropriate messages on sensitive topics that could harm a company\u2019s reputation. In Proceedings of the 8th Workshop on Balto\u2013Slavic Natural Language Processing, pp. 26\u201336, Kiyv. Association for Computational Linguistics."},{"key":"9682_CR2","doi-asserted-by":"crossref","unstructured":"Banko, M., MacKeen, B., & Ray, L. (2020). A unified taxonomy of harmful content. In Proceedings of the Fourth Workshop on Online Abuse and Harms, pp. 125\u2013137, Online. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.alw-1.16"},{"key":"9682_CR3","doi-asserted-by":"crossref","unstructured":"Basile, V., Bosco, C., Fersini, E., Nozza, D., Patti, V., Pardo, F., Manuel, R., Rosso, P., & Sanguinetti, M. (2019). SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter. In Proceedings of the 13th International Workshop on Semantic Evaluation, pp. 54\u201363, Minneapolis, Minnesota. Association for Computational Linguistics.","DOI":"10.18653\/v1\/S19-2007"},{"key":"9682_CR4","doi-asserted-by":"crossref","unstructured":"Bogoradnikova, D., Makhnytkina, O., Matveev, A., Zakharova, A., & Akulov, A. (2021). Multilingual sentiment analysis and toxicity detection for text messages in russian. In 2021 29th Conference of Open Innovations Association (FRUCT), pp 55\u201364.","DOI":"10.23919\/FRUCT52173.2021.9435584"},{"key":"9682_CR5","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1162\/tacl_a_00051","volume":"5","author":"P Bojanowski","year":"2017","unstructured":"Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2017). Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5, 135\u2013146.","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"9682_CR6","doi-asserted-by":"crossref","unstructured":"Breitfeller, L., Ahn, E., Jurgens, D., & Tsvetkov, Y. (2019). Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 1664\u20131674, Hong Kong. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D19-1176"},{"key":"9682_CR7","unstructured":"Cecillon, N., Labatut, V., Dufour, R., & Linar\u00e8s, G. (2020). Wac: A corpus of wikipedia conversations for online abuse detection. In LREC."},{"key":"9682_CR8","doi-asserted-by":"crossref","unstructured":"Chung, Y.-L., Kuzmenko, E., Tekiroglu, S. S., & Guerini, M. (2019). CONAN-COunter NArratives through nichesourcing: A multilingual dataset of responses to fight online hate speech. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 2819\u20132829, Florence. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P19-1271"},{"issue":"1","key":"9682_CR9","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1177\/001316446002000104","volume":"20","author":"J Cohen","year":"1960","unstructured":"Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37\u201346.","journal-title":"Educational and Psychological Measurement"},{"key":"9682_CR10","doi-asserted-by":"crossref","unstructured":"Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzm\u00e1n, F., Grave, E., Ott, M., Zettlemoyer, L., & Stoyanov, V. (2019). Unsupervised cross-lingual representation learning at scale. CoRR, arXiv:1911.02116.","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"9682_CR11","doi-asserted-by":"crossref","unstructured":"Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzm\u00e1n, F., Grave, E., Ott, M., Zettlemoyer, L., & Stoyanov, V. (2020). Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8440\u20138451, Online. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"9682_CR12","doi-asserted-by":"crossref","unstructured":"Davidson, T., Warmsley, D., Macy, M., & Weber, I. (2017). Automated hate speech detection and the problem of offensive language.","DOI":"10.1609\/icwsm.v11i1.14955"},{"key":"9682_CR13","first-page":"20","volume":"28","author":"AP Dawid","year":"1979","unstructured":"Dawid, A. P., & Skene, A. (1979). Maximum likelihood estimation of observer error-rates using the em algorithm. Journal of The Royal Statistical Society Series C, 28, 20\u201328.","journal-title":"Journal of The Royal Statistical Society Series C"},{"key":"9682_CR14","doi-asserted-by":"crossref","unstructured":"de Gibert, O., Perez, N., Garc\u00eda-Pablos, A., & Cuadros, M. (2018). Hate speech dataset from a white supremacy forum. In Proceedings of the 2nd Workshop on Abusive Language Online (ALW2), pp. 11\u201320, Brussels, Belgium. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W18-5102"},{"key":"9682_CR15","unstructured":"Dinan, E., Abercrombie, G., Stevie, B. A., Spruit, S., Hovy, D., Boureau, Y.-L., & Rieser, V. (2021). Anticipating safety issues in e2e conversational ai: Framework and tooling."},{"key":"9682_CR16","doi-asserted-by":"crossref","unstructured":"Dinan, E., Humeau, S., Chintagunta, B., & Weston, J. ( 2019). Build it break it fix it for dialogue safety: Robustness from adversarial human attack. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 4537\u20134546, Hong Kong. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D19-1461"},{"key":"9682_CR17","doi-asserted-by":"crossref","unstructured":"Dixon, L., Li, J., Sorensen, J., Thain, N., & Vasserman, L. (2018). Measuring and mitigating unintended bias in text classification. In Proceedings of the 2018 AAAI\/ACM Conference on AI, Ethics, and Society, AIES \u201918, pp. 67\u201373, New York. Association for Computing Machinery.","DOI":"10.1145\/3278721.3278729"},{"key":"9682_CR18","doi-asserted-by":"crossref","unstructured":"Drisko, J. W., & Maschi, T. (2016). Content analysis. Pocket Guide to Social Work Re.","DOI":"10.1093\/acprof:oso\/9780190215491.001.0001"},{"key":"9682_CR19","doi-asserted-by":"crossref","unstructured":"Fersini, E., Nozza, D., & Rosso, P. (2018). Overview of the evalita 2018 task on automatic misogyny identification (ami). In EVALITA@CLiC-it.","DOI":"10.4000\/books.aaccademia.4497"},{"key":"9682_CR20","doi-asserted-by":"crossref","unstructured":"Gautam, A., Mathur, P., Gosangi, R., Mahata, D., Sawhney, R., & Shah, R. R. (2020). # metooma: Multi-aspect annotations of tweets related to the metoo movement. InProceedings of the International AAAI Conference on Web and Social Media, Vol. 14, pp. 209\u2013216.","DOI":"10.1609\/icwsm.v14i1.7292"},{"key":"9682_CR21","doi-asserted-by":"crossref","unstructured":"Han, X., & Tsvetkov, Y. (2020). Fortifying toxic speech detectors against veiled toxicity. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7732\u20137739, Online. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.emnlp-main.622"},{"key":"9682_CR22","doi-asserted-by":"crossref","unstructured":"Hessel, J., & Lee, L. (2019). Something\u2019s brewing! early prediction of controversy-causing posts from discussion features. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 1648\u20131659, Minneapolis, Minnesota. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N19-1166"},{"key":"9682_CR23","unstructured":"Jigsaw multilingual toxic comment classification. https:\/\/www.kaggle.com\/c\/jigsaw-multilingual-toxic-comment-classification, 2019. Accessed 13 Jan 2021."},{"key":"9682_CR24","unstructured":"Jigsaw unintended bias in toxicity classification. https:\/\/www.kaggle.com\/c\/jigsaw-unintended-bias-in-toxicity-classification, 2019. Accessed 13 Jan 2021."},{"key":"9682_CR25","unstructured":"Jigsaw. (2018). Toxic comment classification challenge. https:\/\/www.kaggle.com\/c\/jigsaw-toxic-comment-classification-challenge. Accessed 01 March 2021."},{"key":"9682_CR26","unstructured":"Jigsaw. (2019). Jigsaw unintended bias in toxicity classification. https:\/\/www.kaggle.com\/c\/jigsaw-unintended-bias-in-toxicity-classification. Accessed 01 March 2021."},{"key":"9682_CR27","unstructured":"Jigsaw. (2020). Jigsaw multilingual toxic comment classification. https:\/\/www.kaggle.com\/c\/jigsaw-multilingual-toxic-comment-classification. Accessed 01 March 2021."},{"key":"9682_CR28","doi-asserted-by":"crossref","unstructured":"Joshi, R., Karnavat, R., Jirapure, K., & Joshi, R. (2021). Evaluation of deep learning models for hostility detection in Hindi text. CoRR, arXiv:2101.04144.","DOI":"10.1109\/I2CT51068.2021.9418073"},{"key":"9682_CR29","unstructured":"Kaggle. (2019). Russian language toxic comments. https:\/\/www.kaggle.com\/blackmoon\/russian-language-toxic-comments. Accessed 01 March 2021."},{"key":"9682_CR30","unstructured":"Kaggle. (2020). Toxic Russian comments. https:\/\/www.kaggle.com\/alexandersemiletov\/toxic-russian-comments. Accessed 01 March 2021."},{"key":"9682_CR31","doi-asserted-by":"crossref","unstructured":"Karan, M., & \u0160najder, J. (2019). Preemptive toxic language detection in Wikipedia comments using thread-level context. In Proceedings of the Third Workshop on Abusive Language Online, pp. 129\u2013134, Florence. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W19-3514"},{"key":"9682_CR32","doi-asserted-by":"crossref","unstructured":"Kim, Y. (2014). Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1746\u20131751, Doha, Qatar. Association for Computational Linguistics.","DOI":"10.3115\/v1\/D14-1181"},{"key":"9682_CR33","unstructured":"Krippendorff, K. (1980). Content analysis: An introduction to its methodolog."},{"key":"9682_CR34","unstructured":"Kuratov, Y., & Arkhipov, M. (2019). Adaptation of deep bidirectional multilingual transformers for Russian language."},{"key":"9682_CR35","volume-title":"Advances in neural information processing systems","author":"Y LeCun","year":"1990","unstructured":"LeCun, Y., Boser, B., Denker, J., Henderson, D., Howard, R., Hubbard, W., & Jackel, L. (1990). Handwritten digit recognition with a back-propagation network. In D. Touretzky (Ed.), Advances in neural information processing systems.  (Vol. 2). Burlington: Morgan-Kaufmann."},{"key":"9682_CR36","unstructured":"Lees, A., Borkan, D., Kivlichan, I., Nario, J., & Goyal, T. (2021). Capturing covertly toxic speech via crowdsourcing. In Proceedings of the First Workshop on Bridging Human-Computer Interaction and Natural Language Processing, pp 14\u201320, Online. Association for Computational Linguistics."},{"key":"9682_CR37","doi-asserted-by":"crossref","unstructured":"Mollas, I., Chrysopoulou, Z., Karlos, S., & Tsoumakas, G. (2021). Ethos: An online hate speech detection dataset.","DOI":"10.1007\/s40747-021-00608-2"},{"key":"9682_CR38","doi-asserted-by":"crossref","unstructured":"Nobata, C., Tetreault, J. R., Thomas, A., Mehdad, Y., & Chang, Y. (2016). Abusive language detection in online user content. In Bourdeau, J., Hendler, J., Nkambou, R., Horrocks, I. & Zhao, B. Y., (eds.) In Proceedings of the 25th International Conference on World Wide Web, WWW 2016, Montreal, April 11\u201315, 2016, pp. 145\u2013153. ACM.","DOI":"10.1145\/2872427.2883062"},{"key":"9682_CR39","doi-asserted-by":"crossref","unstructured":"Ousidhoum, N., Lin, Z., Zhang, H., Song, Y., & Yeung, D.-Y. (2019). Multilingual and multi-aspect hate speech analysis. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 4675\u20134684, Hong Kong. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D19-1474"},{"issue":"5","key":"9682_CR40","doi-asserted-by":"publisher","first-page":"403","DOI":"10.1080\/1364557032000081654","volume":"7","author":"Seth Ovadia","year":"2004","unstructured":"Ovadia, Seth. (2004). Ratings and rankings: Reconsidering the structure of values and their measurement. International Journal of Social Research Methodology, 7(5), 403\u2013414.","journal-title":"International Journal of Social Research Methodology"},{"key":"9682_CR41","doi-asserted-by":"crossref","unstructured":"Pandey, R., Purohit, H., Stabile, B., & Grant, A. (2018). Distributional semantics approach to detect intent in twitter conversations on sexual assaults. In 2018 IEEE\/WIC\/ACM International Conference on Web Intelligence (WI), pp. 270\u2013277.","DOI":"10.1109\/WI.2018.00-80"},{"key":"9682_CR42","doi-asserted-by":"crossref","unstructured":"Park, J. H., Shin, J, & Fung, P. (2018). Reducing gender bias in abusive language detection. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2799\u20132804, Brussels, Belgium, Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1302"},{"key":"9682_CR43","unstructured":"Passonneau, Rebecca, J. (2006). Measuring agreement on set-valued items (masi) for semantic and pragmatic annotation. In LREC."},{"key":"9682_CR44","doi-asserted-by":"crossref","unstructured":"Pavlopoulos, J., Malakasiotis, P., & Androutsopoulos, I. (2017). Deeper attention to abusive user content moderation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 1125\u20131135, Copenhagen. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D17-1117"},{"key":"9682_CR45","doi-asserted-by":"crossref","unstructured":"Pavlopoulos, J., Sorensen, J., Dixon, L., Thain, N., & Androutsopoulos, I. (2020). Toxicity detection: Does context really matter? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4296\u20134305, Online. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.acl-main.396"},{"key":"9682_CR46","doi-asserted-by":"crossref","unstructured":"Qian, J., Bethke, A., Liu, Y., Belding, E., & Wang, W. Y. (2019). A benchmark dataset for learning to intervene in online hate speech.","DOI":"10.18653\/v1\/D19-1482"},{"issue":"8","key":"9682_CR47","first-page":"9","volume":"1","author":"A Radford","year":"2019","unstructured":"Radford, A., Jeffrey, W., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Blog, 1(8), 9.","journal-title":"OpenAI Blog"},{"key":"9682_CR48","doi-asserted-by":"crossref","unstructured":"Rosenthal, S., Atanasova, P., Karadzhov, G., Zampieri, M., & Nakov, P. (2021). SOLID: a large-scale semi-supervised dataset for offensive language identification. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 915\u2013928, Online. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2021.findings-acl.80"},{"key":"9682_CR49","first-page":"15","volume":"2","author":"J Salminen","year":"2020","unstructured":"Salminen, J., Seng\u00fcn, S., Corporan, J., Jung, S., & Jansen, B. J. (2020). Topic-driven toxicity: Exploring the relationship between online toxicity and news topics. PLoS ONE, 2, 15.","journal-title":"PLoS ONE"},{"key":"9682_CR50","doi-asserted-by":"crossref","unstructured":"Schrading, N., Alm, C. O.,, Ptucha, R., & Homan, C. (2015). An analysis of domestic abuse discourse on Reddit. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp 2577\u20132583, Lisbon. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D15-1309"},{"key":"9682_CR51","doi-asserted-by":"crossref","unstructured":"Smetanin, S. (2020). Toxic comments detection in russian. In Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference \u201cDialogue 2020\u201d,","DOI":"10.28995\/2075-7182-2020-19-1149-1159"},{"key":"9682_CR52","unstructured":"Sun, C., Huang, L., & Qiu, X. (2019). Utilizing BERT for aspect-based sentiment analysis via constructing auxiliary sentence. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 380\u2013385, Minneapolis, Minnesota. Association for Computational Linguistics."},{"issue":"1","key":"9682_CR53","doi-asserted-by":"publisher","first-page":"683","DOI":"10.1609\/icwsm.v14i1.7334","volume":"14","author":"A Vaidya","year":"2020","unstructured":"Vaidya, A., Mai, F., & Ning, Y. (2020). Empirical analysis of multi-task learning for reducing identity bias in toxic comment detection. Proceedings of the International AAAI Conference on Web and Social Media, 14(1), 683\u2013693.","journal-title":"Proceedings of the International AAAI Conference on Web and Social Media"},{"key":"9682_CR54","doi-asserted-by":"crossref","unstructured":"Waseem, Z., & Hovy, D. (2016). Hateful symbols or hateful people? predictive features for hate speech detection on Twitter. In Proceedings of the NAACL Student Research Workshop, pp. 88\u201393, San Diego. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N16-2013"},{"key":"9682_CR55","doi-asserted-by":"crossref","unstructured":"Waseem, Z., Davidson, T., Warmsley, D. & Weber, I. (2017). Understanding abuse: A typology of abusive language detection subtasks. In Proceedings of the First Workshop on Abusive Language Online, pp. 78\u201384, Vancouver. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W17-3012"},{"key":"9682_CR56","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1007\/978-3-030-22747-0_7","volume-title":"Computational science-ICCS 2019","author":"X Wu","year":"2019","unstructured":"Wu, X., Lv, S., Zang, L., Han, J., & Hu, S. (2019). Conditional Bert contextual augmentation. In J. M. F. Rodrigues, P. J. S. Cardoso, J. Monteiro, R. Lam, V. V. Krzhizhanovskaya, M. H. Lees, J. J. Dongarra, & P. M. A. Sloot (Eds.), Computational science-ICCS 2019 (pp. 84\u201395). Cham: Springer International Publishing."},{"key":"9682_CR57","doi-asserted-by":"crossref","unstructured":"Xia, M., Field, A., & Tsvetkov, Y. (2020). Demoting racial bias in hate speech detection. In Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media, pp. 7\u201314, Online. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.socialnlp-1.2"},{"key":"9682_CR58","unstructured":"Xia, C., Zhang, C., Nguyen, H., Zhang, J., & Yu, P. S. (2020). Cg-bert: Conditional text generation with Bert for generalized few-shot intent detection. arXiv:2004.01881."},{"key":"9682_CR59","unstructured":"Xu, J., Da, J., Margaret, L., Boureau, Y.-L., Weston, J., & Dinan, E. (2020). Recipes for safety in open-domain chatbots."},{"issue":"4","key":"9682_CR60","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1007\/s41060-017-0088-4","volume":"6","author":"H Yenala","year":"2018","unstructured":"Yenala, H., Jhanwar, A., Chinnakotla, M. K., & Goyal, J. (2018). Deep learning for detecting inappropriate content in text. International Journal of Data Science and Analytics, 6(4), 273\u2013286.","journal-title":"International Journal of Data Science and Analytics"},{"key":"9682_CR61","doi-asserted-by":"publisher","first-page":"176600","DOI":"10.1109\/ACCESS.2019.2953990","volume":"7","author":"S Yu","year":"2019","unstructured":"Yu, S., Su, J., & Luo, D. (2019). Improving Bert-based text classification with auxiliary sentence and domain knowledge. IEEE Access, 7, 176600\u2013176612.","journal-title":"IEEE Access"},{"key":"9682_CR62","doi-asserted-by":"crossref","unstructured":"Zampieri, Marcos, Malmasi, Shervin, Nakov, Preslav, Rosenthal, Sara, F. N., & Kumar, R. (2019). Predicting the type and target of offensive posts in social media. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 1415\u20131420, Minneapolis, Minnesota. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N19-1144"},{"key":"9682_CR63","doi-asserted-by":"crossref","unstructured":"Zhang, G., Bai, B., Zhang, J., Bai, K., Zhu, C., & Zhao, T. (2020). Demographics should not be the reason of toxicity: Mitigating discrimination in text classifications with instance weighting. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4134\u20134145, Online, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.acl-main.380"}],"container-title":["Language Resources and Evaluation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-023-09682-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10579-023-09682-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-023-09682-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,28]],"date-time":"2024-05-28T18:31:44Z","timestamp":1716921104000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10579-023-09682-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,21]]},"references-count":63,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,6]]}},"alternative-id":["9682"],"URL":"https:\/\/doi.org\/10.1007\/s10579-023-09682-z","relation":{},"ISSN":["1574-020X","1574-0218"],"issn-type":[{"value":"1574-020X","type":"print"},{"value":"1574-0218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,21]]},"assertion":[{"value":"13 July 2023","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 October 2023","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"This work has not been submitted to any other journal or conference before. A part of this work has already been published in the Workshop for Balto-Slavic Natural Language Processing (Babakov et al., ). The current work contains new results, descriptions of new methods, and new extended datasets. The differences from the previous work are listed in Sect.\u00a0. This manuscript describes the full conducted work on the collection of datasets of inappropriate messages and sensitive topics and on their use for training the classification models. The results presented in the manuscript can be replicated using the provided datasets, code for training the models, and pre-trained models. All of the above are made available.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Code availability"}},{"value":"The data collection process involved human participants. These were workers hired via Toloka crowdsourcing platform. They were informed and accepted that any data that they produce can be used by employers in the public or private domain. The projects created within Toloka for data annotation fully comply with the rules of this service. Our research was motivated by the need for moderation of automatically generated content by neural models, such as the GPT family of decoder-based Transformers to avoid the reputations risk of companies deploying such models (including PR risks).","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Human and animal rights"}}]}}