{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:14:17Z","timestamp":1750220057277,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":28,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,3,27]],"date-time":"2023-03-27T00:00:00Z","timestamp":1679875200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,3,27]]},"DOI":"10.1145\/3555776.3577856","type":"proceedings-article","created":{"date-parts":[[2023,6,7]],"date-time":"2023-06-07T17:16:29Z","timestamp":1686158189000},"page":"939-941","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Impact of Character n-grams Attention Scores for English and Russian News Articles Authorship Attribution"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2191-4330","authenticated-orcid":false,"given":"Liliya","family":"Makhmutova","sequence":"first","affiliation":[{"name":"School of Computer Science, Technological University Dublin, Dublin, Dublin, Ireland"},{"name":"ML-Labs, SFI Centre for Machine Learning, Dublin, Dublin, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1449-1827","authenticated-orcid":false,"given":"Robert","family":"Ross","sequence":"additional","affiliation":[{"name":"School of Computer Science, Technological University Dublin, Dublin, Dublin, Ireland"},{"name":"ADAPT Centre, Dublin, Dublin, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4301-7000","authenticated-orcid":false,"given":"Giancarlo","family":"Salton","sequence":"additional","affiliation":[{"name":"Unochapec\u00f3 - Universidade Comunit\u00e1ria da Regi\u00e3o de Chapec\u00f3, Chapeco, Santa Catarina, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,6,7]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Universal Language Model Fine-tuning for Text Classification. arXiv:1801.06146","author":"Jeremy Howard S. R.","year":"2018","unstructured":"Jeremy Howard , S. R. ( 2018 ). Universal Language Model Fine-tuning for Text Classification. arXiv:1801.06146 . Jeremy Howard, S. R. (2018). Universal Language Model Fine-tuning for Text Classification. arXiv:1801.06146."},{"key":"e_1_3_2_1_2_1","volume-title":"Attention Is All You Need. 31st Conference on Neural Information Processing Systems (NIPS","author":"Ashish Vaswani N. S.","year":"2017","unstructured":"Ashish Vaswani , N. S. ( 2017 ). Attention Is All You Need. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. Ashish Vaswani, N. S. (2017). Attention Is All You Need. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1007\/11861461_10","article-title":"N-Gram Feature Selection for Authorship Identification","volume":"4183","author":"John Houvardas E. S.","year":"2006","unstructured":"John Houvardas , E. S. ( 2006 ). N-Gram Feature Selection for Authorship Identification . Lecture Notes in Computer Science 4183 , 77 -- 86 . John Houvardas, E. S. (2006). N-Gram Feature Selection for Authorship Identification. Lecture Notes in Computer Science 4183, 77--86.","journal-title":"Lecture Notes in Computer Science"},{"key":"e_1_3_2_1_4_1","volume-title":"Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 93--102","author":"Upendra Sapkota S. B.-y.-G.","year":"2015","unstructured":"Upendra Sapkota , S. B.-y.-G. ( 2015 ). Not All Character N-grams Are Created Equal: A Study in Authorship Attribution . Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 93--102 . Upendra Sapkota, S. B.-y.-G. (2015). Not All Character N-grams Are Created Equal: A Study in Authorship Attribution. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 93--102."},{"key":"e_1_3_2_1_5_1","volume-title":"Proceedings of the 2019 3rd International Conference on Natural Language Processing and Information Retrieval (NLPIR","author":"Tatiana Litvinova O. L.","year":"2019","unstructured":"Tatiana Litvinova , O. L. ( 2019 ). Authorship Attribution of Russian Forum Posts with Different Types of N-gram Features . Proceedings of the 2019 3rd International Conference on Natural Language Processing and Information Retrieval (NLPIR 2019), 9--14. Tatiana Litvinova, O. L. (2019). Authorship Attribution of Russian Forum Posts with Different Types of N-gram Features. Proceedings of the 2019 3rd International Conference on Natural Language Processing and Information Retrieval (NLPIR 2019), 9--14."},{"issue":"3","key":"e_1_3_2_1_6_1","first-page":"59","article-title":"Authorship Attribution in Portuguese Using Character N-grams","volume":"14","author":"Ilia Markov","year":"2017","unstructured":"Ilia Markov 1, J. B.-L. ( 2017 ). Authorship Attribution in Portuguese Using Character N-grams . Acta Polytechnica Hungarica Vol. 14 , No. 3 , 59 -- 78 . Ilia Markov1, J. B.-L. (2017). Authorship Attribution in Portuguese Using Character N-grams. Acta Polytechnica Hungarica Vol. 14, No. 3, 59--78.","journal-title":"Acta Polytechnica Hungarica"},{"key":"e_1_3_2_1_7_1","volume-title":"Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI-16)","author":"Yoon Kim Y. J.","year":"2016","unstructured":"Yoon Kim , Y. J. ( 2016 ). Character-Aware Neural Language Models . Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI-16) , 2741--2749. Yoon Kim, Y. J. (2016). Character-Aware Neural Language Models. Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI-16), 2741--2749."},{"key":"e_1_3_2_1_8_1","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics","volume":"1","author":"Lyan Verwimp J. P.","year":"2017","unstructured":"Lyan Verwimp , J. P. ( 2017 ). Character-Word LSTM Language Models . Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics : Volume 1 , Long Papers, 417--427. Lyan Verwimp, J. P. (2017). Character-Word LSTM Language Models. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, 417--427."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1162\/tacl_a_00051","article-title":"Enriching Word Vectors with Subword Information","volume":"5","author":"Piotr Bojanowski E. G.","year":"2017","unstructured":"Piotr Bojanowski , E. G. ( 2017 ). Enriching Word Vectors with Subword Information . Transactions of the Association for Computational Linguistics , Volume 5 , 135 -- 146 . Piotr Bojanowski, E. G. (2017). Enriching Word Vectors with Subword Information. Transactions of the Association for Computational Linguistics, Volume 5, 135--146.","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"e_1_3_2_1_10_1","unstructured":"Huggingface. (30 09 2022 .). Summary of the tokenizers. Huggingface: https:\/\/huggingface.co\/transformers\/v4.3.0\/tokenizer_summary.html  Huggingface. (30 09 2022 .). Summary of the tokenizers. Huggingface: https:\/\/huggingface.co\/transformers\/v4.3.0\/tokenizer_summary.html"},{"key":"e_1_3_2_1_11_1","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1715--1725","author":"Rico Sennrich B. H.","year":"2016","unstructured":"Rico Sennrich , B. H. ( 2016 ). Neural Machine Translation of Rare Words with Subword Units . Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1715--1725 . Rico Sennrich, B. H. (2016). Neural Machine Translation of Rare Words with Subword Units. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1715--1725."},{"key":"e_1_3_2_1_12_1","volume-title":"International Conference on Acoustics, Speech and Signal Processing, IEEE","author":"Mike Schuster K. N.","year":"2012","unstructured":"Mike Schuster , K. N. ( 2012 ). Japanese and Korean voice search . International Conference on Acoustics, Speech and Signal Processing, IEEE (2012), 5149--5152. Mike Schuster, K. N. (2012). Japanese and Korean voice search. International Conference on Acoustics, Speech and Signal Processing, IEEE (2012), 5149--5152."},{"key":"e_1_3_2_1_13_1","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 66--75","author":"Kudo T.","year":"2018","unstructured":"Kudo , T. ( 2018 ). Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates . Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 66--75 . Kudo, T. (2018). Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 66--75."},{"key":"e_1_3_2_1_14_1","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 66--71","author":"Taku Kudo J. R.","year":"2018","unstructured":"Taku Kudo , J. R. ( 2018 ). SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing . Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 66--71 . Taku Kudo, J. R. (2018). SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 66--71."},{"key":"e_1_3_2_1_15_1","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 1504--1515","author":"John Wieting M. B.","year":"2016","unstructured":"John Wieting , M. B. ( 2016 ). Charagram: Embedding Words and Sentences via Character n-grams . Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 1504--1515 . John Wieting, M. B. (2016). Charagram: Embedding Words and Sentences via Character n-grams. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 1504--1515."},{"key":"e_1_3_2_1_16_1","volume-title":"The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19)","author":"Sho Takase J. S.","year":"2019","unstructured":"Sho Takase , J. S. ( 2019 ). Character n-gram Embeddings to Improve RNN Language Models . The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19) , 5074--5082. Sho Takase, J. S. (2019). Character n-gram Embeddings to Improve RNN Language Models. The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), 5074--5082."},{"key":"e_1_3_2_1_17_1","unstructured":"Tao Shen T. Z. (2019 03 2019 .). Tensorized Self-Attention: Efficiently Modeling Pairwise and Global Dependencies Together. arXiv:1805.00912v4. Retrieved from arxiv: https:\/\/arxiv.org\/pdf\/1805.00912.pdf  Tao Shen T. Z. (2019 03 2019 .). Tensorized Self-Attention: Efficiently Modeling Pairwise and Global Dependencies Together. arXiv:1805.00912v4. Retrieved from arxiv: https:\/\/arxiv.org\/pdf\/1805.00912.pdf"},{"key":"e_1_3_2_1_18_1","volume-title":"Conference: Experimental IR Meets Multilinguality, Multimodality, and Interaction - 8th International Conference of the CLEF Association (CLEF","author":"Miguel A.","year":"2017","unstructured":"Miguel A. Sanchez-Perez , I. M.-A. ( 2017 ). Comparison of Character n-grams and Lexical Features on Author, Gender, and Language Variety Identification on the Same Spanish News Corpus (preprint version) . Conference: Experimental IR Meets Multilinguality, Multimodality, and Interaction - 8th International Conference of the CLEF Association (CLEF 2017). Volume : 10456, 145--151. Miguel A. Sanchez-Perez, I. M.-A. (2017). Comparison of Character n-grams and Lexical Features on Author, Gender, and Language Variety Identification on the Same Spanish News Corpus (preprint version). Conference: Experimental IR Meets Multilinguality, Multimodality, and Interaction - 8th International Conference of the CLEF Association (CLEF 2017). Volume: 10456, 145--151."},{"key":"e_1_3_2_1_19_1","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 3418--3428","author":"Xiaobing Sun W. L.","year":"2020","unstructured":"Xiaobing Sun , W. L. ( 2020 ). Understanding Attention for Text Classification . Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 3418--3428 . Xiaobing Sun, W. L. (2020). Understanding Attention for Text Classification. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 3418--3428."},{"key":"e_1_3_2_1_20_1","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, 1631--1642","author":"Richard Socher A. P.","year":"2013","unstructured":"Richard Socher , A. P. ( 2013 ). Recursive Deep Models for Semantic Compositionality . Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, 1631--1642 . Richard Socher, A. P. (2013). Recursive Deep Models for Semantic Compositionality. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, 1631--1642."},{"key":"e_1_3_2_1_21_1","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, 142--150","author":"Andrew L.","year":"2011","unstructured":"Andrew L. Maas , R. E. ( 2011 ). Learning Word Vectors for Sentiment Analysis . Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, 142--150 . Andrew L. Maas, R. E. (2011). Learning Word Vectors for Sentiment Analysis. Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, 142--150."},{"key":"e_1_3_2_1_22_1","first-page":"361","article-title":"RCV1: A New Benchmark Collection for Text Categorization Research","volume":"5","author":"David D.","year":"2004","unstructured":"David D. Lewis , Y. Y. ( 2004 ). RCV1: A New Benchmark Collection for Text Categorization Research . Journal of Machine Learning Research 5 , 361 -- 397 . David D. Lewis, Y. Y. (2004). RCV1: A New Benchmark Collection for Text Categorization Research. Journal of Machine Learning Research 5, 361--397.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_1_23_1","volume-title":"To the methodology of corpus construction for machine learning:\"Taiga\" syntax tree corpus and parser . PROCEEDINGS OF THE INTERNATIONAL CONFERENCE \u00abCORPUS LINGUISTICS-2017\u00bb, 78--84","author":"Tatiana Shavrina O. S.","year":"2017","unstructured":"Tatiana Shavrina , O. S. ( 2017 ). To the methodology of corpus construction for machine learning:\"Taiga\" syntax tree corpus and parser . PROCEEDINGS OF THE INTERNATIONAL CONFERENCE \u00abCORPUS LINGUISTICS-2017\u00bb, 78--84 . Tatiana Shavrina, O. S. (2017). To the methodology of corpus construction for machine learning:\"Taiga\" syntax tree corpus and parser . PROCEEDINGS OF THE INTERNATIONAL CONFERENCE \u00abCORPUS LINGUISTICS-2017\u00bb, 78--84."},{"key":"e_1_3_2_1_24_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","volume":"1","author":"Jacob Devlin M.-W. C.","year":"2019","unstructured":"Jacob Devlin , M.-W. C. ( 2019 ). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Volume 1 (Long and Short Papers), 4171--4186. Jacob Devlin, M.-W. C. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171--4186."},{"key":"e_1_3_2_1_25_1","volume-title":"DeBERTa: Decoding-enhanced BERT with Disentangled Attention. CoRR abs\/2006.03654","author":"He Pengcheng L. X.","year":"2020","unstructured":"He Pengcheng , L. X. ( 2020 ). DeBERTa: Decoding-enhanced BERT with Disentangled Attention. CoRR abs\/2006.03654 . He Pengcheng, L. X. (2020). DeBERTa: Decoding-enhanced BERT with Disentangled Attention. CoRR abs\/2006.03654."},{"key":"e_1_3_2_1_26_1","unstructured":"Yuri Kuratov M. A. (2019). Adaptation of Deep Bidirectional Multilingual Transformers for Russian Language. Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference \"Dialogue 2019\" 1--7.  Yuri Kuratov M. A. (2019). Adaptation of Deep Bidirectional Multilingual Transformers for Russian Language. Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference \"Dialogue 2019\" 1--7."},{"key":"e_1_3_2_1_27_1","unstructured":"Canziani A. (30 09 2022 .). Attention and the Transformer. Retrieved from atcold: https:\/\/atcold.github.io\/pytorch-Deep-Learning\/en\/week12\/12-3\/  Canziani A. (30 09 2022 .). Attention and the Transformer. Retrieved from atcold: https:\/\/atcold.github.io\/pytorch-Deep-Learning\/en\/week12\/12-3\/"},{"key":"e_1_3_2_1_28_1","volume-title":"Authorship identification of documents with high content similarity. Scientometrics, 223--237","author":"Rexha A, K. M.","year":"2018","unstructured":"Rexha A, K. M. ( 2018 ). Authorship identification of documents with high content similarity. Scientometrics, 223--237 . Rexha A, K. M. (2018). Authorship identification of documents with high content similarity. Scientometrics, 223--237."}],"event":{"name":"SAC '23: 38th ACM\/SIGAPP Symposium on Applied Computing","sponsor":["SIGAPP ACM Special Interest Group on Applied Computing"],"location":"Tallinn Estonia","acronym":"SAC '23"},"container-title":["Proceedings of the 38th ACM\/SIGAPP Symposium on Applied Computing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3555776.3577856","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3555776.3577856","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:08:31Z","timestamp":1750183711000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3555776.3577856"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,27]]},"references-count":28,"alternative-id":["10.1145\/3555776.3577856","10.1145\/3555776"],"URL":"https:\/\/doi.org\/10.1145\/3555776.3577856","relation":{},"subject":[],"published":{"date-parts":[[2023,3,27]]},"assertion":[{"value":"2023-06-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}