{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T22:38:21Z","timestamp":1785883101866,"version":"3.56.0"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2025,12,31]]},"abstract":"<jats:p>The exponential increase in scientific literature and online information necessitates efficient methods for extracting knowledge from textual data. Natural language processing (NLP) plays a crucial role in addressing this challenge, particularly in text classification tasks. While large language models (LLMs) have achieved remarkable success in NLP, their accuracy can suffer in domain-specific contexts due to specialized vocabulary, unique grammatical structures, and imbalanced data distributions. In this systematic literature review (SLR), we investigate the utilization of pre-trained language models (PLMs) for domain-specific text classification. We systematically review 41 articles published between 2018 and January 2024, adhering to the PRISMA statement (preferred reporting items for systematic reviews and meta-analyses). This review methodology involved rigorous inclusion criteria and a multi-step selection process employing AI-powered tools. We delve into the evolution of text classification techniques and differentiate between traditional and modern approaches. We emphasize transformer-based models and explore the challenges and considerations associated with using LLMs for domain-specific text classification. Furthermore, we categorize existing research based on various PLMs and propose a taxonomy of techniques used in the field. To validate our findings, we conducted a comparative experiment involving BERT, SciBERT, and BioBERT in biomedical sentence classification. Finally, we present a comparative study on the performance of LLMs in text classification tasks across different domains. In addition, we examine recent advancements in PLMs for domain-specific text classification and offer insights into future directions and limitations in this rapidly evolving domain.<\/jats:p>","DOI":"10.1145\/3763002","type":"journal-article","created":{"date-parts":[[2025,9,5]],"date-time":"2025-09-05T15:10:58Z","timestamp":1757085058000},"page":"1-41","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9906-6883","authenticated-orcid":false,"given":"Zhyar Rzgar","family":"K. Rostam","sequence":"first","affiliation":[{"name":"Doctoral School of Applied Informatics and Applied Mathematics, Obuda University, Budapest, Hungary"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8845-8301","authenticated-orcid":false,"given":"G\u00e1bor","family":"Kert\u00e9sz","sequence":"additional","affiliation":[{"name":"John von Neumann Faculty of Informatics, Obuda University, Budapest, Hungary and Laboratory of Parallel and Distributed Systems, Institute for Computer Science and Control (SZTAKI), Hungarian Research Network (HUN-REN), Budapest, Hungary"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,17]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3461702.3462624"},{"key":"e_1_3_2_3_2","first-page":"249","volume-title":"Proceedings of the 2022 9th International Conference on Computing for Sustainable Global Development (INDIACom)","author":"Ahanger Mohammad Munzir","year":"2022","unstructured":"Mohammad Munzir Ahanger and M. Arif Wani. 2022. Novel deep learning approach for scientific literature classification. In Proceedings of the 2022 9th International Conference on Computing for Sustainable Global Development (INDIACom). IEEE, 249\u2013254."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3135381"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","unstructured":"Reza Yazdani Aminabadi Samyam Rajbhandari Ammar Ahmad Awan Cheng Li Du Li Elton Zheng Olatunji Ruwase Shaden Smith Minjia Zhang Jeff Rasley et al. 2022. DeepSpeed-inference: Enabling efficient inference of transformer models at unprecedented scale. In Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis (SC \u201922). IEEE 1\u201315.","DOI":"10.1109\/SC41404.2022.00051"},{"key":"e_1_3_2_6_2","unstructured":"Dogu Araci. 2019. Finbert: Financial sentiment analysis with pre-trained language models. arXiv:1908.10063. Retrieved from https:\/\/arxiv.org\/abs\/1908.10063"},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"Iz Beltagy Kyle Lo and Arman Cohan. 2019. SciBERT: A pretrained language model for scientific text. arXiv:1903.10676. Retrieved from https:\/\/arxiv.org\/abs\/1903.10676","DOI":"10.18653\/v1\/D19-1371"},{"key":"e_1_3_2_8_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, Vol. 33, 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_9_2","unstructured":"Lee Burke Karl Pazdernik Daniel Fortin Benjamin Wilson Rustam Goychayev and John Mattingly. 2021. NukeLM: Pre-trained and fine-tuned language models for the nuclear and energy domains. arXiv:2105.12192. Retrieved from https:\/\/arxiv.org\/abs\/2105.12192"},{"key":"e_1_3_2_10_2","article-title":"Large language models for text classification: From zero-shot learning to fine-tuning","author":"Chae Youngjin","year":"2023","unstructured":"Youngjin Chae and Thomas Davidson. 2023. Large language models for text classification: From zero-shot learning to fine-tuning. Open Science Foundation. Retrieved from https:\/\/osf.io\/5t6xz\/","journal-title":"Open Science Foundation"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00492"},{"issue":"2","key":"e_1_3_2_12_2","first-page":"43","article-title":"ResearchRabbit","volume":"44","author":"Cole Victoria","year":"2023","unstructured":"Victoria Cole and Mish Boutet. 2023. ResearchRabbit. The Journal of the Canadian Health Libraries Association 44, 2 (2023), 43.","journal-title":"The Journal of the Canadian Health Libraries Association"},{"key":"e_1_3_2_13_2","first-page":"308","volume-title":"Proceedings of the 8th International Joint Conference on Natural Language ProcessingVol. 2, Short Papers,","author":"Dernoncourt Franck","year":"2017","unstructured":"Franck Dernoncourt and Ji Young Lee. 2017. PubMed 200k RCT: A dataset for sequential sentence classification in medical abstracts. In Proceedings of the 8th International Joint Conference on Natural Language Processing. Greg Kondrak and Taro Watanabe (Eds.), Vol. 2, Short Papers, Asian Federation of Natural Language Processing, Taipei, Taiwan, 308\u2013313. Retrieved from https:\/\/aclanthology.org\/I17-2052\/"},{"key":"e_1_3_2_14_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_2_15_2","unstructured":"Qingxiu Dong Lei Li Damai Dai Ce Zheng Zhiyong Wu Baobao Chang Xu Sun Jingjing Xu and Zhifang Sui. 2022. A survey for in-context learning. arXiv:2301.00234. Retrieved from https:\/\/arxiv.org\/abs\/2301.00234"},{"key":"e_1_3_2_16_2","unstructured":"Alexander Dunn John Dagdelen Nicholas Walker Sanghoon Lee Andrew S. Rosen Gerbrand Ceder Kristin Persson and Anubhav Jain. 2022. Structured information extraction from complex scientific text with fine-tuned large language models. arXiv:2212.05238. Retrieved from https:\/\/arxiv.org\/abs\/2212.05238"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","first-page":"6518\u2013","DOI":"10.1109\/ACCESS.2024.3349952","article-title":"A survey of text classification with transformers: How wide? how large? how long? how accurate? how expensive? how safe?","author":"Fields John","year":"2024","unstructured":"John Fields, Kevin Chovanec, and Praveen Madiraju. 2024. A survey of text classification with transformers: How wide? how large? how long? how accurate? how expensive? how safe? IEEE Access 12 (2024), 6518\u20136531.","journal-title":"IEEE Access"},{"key":"e_1_3_2_18_2","first-page":"1","volume-title":"Proceedings of the 37th Pacific Asia Conference on Language, Information and Computation","author":"Gan Chengguang","year":"2023","unstructured":"Chengguang Gan and Tatsunori Mori. 2023. Sensitivity and robustness of large language models to prompt template in Japanese text classification tasks. In Proceedings of the 37th Pacific Asia Conference on Language, Information and Computation, 1\u201311."},{"issue":"4","key":"e_1_3_2_19_2","doi-asserted-by":"crossref","first-page":"352","DOI":"10.47852\/bonviewJCCE3202838","article-title":"Comparing BERT against traditional machine learning models in text classification","volume":"2","author":"Garrido-Merchan Eduardo C.","year":"2023","unstructured":"Eduardo C. Garrido-Merchan, Roberto Gozalo-Brizuela, and Santiago Gonzalez-Carvajal. 2023. Comparing BERT against traditional machine learning models in text classification. Journal of Computational and Cognitive Engineering 2, 4 (2023), 352\u2013356.","journal-title":"Journal of Computational and Cognitive Engineering"},{"key":"e_1_3_2_20_2","first-page":"117","volume-title":"Proceedings of the 2023 IEEE 3rd International Conference on Data Science and Computer Application (ICDSCA)","author":"Gou Yunhe","year":"2023","unstructured":"Yunhe Gou and Cao Jie. 2023. A lightweight biomedical named entity recognition with pre-trained model. In Proceedings of the 2023 IEEE 3rd International Conference on Data Science and Computer Application (ICDSCA). IEEE, 117\u2013121."},{"issue":"1","key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"102","DOI":"10.1038\/s41524-022-00784-w","article-title":"MatSciBERT: A materials domain language model for text mining and information extraction","volume":"8","author":"Gupta Tanishq","year":"2022","unstructured":"Tanishq Gupta, Mohd Zaki, N. M. Anoop, and Krishnan, Mausam. 2022. MatSciBERT: A materials domain language model for text mining and information extraction. Npj Computational Materials 8, 1 (2022), 102.","journal-title":"Npj Computational Materials"},{"key":"e_1_3_2_22_2","article-title":"Gpipe: Efficient training of giant neural networks using pipeline parallelism","volume":"32","author":"Huang Yanping","year":"2019","unstructured":"Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, et al. 2019. Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Systems, Vol. 32.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_23_2","unstructured":"Ayush Jain N. M. Meenachi Dr. and B. Venkatraman Dr. 2020. NukeBERT: A pre-trained language model for low resource nuclear domain. arXiv:2003.13821. Retrieved from https:\/\/arxiv.org\/abs\/2003.13821"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3180830"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIBA56860.2023.10165621"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","first-page":"143","DOI":"10.18653\/v1\/2021.bionlp-1.16","volume-title":"Proceedings of the 20th Workshop on Biomedical Language Processing","author":"Kanakarajan Kamal Raj","year":"2021","unstructured":"Kamal Raj Kanakarajan, Bhuvana Kundumani, and Malaikannan Sankarasubbu. 2021. BioELECTRA: Pretrained biomedical text encoder using discriminators. In Proceedings of the 20th Workshop on Biomedical Language Processing, 143\u2013154."},{"issue":"24","key":"e_1_3_2_27_2","article-title":"Topic modeling: A comprehensive review","volume":"7","author":"Kherwa Pooja","year":"2019","unstructured":"Pooja Kherwa and Poonam Bansal. 2019. Topic modeling: A comprehensive review. EAI Endorsed Transactions on Scalable Information Systems 7, 24 (Jul. 2019), e2.","journal-title":"EAI Endorsed Transactions on Scalable Information Systems"},{"key":"e_1_3_2_28_2","first-page":"141036","article-title":"MediBioDeBERTa: Biomedical language model with continuous learning and intermediate fine-tuning","author":"Kim Eunhui","year":"2023","unstructured":"Eunhui Kim, Yuna Jeong, and Myung-Seok Choi. 2023. MediBioDeBERTa: Biomedical language model with continuous learning and intermediate fine-tuning. IEEE Access 11 (2023), 141036\u2013141044.","journal-title":"IEEE Access"},{"issue":"5","key":"e_1_3_2_29_2","doi-asserted-by":"crossref","first-page":"109","DOI":"10.12700\/APH.20.5.2023.5.8","article-title":"Sentiment analysis with neural models for hungarian","volume":"20","author":"Laki L\u00e1szl\u00f3 J\u00e1nos","year":"2023","unstructured":"L\u00e1szl\u00f3 J\u00e1nos Laki and Zijian Gy\u0151z\u0151 Yang. 2023. Sentiment analysis with neural models for hungarian. Acta Polytechnica Hungarica 20, 5 (2023), 109\u2013128.","journal-title":"Acta Polytechnica Hungarica"},{"issue":"4","key":"e_1_3_2_30_2","first-page":"1234","article-title":"BioBERT: A pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee Jinhyuk","year":"2020","unstructured":"Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics (Oxford, England) 36, 4 (2020), 1234\u20131240.","journal-title":"Bioinformatics (Oxford, England)"},{"key":"e_1_3_2_31_2","unstructured":"Eric Lehman Evan Hernandez Diwakar Mahajan Jonas Wulff Micah J. Smith Zachary Ziegler Daniel Nadler Peter Szolovits Alistair Johnson and Emily Alsentzer. 2023. Do we still need clinical language models? arXiv:2302.08091. Retrieved from https:\/\/arxiv.org\/abs\/2302.08091"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Yoav Levine Barak Lenz Or Dagan Ori Ram Dan Padnos Or Sharir Shai Shalev-Shwartz Amnon Shashua and Yoav Shoham. 2019. SenseBERT: Driving some sense into BERT. arXiv:1908.05646. Retrieved from https:\/\/arxiv.org\/abs\/1908.05646","DOI":"10.18653\/v1\/2020.acl-main.423"},{"issue":"17","key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"9857","DOI":"10.3390\/app13179857","article-title":"Integrating text classification into topic discovery using semantic embedding models","volume":"13","author":"Lezama-S\u00e1nchez Ana Laura","year":"2023","unstructured":"Ana Laura Lezama-S\u00e1nchez, Mireya Tovar Vidal, and Jos\u00e9 A. Reyes-Ortiz. 2023. Integrating text classification into topic discovery using semantic embedding models. Applied Sciences 13, 17 (2023), 9857.","journal-title":"Applied Sciences"},{"key":"e_1_3_2_34_2","unstructured":"Zhuoyan Li Hangxiao Zhu Zhuoran Lu and Ming Yin. 2023. Synthetic data generation with large language models for text classification: Potential and limitations. arXiv:2310.07849. Retrieved from https:\/\/arxiv.org\/abs\/2310.07849"},{"key":"e_1_3_2_35_2","first-page":"74","article-title":"Rouge: A package for automatic evaluation of summaries","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out, 74\u201381.","journal-title":"Text Summarization Branches Out"},{"key":"e_1_3_2_36_2","unstructured":"Qianchu Liu Stephanie Hyland Shruthi Bannur Kenza Bouzid Daniel C. Castro Maria Teodora Wetscherek Robert Tinn Harshita Sharma Fernando P\u00e9rez-Garc\u00eda Anton Schwaighofer et al. 2023. Exploring the boundaries of GPT-4 in radiology. arXiv:2310.14573. Retrieved from https:\/\/arxiv.org\/abs\/2310.14573"},{"key":"e_1_3_2_37_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_38_2","unstructured":"Hengyu Luo Peng Liu and Stefan Esping. 2023. Exploring small language models with prompt-learning paradigm for efficient domain-specific text classification. arXiv:2309.14779. Retrieved from https:\/\/arxiv.org\/abs\/2309.14779"},{"issue":"6","key":"e_1_3_2_39_2","article-title":"BioGPT: Generative pre-trained transformer for biomedical text generation and mining","volume":"23","author":"Luo Renqian","year":"2022","unstructured":"Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. BioGPT: Generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics 23, 6 (2022), bbac409.","journal-title":"Briefings in Bioinformatics"},{"key":"e_1_3_2_40_2","first-page":"259","volume-title":"Proceedings of the 2022 International Conference on Asian Language Processing (IALP)","author":"Ma Hangchao","year":"2022","unstructured":"Hangchao Ma, You Zhang, and Jin Wang. 2022. Pretrained models with adversarial training for named entity recognition in scientific text. In Proceedings of the 2022 International Conference on Asian Language Processing (IALP). IEEE, 259\u2013264."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3223094"},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"Nikita Nangia Clara Vania Rasika Bhalerao and Samuel R. Bowman. 2020. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. arXiv:2010.00133. Retrieved from https:\/\/arxiv.org\/abs\/2010.00133","DOI":"10.18653\/v1\/2020.emnlp-main.154"},{"issue":"1","key":"e_1_3_2_43_2","first-page":"1","article-title":"Benchmarking for biomedical natural language processing tasks with a domain specific albert","volume":"23","author":"Naseem Usman","year":"2022","unstructured":"Usman Naseem, Adam G. Dunn, Matloob Khushi, and Jinman Kim. 2022. Benchmarking for biomedical natural language processing tasks with a domain specific albert. BMC Bioinformatics 23, 1 (2022), 1\u201315.","journal-title":"BMC Bioinformatics"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1186\/s13643-016-0384-4"},{"key":"e_1_3_2_45_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311\u2013318."},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","first-page":"258","DOI":"10.18653\/v1\/2023.nllp-1.25","volume-title":"Proceedings of the Natural Legal Language Processing Workshop 2023","author":"Parizi Ali Hakimi","year":"2023","unstructured":"Ali Hakimi Parizi, Yuyang Liu, Prudhvi Nokku, Sina Gholamian, and David Emerson. 2023. A comparative study of prompting strategies for legal text classification. In Proceedings of the Natural Legal Language Processing Workshop 2023, 258\u2013265."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.nlp.2023.100033"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11192-022-04602-4"},{"key":"e_1_3_2_49_2","unstructured":"Ying Sheng Lianmin Zheng Binhang Yuan Zhuohan Li Max Ryabinin Beidi Chen Percy Liang Christopher Re Ion Stoica and Ce Zhang. 2023. FlexGen: high-throughput generative inference of large language models with a single GPU. In Proceedings of the 40th International Conference on Machine Learning (ICML'23) Vol. 202 Article 1288 31094\u201331116."},{"key":"e_1_3_2_50_2","unstructured":"Mayank Soni and Vincent Wade. 2023. Comparing abstractive summaries generated by ChatGPT to real summaries through blinded reviewers and text classification algorithms. arXiv:2303.17650. Retrieved from https:\/\/arxiv.org\/abs\/2303.17650"},{"key":"e_1_3_2_51_2","unstructured":"Benjamin Spector and Chris Re. 2023. Accelerating LLM inference with staged speculative decoding. arXiv:2308.04623. Retrieved from https:\/\/arxiv.org\/abs\/2308.04623"},{"key":"e_1_3_2_52_2","unstructured":"Xiaofei Sun Xiaoya Li Jiwei Li Fei Wu Shangwei Guo Tianwei Zhang and Guoyin Wang. 2023. Text classification via large language models. arXiv:2305.08377. Retrieved from https:\/\/arxiv.org\/abs\/2305.08377"},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","unstructured":"Nicol\u00f2 Tamagnone Selim Fekih Ximena Contla Nayid Orozco and Navid Rekabsaz. 2023. Leveraging domain knowledge for inclusive and bias-aware humanitarian response entry classification. arXiv:2305.16756. Retrieved from https:\/\/arxiv.org\/abs\/2305.16756","DOI":"10.24963\/ijcai.2023\/690"},{"issue":"10","key":"e_1_3_2_54_2","doi-asserted-by":"crossref","first-page":"1657","DOI":"10.1093\/jamia\/ocad133","article-title":"Inferring cancer disease response from radiology reports using large language models with data augmentation and prompting","volume":"30","author":"Tan Ryan Shea Ying Cong","year":"2023","unstructured":"Ryan Shea Ying Cong Tan, Qian Lin, Guat Hwa Low, Ruixi Lin, Tzer Chew Goh, Christopher Chu En Chang, Fung Fung Lee, Wei Yin Chan, Wei Chong Tan, Han Jieh Tey, et al. 2023. Inferring cancer disease response from radiology reports using large language models with data augmentation and prompting. Journal of the American Medical Informatics Association: JAMIA 30, 10 (2023), 1657\u20131664.","journal-title":"Journal of the American Medical Informatics Association: JAMIA"},{"key":"e_1_3_2_55_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_2_56_2","first-page":"384","volume-title":"Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics","author":"Turian Joseph","year":"2010","unstructured":"Joseph Turian, Lev Ratinov, and Yoshua Bengio. 2010. Word representations: A simple and general method for semi-supervised learning. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, 384\u2013394."},{"key":"e_1_3_2_57_2","first-page":"460","volume-title":"Proceedings of the 2022 25th International Conference on Computer and Information Technology (ICCIT)","author":"Usha Mahamuda Sultana","year":"2022","unstructured":"Mahamuda Sultana Usha, Afia Mukarrama Smrity, and Sunanda Das. 2022. Named entity recognition using transfer learning with the fusion of pre-trained SciBERT language model and bi-directional long short term memory. In Proceedings of the 2022 25th International Conference on Computer and Information Technology (ICCIT). IEEE, 460\u2013465."},{"key":"e_1_3_2_58_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_59_2","first-page":"1159","volume-title":"Proceedings of the Science and Information Conference","author":"Wahba Yasmen","year":"2023","unstructured":"Yasmen Wahba, Nazim Madhavji, and John Steinbacher. 2023. Attention is not always what you need: Towards efficient classification of domain-specific text: Case-study: IT support tickets. In Proceedings of the Science and Information Conference. Springer, 1159\u20131166."},{"key":"e_1_3_2_60_2","doi-asserted-by":"crossref","unstructured":"Fengjun Wang Moran Beladev Ofri Kleinfeld Elina Frayerman Tal Shachar Eran Fainman Karen Lastmann Assaraf Sarai Mizrachi and Benjamin Wang. 2023. Text2Topic: Multi-label text classification system for efficient topic detection in user generated content with zero-shot capabilities. arXiv:2310.14817. Retrieved from https:\/\/arxiv.org\/abs\/2310.14817","DOI":"10.18653\/v1\/2023.emnlp-industry.10"},{"key":"e_1_3_2_61_2","unstructured":"Junchao Wu Shu Yang Runzhe Zhan Yulin Yuan Derek F. Wong and Lidia S. Chao. 2023. A survey on LLM-gernerated text detection: Necessity methods and future directions. arXiv:2310.14724. Retrieved from https:\/\/arxiv.org\/abs\/2310.14724"},{"key":"e_1_3_2_62_2","unstructured":"Shijie Wu Ozan Irsoy Steven Lu Vadim Dabravolski Mark Dredze Sebastian Gehrmann Prabhanjan Kambadur David Rosenberg and Gideon Mann. 2023. Bloomberggpt: A large language model for finance. arXiv:2303.17564. Retrieved from https:\/\/arxiv.org\/abs\/2303.17564"},{"issue":"4","key":"e_1_3_2_63_2","doi-asserted-by":"crossref","first-page":"e210258","DOI":"10.1148\/ryai.210258","article-title":"RadBERT: Adapting transformer-based language models to radiology","volume":"4","author":"Yan An","year":"2022","unstructured":"An Yan, Julian McAuley, Xing Lu, Jiang Du, Eric Y. Chang, Amilcare Gentili, and Chun-Nan Hsu. 2022. RadBERT: Adapting transformer-based language models to radiology. Radiology. Artificial Intelligence 4, 4 (2022), e210258.","journal-title":"Radiology. Artificial Intelligence"},{"issue":"5","key":"e_1_3_2_64_2","doi-asserted-by":"crossref","first-page":"169","DOI":"10.12700\/APH.20.5.2023.5.11","article-title":"Training experimental language models with low resources, for the Hungarian language","volume":"20","author":"Yang Zijian Gy\u0151z\u0151","year":"2023","unstructured":"Zijian Gy\u0151z\u0151 Yang and Tam\u00e1s V\u00e1radi. 2023. Training experimental language models with low resources, for the Hungarian language. Acta Polytechnica Hungarica 20, 5 (2023), 169\u2013188.","journal-title":"Acta Polytechnica Hungarica"},{"key":"e_1_3_2_65_2","unstructured":"Hao Yu Zachary Yang Kellin Pelrine Jean Francois Godbout and Reihaneh Rabbany. 2023. Open closed or small language models for text classification? arXiv:2308.10092. Retrieved from https:\/\/arxiv.org\/abs\/2308.10092"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"Yue Yu Yuchen Zhuang Rongzhi Zhang Yu Meng Jiaming Shen and Chao Zhang. 2023. ReGen: Zero-shot text classification via training data generation with progressive dense retrieval. arXiv:2305.10703. Retrieved from https:\/\/arxiv.org\/abs\/2305.10703","DOI":"10.18653\/v1\/2023.findings-acl.748"},{"key":"e_1_3_2_67_2","doi-asserted-by":"crossref","unstructured":"Hongyi Yuan Zheng Yuan Ruyi Gan Jiaxing Zhang Yutao Xie and Sheng Yu. 2022. BioBART: Pretraining and evaluation of a biomedical generative language model. arXiv:2204.03905. Retrieved from https:\/\/arxiv.org\/abs\/2204.03905","DOI":"10.18653\/v1\/2022.bionlp-1.9"},{"key":"e_1_3_2_68_2","first-page":"344","volume-title":"Proceedings of the 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS)","author":"Zhai Yujia","year":"2023","unstructured":"Yujia Zhai, Chengquan Jiang, Leyuan Wang, Xiaoying Jia, Shang Zhang, Zizhong Chen, Xin Liu, and Yibo Zhu. 2023. Bytetransformer: A high-performance transformer boosted for variable-length inputs. In Proceedings of the 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 344\u2013355."},{"key":"e_1_3_2_69_2","unstructured":"Tianyi Zhang Varsha Kishore Felix Wu Kilian Q. Weinberger and Yoav Artzi. 2019. Bertscore: Evaluating text generation with BERT. arXiv:1904.09675. Retrieved from https:\/\/arxiv.org\/abs\/1904.09675"},{"key":"e_1_3_2_70_2","unstructured":"Wei Zhao Rahul Singh Tarun Joshi Agus Sudjianto and Vijayan N. Nair. 2021. Self-interpretable convolutional neural networks for text classification. arXiv:2105.08589. Retrieved from https:\/\/arxiv.org\/abs\/2105.08589"},{"key":"e_1_3_2_71_2","unstructured":"Wayne Xin Zhao Kun Zhou Junyi Li Tianyi Tang Xiaolei Wang Yupeng Hou Yingqian Min Beichen Zhang Junjie Zhang Zican Dong et al. 2023. A survey of large language models. arXiv:2303.18223. Retrieved from https:\/\/arxiv.org\/abs\/2303.18223"},{"key":"e_1_3_2_72_2","unstructured":"Xiaoyan Zhao Yang Deng Min Yang Lingzhi Wang Rui Zhang Hong Cheng Wai Lam Ying Shen and Ruifeng Xu. 2023. A comprehensive survey on deep learning for relation extraction: Recent advances and new Frontiers. arXiv:2306.02051. Retrieved from https:\/\/arxiv.org\/abs\/2306.02051"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3763002","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T13:39:20Z","timestamp":1760708360000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3763002"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,17]]},"references-count":71,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,12,31]]}},"alternative-id":["10.1145\/3763002"],"URL":"https:\/\/doi.org\/10.1145\/3763002","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,17]]},"assertion":[{"value":"2024-07-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}