{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T13:00:25Z","timestamp":1780750825737,"version":"3.54.1"},"reference-count":26,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T00:00:00Z","timestamp":1755907200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Information Science"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:p>The advent of Generative Pre-trained Transformer (GPT) has significantly impacted various downstream natural language processing (NLP) tasks, showcasing its remarkable capabilities in language understanding and generation. It has demonstrated state-of-the-art performance in areas such as machine translation, text summarisation and question answering. In this study, we focus on addressing the issue of skewed data sets, which is a challenge across multiple NLP tasks. Conventional balancing techniques including undersampling and oversampling are fraught with limitations and may result in biased predictions and insufficient representation of minority classes. In response, we propose an innovative approach harnessing GPT\u2019s capabilities to generate synthetic samples. We evaluate our approach on a downstream multi-class text classification task and demonstrate significant performance improvements over conventional techniques and state-of-the-art methods, surpassing the accuracy by more than 10%. These findings underscore the potential of GPT to revolutionise data set balancing, thereby augmenting the performance of downstream NLP tasks.<\/jats:p>","DOI":"10.1177\/01655515251362340","type":"journal-article","created":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T09:54:37Z","timestamp":1755942877000},"page":"908-922","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["Addressing class imbalance in text classification with LLMs: A prompt-based GPT-2 approach"],"prefix":"10.1177","volume":"52","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6517-515X","authenticated-orcid":false,"given":"Laouni","family":"Mahmoudi","sequence":"first","affiliation":[{"name":"LISYS Laboratory, University of Mascara, Algeria"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7052-5978","authenticated-orcid":false,"given":"Mohammed","family":"Salem","sequence":"additional","affiliation":[{"name":"LISYS Laboratory, University of Mascara, Algeria"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nawaf R","family":"Alharbe","sequence":"additional","affiliation":[{"name":"Department of Computer Science, College of Computer Science and Engineering, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2025,8,23]]},"reference":[{"key":"e_1_3_3_2_2","volume-title":"Proceedings of the 2022 Ninth International Conference on Social Networks Analysis, Management and Security (SNAMS)","author":"Abdullah M","unstructured":"Abdullah M, Madain A, Jararweh Y. GPT: fundamentals, applications and social impacts. In: Proceedings of the 2022 Ninth International Conference on Social Networks Analysis, Management and Security (SNAMS), Milan, Italy, 28 November\u20131 December 2022."},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/INTELLECT.2017.8277634"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.3390\/info14010054"},{"key":"e_1_3_3_5_2","unstructured":"OpenAI. GPT: optimizing language models for dialogue. Openai 2022 https:\/\/openai.com\/blog\/GPT\/ (accessed 4 February 2023)."},{"key":"e_1_3_3_6_2","first-page":"264","volume-title":"Artificial Intelligence: Theories and Applications. ICAITA 2022. Communications in Computer and Information Science","author":"Mahmoudi L","year":"2023","unstructured":"Mahmoudi L, Salem M. Improving multi-class text classification using balancing techniques. In: Salem M, Merelo JJ, Siarry P, et al. (eds) Artificial Intelligence: Theories and Applications. ICAITA 2022. Communications in Computer and Information Science. Vol. 1769. Cham: Springer, 2023, pp.264\u2013275."},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.3390\/sym14030567"},{"key":"e_1_3_3_8_2","unstructured":"Alharbi B Alamro H Alshehri M et al. ASAD: a Twitter-based benchmark Arabic sentiment analysis dataset. arXiv 2011.00578 2020."},{"key":"e_1_3_3_9_2","unstructured":"Aly MA Atiya AF. LABR: a large scale Arabic book reviews dataset. arXiv 1411.6718 2013."},{"key":"e_1_3_3_10_2","first-page":"4171","volume-title":"Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, Volume 1 (Long and Short Papers) 2\u20137 June 2019","author":"Devlin J","unstructured":"Devlin J, Chang MW, Lee K, et al. BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, Volume 1 (Long and Short Papers) 2\u20137 June 2019, Minneapolis, MN, pp.4171\u20134186. Stroudsburg, PA: ACL Anthology."},{"key":"e_1_3_3_11_2","unstructured":"Jiao W Wang W Huang JT et al. Is GPT a good translator? A preliminary study. arXiv 2301.08745v3 2023."},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.2139\/ssrn.4390455"},{"key":"e_1_3_3_13_2","unstructured":"Luo Z Xie Q Ananiadou S. GPT as a factual inconsistency evaluator for text summarization. arXiv 2303.15621v2 2023."},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.iotcps.2023.04.003"},{"key":"e_1_3_3_15_2","first-page":"102642","article-title":"\u2018So what if GPT wrote it?\u2019 Multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy","volume":"71","author":"Dwivedi YK","year":"2023","unstructured":"Dwivedi YK, Kshetri N, Hughes L, et al. \u2018So what if GPT wrote it?\u2019 Multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy. Int J Inf Manag 2023; 71: 102642.","journal-title":"Int J Inf Manag"},{"key":"e_1_3_3_16_2","unstructured":"Li Y Bonatti R Abdali S et al. Data generation using large language models for text classification: an empirical case study. arXiv 2407.12813 2024."},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.1974"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2008.239"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1177\/01655515221116517"},{"key":"e_1_3_3_20_2","first-page":"248","volume-title":"European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD)","author":"Bellinger C","unstructured":"Bellinger C, Drummond C. Beyond the boundaries of SMOTE \u2013 a framework for manifold-based synthetically oversampling. In: European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD), Riva del Garda, 19\u201323 September 2016, pp. 248\u2013263."},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-16486-1_22"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-59051-2_14"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2164-12-S4-S9"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2014.09.019"},{"key":"e_1_3_3_25_2","unstructured":"Zhang C Zhang C Zheng S et al. A complete survey on generative AI (AIGC): is GPT from GPT-4 to GPT-5 all you need? arXiv 2303.11717 2023."},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2015.2458858"},{"key":"e_1_3_3_27_2","first-page":"9","volume-title":"Proceedings of the 4th workshop on open-source Arabic corpora and processing tools, with a shared task on offensive language detection","author":"Wissam A","unstructured":"Wissam A, Fady B, Hazem H. AraBERT: transformer-based model for Arabic language understanding. In: Proceedings of the 4th workshop on open-source Arabic corpora and processing tools, with a shared task on offensive language detection, Marseille, 11\u201316 May 2020, pp. 9\u201315. Stroudsburg, PA: ACL Anthology."}],"container-title":["Journal of Information Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/01655515251362340","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/01655515251362340","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/01655515251362340","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T12:24:16Z","timestamp":1780748656000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/01655515251362340"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,23]]},"references-count":26,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["10.1177\/01655515251362340"],"URL":"https:\/\/doi.org\/10.1177\/01655515251362340","relation":{},"ISSN":["0165-5515","1741-6485"],"issn-type":[{"value":"0165-5515","type":"print"},{"value":"1741-6485","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,23]]}}}