{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T18:32:04Z","timestamp":1779906724440,"version":"3.53.1"},"reference-count":38,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T00:00:00Z","timestamp":1763683200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["JCP"],"abstract":"<jats:p>Phishing attacks, particularly Smishing (SMS phishing), have become a major cybersecurity threat, with attackers using social engineering tactics to take advantage of human vulnerabilities. Traditional detection models often struggle to keep up with the evolving sophistication of these attacks, especially on devices with constrained computational resources. This research proposes a chain transformer model that integrates GPT-2 for synthetic data generation and BERT for embeddings to detect Smishing within a multiclass dataset, including minority smishing variants. By utilizing compact, open-source transformer models designed to balance accuracy and efficiency, this study explores improved detection of phishing threats on text-based platforms. Experimental results demonstrate an accuracy rate exceeding 97% in detecting phishing attacks across multiple categories. The proposed chained transformer model achieved an F1-score of 0.97, precision of 0.98, and recall of 0.96, indicating strong overall performance.<\/jats:p>","DOI":"10.3390\/jcp5040102","type":"journal-article","created":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T11:04:31Z","timestamp":1763723071000},"page":"102","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Deep Learning Approaches for Multi-Class Classification of Phishing Text Messages"],"prefix":"10.3390","volume":"5","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-9697-2502","authenticated-orcid":false,"given":"Miriam L.","family":"Munoz","sequence":"first","affiliation":[{"name":"Department of Engineering Management and Systems Engineering, The George Washington University, Washington, DC 20052, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Muhammad F.","family":"Islam","sequence":"additional","affiliation":[{"name":"Department of Engineering Management and Systems Engineering, The George Washington University, Washington, DC 20052, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,11,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Alkhalil, Z., Hewage, C., Nawaf, L., and Khan, I. (2021). Phishing attacks: A recent comprehensive study and a new anatomy. Front. Comput. Sci., 3.","DOI":"10.3389\/fcomp.2021.563060"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Gupta, M., Bakliwal, A., Agarwal, S., and Mehndiratta, P. (2018, January 2\u20134). A comparative study of spam SMS detection using machine learning classifiers. Proceedings of the 2018 Eleventh International Conference on Contemporary Computing (IC3), Noida, India.","DOI":"10.1109\/IC3.2018.8530469"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Pant, V.K., Pant, J., Singh, R.K., and Srivastava, S. (2024). Social Engineering in the Digital Age: A Critical Examination of Attack Techniques, Consequences, and Preventative Measures. Effective Strategies for Combatting Social Engineering in Cybersecurity, IGI Global.","DOI":"10.4018\/979-8-3693-6665-3.ch003"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Hummer, D., and Byrne, J. (2023). Phishing for profit. Handbook on Crime and Technology, Edward Elgar Publishing.","DOI":"10.4337\/9781800886643"},{"key":"ref_5","unstructured":"FTC (2025, May 05). New FTC Data Analysis Shows Bank Impersonation is Most-Reported Text Message Scam, Available online: https:\/\/www.ftc.gov\/news-events\/news\/press-releases\/2023\/06\/new-ftc-data-analysis-shows-bank-impersonation-most-reported-text-message-scam."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"232","DOI":"10.1016\/j.jksuci.2019.12.005","article-title":"A predictive model for phishing detection","volume":"34","author":"Orunsolu","year":"2022","journal-title":"J. King Saud Univ. Comput. Inf. Sci."},{"key":"ref_7","unstructured":"Sengupta, P., Zhang, Y., Maharjan, S., and Eliassen, F. (2023). Balancing explainability\u2014Accuracy of complex models. arXiv, Available online: http:\/\/arxiv.org\/abs\/2305.14098."},{"key":"ref_8","unstructured":"Abra-ham, A., Hanne, T., Gandhi, N., Mishra, P.M., Bajaj, A., and Siarry, P. SMS phishing dataset for machine learning and pattern recognition. Proceedings of the 14th International Conference on Soft Computing and Pattern Recognition (SoCPaR 2022)."},{"key":"ref_9","unstructured":"Brownlee, J. (2024, October 17). Random Oversampling and Undersampling for Imbalanced Classification. Available online: https:\/\/machinelearningmastery.com\/random-oversampling-and-undersampling-for-imbalanced-classification\/."},{"key":"ref_10","first-page":"5436","article-title":"Temporal network embedding with high-order nonlinear in-formation [Conference session]","volume":"34","author":"Qiu","year":"2020","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_11","unstructured":"Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv, Available online: https:\/\/arxiv.org\/abs\/1810.04805."},{"key":"ref_12","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. arXiv, Available online: https:\/\/arxiv.org\/abs\/1706.03762."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Indurkhya, N., and Damerau, F.J. (2010). Handbook of Natural Language Processing, Chapman and Hall\/CRC. [2nd ed.].","DOI":"10.1201\/9781420085938"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"803","DOI":"10.1016\/j.future.2020.03.021","article-title":"Smishing detector: A security model to detect Smishing through SMS content analysis and URL behavior analysis","volume":"108","author":"Mishra","year":"2020","journal-title":"Futur. Gener. Comput. Syst."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1007\/s42979-022-01078-0","article-title":"Implementation of \u2018Smishing Detector\u2019: An efficient model for Smishing detection using neural network","volume":"3","author":"Mishra","year":"2022","journal-title":"SN Comput. Sci."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"4975","DOI":"10.1007\/s00521-021-06305-y","article-title":"DSmishSMS-A system to detect Smishing SMS","volume":"35","author":"Mishra","year":"2023","journal-title":"Neural Comput. Appl."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1143","DOI":"10.1093\/comjnl\/bxy039","article-title":"SmiDCA: An anti-Smishing model with machine learning approach","volume":"61","author":"Sonowal","year":"2018","journal-title":"Comput. J."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1007\/s11235-016-0269-9","article-title":"S-Detector: An enhanced security model for detecting Smishing attack for mobile computing","volume":"66","author":"Joo","year":"2017","journal-title":"Telecommun. Syst."},{"key":"ref_19","unstructured":"Harichandana, B.S.S., Kumar, S., Ujjinakoppa, M.B., and Raja, B.R.K. (2024). COPS: A compact on-device pipe-line for real-time Smishing detection. arXiv, Available online: http:\/\/arxiv.org\/abs\/2402.04173."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Verma, S., Ayala-Rivera, V., and Portillo-Dominguez, A.O. (2023, January 6\u201310). Detection of phishing in mobile instant messaging using natural language processing and machine learning. Proceedings of the 2023 11th International Conference in Software Engineering Research and Innovation (CONISOFT), Guanajuato, Mexico.","DOI":"10.1109\/CONISOFT58849.2023.00029"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3605943","article-title":"Recent advances in natural language processing via large pre-trained language models: A survey","volume":"56","author":"Min","year":"2024","journal-title":"ACM Comput. Surv."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"826","DOI":"10.1162\/tacl_a_00577","article-title":"Efficient methods for natural language processing: A survey","volume":"11","author":"Treviso","year":"2023","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"24306","DOI":"10.1109\/ACCESS.2024.3364671","article-title":"Investigating evasive techniques in SMS spam filtering: A com-parative analysis of machine learning models","volume":"12","author":"Salman","year":"2024","journal-title":"IEEE Access"},{"key":"ref_24","unstructured":"Ma, S. (2023). Enhancing NLP Model Performance Through Data Filtering (Technical Report No. UCB\/EECS-2023-170), University of California."},{"key":"ref_25","unstructured":"Uddin, M.A., Islam, M.N., Maglaras, L., Janicke, H., and Sarker, I.H. (2024). ExplainableDetector: Exploring transformer-based language modeling approach for SMS spam detection with explainability analysis. arXiv, Available online: http:\/\/arxiv.org\/abs\/2405.08026."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Tabani, H., Balasubramaniam, A., Marzban, S., Arani, E., and Zonooz, B. (2021, January 1\u20133). Improving the efficiency of transformers for resource-constrained devices [Conference session]. Proceedings of the 2021 24th Euromicro Conference on Digital System Design (DSD), Palermo, Italy.","DOI":"10.1109\/DSD53832.2021.00074"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"15439","DOI":"10.1007\/s00521-024-09707-w","article-title":"Privacy BERT-LSTM: A novel NLP algorithm for sensitive information detection in textual documents","volume":"36","author":"Muralitharan","year":"2024","journal-title":"Neural Comput. Appl."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Khan, M.A., Huang, Y., Feng, J., Prasad, B.K., Ali, Z., Ullah, I., and Kefalas, P. (2023). A multi-attention approach using BERT and stacked bidirectional LSTM for improved dialogue state tracking. Appl. Sci., 13.","DOI":"10.3390\/app13031775"},{"key":"ref_29","first-page":"547","article-title":"A privacy-preserving approach for detecting smishing attacks using federated deep learning","volume":"17","author":"Remmide","year":"2025","journal-title":"Int. J. Inf. Technol."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Almeida, T.A., Hidalgo, J.M.G., and Yamakami, A. (2011, January 19\u201322). Contributions to the study of SMS Spam filtering: New collection and results. Proceedings of the DocEng \u201911: 11th ACM Symposium on Document Engineering, Mountain View, CA, USA.","DOI":"10.1145\/2034691.2034742"},{"key":"ref_31","unstructured":"Alpaydin, E. (2020). Introduction to Machine Learning, The MIT Press. [4th ed.]. Available online: https:\/\/mitpress.mit.edu\/9780262043793\/introduction-to-machine-learning\/."},{"key":"ref_32","unstructured":"Bishop, C.M. (2006). Pattern Recognition and Machine Learning, Springer. Available online: https:\/\/link.springer.com\/book\/9780387310732."},{"key":"ref_33","unstructured":"G\u00e9ron, A. (2022). Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, O\u2019Reilly Media, Inc.. [3rd ed.]. Available online: https:\/\/www.oreilly.com\/library\/view\/hands-on-machine-learning\/9781492032632\/."},{"key":"ref_34","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_35","unstructured":"Montavon, G., Orr, G.B., and M\u00fcller, K.-R. (2012). Practical recommendations for gradient-based training of deep architectures. Neural Networks: Tricks of the Trade, Springer. [3rd ed.]."},{"key":"ref_36","unstructured":"Johnson, R., and Zhang, T. (2015, January 7\u201312). Semi-supervised convolutional neural networks for text categorization via region embedding [Conference presentation]. Proceedings of the NIPS \u201915: 28th International Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_37","unstructured":"Houston, R.A. (2024). Transformer-Enhanced Text Classification in Cybersecurity: GPT Augmented Synthetic Data Generation, BERT-Based Semantic Encoding, and Multiclass Analysis. [Ph.D. Thesis, The George Washington University]. Available online: https:\/\/search.proquest.com\/openview\/fe2a7d3fb1e4ac4426755c3237663c7c\/1?pqorigsite=gscholar&cbl=18750&diss=y."},{"key":"ref_38","unstructured":"Munoz, M., and Islam, M. (2025, July 07). A Balanced Dataset for Spam and Smishing Detection Using Large Language Models (LLMs). Available online: https:\/\/data.mendeley.com\/datasets\/vmg875v4xs."}],"container-title":["Journal of Cybersecurity and Privacy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2624-800X\/5\/4\/102\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T11:29:35Z","timestamp":1763724575000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2624-800X\/5\/4\/102"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,21]]},"references-count":38,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["jcp5040102"],"URL":"https:\/\/doi.org\/10.3390\/jcp5040102","relation":{},"ISSN":["2624-800X"],"issn-type":[{"value":"2624-800X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,21]]}}}