{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T14:18:15Z","timestamp":1776867495266,"version":"3.51.2"},"reference-count":36,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2024,12,19]],"date-time":"2024-12-19T00:00:00Z","timestamp":1734566400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Council of Science and Technology (CONACYT), Mexico","award":["A1-S-47854"],"award-info":[{"award-number":["A1-S-47854"]}]},{"name":"National Council of Science and Technology (CONACYT), Mexico","award":["20241816"],"award-info":[{"award-number":["20241816"]}]},{"name":"National Council of Science and Technology (CONACYT), Mexico","award":["20241819"],"award-info":[{"award-number":["20241819"]}]},{"name":"National Council of Science and Technology (CONACYT), Mexico","award":["20240951"],"award-info":[{"award-number":["20240951"]}]},{"name":"Secretariat of Research and Postgraduate Studies of the Instituto Polit\u00e9cnico Nacional (IPN), Mexico","award":["A1-S-47854"],"award-info":[{"award-number":["A1-S-47854"]}]},{"name":"Secretariat of Research and Postgraduate Studies of the Instituto Polit\u00e9cnico Nacional (IPN), Mexico","award":["20241816"],"award-info":[{"award-number":["20241816"]}]},{"name":"Secretariat of Research and Postgraduate Studies of the Instituto Polit\u00e9cnico Nacional (IPN), Mexico","award":["20241819"],"award-info":[{"award-number":["20241819"]}]},{"name":"Secretariat of Research and Postgraduate Studies of the Instituto Polit\u00e9cnico Nacional (IPN), Mexico","award":["20240951"],"award-info":[{"award-number":["20240951"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Large language models (LLMs) are tools that help us in a variety of activities, from creating well-structured texts to quickly consulting information. But as these new technologies are so easily accessible, many people use them for their own benefit without properly citing the original author, or in other cases the student sector can be heavily compromised because students may opt for a quick answer over understanding and comprehending a specific topic in depth, considerably reducing their basic writing, editing and reading comprehension skills. Therefore, we propose to create a model to identify texts produced by LLM. To do so, we will use natural language processing (NLP) and machine-learning algorithms to recognize texts that mask LLM misuse using different types of adversarial attack, like paraphrasing or translation from one language to another. The main contributions of this work are to identify the texts generated by the large language models, and for this purpose several experiments were developed looking for the best results implementing the f1, accuracy, recall and precision metrics, together with PCA and t-SNE diagrams to see the classification of each one of the texts.<\/jats:p>","DOI":"10.3390\/computers13120346","type":"journal-article","created":{"date-parts":[[2024,12,19]],"date-time":"2024-12-19T03:59:43Z","timestamp":1734580783000},"page":"346","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Identification of Scientific Texts Generated by Large Language Models Using Machine Learning"],"prefix":"10.3390","volume":"13","author":[{"given":"David","family":"Soto-Osorio","sequence":"first","affiliation":[{"name":"Computing Research Center, Instituto Polit\u00e9cnico Nacional, Av. Juan de Dios B\u00e1tiz S\/N, Nueva Industrial Vallejo, Ciudad de M\u00e9xico 07700, Mexico"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3901-3522","authenticated-orcid":false,"given":"Grigori","family":"Sidorov","sequence":"additional","affiliation":[{"name":"Computing Research Center, Instituto Polit\u00e9cnico Nacional, Av. Juan de Dios B\u00e1tiz S\/N, Nueva Industrial Vallejo, Ciudad de M\u00e9xico 07700, Mexico"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liliana","family":"Chanona-Hern\u00e1ndez","sequence":"additional","affiliation":[{"name":"Escuela Superior de Ingeneria Mecanica y Electrica, Unidad Zacatenco, Instituto Polit\u00e9cnico Nacional, Av. Luis Enrique Erro, S\/N, Ciudad de M\u00e9xico 07700, Mexico"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6403-6641","authenticated-orcid":false,"given":"Blanca Cecilia","family":"L\u00f3pez-Ram\u00edrez","sequence":"additional","affiliation":[{"name":"Tecnol\u00f3gico Nacional de M\u00e9xico\/I.T. de Roque, Celaya 38110, Mexico"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,12,19]]},"reference":[{"key":"ref_1","unstructured":"Hugging Face (2024, April 10). Preprocessing Data with Transformers. Available online: https:\/\/huggingface.co\/docs\/transformers\/preprocessing."},{"key":"ref_2","unstructured":"Interactive Chaos (2024, April 22). Machine Learning Tutorial: One Hot Encoding. Available online: https:\/\/interactivechaos.com\/es\/manual\/tutorial-de-machine-learning\/one-hot-encoding."},{"key":"ref_3","unstructured":"IBM (2024, April 24). Bag of Words. Available online: https:\/\/www.ibm.com\/topics\/bag-of-words."},{"key":"ref_4","unstructured":"Towards Data Science (2024, April 24). Understanding Word N-Grams and N-Gram Probability in Natural Language Processing. Available online: https:\/\/towardsdatascience.com\/understanding-word-n-grams-and-n-gram-probability-in-natural-language-processing-9d9eef0fa058."},{"key":"ref_5","unstructured":"Jain, A. (2024, April 25). TF-IDF in NLP: Term Frequency-Inverse Document Frequency. Available online: https:\/\/medium.com\/@abhishekjainindore24\/tf-idf-in-nlp-term-frequency-inverse-document-frequency-e05b65932f1d."},{"key":"ref_6","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). GloVe: Global Vectors for Word Representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar. Available online: https:\/\/aclanthology.org\/D14-1162\/.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_8","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (June, January 2). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the NAACL-HLT, Minneapolis, MN, USA. Available online: https:\/\/arxiv.org\/abs\/1810.04805."},{"key":"ref_9","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv, Available online: https:\/\/arxiv.org\/abs\/1907.11692."},{"key":"ref_10","unstructured":"LlamaIndex (2024, April 10). Ollama Embedding Example. Available online: https:\/\/llamaindex.ai."},{"key":"ref_11","unstructured":"Scikit-Learn (2024, April 23). Logistic Regression. Available online: https:\/\/scikit-learn.org\/stable\/modules\/linear_model.html#logistic-regression."},{"key":"ref_12","unstructured":"Scikit-Learn (2024, April 17). Random Forest. Available online: https:\/\/scikit-learn.org\/stable\/modules\/ensemble.html#random-forests."},{"key":"ref_13","unstructured":"Scikit-Learn (2024, April 15). Support Vector Machines. Available online: https:\/\/scikit-learn.org\/stable\/modules\/svm.html."},{"key":"ref_14","unstructured":"Scikit-Learn (2024, April 23). K-Nearest Neighbors. Available online: https:\/\/scikit-learn.org\/stable\/modules\/neighbors.html."},{"key":"ref_15","unstructured":"BuiltIn (2024, April 23). What Is a Fully Connected Layer in Machine Learning? BuiltIn Machine Learning Topics. Available online: https:\/\/builtin.com\/machine-learning\/fully-connected-layer."},{"key":"ref_16","unstructured":"ScienceDirect (2024, April 23). Long Short-Term Memory Networks. Available online: https:\/\/www.sciencedirect.com\/topics\/computer-science\/long-short-term-memory-networks."},{"key":"ref_17","unstructured":"Hugging Face (2024, April 23). Chapter 1.4\u2013NLP Tasks. Available online: https:\/\/huggingface.co\/course\/chapter1\/4."},{"key":"ref_18","unstructured":"Analytics Vidhya (2024, April 23). Metrics to Evaluate Your Classification Model to Take the Right Decisions. Available online: https:\/\/www.analyticsvidhya.com\/blog\/2021\/07\/metrics-to-evaluate-your-classification-model-to-take-the-right-decisions\/."},{"key":"ref_19","unstructured":"IBM (2024, April 23). Confusion Matrix: What It Is and How to Use It. Available online: https:\/\/www.ibm.com\/mx-es\/topics\/confusion-matrix."},{"key":"ref_20","unstructured":"IBM (2024, April 25). Principal Component Analysis: What It Is and How to Use It. Available online: https:\/\/www.ibm.com\/think\/topics\/principal-component-analysis."},{"key":"ref_21","unstructured":"IBM (2024, April 26). Creaci\u00f3n de gr\u00e1ficos t-SNE en SPSS Statistics. Available online: https:\/\/www.ibm.com\/docs\/es\/spss-statistics\/beta?topic=sslvmb-subs-statistics-mainhelp-ddita-spss-base-chart-creation-tsne-html."},{"key":"ref_22","unstructured":"Google Developers (2024, April 28). ROC and AUC\u2014Machine Learning Crash Course. Google Machine Learning Crash Course. Available online: https:\/\/developers.google.com\/machine-learning\/crash-course\/classification\/roc-and-auc?hl=es-419."},{"key":"ref_23","unstructured":"Pinecone Learning Hub (2024, July 23). LangChain Prompt Templates. Pinecone.io Documentation. Available online: https:\/\/python.langchain.com\/docs\/integrations\/vectorstores\/pinecone\/."},{"key":"ref_24","unstructured":"FreeCodeCamp (2024, July 30). Fine-Tuning LLM Models\u2014FreeCodeCamp. FreeCodeCamp News. Available online: https:\/\/www.freecodecamp.org\/news\/fine-tuning-llm-models-course\/."},{"key":"ref_25","unstructured":"Amazon Web Services (2024, August 01). What Is Retrieval-Augmented Generation (RAG)? AWS Documentation. Available online: https:\/\/aws.amazon.com\/what-is\/retrieval-augmented-generation\/."},{"key":"ref_26","unstructured":"Sadasivaan, V.S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S. (2023). Can AI-Generated Text be Reliably Detected?. arXiv."},{"key":"ref_27","unstructured":"Wu, J., Yang, S., Zhan, R., Yuan, Y., Wong, D.F., and Chao, L.S. (2023). A Survey on LLM-generated Text Detection: Necessity, Methods, and Future Directions. arXiv."},{"key":"ref_28","unstructured":"Bv, P., Ahmed, S., and Sadanandam, M. (2024). DistilBERT: A Novel Approach to Detect Text Generated by Large Language Models (LLM). arXiv."},{"key":"ref_29","unstructured":"Major, A., Capobianco, M., Reynolds, M., Phelan, C., Shah-Nathwani, K., Luong, D., Lee, K., and Kumaravel, M. (2023). Supervised Machine Generated Text Detection Using LLM Encoders in Various Data Resource Scenarios. [Doctoral Dissertation, Worcester Polytechnic Institute]. Available online: https:\/\/www.semanticscholar.org\/paper\/Supervised-Machine-Generated-Text-Detection-Using-Major-Capobianco\/a79561bad0a5a3f5b0cb3ba9750ad7851369ff2a."},{"key":"ref_30","unstructured":"Blecher, L., Cucurull, G., Scialom, T., and Stojnic, R. (2023). Nougat: Neural Optical Understanding for Academic Documents. arXiv, Available online: https:\/\/arxiv.org\/abs\/2308.13418."},{"key":"ref_31","unstructured":"Ollama-Llama3 (2024, February 25). LLaMA3. Ollama Library. Available online: https:\/\/ollama.com\/library\/llama3."},{"key":"ref_32","unstructured":"Ollama-Llama2 (2024, February 25). Llama2. Ollama Library. Available online: https:\/\/ollama.com\/library\/llama2."},{"key":"ref_33","unstructured":"Ollama-Gemma (2024, February 25). Gemma. Ollama Library. Available online: https:\/\/ollama.com\/library\/gemma."},{"key":"ref_34","unstructured":"Ollama-Llava (2024, February 25). LLaVA. Ollama Library. Available online: https:\/\/ollama.com\/library\/llava."},{"key":"ref_35","unstructured":"Hugging Face-Distilbert (2024, April 20). DistilBERT Model Documentation. Hugging Face Transformers. Available online: https:\/\/huggingface.co\/distilbert."},{"key":"ref_36","unstructured":"Hugging Face (2024, April 25). DistilRoBERTa-Base-Distilroberta. Hugging Face Transformers. Available online: https:\/\/huggingface.co\/distilroberta-base."}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/13\/12\/346\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T16:55:14Z","timestamp":1760115314000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/13\/12\/346"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,19]]},"references-count":36,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["computers13120346"],"URL":"https:\/\/doi.org\/10.3390\/computers13120346","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,19]]}}}