{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,15]],"date-time":"2025-11-15T05:19:57Z","timestamp":1763183997593,"version":"3.45.0"},"reference-count":30,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2025,11,13]],"date-time":"2025-11-13T00:00:00Z","timestamp":1762992000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100013185","name":"Florida Institute of\nTechnology","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100013185","id-type":"DOI","asserted-by":"crossref"}]},{"name":"AI Pipeline for\nData Assimilation"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Large language models (LLMs) such as GPT-4 are increasingly integrated into research, industry, and enterprise workflows, yet little is known about how input file formats shape their outputs. While prior work has shown that formats can influence response time, the effects on readability, complexity, and semantic stability remain underexplored. This study systematically evaluates GPT-4\u2019s responses to 100 queries drawn from 50 academic papers, each tested across four formats, TXT, DOCX, PDF, and XML, yielding 400 question\u2013answer pairs. We have assessed two aspects of the responses to the queries: first, efficiency quantified by response time and answer length, and second, linguistic style measured by readability indices, sentence length, word length, and lexical diversity where semantic similarity was considered to control for preservation of semantic context. Results show that readability and semantic content remain stable across formats, with no significant differences in Flesch\u2013Kincaid or Dale\u2013Chall scores, but response time is sensitive to document encoding, with XML consistently outperforming PDF, DOCX, and TXT in the initial experiments conducted in February 2025. Verbosity, rather than input size, emerged as the main driver of latency. However, follow-up replications conducted several months later (October 2025) under the updated Microsoft Copilot Studio (GPT-4) environment showed that these latency differences had largely converged, indicating that backend improvements, particularly in GPT-4o\u2019s document-ingestion and parsing pipelines, have reduced the earlier disparities. These findings suggest that the file format matters and affects how fast the LLMs respond, although its influence may diminish as enterprise-level AI systems continue to evolve. Overall, the content and semantics of the responses are fairly similar and consistent across different file formats, demonstrating that LLMs can handle diverse encodings without compromising response quality. For large-scale applications, adopting structured formats such as XML or semantically tagged HTML can still yield measurable throughput gains in earlier system versions, whereas in more optimized environments, such differences may become minimal.<\/jats:p>","DOI":"10.3390\/computers14110493","type":"journal-article","created":{"date-parts":[[2025,11,13]],"date-time":"2025-11-13T12:57:13Z","timestamp":1763038633000},"page":"493","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Document Encoding Effects on Large Language Model Response Time and Consistency"],"prefix":"10.3390","volume":"14","author":[{"given":"Dianeliz","family":"Ortiz Martes","sequence":"first","affiliation":[{"name":"Department of Mathematics and Systems Engineering, Florida Institute of Technology, Melbourne, FL 32901, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9397-1807","authenticated-orcid":false,"given":"Nezamoddin N.","family":"Kachouie","sequence":"additional","affiliation":[{"name":"Department of Mathematics and Systems Engineering, Florida Institute of Technology, Melbourne, FL 32901, USA"},{"name":"Department of Electrical Engineering and Computer Science, Florida Institute of Technology, Melbourne, FL 32901, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,11,13]]},"reference":[{"key":"ref_1","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., and Altman, S. (2023). GPT-4 technical report. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Li, W., Duan, M., An, D., and Shao, Y. (2024). Large language models understand layout. arXiv.","DOI":"10.3233\/FAIA240700"},{"key":"ref_3","unstructured":"Microsoft (2025, July 12). Format Guidelines for Custom Question Answering. Microsoft Learn. 21 November 2024., Available online: https:\/\/learn.microsoft.com\/en-us\/azure\/ai-services\/language-service\/question-answering\/reference\/document-format-guidelines."},{"key":"ref_4","unstructured":"Ortiz Martes, D., and Kachouie, N.N. (2025, January 21\u201324). File-format effects on LLMs for local knowledge: A multi-metric study. Proceedings of the CSCE 2025 Congress, Las Vegas, NC, USA."},{"key":"ref_5","unstructured":"Yang, L., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. (2023). G-Eval: NLG evaluation using GPT-4 with better human alignment. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Sellam, T., Das, D., and Parikh, A. (2020, January 5\u201310). BLEURT: Learning robust metrics for text generation. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.704"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Rei, R., Stewart, C., Farinha, A.C., and Lavie, A. (2020, January 16\u201320). COMET: A neural framework for MT evaluation. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online.","DOI":"10.18653\/v1\/2020.emnlp-main.213"},{"key":"ref_8","unstructured":"Kincaid, J.P., Fishburne, R.P., Rogers, R.L., and Chissom, B.S. (2025, May 03). Derivation of New Readability Formulas for Navy Enlisted Personnel. Research Branch Report 8-75. Available online: https:\/\/apps.dtic.mil\/sti\/pdfs\/ADA006655.pdf."},{"key":"ref_9","first-page":"37","article-title":"A formula for predicting readability: Instructions","volume":"27","author":"Dale","year":"1948","journal-title":"Educ. Res. Bull."},{"key":"ref_10","first-page":"206","article-title":"Towards predicting post-editing effort with source text readability: An investigation for English\u2013Chinese machine translation","volume":"38","author":"Dai","year":"2024","journal-title":"Mach. Transl."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"491","DOI":"10.3758\/s13428-022-01802-x","article-title":"A large-scaled corpus for assessing text readability","volume":"55","author":"Crossley","year":"2023","journal-title":"Behav. Res. Methods"},{"key":"ref_12","unstructured":"(2025, August 02). Textstat. (n.d.). Readability Scores Documentation. Available online: https:\/\/pypi.org\/project\/textstat\/."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"323","DOI":"10.1023\/A:1001749303137","article-title":"How Variable May a Constant Be? Measures of lexical richness in perspective","volume":"32","author":"Tweedie","year":"1998","journal-title":"Comput. Humanit."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1016\/j.jslw.2015.06.003","article-title":"Syntactic complexity in college-level English writing","volume":"29","author":"Lu","year":"2015","journal-title":"J. Second. Lang. Writ."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Reimers, N., and Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. arXiv.","DOI":"10.18653\/v1\/D19-1410"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ortiz Martes, D., Gunderson, E., Neuman, C., and Kachouie, N.N. (2025). Transformer models for paraphrase detection: A comprehensive semantic similarity study. Computers, 14.","DOI":"10.3390\/computers14090385"},{"key":"ref_17","unstructured":"Tyagi, S. (2025, June 15). How File Formats can Impact the Performance of LLM-Powered Applications. LinkedIn Articles. 18 January 2025. Available online: https:\/\/www.linkedin.com\/pulse\/how-file-formats-can-impact-performance-llm-powered-applications-saurabh-tyagi\/."},{"key":"ref_18","unstructured":"Gong, C., Ren, X., and Wang, G. (2023). Bytes are all you need: Transformers operating directly on file bytes. arXiv."},{"key":"ref_19","unstructured":"Yamaguchi, F., Gros, A., and Schreck, T. (2022). Toward the detection of polyglot files. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1038\/s41746-023-00773-3","article-title":"The impact of inconsistent human annotations on AI driven clinical decision making","volume":"6","author":"Sylolypavan","year":"2023","journal-title":"npj Digit. Med."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Klie, J.-C., Webber, B., and Gurevych, I. (2022). Annotation error detection: Analyzing the past and present for a more coherent future [Preprint]. arXiv.","DOI":"10.1162\/coli_a_00464"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"876","DOI":"10.1136\/amiajnl-2012-001173","article-title":"Using rule-based natural language processing to improve disease normalization in biomedical text","volume":"20","author":"Kang","year":"2013","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_23","unstructured":"Chalkidis, I., Fergadiotis, M., Malakasiotis, P., and Androutsopoulos, I. (August, January 28). Neural legal judgment prediction in English. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"591","DOI":"10.1093\/biomet\/52.3-4.591","article-title":"An analysis of variance test for normality","volume":"52","author":"Shapiro","year":"1965","journal-title":"Biometrika"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"765","DOI":"10.1080\/01621459.1954.10501232","article-title":"A test of goodness of fit","volume":"49","author":"Anderson","year":"1954","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"583","DOI":"10.1080\/01621459.1952.10483441","article-title":"Use of ranks in one-criterion variance analysis","volume":"47","author":"Kruskal","year":"1952","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"241","DOI":"10.1080\/00401706.1964.10490181","article-title":"Multiple comparisons using rank sums","volume":"6","author":"Dunn","year":"1964","journal-title":"Technometrics"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"50","DOI":"10.1214\/aoms\/1177730491","article-title":"A test whether one of two random variables is stochastically larger","volume":"18","author":"Mann","year":"1947","journal-title":"Ann. Math. Stat."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"72","DOI":"10.2307\/1412159","article-title":"The proof and measurement of association between two things","volume":"15","author":"Spearman","year":"1904","journal-title":"Am. J. Psychol."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"420","DOI":"10.1037\/0033-2909.86.2.420","article-title":"Intraclass correlations: Uses in assessing rater reliability","volume":"86","author":"Shrout","year":"1979","journal-title":"Psychol. Bull."}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/14\/11\/493\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,15]],"date-time":"2025-11-15T05:17:30Z","timestamp":1763183850000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/14\/11\/493"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,13]]},"references-count":30,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2025,11]]}},"alternative-id":["computers14110493"],"URL":"https:\/\/doi.org\/10.3390\/computers14110493","relation":{},"ISSN":["2073-431X"],"issn-type":[{"type":"electronic","value":"2073-431X"}],"subject":[],"published":{"date-parts":[[2025,11,13]]}}}