{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T03:17:16Z","timestamp":1783048636467,"version":"3.54.6"},"reference-count":27,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T00:00:00Z","timestamp":1773360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Clinical discharge summaries contain rich patient information but remain difficult to convert into structured representations for downstream analysis. Recent advances in large language models (LLMs) have introduced new approaches for clinical text extraction, yet their relative strengths compared with supervised methods remain unclear. This study presents a controlled evaluation of three dominant strategies for structured clinical information extraction from electronic health records: prompting-based extraction using LLMs, retrieval-augmented generation for terminology canonicalization, and supervised fine-tuning of domain-specific transformer models. Using discharge summaries from the MIMIC-IV dataset, we compare zero-shot, few-shot, and verification-based prompting across closed-source and open-source LLMs, evaluate retrieval-augmented canonicalization as a post-processing mechanism, and benchmark these methods against a fine-tuned BioClinicalBERT model. Performance is assessed using a multi-level evaluation framework that combines exact matching, fuzzy lexical matching, and semantic assessment via an LLM-based judge. The results reveal clear tradeoffs across approaches: prompting achieves strong semantic correctness with minimal supervision, retrieval augmentation improves terminology consistency without expanding extraction coverage, and supervised fine-tuning yields the highest overall accuracy when labeled data are available. Across all methods, we observe a consistent 40\u221250% gap between exact-match and semantic correctness, highlighting the limitations of string-based metrics for clinical Natural Language Processing (NLP). These findings provide practical guidance for selecting extraction strategies under varying resource constraints and emphasize the importance of evaluation methodologies that reflect clinical equivalence rather than surface-form similarity.<\/jats:p>","DOI":"10.3390\/a19030215","type":"journal-article","created":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T09:26:38Z","timestamp":1773393998000},"page":"215","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Understanding Tradeoffs in Clinical Text Extraction: Prompting, Retrieval-Augmented Generation, and Supervised Learning on Electronic Health Records"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-9544-5834","authenticated-orcid":false,"given":"Tanya","family":"Yadav","sequence":"first","affiliation":[{"name":"Department of Applied Data Science, San Jose State University, San Jose, CA 95192, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9389-3723","authenticated-orcid":false,"given":"Aditya","family":"Tekale","sequence":"additional","affiliation":[{"name":"Department of Applied Data Science, San Jose State University, San Jose, CA 95192, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeff","family":"Chong","sequence":"additional","affiliation":[{"name":"Department of Applied Data Science, San Jose State University, San Jose, CA 95192, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9974-6950","authenticated-orcid":false,"given":"Mohammad","family":"Masum","sequence":"additional","affiliation":[{"name":"Department of Applied Data Science, San Jose State University, San Jose, CA 95192, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,3,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1038\/s41597-022-01899-x","article-title":"MIMIC-IV, a freely accessible electronic health record dataset","volume":"10","author":"Johnson","year":"2023","journal-title":"Sci. Data"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Kundeti, S.R., Vijayananda, J., Mujjiga, S., and Kalyan, M. (2016, January 5\u20138). Clinical named entity recognition: Challenges and opportunities. Proceedings of the 2016 IEEE International Conference on Big Data (Big Data), Washington, DC, USA.","DOI":"10.1109\/BigData.2016.7840814"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"296","DOI":"10.1197\/jamia.M1733","article-title":"Agreement, the f-measure, and reliability in information retrieval","volume":"12","author":"Hripcsak","year":"2005","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_4","unstructured":"Koroteev, M.V. (2021). BERT: A review of applications in natural language processing and understanding. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"BioBERT: A pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2020","journal-title":"Bioinformatics"},{"key":"ref_6","first-page":"2","article-title":"Domain-specific language model pretraining for biomedical natural language processing","volume":"3","author":"Gu","year":"2021","journal-title":"ACM Trans. Comput. Healthc."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1935","DOI":"10.1093\/jamia\/ocaa189","article-title":"Clinical concept extraction using transformers","volume":"27","author":"Yang","year":"2020","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_8","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_9","unstructured":"Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., and Metzler, D. (2022). Emergent abilities of large language models. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"e55318","DOI":"10.2196\/55318","article-title":"An empirical evaluation of prompting strategies for large language models in zero-shot clinical natural language processing: Algorithm development and validation study","volume":"12","author":"Sivarajkumar","year":"2024","journal-title":"JMIR Med. Inform."},{"key":"ref_11","unstructured":"Monajatipoor, M., Yang, J., Stremmel, J., Emami, M., Mohaghegh, F., Rouhsedaghat, M., and Chang, K.W. (2024). LLMs in biomedicine: A study on clinical named entity recognition. arXiv."},{"key":"ref_12","first-page":"9459","article-title":"Retrieval-augmented generation for knowledge-intensive NLP tasks","volume":"33","author":"Lewis","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"20552076251337177","DOI":"10.1177\/20552076251337177","article-title":"Enhancing medical AI with retrieval-augmented generation: A mini narrative review","volume":"11","author":"Gargari","year":"2025","journal-title":"Digit. Health"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. (2023). G-Eval: NLG evaluation using GPT-4 with better human alignment. arXiv.","DOI":"10.18653\/v1\/2023.emnlp-main.153"},{"key":"ref_15","unstructured":"Gera, A., Boni, O., Perlitz, Y., Bar-Haim, R., Eden, L., and Yehudai, A. (August, January 27). JuStRank: Benchmarking LLM judges for system ranking. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria."},{"key":"ref_16","first-page":"685","article-title":"Detection of medication mentions and medication change events in clinical notes using transformer-based models","volume":"310","author":"Guo","year":"2024","journal-title":"Stud. Health Technol. Inform."},{"key":"ref_17","first-page":"22199","article-title":"Large language models are zero-shot reasoners","volume":"35","author":"Kojima","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_18","unstructured":"Zhou, D., Sch\u00e4rli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., and Le, Q. (2022). Least-to-most prompting enables complex reasoning in large language models. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., and Weston, J. (2024, January 11\u201316). Chain-of-verification reduces hallucination in large language models. Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand. Available online: https:\/\/aclanthology.org\/2024.findings-acl.212.pdf.","DOI":"10.18653\/v1\/2024.findings-acl.212"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Alsentzer, E., Murphy, J., Boag, W., Weng, W.H., Jindi, D., Naumann, T., and McDermott, M. (2019, January 7). Publicly available clinical BERT embeddings. Proceedings of the 2nd Clinical Natural Language Processing Workshop, Minneapolis, MN, USA.","DOI":"10.18653\/v1\/W19-1909"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Rezaei-Dastjerdehei, M.R., Mijani, A., and Fatemizadeh, E. (2020, January 26\u201327). Addressing imbalance in multi-label classification using weighted cross-entropy loss function. Proceedings of the 2020 27th National and 5th International Iranian Conference on Biomedical Engineering (ICBME), Tehran, Iran.","DOI":"10.1109\/ICBME51989.2020.9319440"},{"key":"ref_22","unstructured":"Mosbach, M., Andriushchenko, M., and Klakow, D. (2020). On the stability of fine-tuning BERT: Misconceptions, explanations, and strong baselines. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"e215","DOI":"10.1161\/01.CIR.101.23.e215","article-title":"PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals","volume":"101","author":"Goldberger","year":"2000","journal-title":"Circulation"},{"key":"ref_24","unstructured":"Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning, MIT Press."},{"key":"ref_25","unstructured":"White, J. (2023). A prompt pattern catalog to enhance prompt engineering with ChatGPT. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1186\/s40537-025-01077-x","article-title":"Survey on terminology extraction from texts","volume":"12","author":"Xu","year":"2025","journal-title":"J. Big Data"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Amugongo, L.M., Mascheroni, P., Brooks, S., Doering, S., and Seidel, J. (2025). Retrieval-augmented generation for large language models in healthcare: A systematic review. PLoS Digit. Health, 4.","DOI":"10.1371\/journal.pdig.0000877"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/3\/215\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T09:28:19Z","timestamp":1773394099000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/3\/215"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,13]]},"references-count":27,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2026,3]]}},"alternative-id":["a19030215"],"URL":"https:\/\/doi.org\/10.3390\/a19030215","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,13]]}}}