{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,22]],"date-time":"2026-06-22T14:49:54Z","timestamp":1782139794832,"version":"3.54.5"},"reference-count":47,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2025,9,7]],"date-time":"2025-09-07T00:00:00Z","timestamp":1757203200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2423235"],"award-info":[{"award-number":["2423235"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Large language models (LLMs) often fail to correctly associate biomedical terms with their standardized ontology identifiers, posing challenges for downstream applications that rely on accurate, machine-readable codes. These linking failures can compromise the integrity of data used in precision medicine, clinical decision support, and population health. Fine-tuning can partially remedy these issues, but the degree of improvement varies across terms and terminologies. Focusing on the Human Phenotype Ontology (HPO), we show that a model\u2019s prior knowledge of term\u2013identifier pairs, acquired during pre-training, strongly predicts whether fine-tuning will enhance its linking accuracy. We evaluate prior knowledge in three complementary ways: (1) latent probabilistic knowledge, revealed through stochastic prompting, captures hidden associations not evident in deterministic output; (2) partial subtoken knowledge, reflected in incomplete but non-random generation of identifier components; and (3) term familiarity, inferred from annotation frequencies in the biomedical literature, which serve as a proxy for training exposure. We then assess how these forms of prior knowledge influence the accuracy of deterministic identifier linking. Fine-tuning performance varies most for terms in what we call the reactive middle zone of the ontology\u2014terms with intermediate levels of prior knowledge that are neither absent nor fully consolidated. Fine-tuning was most successful when prior knowledge as measured by partial subtoken knowledge, was \u2018weak\u2019 or \u2018medium\u2019 or when prior knowledge as measured by latent probabilistic knowledge was \u2018unknown\u2019 or \u2018weak\u2019 (p&lt;0.001). These terms from the \u2018reactive middle\u2019 exhibited the largest gains or losses in accuracy during fine-tuning, suggesting that the success of knowledge injection critically depends on the level of term\u2013identifier pair knowledge in the LLM before fine-tuning.<\/jats:p>","DOI":"10.3390\/info16090776","type":"journal-article","created":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T09:32:01Z","timestamp":1757496721000},"page":"776","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Prior Knowledge Shapes Success When Large Language Models Are Fine-Tuned for Biomedical Term Normalization"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6179-0793","authenticated-orcid":false,"given":"Daniel B.","family":"Hier","sequence":"first","affiliation":[{"name":"Department of Neurology & Rehabilitation, University of Illinois at Chicago, Chicago, IL 60612, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8270-1932","authenticated-orcid":false,"given":"Steven K.","family":"Platt","sequence":"additional","affiliation":[{"name":"Laboratory for Applied Artificial Intelligence, Loyola University Chicago, Chicago, IL 60611, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8836-7944","authenticated-orcid":false,"given":"Anh","family":"Nguyen","sequence":"additional","affiliation":[{"name":"Laboratory for Applied Artificial Intelligence, Loyola University Chicago, Chicago, IL 60611, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,9,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"bav005","DOI":"10.1093\/database\/bav005","article-title":"Automatic concept recognition using the human phenotype ontology reference and test suite corpora","volume":"2015","author":"Groza","year":"2015","journal-title":"Database"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Fu, S., Chen, D., He, H., Liu, S., Moon, S., Peterson, K.J., Shen, F., Wang, L., Wang, Y., and Wen, A. (2020). Clinical concept extraction: A methodology review. J. Biomed. Inform., 109.","DOI":"10.1016\/j.jbi.2020.103526"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1529-e1","DOI":"10.1093\/jamia\/ocaa106","article-title":"The 2019 n2c2\/UMass Lowell shared task on clinical concept normalization","volume":"27","author":"Luo","year":"2020","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1016\/j.jbi.2017.11.011","article-title":"Clinical information extraction applications: A literature review","volume":"77","author":"Wang","year":"2018","journal-title":"J. Biomed. Inform."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/1472-6947-15-S1-S4","article-title":"Entity linking for biomedical literature","volume":"15","author":"Zheng","year":"2015","journal-title":"BMC Med. Inform. Decis. Mak."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"512","DOI":"10.1016\/j.jbi.2004.08.004","article-title":"Term identification in the biomedical literature","volume":"37","author":"Krauthammer","year":"2004","journal-title":"J. Biomed. Inform."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"777","DOI":"10.1002\/humu.22080","article-title":"Deep phenotyping for precision medicine","volume":"33","author":"Robinson","year":"2012","journal-title":"Hum. Mutat."},{"key":"ref_8","first-page":"223","article-title":"Ergonomic adequacy of university tablet armchairs for male and female: A multigroup item response theory analysis","volume":"1","author":"Bispo","year":"2024","journal-title":"J. Saf. Sustain."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1686","DOI":"10.1002\/int.22357","article-title":"Knowledge modeling: A survey of processes and techniques","volume":"36","author":"Yun","year":"2021","journal-title":"Int. J. Intell. Syst."},{"key":"ref_10","first-page":"9","article-title":"A review of knowledge management models","volume":"2","author":"Haslinda","year":"2009","journal-title":"J. Int. Soc. Res."},{"key":"ref_11","unstructured":"Pan, J.Z., Razniewski, S., Kalo, J.C., Singhania, S., Chen, J., Dietze, S., Jabeen, H., Omeliyanenko, J., Zhang, W., and Lissandrini, M. (2023). Large language models and knowledge graphs: Opportunities and challenges. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Abellanosa, A.D., Pereira, E., Lefsrud, L., and Mohamed, Y. (J. Saf. Sustain., 2025). Integrating Knowledge Management and Large Language Models to Advance Construction Job Hazard Analysis: A Systematic Review and Conceptual Framework, J. Saf. Sustain., in press, corrected proof.","DOI":"10.1016\/j.jsasus.2025.05.004"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1038\/s41597-023-01960-3","article-title":"Building a knowledge graph to enable precision medicine","volume":"10","author":"Chandak","year":"2023","journal-title":"Sci. Data"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"6878","DOI":"10.1109\/ACCESS.2024.3524588","article-title":"A Review of Applying Large Language Models in Healthcare","volume":"13","author":"Liu","year":"2024","journal-title":"IEEE Access"},{"key":"ref_15","unstructured":"Wang, Y., Zhao, Y., and Petzold, L. (2023). Are Large Language Models Ready for Healthcare? A Comparative Study on Clinical Language Understanding. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2017","DOI":"10.1093\/jamia\/ocab084","article-title":"The use of SNOMED CT, 2013\u20132020: A literature review","volume":"28","author":"Chang","year":"2021","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"The Gene Ontology Consortium (2019). The gene ontology resource: 20 years and still GOing strong. Nucleic Acids Res., 47, D330\u2013D338.","DOI":"10.1093\/nar\/gky1055"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1178","DOI":"10.1093\/bioinformatics\/bth060","article-title":"Recognizing names in biomedical texts: A machine learning approach","volume":"20","author":"Zhou","year":"2004","journal-title":"Bioinformatics"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"D865","DOI":"10.1093\/nar\/gkw1039","article-title":"The human phenotype ontology in 2017","volume":"45","author":"Vasilevsky","year":"2017","journal-title":"Nucleic Acids Res."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"610","DOI":"10.1016\/j.ajhg.2008.09.017","article-title":"The Human Phenotype Ontology: A tool for annotating and analyzing human hereditary disease","volume":"83","author":"Robinson","year":"2008","journal-title":"Am. J. Hum. Genet."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Jahan, I., Laskar, M.T.R., Peng, C., and Huang, J.X. (2024). A comprehensive evaluation of large language models on benchmark biomedical text processing tasks. Comput. Biol. Med., 171.","DOI":"10.1016\/j.compbiomed.2024.108189"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Do, T.S., Hier, D.B., and Obafemi-Ajayi, T. (2025, January 5\u20137). Mapping Biomedical Ontology Terms to IDs: Effect of Domain Prevalence on Prediction Accuracy. Proceedings of the 2025 IEEE Conference on Artificial Intelligence (CAI), Santa Clara, CA, USA.","DOI":"10.1109\/CAI64502.2025.00101"},{"key":"ref_23","unstructured":"Kandpal, N., Deng, H., Roberts, A., Wallace, E., and Raffel, C. (2023, January 23\u201329). Large language models struggle to learn long-tail knowledge. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA."},{"key":"ref_24","unstructured":"Wu, E., Wu, K., and Zou, J. (2024). FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Braga, M. (2024, January 14\u201318). Personalized Large Language Models through Parameter Efficient Fine-Tuning Techniques. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, Washington, DC, USA.","DOI":"10.1145\/3626772.3657657"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"100729","DOI":"10.1016\/j.patter.2023.100729","article-title":"Fine-tuning large neural language models for biomedical natural language processing","volume":"4","author":"Tinn","year":"2023","journal-title":"Patterns"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"220","DOI":"10.1038\/s42256-023-00626-4","article-title":"Parameter-efficient fine-tuning of large-scale pre-trained language models","volume":"5","author":"Ding","year":"2023","journal-title":"Nat. Mach. Intell."},{"key":"ref_28","unstructured":"Wang, C., Yan, J., Zhang, W., and Huang, J. (2023). Towards Better Parameter-Efficient Fine-Tuning for Large Language Models: A Position Paper. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"AIcs2401155","DOI":"10.1056\/AIcs2401155","article-title":"Limitations of Learning New and Updated Medical Knowledge with Commercial Fine-Tuning Large Language Models","volume":"2","author":"Wu","year":"2025","journal-title":"NEJM AI"},{"key":"ref_30","unstructured":"Mecklenburg, N., Lin, Y., Li, X., Holstein, D., Nunes, L., Malvar, S., Silva, B., Chandra, R., Aski, V., and Yannam, P.K.R. (2024). Injecting New Knowledge Into Large Language Models Via Supervised Fine-Tuning. arXiv."},{"key":"ref_31","unstructured":"Chu, T., Zhai, Y., Yang, J., Tong, S., Xie, S., Schuurmans, D., Le, Q.V., Levine, S., and Ma, Y. (2025). SFT memorizes, RL generalizes: A comparative study of foundation model post-training. arXiv."},{"key":"ref_32","unstructured":"Pan, X., Hahami, E., Zhang, Z., and Sompolinsky, H. (2025). Memorization and Knowledge Injection in Gated LLMs. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Pletenev, S., Marina, M., Moskovskiy, D., Konovalov, V., Braslavski, P., Panchenko, A., and Salnikov, M. (2025). How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?. arXiv.","DOI":"10.18653\/v1\/2025.findings-naacl.243"},{"key":"ref_34","unstructured":"Orgad, H., Toker, M., Gekhman, Z., Reichart, R., Szpektor, I., Kotek, H., and Belinkov, Y. (2025). LLMS Know More Than They Show: On The Intrinsic Representation of LLM Hallucinations. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Gekhman, Z., Yona, G., Aharoni, R., Eyal, M., Feder, A., Reichart, R., and Herzig, J. (2024, January 12\u201316). Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA.","DOI":"10.18653\/v1\/2024.emnlp-main.444"},{"key":"ref_36","unstructured":"Gekhman, Z., David, E.B., Orgad, H., Ofek, E., Belinkov, Y., Szpektor, I., Herzig, J., and Reichart, R. (2025). Inside-out: Hidden factual knowledge in LLMs. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"D789","DOI":"10.1093\/nar\/gku1205","article-title":"OMIM. org: Online Mendelian Inheritance in Man (OMIM\u00ae), an online catalog of human genes and genetic disorders","volume":"43","author":"Amberger","year":"2015","journal-title":"Nucleic Acids Res."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"S3","DOI":"10.1016\/S0035-3787(13)70052-3","article-title":"Orphanet and its consortium: Where to find expert-validated information on rare diseases","volume":"169","author":"Maiella","year":"2013","journal-title":"Rev. Neurol."},{"key":"ref_39","unstructured":"McEntyre, J., and Ostell, J. (2025, September 03). PubMed Central (PMC): An Archive for Literature from Life Sciences Journals, The NCBI Handbook [Internet], Available online: https:\/\/www.ncbi.nlm.nih.gov\/sites\/books\/NBK21087\/pdf\/Bookshelf_NBK21087.pdf."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Hier, D.B., Carrithers, M.A., Platt, S.K., Nguyen, A., Giannopoulos, I., and Obafemi-Ajayi, T. (2025). Preprocessing of Physician Notes by LLMs Improves Clinical Concept Extraction Without Information Loss. Information, 16.","DOI":"10.20944\/preprints202504.1452.v1"},{"key":"ref_41","unstructured":"Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., and Vaughan, A. (2024). The llama 3 herd of models. arXiv."},{"key":"ref_42","first-page":"5998","article-title":"Attention is all you need","volume":"Volume 30","author":"Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_43","first-page":"3","article-title":"Lora: Low-rank adaptation of large language models","volume":"1","author":"Hu","year":"2022","journal-title":"ICLR"},{"key":"ref_44","unstructured":"Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y. (2023). An empirical study of catastrophic forgetting in large language models during continual fine-tuning. arXiv."},{"key":"ref_45","unstructured":"Kalajdzievski, D. (2024). Scaling laws for forgetting when fine-tuning large language models. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"2076","DOI":"10.1093\/jamia\/ocae133","article-title":"Fine-tuning large language models for rare disease concept normalization","volume":"31","author":"Wang","year":"2024","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_47","first-page":"10088","article-title":"Qlora: Efficient finetuning of quantized llms","volume":"36","author":"Dettmers","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/9\/776\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:41:29Z","timestamp":1760035289000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/9\/776"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,7]]},"references-count":47,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2025,9]]}},"alternative-id":["info16090776"],"URL":"https:\/\/doi.org\/10.3390\/info16090776","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,7]]}}}