{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T11:18:39Z","timestamp":1776424719070,"version":"3.51.2"},"reference-count":31,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T00:00:00Z","timestamp":1776384000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Large Language Models (LLMs) have witnessed significant adoption across numerous domains since 2020, but their proclivity to hallucinate creates unacceptable dangers in high-risk environments like healthcare, where wrong outputs can directly jeopardize human safety. While present systems focus on pre-generation mitigation strategies, they cannot ensure the safety of individual outputs during inference. We provide a post hoc Hallucination Risk Scoring (HRS) methodology that intercepts questionable outputs before they reach patients via an agentic pipeline. Given a medical question, a domain-specific LLM generates an initial response from which five complimentary uncertainty signals are computed, which are then separated into a decision layer that governs escalation and a guidance layer that directs clinical knowledge injection by a GPT. The framework is tested using three biological question-answering datasets of various complexity: PubMedQA-Labeled, PubMedQA-Artificial, and BioASQ Task B. The results show an up to 38% safety increase at the most sensitive threshold configuration, zero deterioration across all experimental configurations enforced by the Revert Baseline method, and complexity-aware escalation rates that scale organically with dataset difficulty. Tunable thresholds allow physicians to calibrate system behavior based on deployment requirements, providing a practical safety\u2013accuracy trade-off. Statistical research finds entropy as the primary uncertainty signal separating escalated from non-escalated situations across all datasets. These findings provide a deployable, interpretable, and configurable post hoc safety paradigm for reliable medical AI implementation.<\/jats:p>","DOI":"10.3390\/a19040315","type":"journal-article","created":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T10:11:04Z","timestamp":1776420664000},"page":"315","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Agentic Hallucination Risk Scoring for Medical LLMs via Uncertainty Quantification and Clinical Knowledge Injection"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-6777-9227","authenticated-orcid":false,"given":"Mayank","family":"Kapadia","sequence":"first","affiliation":[{"name":"Department of Applied Data Science, San Jos\u00e9 State University, San Jos\u00e9, CA 95192, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9974-6950","authenticated-orcid":false,"given":"Mohammad","family":"Masum","sequence":"additional","affiliation":[{"name":"Department of Applied Data Science, San Jos\u00e9 State University, San Jos\u00e9, CA 95192, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,17]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3641289","article-title":"A survey on evaluation of large language models","volume":"15","author":"Chang","year":"2024","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"ref_2","first-page":"1","article-title":"A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions","volume":"43","author":"Huang","year":"2025","journal-title":"ACM Trans. Inf. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"459","DOI":"10.1111\/1468-0009.12023","article-title":"High-reliability health care: Getting there from here","volume":"91","author":"Chassin","year":"2013","journal-title":"Milbank Q."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Kim, Y., Jeong, H., Chen, S., Li, S.S., Park, C., Lu, M., and Alhamoud, K. (2025). Medical hallucinations in foundation models and their impact on healthcare. arXiv.","DOI":"10.1101\/2025.02.28.25323115"},{"key":"ref_5","unstructured":"Tonmoy, S.M.T.I., Zaman, S.M., Jain, V., Rani, A., Rawte, V., Chadha, A., and Das, A. (2024). A comprehensive survey of hallucination mitigation techniques in large language models. arXiv."},{"key":"ref_6","unstructured":"Zhou, Y., Muresanu, A.I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. (2023, January 25\u201329). Large language models are human-level prompt engineers. Proceedings of the Eleventh International Conference on Learning Representations (ICLR), Online."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Li, Y., Zhang, J., and Xu, H. (2024, January 16\u201321). LLM-driven knowledge injection advances zero-shot and cross-target stance detection. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), Mexico City, Mexico.","DOI":"10.18653\/v1\/2024.naacl-short.32"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"MacFarland, T.W., and Yates, J.M. (2016). Mann\u2013Whitney U test. Introduction to Nonparametric Statistics for the Biological Sciences Using R, Springer International Publishing.","DOI":"10.1007\/978-3-319-30634-6_4"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Pal, A., Umapathi, L.K., and Sankarasubbu, M. (2023, January 6\u20137). Med-HALT: Medical domain hallucination test for large language models. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), Singapore.","DOI":"10.18653\/v1\/2023.conll-1.21"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Kozlakidis, Z., Wootton, T., and Mayrhofer, M.T. (2025). Through the looking glass: Ethical considerations regarding LLM-induced hallucinations to medical questions. Front. Digit. Health, 8.","DOI":"10.3389\/fdgth.2026.1736616"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zhu, Z., Zhang, Y., Zhuang, X., Zhang, F., Wan, Z., Chen, Y., Long, Q., Zheng, Y., and Wu, X. (2025). Can we trust AI doctors? A survey of medical hallucination in large language and large vision-language models. Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria, 27 July\u20131 August 2025, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2025.findings-acl.350"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"393","DOI":"10.1162\/coli_a_00322","article-title":"A structured review of the validity of BLEU","volume":"44","author":"Reiter","year":"2018","journal-title":"Comput. Linguist."},{"key":"ref_13","unstructured":"Lin, C.Y. (2004, January 25\u201326). ROUGE: A package for automatic evaluation of summaries. Proceedings of the Workshop on Text Summarization Branches Out, Barcelona, Spain. Available online: https:\/\/aclanthology.org\/W04-1013\/."},{"key":"ref_14","first-page":"15356","article-title":"Benchmarking LLMs via uncertainty quantification","volume":"37","author":"Ye","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Liu, X., Chen, T., Da, L., Chen, C., Lin, Z., and Wei, H. (2025). Uncertainty quantification and confidence calibration in large language models: A survey. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Barcelona, Spain, 3\u20137 August 2025, Association for Computing Machinery.","DOI":"10.1145\/3711896.3736569"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1016\/S0016-0032(96)00063-4","article-title":"The Jensen-Shannon divergence","volume":"334","author":"Pardo","year":"1997","journal-title":"J. Frankl. Inst."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Sah, C.K. (2025). Understanding uncertainty in LLMs. Proceedings of the 40th IEEE\/ACM International Conference on Automated Software Engineering (ASE), Sacramento, CA, USA, 16\u201320 November 2025, IEEE.","DOI":"10.1109\/ASE63991.2025.00392"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chen, Z., Li, P., Dong, X., and Hong, P. (2025). Uncertainty quantification for clinical outcome predictions with (large) language models. Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, 29 April\u20134 May 2025, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2025.findings-naacl.419"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Henard, C., Papadakis, M., Harman, M., Jia, Y., and Le Traon, Y. (2016). Comparing white-box and black-box test prioritization. Proceedings of the 38th International Conference on Software Engineering (ICSE), Austin, TX, USA, 14\u201322 May 2016, IEEE.","DOI":"10.1145\/2884781.2884791"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Neha, F., and Bhati, D. (2025). A survey of DeepSeek models. TechRxiv.","DOI":"10.36227\/techrxiv.173896582.25938392\/v1"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Bhati, D., Neha, F., Bandaru, D.S., Weber, M., and Gajera, I.D. (2026). Large language models: A survey of architectures, training paradigms, and alignment methods. Preprints, 2026011138.","DOI":"10.20944\/preprints202601.1138.v1"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Jin, Q., Dhingra, B., Liu, Z., Cohen, W., and Lu, X. (2019, January 3\u20137). PubMedQA: A dataset for biomedical research question answering. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China.","DOI":"10.18653\/v1\/D19-1259"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1038\/s41597-023-02068-4","article-title":"BioASQ-QA: A manually curated corpus for biomedical question answering","volume":"10","author":"Krithara","year":"2023","journal-title":"Sci. Data"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Singh, R., and Mangat, N.S. (1996). Stratified sampling. Elements of Survey Sampling, Springer.","DOI":"10.1007\/978-94-017-1404-4"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1238","DOI":"10.1038\/s41467-025-67998-6","article-title":"Bayesian teaching enables probabilistic reasoning in large language models","volume":"17","author":"Qiu","year":"2026","journal-title":"Nat. Commun."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Xu, J., Zhang, H., Yang, Y., Yang, L., Cheng, Z., Lyu, J., Liu, B., Zhou, X., Bacchelli, A., and Chiam, Y.K. (2025). One size does not fit all: Investigating efficacy of perplexity in detecting LLM-generated code. ACM Trans. Softw. Eng. Methodol., in press.","DOI":"10.1145\/3748506"},{"key":"ref_27","unstructured":"Kuhn, L., Gal, Y., and Farquhar, S. (2023). Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Hershey, J.R., and Olsen, P.A. (2007). Approximating the Kullback Leibler divergence between Gaussian mixture models. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Honolulu, HI, USA, 15\u201320 April 2007, IEEE.","DOI":"10.1109\/ICASSP.2007.366913"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1016\/j.ins.2015.02.024","article-title":"Learning similarity with cosine similarity ensemble","volume":"307","author":"Xia","year":"2015","journal-title":"Inf. Sci."},{"key":"ref_30","unstructured":"Zhang, T., Shi, H., Wang, Y., Wang, H., He, X., Li, Z., and Chen, H. (2025). TokUR: Token-level uncertainty estimation for large language model reasoning. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lee, J., Park, S., Shin, J., and Cho, B. (2024). Analyzing evaluation methods for large language models in the medical field: A scoping review. BMC Med. Inform. Decis. Mak., 24.","DOI":"10.1186\/s12911-024-02709-7"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/4\/315\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T10:17:42Z","timestamp":1776421062000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/4\/315"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,17]]},"references-count":31,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["a19040315"],"URL":"https:\/\/doi.org\/10.3390\/a19040315","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,17]]}}}