{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T03:53:42Z","timestamp":1782186822711,"version":"3.54.5"},"reference-count":51,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T00:00:00Z","timestamp":1760313600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T00:00:00Z","timestamp":1760313600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Universit\u00e4tsklinikum RWTH Aachen"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Med Inform Decis Mak"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Background<\/jats:title>\n                    <jats:p>Clinical utilization of machine learning is hampered by the lack of interpretability inherent in most non-linear black box modeling approaches, reducing trust among clinicians and regulators. Advanced large language models offer a potential framework for integrating medical knowledge into these models, potentially enhancing their interpretability.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods<\/jats:title>\n                    <jats:p>A hybrid mechanistic\/data-driven modeling framework is presented for developing an ICU risk of death prediction model for mechanically ventilated patients. In the mechanistic modeling part, GPT-4o is used to generate detailed medical feature descriptions, which are then aggregated into a comprehensive corpus and processed with TF-I DF vectorization. Fuzzy C-means clustering is subsequently applied to these vectorized features to identify significant mortality cause-specific feature clusters, and a physician reviewed the resulting clusters to validate their relevance to actionable insights for clinical decision support. In the data-driven part, the identified clusters inform the creation of XGBoost-based weak classifiers, whose outcomes are combined into a single XGBoost-based strong classifier through a hierarchically structured feed-forward network. This process results in a novel GPT hybrid model for ICU risk of death prediction.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>This study enrolled 16,018 mechanically ventilated ICU patients, divided into derivation (12,758) and validation (3,260) cohorts, to develop and evaluate a GPT hybrid model for predicting in-ICU death. Leveraging GPT-4o, we implemented an automated process for clustering mortality cause-specific features, resulting in six feature clusters: Liver Failure, Infection, Renal Failure, Hypoxia, Cardiac Failure, and Mechanical Ventilation. This approach significantly improved upon previous manual methods, automating the reconstruction of structured hybrid models. While the GPT hybrid model showed similar predictive accuracy to a Global XGBoost model, it demonstrated superior interpretability and clinical relevance by incorporating a wider array of features and providing a hierarchical structure of feature importance aligned with medical knowledge.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusion<\/jats:title>\n                    <jats:p>We introduce a novel approach to predicting in-ICU risk of death for mechanically ventilated patients using a GPT hybrid model. Our methodology demonstrates the potential of integrating large language models with traditional machine learning techniques to create interpretable and clinically relevant predictive models.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1186\/s12911-025-03224-z","type":"journal-article","created":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T09:27:24Z","timestamp":1760347644000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["GPT-4o and the quest for machine learning interpretability in ICU risk of death prediction"],"prefix":"10.1186","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3916-1085","authenticated-orcid":false,"given":"Moein","family":"E. Samadi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kateryna","family":"Nikulina","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8350-8584","authenticated-orcid":false,"given":"Sebastian Johannes","family":"Fritsch","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3783-6605","authenticated-orcid":false,"given":"Andreas","family":"Schuppert","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,10,13]]},"reference":[{"key":"3224_CR1","doi-asserted-by":"crossref","unstructured":"Holmes J, Sacchi L, Bellazzi R, et al. Artificial intelligence in medicine. Ann R Coll surgengl. 2004;86:334\u201338.","DOI":"10.1308\/147870804290"},{"key":"3224_CR2","doi-asserted-by":"crossref","unstructured":"Hamet P, Tremblay J. Artificial intelligence in medicine. Metabolism. 2017;69:S36\u2013S40.","DOI":"10.1016\/j.metabol.2017.01.011"},{"key":"3224_CR3","doi-asserted-by":"crossref","unstructured":"Rajpurkar P, Chen E, Banerjee O, Topol EJ. AI in health and medicine. NatMed. 2022;28(1):31\u201338.","DOI":"10.1038\/s41591-021-01614-0"},{"key":"3224_CR4","doi-asserted-by":"crossref","unstructured":"Vellido A. The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural ComputAppl. 2020;32(24):18069\u201383.","DOI":"10.1007\/s00521-019-04051-w"},{"key":"3224_CR5","doi-asserted-by":"crossref","unstructured":"Holzinger A, Langs G, Denk H, Zatloukal K, M\u00fcller H. Causability and explainability of artificial intelligence in medicine. Wiley Interdiscip Rev Data Min knowldiscov. 2019;9(4):e1312.","DOI":"10.1002\/widm.1312"},{"key":"3224_CR6","doi-asserted-by":"crossref","unstructured":"Yoon CH, Torrance R, Scheinerman N. Machine learning in medicine: Should the pursuit of enhanced interpretability be abandoned? J Med Ethics. 2022;48(9):581\u201385.","DOI":"10.1136\/medethics-2020-107102"},{"key":"3224_CR7","doi-asserted-by":"crossref","unstructured":"Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat MachIntell. 2019;1(5):206\u201315.","DOI":"10.1038\/s42256-019-0048-x"},{"key":"3224_CR8","doi-asserted-by":"crossref","unstructured":"Rudin C, Chen C, Chen, Huang H, Semenova L, Zhong C. Interpretable machine learning: Fundamental principles and 10 grand challenges. Statistic Surv. 2022;16:1\u201385.","DOI":"10.1214\/21-SS133"},{"issue":"4","key":"3224_CR9","first-page":"958","volume":"30","author":"S Feuerriegel","year":"2024","unstructured":"Feuerriegel S, Frauen D, Melnychuk V, Schweisthal J, Hess K, Curth A, et al. Causal machine learning for predicting treatment outcomes. NatMed. 2024;30(4):958\u201368.","journal-title":"NatMed"},{"issue":"3","key":"3224_CR10","doi-asserted-by":"publisher","first-page":"1626","DOI":"10.1287\/ijoc.2021.1143","volume":"34","author":"T Wang","year":"2022","unstructured":"Wang T, Rudin C. Causal rule sets for identifying subgroups with enhanced treatment effects. INFORMS JComput. 2022;34(3):1626\u201343.","journal-title":"INFORMS JComput"},{"key":"3224_CR11","doi-asserted-by":"crossref","unstructured":"Piccininni M, Konigorski S, Rohmann JL, Kurth T. Directed acyclic graphs and causal thinking in clinical risk prediction modeling. BMC Med resmethodol. 2020;20:1\u20139.","DOI":"10.1186\/s12874-020-01058-z"},{"key":"3224_CR12","unstructured":"Liu Z, Wang Y, Vaidya S, Ruehle F, Halverson J, Solja\u010di\u0107 M, et al. Kan: Kolmogorov-arnold networks. arXiv preprint arXiv:240419756. 2024."},{"key":"3224_CR13","unstructured":"Samadi ME, M\u00fcller Y, Schuppert A. Smooth kolmogorov arnold networks enabling structural knowledge representation. arXiv preprint arXiv:240511318. 2024."},{"key":"3224_CR14","first-page":"119","volume":"137","author":"J Schmidt-Hieber","year":"2021","unstructured":"Schmidt-Hieber J. The kolmogorov\u2013Arnold representation theorem revisited. NeuralNetw. 2021;137:119\u201326.","journal-title":"NeuralNetw"},{"key":"3224_CR15","doi-asserted-by":"crossref","unstructured":"Fiedler B, Schuppert A. Local identification of scalar hybrid models with tree structure. IMA J applmath. 2008;73(3):449\u201376.","DOI":"10.1093\/imamat\/hxn011"},{"key":"3224_CR16","doi-asserted-by":"crossref","unstructured":"Schuppert AA. Extrapolability of structured hybrid models: a key to optimization of complex processes. In: Equadiff 99: (in 2. Vol. Volumes. World Scientific; 2000. p. 1135\u201351.","DOI":"10.1142\/9789812792617_0218"},{"key":"3224_CR17","doi-asserted-by":"crossref","unstructured":"Samadi E, Kiefer M, S FS, Bickenbach J, Schuppert A. A training strategy for hybrid models to break the curse of dimensionality. PLoS One. 2022;17(9):e0274569.","DOI":"10.1371\/journal.pone.0274569"},{"issue":"1","key":"3224_CR18","doi-asserted-by":"publisher","first-page":"5725","DOI":"10.1038\/s41598-024-55577-6","volume":"14","author":"ME Samadi","year":"2024","unstructured":"Samadi ME, Guzman-Maldonado J, Nikulina K, Mirzaieazar H, Sharafutdinov K, Fritsch SJ, et al. A hybrid modeling framework for generalizable and interpretable predictions of ICU mortality across multiple hospitals. Sci Rep. 2024;14(1):5725.","journal-title":"Sci Rep"},{"issue":"8","key":"3224_CR19","first-page":"1930","volume":"29","author":"AJ Thirunavukarasu","year":"2023","unstructured":"Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. NatMed. 2023;29(8):1930\u201340.","journal-title":"NatMed"},{"issue":"13","key":"3224_CR20","doi-asserted-by":"publisher","first-page":"1233","DOI":"10.1056\/NEJMsr2214184","volume":"388","author":"P Lee","year":"2023","unstructured":"Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. NEJM Evid. 2023;388(13):1233\u201339.","journal-title":"NEJM Evid"},{"issue":"1","key":"3224_CR21","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1038\/s43856-023-00370-1","volume":"3","author":"J Clusmann","year":"2023","unstructured":"Clusmann J, Kolbinger FR, Muti HS, Carrero ZI, Eckardt JN, Laleh NG, et al. The future landscape of large language models in medicine. Commun Med (lond). 2023;3(1):141.","journal-title":"Commun Med (lond)"},{"issue":"7","key":"3224_CR22","doi-asserted-by":"publisher","first-page":"1151","DOI":"10.1016\/j.jacr.2023.11.021","volume":"21","author":"MP Lungren","year":"2024","unstructured":"Lungren MP, Fishman EK, Chu LC, Rizk RC, Rowe SP. More is different: Large language models in health care. J Am Coll ofradiol. 2024;21(7):1151\u201354.","journal-title":"J Am Coll ofradiol"},{"issue":"4","key":"3224_CR23","doi-asserted-by":"publisher","first-page":"2773","DOI":"10.3390\/ai5040134","volume":"5","author":"G Feretzakis","year":"2024","unstructured":"Feretzakis G, Verykios VS. Trustworthy AI: Securing sensitive data in large language models. AI. 2024;5(4):2773\u2013800.","journal-title":"AI"},{"key":"3224_CR24","doi-asserted-by":"crossref","unstructured":"Feretzakis G, Papaspyridis K, Gkoulalas-Divanis A, Verykios VS. Privacy-preserving techniques in generative AI and Large language models: A narrative review. Information. 2024;15(11):697.","DOI":"10.3390\/info15110697"},{"key":"3224_CR25","doi-asserted-by":"crossref","unstructured":"Van Veen D, Van Uden C, Blankemeier L, Delbrouck JB, Aali A, Bluethgen C, et al. Adapted large language models can outperform medical experts in clinical text summarization. NatMed. 2024;30(4):1134\u201342.","DOI":"10.1038\/s41591-024-02855-5"},{"key":"3224_CR26","unstructured":"Eriksen AV, M\u00f6ller S, Ryg J. Use of GPT-4 to diagnose complex clinical cases. Massachusetts Medical Society."},{"key":"3224_CR27","doi-asserted-by":"crossref","unstructured":"Huang J, Yang DM, Rong R, Nezafati K, Treager C, Chi Z, et al. A critical assessment of using ChatGPT for extracting structured data from clinical notes. NPJ digitmed. 2024;7(1):106.","DOI":"10.1038\/s41746-024-01079-8"},{"key":"3224_CR28","doi-asserted-by":"crossref","unstructured":"Hager P, Jungmann F, Holland R, Bhagat K, Hubrecht I, Knauer M, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. NatMed. 2024;1\u201310.","DOI":"10.1101\/2024.01.26.24301810"},{"key":"3224_CR29","doi-asserted-by":"crossref","unstructured":"Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172\u2013180.","DOI":"10.1038\/s41586-023-06291-2"},{"key":"3224_CR30","doi-asserted-by":"crossref","unstructured":"Chen T, Guestrin C. Xgboost: A scalable tree boosting system. Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining. 2016. p. 785\u201394.","DOI":"10.1145\/2939672.2939785"},{"key":"3224_CR31","doi-asserted-by":"crossref","unstructured":"Li J, Liu S, Hu Y, Zhu L, Mao Y, Liu J. Predicting mortality in intensive care unit patients with heart failure using an interpretable machine learning model: Retrospective cohort study. J Med InternetRes. 2022;24(8):e38082.","DOI":"10.2196\/38082"},{"key":"3224_CR32","doi-asserted-by":"crossref","unstructured":"Deshmukh F, Merchant SS. Explainable machine learning model for predicting GI bleed mortality in the intensive care unit. Off J Am Coll Of Gastroenterol| ACG. 2020;115(10):1657\u201368.","DOI":"10.14309\/ajg.0000000000000632"},{"key":"3224_CR33","doi-asserted-by":"crossref","unstructured":"Schweidtmann AM, Zhang D, von Stosch M. A review and perspective on hybrid modeling methodologies. Digit ChemEng. 2024;10:100136.","DOI":"10.1016\/j.dche.2023.100136"},{"key":"3224_CR34","doi-asserted-by":"crossref","unstructured":"Samadi ME, Mirzaieazar H, Mitsos A, Schuppert A. Noisecut: A python package for noise-tolerant classification of binary data using prior knowledge integration and max-cut solutions. BMCBioinf. 2024;25(1):155.","DOI":"10.1186\/s12859-024-05769-8"},{"issue":"3","key":"3224_CR35","doi-asserted-by":"publisher","first-page":"925","DOI":"10.1007\/s10957-018-1396-0","volume":"180","author":"AM Schweidtmann","year":"2019","unstructured":"Schweidtmann AM, Mitsos A. Deterministic global optimization with artificial neural networks embedded. J Optim theoryappl. 2019;180(3):925\u201348.","journal-title":"J Optim theoryappl"},{"key":"3224_CR36","doi-asserted-by":"crossref","unstructured":"Marx G, Bickenbach J, Fritsch SJ, Kunze JB, Maassen O, Deffge S, et al. Algorithmic surveillance of ICU patients with acute respiratory distress syndrome (ASIC): Protocol for a multicentre stepped-wedge cluster randomised quality improvement strategy. BMJ Open. 2021;11(4):e045589.","DOI":"10.1136\/bmjopen-2020-045589"},{"key":"3224_CR37","doi-asserted-by":"crossref","unstructured":"Winter A, St\u00e4ubert S, Ammon D, Aiche S, Beyan O, Bischoff V, et al. Smart medical information technology for healthcare (SMITH). Methods infmed. 2018;57:01):e92\u2013e105.","DOI":"10.3414\/ME18-02-0004"},{"key":"3224_CR38","doi-asserted-by":"crossref","unstructured":"Collins GS, Moons KG, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+ AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385.","DOI":"10.1136\/bmj-2023-078378"},{"key":"3224_CR39","doi-asserted-by":"crossref","unstructured":"Levey AS, Stevens LA. Estimating GFR using the CKD epidemiology collaboration (CKD-EPI) creatinine equation: More accurate GFR estimates, lower CKD prevalence estimates, and better risk predictions. Am J kidneydis. 2010;55(4):622\u201327.","DOI":"10.1053\/j.ajkd.2010.02.337"},{"key":"3224_CR40","doi-asserted-by":"crossref","unstructured":"Ranieri VM, Rubenfeld GD, Taylor Thompson B, Ferguson ND, Caldwell E, Fan E, et al. Acute respiratory distress syndrome: The Berlin definition. JAMA: J Am Med Assoc. 2012;307(23).","DOI":"10.1001\/jama.2012.5669"},{"key":"3224_CR41","doi-asserted-by":"crossref","unstructured":"Levy MM, Dellinger RP, Townsend SR, Linde-Zwirble WT, Marshall JC, Bion J, et al. The surviving sepsis campaign: Results of an international guideline-based performance improvement program targeting severe sepsis. Intensive caremed. 2010;36:222\u201331.","DOI":"10.1007\/s00134-009-1738-3"},{"key":"3224_CR42","unstructured":"Ramos J, et al. Using tf-idf to determine word relevance in document queries. Proceedings of the first instructional conference on machine learning. Citeseer; 2003. p. 29\u201348. vol. 242."},{"key":"3224_CR43","doi-asserted-by":"crossref","unstructured":"Wu HC, Luk RWP, Wong KF, Kwok KL. Interpreting TF-IDF term weights as making relevance decisions. ACM Trans On Inf Syst (TOIS). 2008;26(3):1\u201337.","DOI":"10.1145\/1361684.1361686"},{"key":"3224_CR44","doi-asserted-by":"crossref","unstructured":"Bezdek JC, Ehrlich R, Full W. FCM: The fuzzy c-means clustering algorithm. ComputGeosci. 1984;10(2\u20133):191\u2013203.","DOI":"10.1016\/0098-3004(84)90020-7"},{"key":"3224_CR45","unstructured":"Freund Y, Schapire R, Abe N. A short introduction to boosting. j-Jpn Soc For Artif Intel. 1999;14(771\u2013780):1612."},{"key":"3224_CR46","doi-asserted-by":"crossref","unstructured":"Meir R, R\u00e4tsch G. An introduction to boosting and leveraging. In: Advanced lectures on machine learning: Machine learning Summer school 2002. Canberra, Australia: Springer; 2003. p. 118\u201383. February 11\u201322, 2002 Revised Lectures.","DOI":"10.1007\/3-540-36434-X_4"},{"key":"3224_CR47","unstructured":"Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine learning in Python. J Mach learnres. 2011;12:2825\u201330."},{"key":"3224_CR48","unstructured":"Dem\u0161ar J. Statistical comparisons of classifiers over multiple data sets. J Mach learnres. 2006;7:1\u201330."},{"key":"3224_CR49","unstructured":"Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Advances in neural information processing systems. 2017;30."},{"key":"3224_CR50","doi-asserted-by":"crossref","unstructured":"Sinno ZC, Shay D, Kruppa J, Klopfenstein SA, Giesa N, Flint AR, et al. The influence of patient characteristics on the alarm rate in intensive care units: A retrospective cohort study. Sci Rep. 2022;12(1):21801.","DOI":"10.1038\/s41598-022-26261-4"},{"key":"3224_CR51","doi-asserted-by":"crossref","unstructured":"Valik JK, Ward L, Tanushi H, Johansson AF, F\u00e4rnert A, Mogensen ML, et al. Predicting sepsis onset using a machine learned causal probabilistic network algorithm based on electronic health records data. Sci Rep. 2023;13(1):11760.","DOI":"10.1038\/s41598-023-38858-4"}],"container-title":["BMC Medical Informatics and Decision Making"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-025-03224-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12911-025-03224-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-025-03224-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T09:28:01Z","timestamp":1760347681000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcmedinformdecismak.biomedcentral.com\/articles\/10.1186\/s12911-025-03224-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,13]]},"references-count":51,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["3224"],"URL":"https:\/\/doi.org\/10.1186\/s12911-025-03224-z","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-4816139\/v1","asserted-by":"object"}]},"ISSN":["1472-6947"],"issn-type":[{"value":"1472-6947","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,13]]},"assertion":[{"value":"28 July 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 September 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 October 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"This study was conducted in accordance with the principles of the Declaration of Helsinki. All experimental protocols were approved by the Ethics Committee of the RWTH Aachen Faculty of Medicine (local Ethics Committee reference number: EK 102\/19, date of approval: 26.03.2019). As well, the Ethics Committee of the RWTH Aachen Faculty of Medicine (local Ethics Committee reference number: EK 102\/19, date of approval: 26.03.2019) waived the need to obtain Informed consent for the collection and retrospective analysis of the de-identified data as well as the publication of the results of the analysis. All methods were performed in accordance with the relevant guidelines and regulations.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"373"}}