{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T09:00:26Z","timestamp":1783069226178,"version":"3.54.6"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,7,7]],"date-time":"2025-07-07T00:00:00Z","timestamp":1751846400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,7,7]],"date-time":"2025-07-07T00:00:00Z","timestamp":1751846400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005713","name":"Technische Universit\u00e4t M\u00fcnchen","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005713","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Digit Imaging. Inform. med."],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Large language models (LLMs) have shown promising potential in analyzing complex textual data, including radiological reports. These models can assist clinicians, particularly those with limited experience, by integrating and presenting diagnostic criteria within radiological classifications. However, before clinical adoption, LLMs must be rigorously validated by medical professionals to ensure accuracy, especially in the context of advanced radiological classification systems. This study evaluates the performance of four LLMs\u2014ChatGPT-4o, AmbossGPT, Claude 3.5 Sonnet, and Gemini 2.0 Flash\u2014in classifying fractures based on the AO classification system using CT reports. A dataset of 292 fictitious physician-generated CT reports, representing 310 fractures, was used to assess the accuracy of each LLM in AO fracture classification retrospectively. Performance was evaluated by comparing the models\u2019 classifications to ground truth labels, with accuracy rates analyzed across different fracture types and subtypes. ChatGPT-4o and AmbossGPT achieved the highest overall accuracy (74.6 and 74.3%, respectively), outperforming Claude 3.5 Sonnet (69.5%) and Gemini 2.0 Flash (62.7%). Statistically significant differences were observed in fracture type classification, particularly between ChatGPT-4o and Gemini 2.0 Flash (\u039412%,\n                    <jats:italic>p<\/jats:italic>\n                    \u2009&lt;\u20090.001). While all models demonstrated strong bone recognition rates (90\u201399%), their accuracy in fracture subtype classification remained lower (71\u201377%), indicating limitations in nuanced diagnostic categorization. LLMs show potential in assisting radiologists with initial fracture classification, particularly in high-volume or resource-limited settings. However, their performance remains inconsistent for detailed subtype classification, highlighting the need for further refinement and validation before clinical integration in advanced diagnostic workflows.\n                  <\/jats:p>","DOI":"10.1007\/s10278-025-01603-6","type":"journal-article","created":{"date-parts":[[2025,7,7]],"date-time":"2025-07-07T15:37:20Z","timestamp":1751902640000},"page":"1861-1867","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Leveraging Large Language Models for Accurate AO Fracture Classification from CT Text Reports"],"prefix":"10.1007","volume":"39","author":[{"given":"Markus","family":"Mergen","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5878-3189","authenticated-orcid":false,"given":"Daniel","family":"Spitzl","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Conrad","family":"Ketzer","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Maximilian","family":"Strenzke","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexander W.","family":"Marka","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marcus R.","family":"Makowski","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Keno K.","family":"Bressem","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lisa C.","family":"Adams","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Florian T.","family":"Gassert","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,7,7]]},"reference":[{"key":"1603_CR1","doi-asserted-by":"crossref","unstructured":"Meinberg EG, Agel J, Roberts CS, Karam MD, Kellam JF. Fracture and Dislocation Classification Compendium-2018. J Orthop Trauma. 2018;32 Suppl 1:S1-S170.","DOI":"10.1097\/BOT.0000000000001063"},{"key":"1603_CR2","doi-asserted-by":"crossref","unstructured":"Park JW, Jo WL, Park BK, Go JJ, Han M, Chun S, Lee YK. Reliability of the 2018 Revised Version of AO\/OTA Classification for Femoral Shaft Fractures. Clin Orthop Surg. 2024;16(5):688-93.","DOI":"10.4055\/cios23292"},{"key":"1603_CR3","doi-asserted-by":"crossref","unstructured":"Cheng CT, Wang Y, Chen HW, Hsiao PM, Yeh CN, Hsieh CH, et al. A scalable physician-level deep learning algorithm detects universal trauma on pelvic radiographs. Nat Commun. 2021;12(1):1066.","DOI":"10.1038\/s41467-021-21311-3"},{"key":"1603_CR4","doi-asserted-by":"crossref","unstructured":"Nobel JM, van Geel K, Robben SGF. Structured reporting in radiology: a systematic review to explore its potential. Eur Radiol. 2022;32(4):2837-54.","DOI":"10.1007\/s00330-021-08327-5"},{"key":"1603_CR5","doi-asserted-by":"crossref","unstructured":"Menezes MCS, Hoffmann AF, Tan ALM, Nalbandyan M, Omenn GS, Mazzotti DR, et al. The potential of Generative Pre-trained Transformer 4 (GPT-4) to analyse medical notes in three different languages: a retrospective model-evaluation study. Lancet Digit Health. 2025;7(1):e35-e43.","DOI":"10.1016\/S2589-7500(24)00246-2"},{"key":"1603_CR6","doi-asserted-by":"crossref","unstructured":"Bhattarai K, Oh IY, Sierra JM, Tang J, Payne PRO, Abrams Z, Lai AM. Leveraging GPT-4 for identifying cancer phenotypes in electronic health records: a performance comparison between GPT-4, GPT-3.5-turbo, Flan-T5, Llama-3\u20138B, and spaCy\u2019s rule-based and machine learning-based methods. JAMIA Open. 2024;7(3):ooae060.","DOI":"10.1093\/jamiaopen\/ooae060"},{"key":"1603_CR7","doi-asserted-by":"crossref","unstructured":"Ford E, Carroll JA, Smith HE, Scott D, Cassell JA. Extracting information from the text of electronic medical records to improve case detection: a systematic review. J Am Med Inform Assoc. 2016;23(5):1007-15.","DOI":"10.1093\/jamia\/ocv180"},{"key":"1603_CR8","doi-asserted-by":"crossref","unstructured":"Wu J, Li C, Gensheimer M, Padda S, Kato F, Shirato H, et al. Radiological tumor classification across imaging modality and histology. Nat Mach Intell. 2021;3:787-98.","DOI":"10.1038\/s42256-021-00377-0"},{"key":"1603_CR9","doi-asserted-by":"crossref","unstructured":"Oh Y, Park S, Byun HK, Cho Y, Lee IJ, Kim JS, Ye JC. LLM-driven multimodal target volume contouring in radiation oncology. Nat Commun. 2024;15(1):9186.","DOI":"10.1038\/s41467-024-53387-y"},{"key":"1603_CR10","doi-asserted-by":"crossref","unstructured":"Moura Cunha G, Chernyak V, Fowler KJ, Sirlin CB (2021) Up-to-Date Role of CT\/MRI LI-RADS in Hepatocellular Carcinoma. J Hepatocell Carcinoma. 8:513\u2013527","DOI":"10.2147\/JHC.S268288"},{"key":"1603_CR11","doi-asserted-by":"crossref","unstructured":"Yasaka K, Akai H, Abe O, Kiryu S. Deep Learning with Convolutional Neural Network for Differentiation of Liver Masses at Dynamic Contrast-enhanced CT: A Preliminary Study. Radiology. 2018;286(3):887-96.","DOI":"10.1148\/radiol.2017170706"},{"key":"1603_CR12","doi-asserted-by":"crossref","unstructured":"Fervers P, Hahnfeldt R, Kottlors J, Wagner A, Maintz D, Pinto Dos Santos D, et al. ChatGPT yields low accuracy in determining LI-RADS scores based on free-text and structured radiology reports in German language. Front Radiol. 2024;4:1390774.","DOI":"10.3389\/fradi.2024.1390774"},{"key":"1603_CR13","doi-asserted-by":"crossref","unstructured":"Matute-Gonzalez M, Darnell A, Comas-Cufi M, Pazo J, Soler A, Saborido B, et al. Utilizing a domain-specific large language model for LI-RADS v2018 categorization of free-text MRI reports: a feasibility study. Insights Imaging. 2024;15(1):280.","DOI":"10.1186\/s13244-024-01850-1"},{"key":"1603_CR14","doi-asserted-by":"crossref","unstructured":"Wagholikar A, Zuccon G, Nguyen A, Chu K, Martin S, Lai K, Greenslade J. Automated classification of limb fractures from free-text radiology reports using a clinician-informed gazetteer methodology. Australas Med J. 2013;6(5):301-7.","DOI":"10.4066\/AMJ.2013.1651"},{"key":"1603_CR15","doi-asserted-by":"crossref","unstructured":"Lee JE, Park KS, Kim YH, Song HC, Park B, Jeong YJ. Lung Cancer Staging Using Chest CT and FDG PET\/CT Free-Text Reports: Comparison Among Three ChatGPT Large Language Models and Six Human Readers of Varying Experience. AJR Am J Roentgenol. 2024;223(6):e2431696.","DOI":"10.2214\/AJR.24.31696"},{"key":"1603_CR16","doi-asserted-by":"crossref","unstructured":"Singhal K, Tu T, Gottweis J, Sayres R, Wulczyn E, Amin M, et al. Toward expert-level medical question answering with large language models. Nat Med. 2025.","DOI":"10.1038\/s41591-024-03423-7"},{"key":"1603_CR17","doi-asserted-by":"publisher","unstructured":"McKinney, W. (2010) Data Structures for Statistical Computing in Python. Proceedings of the 9th Python in Science Conference, Austin, 28 June-3 July 2010, 56-61. https:\/\/doi.org\/10.25080\/Majora-92bf1922-00a","DOI":"10.25080\/Majora-92bf1922-00a"},{"key":"1603_CR18","doi-asserted-by":"crossref","unstructured":"Seabold, Skipper, and Josef Perktold. \"Statsmodels: econometric and statistical modeling with python.\" SciPy 7.1 (2010): 92-96.","DOI":"10.25080\/Majora-92bf1922-011"},{"key":"1603_CR19","doi-asserted-by":"crossref","unstructured":"Waskom, Michael L. \"Seaborn: statistical data visualization.\" Journal of Open Source Software 6.60 (2021): 3021.","DOI":"10.21105\/joss.03021"},{"key":"1603_CR20","doi-asserted-by":"crossref","unstructured":"Audige L, Bhandari M, Hanson B, Kellam J. A concept for the validation of fracture classifications. J Orthop Trauma. 2005;19(6):401-6.","DOI":"10.1097\/01.bot.0000155310.04886.37"},{"key":"1603_CR21","doi-asserted-by":"crossref","unstructured":"Curtis EM, van der Velde R, Moon RJ, van den Bergh JP, Geusens P, de Vries F, et al. Epidemiology of fractures in the United Kingdom 1988-2012: Variation with age, sex, geography, ethnicity and socioeconomic status. Bone. 2016;87:19-26.","DOI":"10.1016\/j.bone.2016.03.006"},{"key":"1603_CR22","doi-asserted-by":"crossref","unstructured":"Gao L, Zhang L, Liu C, Wu S. Handling imbalanced medical image data: A deep-learning-based one-class classification approach. Artif Intell Med. 2020;108:101935.","DOI":"10.1016\/j.artmed.2020.101935"},{"key":"1603_CR23","doi-asserted-by":"crossref","unstructured":"Court-Brown CM, Caesar B. Epidemiology of adult fractures: A review. Injury. 2006;37(8):691-7.","DOI":"10.1016\/j.injury.2006.04.130"},{"key":"1603_CR24","doi-asserted-by":"crossref","unstructured":"Pape HC, Tornetta P, 3rd, Tarkin I, Tzioupis C, Sabeson V, Olson SA. Timing of fracture fixation in multitrauma patients: the role of early total care and damage control surgery. J Am Acad Orthop Surg. 2009;17(9):541-9.","DOI":"10.5435\/00124635-200909000-00001"},{"key":"1603_CR25","doi-asserted-by":"crossref","unstructured":"Berger-Groch J, Thiesen DM, Grossterlinden LG, Schaewel J, Fensky F, Hartel MJ. The intra- and interobserver reliability of the Tile AO, the Young and Burgess, and FFP classifications in pelvic trauma. Arch Orthop Trauma Surg. 2019;139(5):645-50.","DOI":"10.1007\/s00402-019-03123-9"},{"key":"1603_CR26","doi-asserted-by":"crossref","unstructured":"Driver CN, Bowles BS, Bartholmai BJ, Greenberg-Worisek AJ. Artificial Intelligence in Radiology: A Call for Thoughtful Application. Clin Transl Sci. 2020;13(2):216-8.","DOI":"10.1111\/cts.12704"},{"key":"1603_CR27","doi-asserted-by":"crossref","unstructured":"Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44-56.","DOI":"10.1038\/s41591-018-0300-7"},{"key":"1603_CR28","doi-asserted-by":"crossref","unstructured":"Schwartz LH, Panicek DM, Berk AR, Li Y, Hricak H. Improving communication of diagnostic radiology findings through structured reporting. Radiology. 2011;260(1):174-81.","DOI":"10.1148\/radiol.11101913"},{"key":"1603_CR29","doi-asserted-by":"crossref","unstructured":"Rahsepar AA, Tavakoli N, Kim GHJ, Hassani C, Abtin F, Bedayat A. How AI Responds to Common Lung Cancer Questions: ChatGPT vs Google Bard. Radiology. 2023;307(5):e230922.","DOI":"10.1148\/radiol.230922"},{"key":"1603_CR30","doi-asserted-by":"crossref","unstructured":"Sun Z, Ong H, Kennedy P, Tang L, Chen S, Elias J, et al. Evaluating GPT4 on Impressions Generation in Radiology Reports. Radiology. 2023;307(5):e231259.","DOI":"10.1148\/radiol.231259"},{"key":"1603_CR31","doi-asserted-by":"crossref","unstructured":"Fink MA, Bischoff A, Fink CA, Moll M, Kroschke J, Dulz L, et al. Potential of ChatGPT and GPT-4 for Data Mining of Free-Text CT Reports on Lung Cancer. Radiology. 2023;308(3):e231362.","DOI":"10.1148\/radiol.231362"},{"key":"1603_CR32","doi-asserted-by":"crossref","unstructured":"Le Guellec B, Lefevre A, Geay C, Shorten L, Bruge C, Hacein-Bey L, et al. Performance of an Open-Source Large Language Model in Extracting Information from Free-Text Radiology Reports. Radiol Artif Intell. 2024;6(4):e230364.","DOI":"10.1148\/ryai.230364"},{"key":"1603_CR33","doi-asserted-by":"crossref","unstructured":"Adams LC, Truhn D, Busch F, Kader A, Niehues SM, Makowski MR, Bressem KK. Leveraging GPT-4 for Post Hoc Transformation of Free-text Radiology Reports into Structured Reporting: A Multilingual Feasibility Study. Radiology. 2023;307(4):e230725.","DOI":"10.1148\/radiol.230725"},{"key":"1603_CR34","doi-asserted-by":"crossref","unstructured":"Clusmann J, Kolbinger FR, Muti HS, Carrero ZI, Eckardt JN, Laleh NG, et al. The future landscape of large language models in medicine. Commun Med (Lond). 2023;3(1):141.","DOI":"10.1038\/s43856-023-00370-1"},{"key":"1603_CR35","doi-asserted-by":"crossref","unstructured":"Zhao Z, Wang S, Gu J, Zhu Y, Mei L, Zhuang Z, et al. ChatCAD+: Toward a Universal and Reliable Interactive CAD Using LLMs. IEEE Trans Med Imaging. 2024;43(11):3755-66.","DOI":"10.1109\/TMI.2024.3398350"},{"key":"1603_CR36","doi-asserted-by":"crossref","unstructured":"Tan R, Lin Q, Low GH, Lin R, Goh TC, Chang CCE, et al. Inferring cancer disease response from radiology reports using large language models with data augmentation and prompting. J Am Med Inform Assoc. 2023;30(10):1657-64.","DOI":"10.1093\/jamia\/ocad133"},{"key":"1603_CR37","doi-asserted-by":"crossref","unstructured":"Rau A, Rau S, Zoeller D, Fink A, Tran H, Wilpert C, et al. A Context-based Chatbot Surpasses Trained Radiologists and Generic ChatGPT in Following the ACR Appropriateness Guidelines. Radiology. 2023;308(1):e230970.","DOI":"10.1148\/radiol.230970"},{"key":"1603_CR38","doi-asserted-by":"crossref","unstructured":"Benary M, Wang XD, Schmidt M, Soll D, Hilfenhaus G, Nassir M, et al. Leveraging Large Language Models for Decision Support in Personalized Oncology. JAMA Netw Open. 2023;6(11):e2343689.","DOI":"10.1001\/jamanetworkopen.2023.43689"},{"key":"1603_CR39","doi-asserted-by":"crossref","unstructured":"Kresevic S, Giuffre M, Ajcevic M, Accardo A, Croce LS, Shung DL. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework. NPJ Digit Med. 2024;7(1):102.","DOI":"10.1038\/s41746-024-01091-y"},{"key":"1603_CR40","doi-asserted-by":"crossref","unstructured":"Kanzawa J, Yasaka K, Fujita N, Fujiwara S, Abe O Automated classification of brain MRI reports using fine-tuned large language models. Neuroradiology. 2024;66(12):2177\u20132183","DOI":"10.1007\/s00234-024-03427-7"}],"container-title":["Journal of Imaging Informatics in Medicine"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10278-025-01603-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10278-025-01603-6","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10278-025-01603-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T16:22:29Z","timestamp":1776874949000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10278-025-01603-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,7]]},"references-count":40,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["1603"],"URL":"https:\/\/doi.org\/10.1007\/s10278-025-01603-6","relation":{},"ISSN":["2948-2933"],"issn-type":[{"value":"2948-2933","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,7]]},"assertion":[{"value":"17 March 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 May 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 June 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 July 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Approval from an institutional review board was not required for the fictitious cohort due to the use of nonidentifiable data. The real-world validation was conducted in accordance with local ethical guidelines under ethics approval number\n                      2024\u2013590-S-CB\n                      .","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics Approval"}},{"value":"The authors declare no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}