{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,6]],"date-time":"2026-02-06T00:39:53Z","timestamp":1770338393469,"version":"3.49.0"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,5,17]],"date-time":"2024-05-17T00:00:00Z","timestamp":1715904000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,5,17]],"date-time":"2024-05-17T00:00:00Z","timestamp":1715904000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Healthc Inform Res"],"published-print":{"date-parts":[[2024,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>\u00a0Pulmonary nodules and nodule characteristics are important indicators of lung nodule malignancy. However, nodule information is often documented as free text in clinical narratives such as radiology reports in electronic health record systems. Natural language processing (NLP) is the key technology to extract and standardize patient information from radiology reports into structured data elements. This study aimed to develop an NLP system using state-of-the-art transformer models to extract pulmonary nodules and associated nodule characteristics from radiology reports. We identified a cohort of 3080 patients who underwent LDCT at the University of Florida health system and collected their radiology reports. We manually annotated 394 reports as the gold standard. We explored eight pretrained transformer models from three transformer architectures including bidirectional encoder representations from transformers (BERT), robustly optimized BERT approach (RoBERTa), and A Lite BERT (ALBERT), for clinical concept extraction, relation identification, and negation detection. We examined general transformer models pretrained using general English corpora, transformer models fine-tuned using a clinical corpus, and a large clinical transformer model, GatorTron, which was trained from scratch using 90 billion words of clinical text. We compared transformer models with two baseline models including a recurrent neural network implemented using bidirectional long short-term memory with a conditional random fields layer and support vector machines. RoBERTa-mimic achieved the best <jats:italic>F<\/jats:italic>1-score of 0.9279 for nodule concept and nodule characteristics extraction. ALBERT-base and GatorTron achieved the best <jats:italic>F<\/jats:italic>1-score of 0.9737 in linking nodule characteristics to pulmonary nodules. Seven out of eight transformers achieved the best <jats:italic>F<\/jats:italic>1-score of 1.0000 for negation detection. Our end-to-end system achieved an overall <jats:italic>F<\/jats:italic>1-score of 0.8869. This study demonstrated the advantage of state-of-the-art transformer models for pulmonary nodule information extraction from radiology reports.<\/jats:p>","DOI":"10.1007\/s41666-024-00166-5","type":"journal-article","created":{"date-parts":[[2024,5,17]],"date-time":"2024-05-17T15:02:29Z","timestamp":1715958149000},"page":"463-477","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Extracting Pulmonary Nodules and Nodule Characteristics from Radiology Reports of Lung Cancer Screening Patients Using Transformer Models"],"prefix":"10.1007","volume":"8","author":[{"given":"Shuang","family":"Yang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xi","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tianchen","family":"Lyu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James L.","family":"Huang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aokun","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xing","family":"He","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dejana","family":"Braithwaite","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiren J.","family":"Mehta","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yonghui","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Guo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiang","family":"Bian","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,5,17]]},"reference":[{"key":"166_CR1","doi-asserted-by":"publisher","first-page":"7","DOI":"10.3322\/caac.21654","volume":"71","author":"RL Siegel","year":"2021","unstructured":"Siegel RL, Miller KD, Fuchs HE et al (2021) Cancer statistics, 2021. CA Cancer J Clin 71:7\u201333","journal-title":"CA Cancer J Clin"},{"key":"166_CR2","doi-asserted-by":"publisher","first-page":"395","DOI":"10.1056\/NEJMoa1102873","volume":"365","author":"DR Aberle","year":"2011","unstructured":"National Lung Screening Trial Research Team, Aberle DR, Adams AM et al (2011) Reduced lung-cancer mortality with low-dose computed tomographic screening. N Engl J Med 365:395\u2013409","journal-title":"N Engl J Med"},{"key":"166_CR3","doi-asserted-by":"publisher","first-page":"971","DOI":"10.1001\/jama.2021.0377","volume":"325","author":"DE Jonas","year":"2021","unstructured":"Jonas DE, Reuland DS, Reddy SM et al (2021) Screening for lung cancer with low-dose computed tomography: updated evidence report and systematic review for the US Preventive Services Task Force. JAMA 325:971\u2013987","journal-title":"JAMA"},{"key":"166_CR4","unstructured":"Centers for Medicare & Medicaid Services. Decision memo for screening for lung cancer with low dose computed tomography (LDCT)(CAG-00439N). https:\/\/www.cms.gov\/medicare-coverage-database\/details\/nca-decision-memo.aspx"},{"key":"166_CR5","doi-asserted-by":"publisher","first-page":"1587","DOI":"10.1016\/j.jacr.2019.04.026","volume":"16","author":"SK Kang","year":"2019","unstructured":"Kang SK, Garry K, Chung R et al (2019) Natural language processing for identification of incidental pulmonary nodules in radiology reports. J Am Coll Radiol 16:1587\u20131594","journal-title":"J Am Coll Radiol"},{"key":"166_CR6","doi-asserted-by":"publisher","first-page":"1902","DOI":"10.1016\/j.chest.2021.05.048","volume":"160","author":"C Zheng","year":"2021","unstructured":"Zheng C, Huang BZ, Agazaryan AA et al (2021) Natural language processing to identify pulmonary nodules and extract nodule characteristics from radiology reports. Chest 160:1902\u20131914","journal-title":"Chest"},{"key":"166_CR7","doi-asserted-by":"publisher","first-page":"3114","DOI":"10.21037\/jtd.2017.08.13","volume":"9","author":"SE Beyer","year":"2017","unstructured":"Beyer SE, McKee BJ, Regis SM et al (2017) Automatic Lung-RADSTM classification with a natural language processing system. J Thorac Dis 9:3114\u20133122","journal-title":"J Thorac Dis"},{"key":"166_CR8","doi-asserted-by":"publisher","first-page":"80","DOI":"10.1093\/jamia\/ocaa209","volume":"28","author":"R Lacson","year":"2021","unstructured":"Lacson R, Cochon L, Ching PR et al (2021) Integrity of clinical information in radiology reports documenting pulmonary nodules. J Am Med Inform Assoc 28:80\u201385","journal-title":"J Am Med Inform Assoc"},{"key":"166_CR9","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1016\/j.cosrev.2018.06.001","volume":"29","author":"A Goyal","year":"2018","unstructured":"Goyal A, Gupta V, Kumar M (2018) Recent named entity recognition and classification techniques: a systematic review. Comput Sci Rev 29:21\u201343","journal-title":"Comput Sci Rev"},{"key":"166_CR10","unstructured":"Devlin J, Chang MW, Lee K, Toutanova K (2019) BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp 4171\u20134186), Minneapolis, Minnesota. Association for Computational Linguistics"},{"key":"166_CR11","first-page":"1812","volume":"2018","author":"Y Wu","year":"2017","unstructured":"Wu Y, Jiang M, Xu J et al (2017) Clinical named entity recognition using deep learning models. AMIA Annu Symp Proc 2018:1812\u20131819","journal-title":"AMIA Annu Symp Proc"},{"key":"166_CR12","doi-asserted-by":"crossref","unstructured":"Lample G, Ballesteros M, Subramanian S, Kawakami K, Dyer C (2016) Neural architectures for named entity recognition. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp 260\u2013270), San Diego, California. Association for Computational Linguistics","DOI":"10.18653\/v1\/N16-1030"},{"key":"166_CR13","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1186\/s12911-017-0468-7","volume":"17","author":"Z Liu","year":"2017","unstructured":"Liu Z, Yang M, Wang X et al (2017) Entity recognition from clinical texts via recurrent neural network. BMC Med Inform Decis Mak 17:67","journal-title":"BMC Med Inform Decis Mak"},{"key":"166_CR14","first-page":"455","volume":"2016","author":"W Yim","year":"2016","unstructured":"Yim W, Denman T, Kwan SW et al (2016) Tumor information extraction in radiology reports for hepatocellular carcinoma patients. AMIA Jt Summits Transl Sci Proc 2016:455\u2013464","journal-title":"AMIA Jt Summits Transl Sci Proc"},{"key":"166_CR15","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1016\/j.artmed.2015.09.007","volume":"66","author":"S Hassanpour","year":"2016","unstructured":"Hassanpour S, Langlotz CP (2016) Information extraction from multi-institutional radiology reports. Artif Intell Med 66:29\u201339","journal-title":"Artif Intell Med"},{"key":"166_CR16","first-page":"1079","volume-title":"AMIA Annual Symposium Proceedings","author":"T Santos","year":"2021","unstructured":"Santos T, Kallas ON, Newsome J et al (2021) A fusion NLP model for the inference of standardized thyroid nodule malignancy scores from radiology report text. In: AMIA Annual Symposium Proceedings. American Medical Informatics Association, p 1079"},{"key":"166_CR17","doi-asserted-by":"publisher","first-page":"103985","DOI":"10.1016\/j.ijmedinf.2019.103985","volume":"132","author":"X Zhang","year":"2019","unstructured":"Zhang X, Zhang Y, Zhang Q et al (2019) Extracting comprehensive clinical information for breast cancer using deep learning methods. Int J Med Inform 132:103985","journal-title":"Int J Med Inform"},{"key":"166_CR18","doi-asserted-by":"publisher","first-page":"544","DOI":"10.1136\/amiajnl-2011-000464","volume":"18","author":"PM Nadkarni","year":"2011","unstructured":"Nadkarni PM, Ohno-Machado L, Chapman WW (2011) Natural language processing: an introduction. J Am Med Inform Assoc 18:544\u2013551","journal-title":"J Am Med Inform Assoc"},{"key":"166_CR19","unstructured":"Kumar S (2017) A survey of deep learning methods for relation extraction. ArXiv [Cs.CL]. arXiv.\u00a0https:\/\/arxiv.org\/abs\/1705.03645"},{"key":"166_CR20","volume-title":"Learning to detect negation with \u2018not\u2019 in medical texts","author":"I Goldin","year":"2003","unstructured":"Goldin I, Chapman WW (2003) Learning to detect negation with \u2018not\u2019 in medical texts. Proc Workshop on Text Analysis and Search for Bioinformatics, ACM SIGIR"},{"key":"166_CR21","unstructured":"Zhuang L, Wayne L, Ya S, Jun Z (2021) A robustly optimized BERT pre-training approach with post-training. In Proceedings of the 20th Chinese National Conference on Computational Linguistics (pp 1218\u20131227), Huhhot, China. Chinese Information Processing Society of China"},{"key":"166_CR22","doi-asserted-by":"crossref","unstructured":"Lan Z, Chen M, Goodman S, Gimpel K, Sharma P, Soricut R (2020) ALBERT: A Lite BERT for self-supervised learning of language representations. Paper presented at the meeting of the ICLR, 2020.","DOI":"10.1109\/SLT48900.2021.9383575"},{"key":"166_CR23","unstructured":"Stenetorp P, Pyysalo S, Topi\u0107 G, Ohta T, Ananiadou S, Tsujii JI (2012) BRAT: a web-based tool for nlp-assisted text annotation. In Proceedings of the Demonstrations at the 13th Conference of the European Chapter of the Association for Computational Linguistics (pp 102\u2013107), Avignon, France. Association for Computational Linguistics"},{"key":"166_CR24","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1177\/001316446002000104","volume":"20","author":"J Cohen","year":"1960","unstructured":"Cohen J (1960) A coefficient of agreement for nominal scales. Educ Psychol Meas 20:37\u201346","journal-title":"Educ Psychol Meas"},{"key":"166_CR25","doi-asserted-by":"publisher","first-page":"1935","DOI":"10.1093\/jamia\/ocaa189","volume":"27","author":"X Yang","year":"2020","unstructured":"Yang X, Bian J, Hogan WR et al (2020) Clinical concept extraction using transformers. J Am Med Inform Assoc 27:1935\u20131942","journal-title":"J Am Med Inform Assoc"},{"key":"166_CR26","doi-asserted-by":"publisher","first-page":"160035","DOI":"10.1038\/sdata.2016.35","volume":"3","author":"AEW Johnson","year":"2016","unstructured":"Johnson AEW, Pollard TJ, Shen L et al (2016) MIMIC-III, a freely accessible critical care database. Sci Data 3:160035. https:\/\/doi.org\/10.1038\/sdata.2016.35","journal-title":"Sci Data"},{"key":"166_CR27","unstructured":"Yang X, Yu Z, Guo Y, Bian J, Wu Y (2021) Clinical relation extraction using transformer-based models. ArXiv [Cs.CL]. arXiv. http:\/\/arxiv.org\/abs\/2107.08957"},{"key":"166_CR28","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1093\/jamia\/ocz144","volume":"27","author":"X Yang","year":"2020","unstructured":"Yang X, Bian J, Fang R et al (2020) Identifying relations of medications with adverse drug events using recurrent convolutional neural networks and gradient boosting. J Am Med Inform Assoc 27:65\u201372","journal-title":"J Am Med Inform Assoc"},{"key":"166_CR29","doi-asserted-by":"publisher","first-page":"e22982","DOI":"10.2196\/22982","volume":"8","author":"X Yang","year":"2020","unstructured":"Yang X, Zhang H, He X et al (2020) Extracting family history of patients from clinical narratives: exploring an end-to-end solution with deep learning models. JMIR Med Inform 8:e22982","journal-title":"JMIR Med Inform"},{"key":"166_CR30","doi-asserted-by":"publisher","first-page":"189","DOI":"10.1016\/j.neucom.2019.10.118","volume":"408","author":"J Cervantes","year":"2020","unstructured":"Cervantes J, Garcia-Lamont F, Rodr\u00edguez-Mazahua L, Lopez A (2020) A comprehensive survey on support vector machine classification: applications, challenges and trends. Neurocomputing 408:189\u2013215","journal-title":"Neurocomputing"},{"key":"166_CR31","unstructured":"LIBSVM: A library for support vector machines: ACM Transactions on Intelligent Systems and Technology: Vol 2, No 3. https:\/\/dl.acm.org\/doi\/abs\/10.1145\/1961189.1961199?casa_token=Qs6g7IO8tZYAAAAA:5tlZ57sdN_78cebeKSjO-5X71ruAlyiE1h5xzAKTIzWemYxONtT4-Fy1W8ZvBJ-qn4MzbHXwCXGc (accessed 29 September 2022)"},{"key":"166_CR32","doi-asserted-by":"publisher","unstructured":"Alsentzer E, Murphy JR, Boag W et al (2019) Publicly available clinical BERT embeddings. https:\/\/doi.org\/10.48550\/arXiv.1904.03323","DOI":"10.48550\/arXiv.1904.03323"},{"key":"166_CR33","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41746-022-00742-2","volume":"5","author":"X Yang","year":"2022","unstructured":"Yang X, Chen A, PourNejatian N et al (2022) A large language model for electronic health records. npj Digit Med 5:1\u20139. https:\/\/doi.org\/10.1038\/s41746-022-00742-2","journal-title":"npj Digit Med"},{"key":"166_CR34","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1093\/jamia\/ocz166","volume":"27","author":"S Henry","year":"2020","unstructured":"Henry S, Buchan K, Filannino M et al (2020) 2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records. J Am Med Inform Assoc 27:3\u201312","journal-title":"J Am Med Inform Assoc"},{"key":"166_CR35","unstructured":"Please, don\u2019t forget the difference and the confidence interval when seeking for the state-of-the-art status - ACL Anthology. https:\/\/aclanthology.org\/2022.lrec-1.640\/ (accessed 1 April 2024)"},{"key":"166_CR36","doi-asserted-by":"publisher","unstructured":"Bommasani R, Hudson DA, Adeli E et al (2022) On the opportunities and risks of foundation models. https:\/\/doi.org\/10.48550\/arXiv.2108.07258","DOI":"10.48550\/arXiv.2108.07258"},{"key":"166_CR37","doi-asserted-by":"publisher","first-page":"1486","DOI":"10.1093\/jamia\/ocad107","volume":"30","author":"C Peng","year":"2023","unstructured":"Peng C, Yang X, Yu Z et al (2023) Clinical concept and relation extraction using prompt-based machine reading comprehension. J Am Med Inform Assoc 30:1486\u20131493","journal-title":"J Am Med Inform Assoc"},{"key":"166_CR38","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2110.07602","volume-title":"P-Tuning v2: prompt tuning can be comparable to fine-tuning universally across scales and tasks","author":"X Liu","year":"2022","unstructured":"Liu X, Ji K, Fu Y et al (2022) P-Tuning v2: prompt tuning can be comparable to fine-tuning universally across scales and tasks. https:\/\/doi.org\/10.48550\/arXiv.2110.07602"}],"container-title":["Journal of Healthcare Informatics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41666-024-00166-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s41666-024-00166-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41666-024-00166-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,8]],"date-time":"2024-08-08T14:25:30Z","timestamp":1723127130000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s41666-024-00166-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,17]]},"references-count":38,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,9]]}},"alternative-id":["166"],"URL":"https:\/\/doi.org\/10.1007\/s41666-024-00166-5","relation":{},"ISSN":["2509-4971","2509-498X"],"issn-type":[{"value":"2509-4971","type":"print"},{"value":"2509-498X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,17]]},"assertion":[{"value":"14 October 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 April 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 May 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 May 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of Interest"}},{"value":"The content is solely the responsibility of the authors and does not necessarily represent the official views of the funding institutions.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Disclaimer"}}]}}