{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T10:46:20Z","timestamp":1767091580602},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,4,15]],"date-time":"2022-04-15T00:00:00Z","timestamp":1649980800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,4,15]],"date-time":"2022-04-15T00:00:00Z","timestamp":1649980800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"National Institutes of Health\/National Cancer Institute","award":["P30-CA014236","P30-CA014236","P30-CA014236","P30-CA014236"],"award-info":[{"award-number":["P30-CA014236","P30-CA014236","P30-CA014236","P30-CA014236"]}]},{"name":"National Institutes of Health\/National Institute of Biomedical Imaging and Bioengineering","award":["P41-EB028744","P41-EB028744","P41-EB028744","P41-EB028744"],"award-info":[{"award-number":["P41-EB028744","P41-EB028744","P41-EB028744","P41-EB028744"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Med Inform Decis Mak"],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>There is progress to be made in building artificially intelligent systems to detect abnormalities that are not only accurate but can handle the true breadth of findings that radiologists encounter in body (chest, abdomen, and pelvis) computed tomography (CT). Currently, the major bottleneck for developing multi-disease classifiers is a lack of manually annotated data. The purpose of this work was to develop high throughput multi-label annotators for body CT reports that can be applied across a variety of abnormalities, organs, and disease states thereby mitigating the need for human annotation.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Methods<\/jats:title>\n                <jats:p>We used a dictionary approach to develop rule-based algorithms (RBA) for extraction of disease labels from radiology text reports. We targeted three organ systems (lungs\/pleura, liver\/gallbladder, kidneys\/ureters) with four diseases per system based on their prevalence in our dataset. To expand the algorithms beyond pre-defined keywords, attention-guided recurrent neural networks (RNN) were trained using the RBA-extracted labels to classify reports as being positive for one or more diseases or normal for each organ system. Alternative effects on disease classification performance were evaluated using random initialization or pre-trained embedding as well as different sizes of training datasets. The RBA was tested on a subset of 2158 manually labeled reports and performance was reported as accuracy and F-score. The RNN was tested against a test set of 48,758 reports labeled by RBA and performance was reported as area under the receiver operating characteristic curve (AUC), with 95% CIs calculated using the DeLong method.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>Manual validation of the RBA confirmed 91\u201399% accuracy across the 15 different labels. Our models extracted disease labels from 261,229 radiology reports of 112,501 unique subjects. Pre-trained models outperformed random initialization across all diseases. As the training dataset size was reduced, performance was robust except for a few diseases with a relatively small number of cases. Pre-trained classification AUCs reached &gt;\u20090.95 for all four disease outcomes and normality across all three organ systems.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusions<\/jats:title>\n                <jats:p>Our label-extracting pipeline was able to encompass a variety of cases and diseases in body CT reports by generalizing beyond strict rules with exceptional accuracy. The method described can be easily adapted to enable automated labeling of hospital-scale medical data sets for training image-based disease classifiers.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s12911-022-01843-4","type":"journal-article","created":{"date-parts":[[2022,4,15]],"date-time":"2022-04-15T08:02:54Z","timestamp":1650009774000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Multi-label annotation of text reports from computed tomography of the chest, abdomen, and pelvis using deep learning"],"prefix":"10.1186","volume":"22","author":[{"given":"Vincent M.","family":"D\u2019Anniballe","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fakrul Islam","family":"Tushar","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Khrystyna","family":"Faryna","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Songyue","family":"Han","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maciej A.","family":"Mazurowski","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Geoffrey D.","family":"Rubin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joseph Y.","family":"Lo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,4,15]]},"reference":[{"issue":"2","key":"1843_CR1","doi-asserted-by":"publisher","first-page":"329","DOI":"10.1148\/radiol.16142770","volume":"279","author":"E Pons","year":"2016","unstructured":"Pons E, Braun LM, Hunink MM, Kors JA. Natural language processing in radiology: a systematic review. Radiology. 2016;279(2):329\u201343.","journal-title":"Radiology"},{"issue":"2","key":"1843_CR2","doi-asserted-by":"publisher","first-page":"323","DOI":"10.1148\/radiol.2341040049","volume":"234","author":"KJ Dreyer","year":"2005","unstructured":"Dreyer KJ, Kalra MK, Maher MM, Hurier AM, Asfaw BA, Schultz T, et al. Application of recently developed computer algorithm for automatic classification of unstructured radiology reports: validation study. Radiology. 2005;234(2):323\u20139.","journal-title":"Radiology"},{"key":"1843_CR3","doi-asserted-by":"crossref","unstructured":"Solti I, Cooke CR, Xia F, Wurfel MM, editors. Automated classification of radiology reports for acute lung injury: comparison of keyword and machine learning based natural language processing approaches. In 2009 IEEE international conference on bioinformatics and biomedicine workshop; 2009: IEEE.","DOI":"10.1109\/BIBMW.2009.5332081"},{"issue":"5","key":"1843_CR4","doi-asserted-by":"publisher","first-page":"913","DOI":"10.1136\/amiajnl-2011-000607","volume":"19","author":"B Percha","year":"2012","unstructured":"Percha B, Nassif H, Lipson J, Burnside E, Rubin D. Automatic classification of mammography reports by BI-RADS breast tissue composition class. J Am Med Inform Assoc. 2012;19(5):913\u20136.","journal-title":"J Am Med Inform Assoc"},{"key":"1843_CR5","doi-asserted-by":"crossref","unstructured":"Wang X, Peng Y, Lu L, Lu Z, Bagheri M, Summers RM. ChestX-Ray8: hospital-scale chest X-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In: 2017 IEEE conference on computer vision and pattern recognition (CVPR), 2017, pp. 3462\u201371.","DOI":"10.1109\/CVPR.2017.369"},{"key":"1843_CR6","doi-asserted-by":"publisher","first-page":"101857","DOI":"10.1016\/j.media.2020.101857","volume":"67","author":"RL Draelos","year":"2021","unstructured":"Draelos RL, Dov D, Mazurowski MA, Lo JY, Henao R, Rubin GD, et al. Machine-learning-based multiple abnormality prediction with large-scale chest computed tomography volumes. Med Image Anal. 2021;67:101857.","journal-title":"Med Image Anal"},{"issue":"1","key":"1843_CR7","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1016\/j.acra.2017.08.005","volume":"25","author":"D Ganeshan","year":"2018","unstructured":"Ganeshan D, Duong P-AT, Probyn L, Lenchik L, McArthur TA, Retrouvey M, et al. Structured reporting in radiology. Acad Radiol. 2018;25(1):66\u201373.","journal-title":"Acad Radiol"},{"key":"1843_CR8","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-10-5209-5","volume-title":"Deep learning in natural language processing","author":"L Deng","year":"2018","unstructured":"Deng L, Liu Y. Deep learning in natural language processing. Berlin: Springer; 2018."},{"key":"1843_CR9","unstructured":"Bahdanau D, Cho K, Bengio Y. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:14090473. 2014."},{"issue":"5","key":"1843_CR10","doi-asserted-by":"publisher","first-page":"e180052","DOI":"10.1148\/ryai.2019180052","volume":"1","author":"JM Steinkamp","year":"2019","unstructured":"Steinkamp JM, Chambers CM, Lalevic D, Zafar HM, Cook TS. Automated organ-level classification of free-text pathology reports to support a radiology follow-up tracking engine. Radiol Artificial Intell. 2019;1(5):e180052.","journal-title":"Radiol Artificial Intell"},{"key":"1843_CR11","first-page":"285","volume":"2019","author":"J Yuan","year":"2019","unstructured":"Yuan J, Zhu H, Tahmasebi A. Classification of pulmonary nodular findings based on characterization of change using radiology reports. AMIA Jt Summits Transl Sci Proc. 2019;2019:285\u201394.","journal-title":"AMIA Jt Summits Transl Sci Proc"},{"key":"1843_CR12","unstructured":"Raffel C, Ellis DP. Feed-forward networks with attention can solve some long-term memory problems. arXiv preprint arXiv:151208756. 2015."},{"key":"1843_CR13","doi-asserted-by":"crossref","unstructured":"Wang Y, Huang M, Zhu X, Zhao L. Attention-based LSTM for aspect-level sentiment classification. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing; 2016.","DOI":"10.18653\/v1\/D16-1058"},{"key":"1843_CR14","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1016\/j.artmed.2018.11.004","volume":"97","author":"I Banerjee","year":"2019","unstructured":"Banerjee I, Ling Y, Chen MC, Hasan SA, Langlotz CP, Moradzadeh N, et al. Comparative effectiveness of convolutional neural network (CNN) and recurrent neural network (RNN) architectures for radiology text report classification. Artif Intell Med. 2019;97:79\u201388.","journal-title":"Artif Intell Med"},{"key":"1843_CR15","doi-asserted-by":"crossref","unstructured":"Han S, Tian J, Kelly M, Selvakumaran V, Henao R, Rubin GD, Lo JY. Classifying abnormalities in\ncomputed tomography radiology reports with rule-based and natural language \nprocessing models. Proc. SPIE 10950, Medical Imaging 2019: Computer-Aided Diagnosis, 109504H.","DOI":"10.1117\/12.2513577"},{"key":"1843_CR16","doi-asserted-by":"crossref","unstructured":"Faryna K, Tushar FI, D'Anniballe VM, Hou R, Rubin GD, Lo JY. Attention-guided classification of abnormalities in\nsemi-structured computed tomography reports. Proc. SPIE 11314, Medical \nImaging 2020: Computer-Aided Diagnosis, 113141P.","DOI":"10.1117\/12.2551370"},{"key":"1843_CR17","doi-asserted-by":"crossref","unstructured":"Wu HC, Luk RWP, Wong KF, Kwok KL. Interpreting TF-IDF term weights as making relevance decisions. ACM Trans Inf Syst. 2008;26(3):Article 13.","DOI":"10.1145\/1361684.1361686"},{"issue":"8","key":"1843_CR18","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735\u201380.","journal-title":"Neural Comput"},{"key":"1843_CR19","doi-asserted-by":"crossref","unstructured":"Zhou Q, Wu H, editors. NLP at IEST 2018: BiLSTM-attention and LSTM-attention via soft voting in emotion classification. In: Proceedings of the 9th workshop on computational approaches to subjectivity, sentiment and social media analysis; 2018.","DOI":"10.18653\/v1\/W18-6226"},{"issue":"1","key":"1843_CR20","doi-asserted-by":"publisher","first-page":"52","DOI":"10.1038\/s41597-019-0055-0","volume":"6","author":"Y Zhang","year":"2019","unstructured":"Zhang Y, Chen Q, Yang Z, Lin H, Lu Z. BioWordVec, improving biomedical word embeddings with subword information and MeSH. Sci Data. 2019;6(1):52.","journal-title":"Sci Data"},{"issue":"3","key":"1843_CR21","doi-asserted-by":"publisher","first-page":"837","DOI":"10.2307\/2531595","volume":"44","author":"ER DeLong","year":"1988","unstructured":"DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach. Biometrics. 1988;44(3):837\u201345.","journal-title":"Biometrics"},{"key":"1843_CR22","doi-asserted-by":"crossref","unstructured":"Tushar FI, D'Anniballe VM, Hou R, Mazurowski MA, Fu W, Samei E, et al. Classification of multiple diseases on body CT scans using weakly supervised deep learning. Radiol Artif Intell 2021:e210026.","DOI":"10.1148\/ryai.210026"},{"issue":"1","key":"1843_CR23","first-page":"3","volume":"81","author":"A Brady","year":"2012","unstructured":"Brady A, Laoide R, McCarthy P, McDermott R. Discrepancy and error in radiology: concepts, causes and consequences. Ulster Med J. 2012;81(1):3\u20139.","journal-title":"Ulster Med J"},{"issue":"5","key":"1843_CR24","doi-asserted-by":"publisher","first-page":"639","DOI":"10.1016\/j.jacr.2019.12.026","volume":"17","author":"V Sorin","year":"2020","unstructured":"Sorin V, Barash Y, Konen E, Klang E. Deep learning for natural language processing in radiology\u2014fundamentals and a systematic review. J Am Coll Radiol. 2020;17(5):639\u201348.","journal-title":"J Am Coll Radiol"},{"issue":"5","key":"1843_CR25","doi-asserted-by":"publisher","first-page":"685","DOI":"10.1007\/s10278-018-0141-4","volume":"32","author":"RG Short","year":"2019","unstructured":"Short RG, Bralich J, Bogaty D, Befera NT. Comprehensive word-level classification of screening mammography reports using a neural network sequence labeling approach. J Digit Imaging. 2019;32(5):685\u201392.","journal-title":"J Digit Imaging"},{"issue":"1","key":"1843_CR26","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12911-018-0723-6","volume":"19","author":"Y Wang","year":"2019","unstructured":"Wang Y, Sohn S, Liu S, Shen F, Wang L, Atkinson EJ, et al. A clinical text classification paradigm using weak supervision and deep representation. BMC Med Inform Decis Mak. 2019;19(1):1\u201313.","journal-title":"BMC Med Inform Decis Mak"},{"issue":"1","key":"1843_CR27","doi-asserted-by":"publisher","first-page":"155","DOI":"10.1186\/s12911-017-0556-8","volume":"17","author":"WH Weng","year":"2017","unstructured":"Weng WH, Wagholikar KB, McCray AT, Szolovits P, Chueh HC. Medical subdomain classification of clinical notes using a machine learning-based natural language processing approach. BMC Med Inform Decis Mak. 2017;17(1):155.","journal-title":"BMC Med Inform Decis Mak"},{"issue":"1","key":"1843_CR28","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1186\/s41747-019-0118-1","volume":"3","author":"A Spandorfer","year":"2019","unstructured":"Spandorfer A, Branch C, Sharma P, Sahbaee P, Schoepf UJ, Ravenel JG, et al. Deep learning to convert unstructured CT pulmonary angiography reports into structured reports. Eur Radiol Exp. 2019;3(1):37.","journal-title":"Eur Radiol Exp"},{"key":"1843_CR29","doi-asserted-by":"crossref","unstructured":"Shin H-C, Roberts K, Lu L, Demner-Fushman D, Yao J, Summers RM. Learning to read chest X-rays: recurrent neural cascade model for automated image annotation. Proc CVPR IEEE 2016. p. 2497\u2013506.","DOI":"10.1109\/CVPR.2016.274"},{"key":"1843_CR30","doi-asserted-by":"crossref","unstructured":"Laserson J, Lantsman CD, Cohen-Sfady M, Tamir I, Goz E, Brestel C, et al. TextRay: mining clinical reports to gain a broad understanding of chest X-rays. In: Medical image computing and computer assisted intervention\u2014MICCAI 2018. Lecture Notes in Computer Science, 2018. p. 553\u201361.","DOI":"10.1007\/978-3-030-00934-2_62"},{"issue":"2","key":"1843_CR31","first-page":"1021","volume":"14","author":"C Kim","year":"2019","unstructured":"Kim C, Zhu V, Obeid J, Lenert L. Natural language processing and machine learning algorithm to identify brain MRI reports with acute ischemic stroke. PLoS ONE. 2019;14(2):1021.","journal-title":"PLoS ONE"},{"issue":"2","key":"1843_CR32","doi-asserted-by":"publisher","first-page":"361","DOI":"10.1093\/jamia\/ocw112","volume":"24","author":"E Choi","year":"2017","unstructured":"Choi E, Schuetz A, Stewart WF, Sun J. Using recurrent neural network models for early detection of heart failure onset. J Am Med Inform Assoc. 2017;24(2):361\u201370.","journal-title":"J Am Med Inform Assoc"},{"issue":"1","key":"1843_CR33","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1148\/radiol.2020192224","volume":"295","author":"MJ Willemink","year":"2020","unstructured":"Willemink MJ, Koszek WA, Hardell C, Wu J, Fleischmann D, Harvey H, et al. Preparing medical imaging data for machine learning. Radiology. 2020;295(1):4\u201315.","journal-title":"Radiology"},{"issue":"5","key":"1843_CR34","doi-asserted-by":"publisher","first-page":"828","DOI":"10.1093\/bioinformatics\/btx659","volume":"34","author":"Y Zhang","year":"2018","unstructured":"Zhang Y, Zheng W, Lin H, Wang J, Yang Z, Dumontier M. Drug-drug interaction extraction via hierarchical RNNs on sequence and shortest dependency paths. Bioinformatics. 2018;34(5):828\u201335.","journal-title":"Bioinformatics"},{"key":"1843_CR35","doi-asserted-by":"crossref","unstructured":"Tang D, Wei F, Yang N, Zhou M, Liu T, Qin B, editors. Learning sentiment-specific word embedding for twitter sentiment classification. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); 2014.","DOI":"10.3115\/v1\/P14-1146"},{"key":"1843_CR36","doi-asserted-by":"publisher","first-page":"12","DOI":"10.1016\/j.jbi.2018.09.008","volume":"87","author":"Y Wang","year":"2018","unstructured":"Wang Y, Liu S, Afzal N, Rastegar-Mojarad M, Wang L, Shen F, et al. A comparison of word embeddings for the biomedical natural language processing. J Biomed Inform. 2018;87:12\u201320.","journal-title":"J Biomed Inform"}],"container-title":["BMC Medical Informatics and Decision Making"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-022-01843-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12911-022-01843-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-022-01843-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,4,16]],"date-time":"2022-04-16T05:05:33Z","timestamp":1650085533000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcmedinformdecismak.biomedcentral.com\/articles\/10.1186\/s12911-022-01843-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,15]]},"references-count":36,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["1843"],"URL":"https:\/\/doi.org\/10.1186\/s12911-022-01843-4","relation":{},"ISSN":["1472-6947"],"issn-type":[{"value":"1472-6947","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,4,15]]},"assertion":[{"value":"20 April 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 April 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 April 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"This study was approved by the IRB at Duke University under protocol # Pro00082329. Informed consent was waived by the IRB at Duke University for this retrospective study that was compliant with the Health Insurance Portability and Accountability Act. IRB approval included permission to access the raw data. All experiments were performed in accordance with relevant guidelines and regulations.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"102"}}