{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T05:54:20Z","timestamp":1783576460022,"version":"3.55.0"},"reference-count":116,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2022,6,16]],"date-time":"2022-06-16T00:00:00Z","timestamp":1655337600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,6,16]],"date-time":"2022-06-16T00:00:00Z","timestamp":1655337600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100009244","name":"Stockholm University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100009244","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Knowl Inf Syst"],"published-print":{"date-parts":[[2022,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Diagnostic captioning (DC) concerns the automatic generation of a diagnostic text from a set of medical images of a patient collected during an examination. DC can assist inexperienced physicians, reducing clinical errors. It can also help experienced physicians produce diagnostic reports faster. Following the advances of deep learning, especially in generic image captioning, DC has recently attracted more attention, leading to several systems and datasets. This article is an extensive overview of DC. It presents relevant datasets, evaluation measures, and up-to-date systems. It also highlights shortcomings that hinder DC\u2019s progress and proposes future directions.<\/jats:p>","DOI":"10.1007\/s10115-022-01684-7","type":"journal-article","created":{"date-parts":[[2022,6,16]],"date-time":"2022-06-16T06:02:36Z","timestamp":1655359356000},"page":"1691-1722","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":24,"title":["Diagnostic captioning: a survey"],"prefix":"10.1007","volume":"64","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9188-7425","authenticated-orcid":false,"given":"John","family":"Pavlopoulos","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vasiliki","family":"Kougia","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ion","family":"Androutsopoulos","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dimitris","family":"Papamichail","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,6,16]]},"reference":[{"key":"1684_CR1","first-page":"1","volume":"5","author":"HJ Aerts","year":"2014","unstructured":"Aerts HJ, Velazquez ER, Leijenaar RT, Parmar C, Grossmann P, Carvalho S, Bussink J, Monshouwer R, Haibe-Kains B, Rietveld D et al (2014) Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach. Nat Commun 5:1\u20139","journal-title":"Nat Commun"},{"key":"1684_CR2","doi-asserted-by":"crossref","unstructured":"Agrawal H, Desai K, Wang Y, Chen X, Jain R, Johnson M, Batra D, Parikh D, Lee S, Anderson P (2019) nocaps: novel object captioning at scale. In: Proceedings of the IEEE international conference on computer vision, Seoul, Korea, pp 8948\u20138957","DOI":"10.1109\/ICCV.2019.00904"},{"key":"1684_CR3","doi-asserted-by":"crossref","unstructured":"Anderson P, Fernando B, Johnson M, Gould S (2016) SPICE: semantic propositional image caption evaluation. In: Proceedings of the European conference on computer vision, Amsterdam, Netherlands, pp 382\u2013398","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"1684_CR4","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1016\/j.neucom.2018.05.080","volume":"311","author":"S Bai","year":"2018","unstructured":"Bai S, An S (2018) A survey on automatic image caption generation. Neurocomputing 311:291\u2013304","journal-title":"Neurocomputing"},{"key":"1684_CR5","unstructured":"Banerjee S, Lavie A (2005) METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In: Proceedings of the workshop on intrinsic and extrinsic evaluation measures for machine translation and\/or summarization of the annual conference of the association for computational linguistics, Ann Arbor, MI, USA, pp 65\u201372"},{"key":"1684_CR6","doi-asserted-by":"publisher","first-page":"1173","DOI":"10.2214\/AJR.06.1270","volume":"188","author":"L Berlin","year":"2007","unstructured":"Berlin L (2007) Accuracy of diagnostic procedures: has it improved over the past five decades? Am J Roentgenol 188:1173\u20131178","journal-title":"Am J Roentgenol"},{"key":"1684_CR7","doi-asserted-by":"publisher","first-page":"409","DOI":"10.1613\/jair.4900","volume":"55","author":"R Bernardi","year":"2016","unstructured":"Bernardi R, Cakici R, Elliott D, Erdem A, Erdem E, Ikizler-Cinbis N, Keller F, Muscat A, Plank B (2016) Automatic description generation from images: a survey of models, datasets, and evaluation measures. J Artif Intell Res 55:409\u2013442","journal-title":"J Artif Intell Res"},{"key":"1684_CR8","unstructured":"Boag W, Hsu T-MH, McDermott M, Berner G, Alesentzer E, Szolovits P (2020) Baselines for chest x-ray report generation. In: Machine learning for health workshop, pp 126\u2013140"},{"key":"1684_CR9","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1007\/s13244-016-0534-1","volume":"8","author":"AP Brady","year":"2017","unstructured":"Brady AP (2017) Error and discrepancy in radiology: inevitable or avoidable? Insights Imaging 8:171\u2013182","journal-title":"Insights Imaging"},{"key":"1684_CR10","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, Agarwal S, Herbert-Voss A, Krueger G, Henighan T, Child R, Ramesh A, Ziegler D, Wu J, Winter C, Hesse C, Chen M, Sigler E, Litwin M, Gray S, Chess B, Clark J, Berner C, McCandlish S, Radford A, Sutskever I, Amodei D (2020) Language models are few-shot learners. In: Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H (eds) Advances in neural information processing systems, vol 33. Curran Associates Inc, pp 1877\u20131901"},{"key":"1684_CR11","doi-asserted-by":"publisher","first-page":"101797","DOI":"10.1016\/j.media.2020.101797","volume":"66","author":"A Bustos","year":"2020","unstructured":"Bustos A, Pertusa A, Salinas J-M, de la Iglesia-Vay\u00e1 M (2020) Padchest: a large chest X-ray image dataset with multi-label annotated reports. Med Image Anal 66:101797","journal-title":"Med Image Anal"},{"issue":"1","key":"1684_CR12","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1177\/0846537120938328","volume":"72","author":"D Byrne","year":"2021","unstructured":"Byrne D, Neill SBO, M\u00fcller NL, M\u00fcller CIS, Walsh JP, Jalal S, Parker W, Bilawich A-M, Nicolaou S (2021) RSNA expert consensus statement on reporting chest CT findings related to COVID-19: interobserver agreement between chest radiologists. Can Assoc Radiol J 72(1):159\u2013166","journal-title":"Can Assoc Radiol J"},{"issue":"5","key":"1684_CR13","doi-asserted-by":"publisher","first-page":"1626","DOI":"10.1007\/s00259-021-05245-y","volume":"48","author":"F Ceci","year":"2021","unstructured":"Ceci F, Oprea-Lager DE, Emmett L, Adam JA, Bomanji J, Czernin J, Eiber M, Haberkorn U, Hofman MS, Hope TA et al (2021) E-PSMA: the EANM standardized reporting guidelines v1. 0 for PSMA-PET. Eur J Nucl Med Mol Imaging 48(5):1626\u20131638","journal-title":"Eur J Nucl Med Mol Imaging"},{"key":"1684_CR14","volume-title":"Introduction to deep learning","author":"E Charniak","year":"2018","unstructured":"Charniak E (2018) Introduction to deep learning. MIT Press, Cambridge"},{"key":"1684_CR15","unstructured":"Chen X, Fang H, Lin T-Y, Vedantam R, Gupta S, Doll\u00e1r P, Zitnick CL (2015) Microsoft COCO captions: data collection and evaluation server. arXiv:1504.00325"},{"key":"1684_CR16","doi-asserted-by":"crossref","unstructured":"Chen Z, Song Y, Chang T-H, Wan X (2020) Generating radiology reports via memory-driven transformer. In: Proceedings of the 2020 conference on empirical methods in natural language processing","DOI":"10.18653\/v1\/2020.emnlp-main.112"},{"key":"1684_CR17","doi-asserted-by":"crossref","unstructured":"Cho K, van Merrienboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, Bengio Y (2014) Learning phrase representations using RNN encoder\u2013decoder for statistical machine translation. In: Proceedings of the conference on empirical methods in natural language processing, Doha, Qatar, pp 1724\u20131734","DOI":"10.3115\/v1\/D14-1179"},{"key":"1684_CR18","doi-asserted-by":"publisher","first-page":"664","DOI":"10.1016\/j.jacr.2015.02.009","volume":"12","author":"FH Chokshi","year":"2015","unstructured":"Chokshi FH, Hughes DR, Wang JM, Mullins ME, Hawkins CM, Duszak R Jr (2015) Diagnostic radiology resident and fellow workloads: a 12-year longitudinal trend analysis using national medicare aggregate claims data. J Am Coll Radiol 12:664\u2013669","journal-title":"J Am Coll Radiol"},{"issue":"2","key":"1684_CR19","doi-asserted-by":"publisher","first-page":"318","DOI":"10.1148\/radiol.2018171820","volume":"288","author":"G Choy","year":"2018","unstructured":"Choy G, Khalilzadeh O, Michalski M, Do S, Samir AE, Pianykh OS, Geis JR, Pandharipande PV, Brink JA, Dreyer KJ (2018) Current applications and future impact of machine learning in radiology. Radiology 288(2):318\u2013328","journal-title":"Radiology"},{"key":"1684_CR20","unstructured":"de\u00a0Herrera AGS, Eickhoff C, Andrearczyk V, M\u00fcller H (2018) Overview of the ImageCLEF 2018 caption prediction tasks. In: Proceedings of the CEUR workshop, CLEF2018 working notes, Avignon, France"},{"key":"1684_CR21","doi-asserted-by":"publisher","first-page":"304","DOI":"10.1093\/jamia\/ocv080","volume":"23","author":"D Demner-Fushman","year":"2015","unstructured":"Demner-Fushman D, Kohli MD, Rosenman MB, Shooshan SE, Rodriguez L, Antani S, Thoma GR, McDonald CJ (2015) Preparing a collection of radiology examinations for distribution and retrieval. J Am Med Inform Assoc 23:304\u2013310","journal-title":"J Am Med Inform Assoc"},{"key":"1684_CR22","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2018) BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the North American chapter of the association for computational linguistics, Minneapolis, MN, USA, pp 4171\u20144186"},{"key":"1684_CR23","doi-asserted-by":"crossref","unstructured":"Donahue J, Hendricks LA, Guadarrama S, Rohrbach M, Venugopalan S, Saenko K, Darrell T (2015) Long-term recurrent convolutional networks for visual recognition and description. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Boston, MA, USA, pp 2625\u20132634","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"1684_CR24","unstructured":"Eickhoff C, Schwall I, de\u00a0Herrera AGS, M\u00fcller H (2017) Overview of ImageCLEFcaption 2017\u2014the image caption prediction and concept extraction tasks to understand biomedical images. In: Proceeding of the CEUR workshop, CLEF2017 working notes, Dublin, Ireland"},{"key":"1684_CR25","doi-asserted-by":"crossref","unstructured":"European Society of Radiology (ESR) (2018) ESR paper on structured reporting in radiology. Insights Imaging 9:1\u20137","DOI":"10.1007\/s13244-017-0588-8"},{"key":"1684_CR26","doi-asserted-by":"publisher","first-page":"246","DOI":"10.1016\/j.ejrad.2018.06.020","volume":"105","author":"MI Fazal","year":"2018","unstructured":"Fazal MI, Patel ME, Tye J, Gupta Y (2018) The past, present and future role of artificial intelligence in imaging. Eur J Radiol 105:246\u2013250","journal-title":"Eur J Radiol"},{"key":"1684_CR27","doi-asserted-by":"crossref","unstructured":"Fellbaum C (2012) WordNet. The encyclopedia of applied linguistics","DOI":"10.1002\/9781405198431.wbeal1285"},{"key":"1684_CR28","doi-asserted-by":"publisher","first-page":"601","DOI":"10.1197\/jamia.M2702","volume":"15","author":"FJ Friedlin","year":"2008","unstructured":"Friedlin FJ, McDonald CJ (2008) A software tool for removing patient identifying information from clinical documents. J Am Med Inform Assoc 15:601\u2013610","journal-title":"J Am Med Inform Assoc"},{"key":"1684_CR29","doi-asserted-by":"crossref","unstructured":"Gale W, Oakden-Rayner L, Carneiro G, Bradley AP, Palmer LJ (2018) Producing radiologist-quality reports for interpretable artificial intelligence. arXiv:1806.00340","DOI":"10.1109\/ISBI.2019.8759236"},{"key":"1684_CR30","doi-asserted-by":"crossref","unstructured":"Gasimova A, Seegoolam G, Chen L, Bentley P, Rueckert D (2020) Spatial semantic-preserving latent space learning for accelerated DWI diagnostic report generation. In: International conference on medical image computing and computer-assisted intervention, Springer, Berlin, pp 333\u2013342","DOI":"10.1007\/978-3-030-59728-3_33"},{"key":"1684_CR31","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1613\/jair.5477","volume":"61","author":"A Gatt","year":"2018","unstructured":"Gatt A, Krahmer E (2018) Survey of the state of the art in natural language generation: core tasks, applications and evaluation. J Artif Intell Res 61:65\u2013170","journal-title":"J Artif Intell Res"},{"key":"1684_CR32","doi-asserted-by":"crossref","unstructured":"Goldberg Y (2017) Neural network methods in natural language processing. Morgan and Claypool Publishers","DOI":"10.1007\/978-3-031-02165-7"},{"key":"1684_CR33","unstructured":"Goodfellow I, Bengio Y, Courville A (2016) Deep learning. MIT press, Cambridge"},{"key":"1684_CR34","doi-asserted-by":"crossref","unstructured":"Graham Y (2015) Re-evaluating automatic summarization with BLEU and 192 shades of ROUGE. In: Proceedings of the conference on empirical methods in natural language processing, Lisbon, Portugal, pp 128\u2013137","DOI":"10.18653\/v1\/D15-1013"},{"issue":"1108","key":"1684_CR35","doi-asserted-by":"publisher","first-page":"20190840","DOI":"10.1259\/bjr.20190840","volume":"93","author":"M Hardy","year":"2020","unstructured":"Hardy M, Harvey H (2020) Artificial intelligence in diagnostic imaging: impact on the radiography profession. Br J Radiol 93(1108):20190840","journal-title":"Br J Radiol"},{"key":"1684_CR36","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Las Vegas, NV, USA, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"1684_CR37","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9:1735\u20131780","journal-title":"Neural Comput"},{"key":"1684_CR38","doi-asserted-by":"publisher","first-page":"500","DOI":"10.1038\/s41568-018-0016-5","volume":"18","author":"A Hosny","year":"2018","unstructured":"Hosny A, Parmar C, Quackenbush J, Schwartz LH, Aerts HJ (2018) Artificial intelligence in radiology. Nat Rev Cancer 18:500\u2013510","journal-title":"Nat Rev Cancer"},{"key":"1684_CR39","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3295748","volume":"51","author":"M Hossain","year":"2019","unstructured":"Hossain M, Sohel F, Shiratuddin MF, Laga H (2019) A comprehensive survey of deep learning for image captioning. ACM Comput Surv 51:1\u201336","journal-title":"ACM Comput Surv"},{"key":"1684_CR40","doi-asserted-by":"crossref","unstructured":"Huang G, Liu Z, Maaten LVD, Weinberger KQ (2017) Densely connected convolutional networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA, pp 4700\u20134708","DOI":"10.1109\/CVPR.2017.243"},{"key":"1684_CR41","doi-asserted-by":"publisher","first-page":"154808","DOI":"10.1109\/ACCESS.2019.2947134","volume":"7","author":"X Huang","year":"2019","unstructured":"Huang X, Yan F, Xu W, Li M (2019) Multi-attention and incorporating background information model for chest X-ray image report generation. IEEE Access 7:154808\u2013154817","journal-title":"IEEE Access"},{"key":"1684_CR42","unstructured":"Iandola F, Moskewicz M, Karayev S, Girshick R, Darrell T, Keutzer K (2014) DenseNet: implementing efficient ConvNet descriptor pyramids. arXiv:1404.1869"},{"key":"1684_CR43","doi-asserted-by":"crossref","unstructured":"Irvin J, Rajpurkar P, Ko M, Yu Y, Ciurea-Ilcus S, Chute C, Marklund H, Haghgoo B, Ball R, Shpanskaya K et al (2019) CheXpert: a large chest radiograph dataset with uncertainty labels and expert comparison. In: Proceedings of the AAAI conference on artificial intelligence, Honolulu, HI, USA, pp 590\u2013597","DOI":"10.1609\/aaai.v33i01.3301590"},{"key":"1684_CR44","doi-asserted-by":"crossref","unstructured":"Jia Y, Shelhamer E, Donahue J, Karayev S, Long J, Girshick R, Guadarrama S, Darrell T (2014) Caffe: convolutional architecture for fast feature embedding. In: Proceedings of the 22nd ACM international conference on multimedia, Orlando, FL, USA, pp 675\u2013678","DOI":"10.1145\/2647868.2654889"},{"key":"1684_CR45","doi-asserted-by":"crossref","unstructured":"Jing B, Xie P, Xing E (2018) On the automatic generation of medical imaging reports. In: Proceedings of the 56th annual meeting of the association for computational linguistics, Melbourne, Australia, pp 2577\u20132586","DOI":"10.18653\/v1\/P18-1240"},{"key":"1684_CR46","doi-asserted-by":"crossref","unstructured":"Johnson AE, Pollard TJ, Berkowitz S, Greenbaum NR, Lungren MP, Deng C-Y, Mark RG, Horng S (2019) MIMIC-CXR: a large publicly available database of labeled chest radiographs. arXiv:1901.07042","DOI":"10.1038\/s41597-019-0322-0"},{"key":"1684_CR47","doi-asserted-by":"crossref","unstructured":"Karpathy A, Fei-Fei L (2015) Deep visual-semantic alignments for generating image descriptions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Boston, MA, USA, pp 3128\u20133137","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"1684_CR48","doi-asserted-by":"crossref","unstructured":"Kassner N, Sch\u00fctze H (2020) Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly. In: Proceedings of the 58th annual meeting of the association for computational linguistics, pp 7811\u20137818, held on-line","DOI":"10.18653\/v1\/2020.acl-main.698"},{"key":"1684_CR49","unstructured":"Khandelwal U, Levy O, Jurafsky D, Zettlemoyer L, Lewis M (2020) Generalization through memorization: nearest neighbor language models. In: Proceedings of the international conference on learning representations, pp 1\u201320, held on-line"},{"key":"1684_CR50","doi-asserted-by":"crossref","unstructured":"Kilickaya M, Erdem A, Ikizler-Cinbis N, Erdem E (2016) Re-evaluating automatic metrics for image captioning. In: Proceedings of the conference of the European chapter of the association for computational linguistics, Valencia, Spain, pp 199\u2013209","DOI":"10.18653\/v1\/E17-1019"},{"issue":"3","key":"1684_CR51","doi-asserted-by":"publisher","first-page":"405","DOI":"10.3348\/kjr.2019.0025","volume":"20","author":"DW Kim","year":"2019","unstructured":"Kim DW, Jang HY, Kim KW, Shin Y, Park SH (2019) Design characteristics of studies reporting the performance of artificial intelligence algorithms for diagnostic analysis of medical images: results from recently published papers. Korean J Radiol 20(3):405\u2013410","journal-title":"Korean J Radiol"},{"key":"1684_CR52","doi-asserted-by":"crossref","unstructured":"Kisilev P, Sason E, Barkan E, Hashoul S (2016) Medical image captioning: learning to describe medical image findings using multi-task-loss CNN. In: Proceedings of the 1st international workshop on deep learning for precision medicine, Riva del Garda, Italy","DOI":"10.1007\/978-3-319-46976-8_13"},{"key":"1684_CR53","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1147\/JRD.2015.2393193","volume":"59","author":"P Kisilev","year":"2015","unstructured":"Kisilev P, Walach E, Barkan E, Ophir B, Alpert S, Hashoul SY (2015a) From medical image to automatic medical report generation. IBM J Res Dev 59:1\u20137","journal-title":"IBM J Res Dev"},{"key":"1684_CR54","doi-asserted-by":"crossref","unstructured":"Kisilev P, Walach E, Hashoul SY, Barkan E, Ophir B, Alpert S (2015b) Semantic description of medical image findings: structured learning approach. In: Proceedings of the British machine vision conference, Swansea, UK, pp 171.1\u2013171.11","DOI":"10.5244\/C.29.171"},{"key":"1684_CR55","doi-asserted-by":"crossref","unstructured":"Kougia V, Pavlopoulos J, Androutsopoulos I (2019) A survey on biomedical image captioning. In: Proceedings of the workshop on shortcomings in vision and language of the annual conference of the North American chapter of the association for computational linguistics, Minneapolis, MN, USA, pp 26\u201336","DOI":"10.18653\/v1\/W19-1803"},{"key":"1684_CR56","doi-asserted-by":"publisher","first-page":"1205","DOI":"10.3758\/APP.72.5.1205","volume":"72","author":"EA Krupinski","year":"2010","unstructured":"Krupinski EA (2010) Current perspectives in medical image perception. Attention, Perception, & Psychophysics 72:1205\u20131217","journal-title":"Attention, Perception, & Psychophysics"},{"issue":"3","key":"1684_CR57","doi-asserted-by":"publisher","first-page":"e190058","DOI":"10.1148\/ryai.2019190058","volume":"1","author":"CP Langlotz","year":"2019","unstructured":"Langlotz CP (2019) Will artificial intelligence replace radiologists? Radiol Artif Intell 1(3):e190058","journal-title":"Radiol Artif Intell"},{"key":"1684_CR58","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun Y, Bengio Y, Hinton G (2015) Deep learning. Nature 521:436\u2013444","journal-title":"Nature"},{"key":"1684_CR59","unstructured":"Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, K\u00fcttler H, Lewis M, Yih W-T, Rockt\u00e4schel T et al (2020) Retrieval-augmented generation for knowledge-intensive NLP tasks. In: NIPS, Vancouver, Canada"},{"key":"1684_CR60","unstructured":"Li Y, Liang X, Hu Z, Xing E (2018) Hybrid retrieval-generation reinforced agent for medical image report generation. In: Proceedings of the 32nd international conference on neural information processing systems, Montreal, Canada, pp 1537\u20131547"},{"key":"1684_CR61","doi-asserted-by":"crossref","unstructured":"Li Y, Liang X, Hu Z, Xing E (2019) Knowledge-driven encode, retrieve, paraphrase for medical image report generation. In: Proceedings of the AAAI conference on artificial intelligence, Honolulu, HI, USA, pp 6666\u20136673","DOI":"10.1609\/aaai.v33i01.33016666"},{"key":"1684_CR62","unstructured":"Liang S, Li X, Zhu Y, Li X, Jiang S (2017) ISIA at the ImageCLEF 2017 image caption task. In: Proceedings of the CEUR workshop, CLEF2017 working notes, Dublin, Ireland"},{"key":"1684_CR63","doi-asserted-by":"publisher","first-page":"152","DOI":"10.1016\/j.ejrad.2018.03.019","volume":"102","author":"C Liew","year":"2018","unstructured":"Liew C (2018) The future of radiology augmented with artificial intelligence: a strategy for success. Eur J Radiol 102:152\u2013156","journal-title":"Eur J Radiol"},{"key":"1684_CR64","unstructured":"Lin C-Y (2004) ROUGE: A package for automatic evaluation of summaries. In: Proceedings of the workshop on text summarization branches out of the annual conference of the association for computational linguistics, Barcelona, Spain, pp 74\u201381"},{"key":"1684_CR65","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft COCO: common objects in context. In: Proceedings of the European conference on computer vision, Zurich, Switzerland, pp 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"1684_CR66","doi-asserted-by":"crossref","unstructured":"Liu F, Wu X, Ge S, Fan W, Zou Y (2021) Exploring and distilling posterior and prior knowledge for radiology report generation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 13753\u201313762, held on-line","DOI":"10.1109\/CVPR46437.2021.01354"},{"key":"1684_CR67","unstructured":"Liu G, Hsu T-MH, McDermott M, Boag W, Weng W-H, Szolovits P, Ghassemi M (2019a) Clinically accurate chest X-ray report generation. In: Proceedings of the machine learning for healthcare conference, Ann Arbor, MI, USA, pp 1\u201320"},{"key":"1684_CR68","doi-asserted-by":"publisher","first-page":"445","DOI":"10.1007\/s00371-018-1566-y","volume":"35","author":"X Liu","year":"2019","unstructured":"Liu X, Xu Q, Wang N (2019) A survey on deep neural network-based image captioning. Vis Comput 35:445\u2013470","journal-title":"Vis Comput"},{"key":"1684_CR69","doi-asserted-by":"crossref","unstructured":"Lu J, Xiong C, Parikh D, Socher R (2017) Knowing when to look: adaptive attention via a visual sentinel for image captioning. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA, pp 375\u2013383","DOI":"10.1109\/CVPR.2017.345"},{"key":"1684_CR70","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511809071","volume-title":"Introduction to information retrieval","author":"CD Manning","year":"2008","unstructured":"Manning CD, Raghavan P, Sch\u00fctze H (2008) Introduction to information retrieval. Cambridge University Press, Cambridge"},{"issue":"1","key":"1684_CR71","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1007\/s12553-020-00515-5","volume":"11","author":"LG Marcu","year":"2021","unstructured":"Marcu LG, Marcu D (2021) Points of view on artificial intelligence in medical imaging-one good, one bad, one fuzzy. Heal Technol 11(1):17\u201322","journal-title":"Heal Technol"},{"key":"1684_CR72","doi-asserted-by":"publisher","first-page":"101878","DOI":"10.1016\/j.artmed.2020.101878","volume":"106","author":"MMA Monshi","year":"2020","unstructured":"Monshi MMA, Poon J, Chung V (2020) Deep learning in generating radiology reports: a survey. Artif Intell Med 106:101878","journal-title":"Artif Intell Med"},{"key":"1684_CR73","unstructured":"Mork JG, Jimeno-Yepes A, Aronson AR (2013) The NLM medical text indexer system for indexing biomedical literature. In: Proceedings of BioASQ, Valencia, Spain"},{"key":"1684_CR74","volume-title":"Machine learning: a probabilistic perspective","author":"KP Murphy","year":"2012","unstructured":"Murphy KP (2012) Machine learning: a probabilistic perspective. MIT press, Cambridge"},{"key":"1684_CR75","doi-asserted-by":"publisher","first-page":"661","DOI":"10.1613\/jair.1.12025","volume":"68","author":"OM Nezami","year":"2020","unstructured":"Nezami OM, Dras M, Wan S, Paris C (2020) Image captioning using facial expression and attention. J Artif Intell Res 68:661\u2013689","journal-title":"J Artif Intell Res"},{"key":"1684_CR76","doi-asserted-by":"crossref","unstructured":"Papineni K, Roukos S, Ward T, Zhu W-J (2002) BLEU: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting on association for computational linguistics, Philadelphia, PA, USA, pp 311\u2013318","DOI":"10.3115\/1073083.1073135"},{"key":"1684_CR77","unstructured":"Pelka O, Friedrich CM, de\u00a0Herrera AGS, M\u00fcller H (2019) Overview of the ImageCLEFmed 2019 concept prediction task. In: Proceedings of the CEUR workshop, CLEF2019 working notes, Lugano, Switzerland"},{"key":"1684_CR78","unstructured":"Pelka O, Friedrich CM, Garc\u0131a Seco\u00a0de Herrera A, M\u00fcller H (2020) Overview of the imageclefmed 2020 concept prediction task: medical image understanding. In: Proceedings of the CEUR workshop, CLEF2020 working notes, Thessaloniki, Greece"},{"key":"1684_CR79","first-page":"9","volume":"1","author":"A Radford","year":"2019","unstructured":"Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I (2019) Language models are unsupervised multitask learners. OpenAI Blog 1:9","journal-title":"OpenAI Blog"},{"key":"1684_CR80","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511519857","volume-title":"Building natural language generation systems","author":"E Reiter","year":"2000","unstructured":"Reiter E, Dale R (2000) Building natural language generation systems. Cambridge University Press, Cambridge"},{"key":"1684_CR81","doi-asserted-by":"crossref","unstructured":"Rennie SJ, Marcheret E, Mroueh Y, Ross J, Goel V (2017) Self-critical sequence training for image captioning. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA, pp 7008\u20137024","DOI":"10.1109\/CVPR.2017.131"},{"key":"1684_CR82","doi-asserted-by":"crossref","unstructured":"Schlegl T, Waldstein SM, Vogl W-D, Schmidt-Erfurth U, Langs G (2015) Predicting semantic descriptions from medical images with convolutional neural networks. In: Proceedings of the international conference on information processing in medical imaging, Isle of Skye, UK, pp 437\u2013448","DOI":"10.1007\/978-3-319-19992-4_34"},{"key":"1684_CR83","doi-asserted-by":"crossref","unstructured":"Sellam T, Das D, Parikh AP (2020) Bleurt: learning robust metrics for text generation. In: Proceedings of the 58th annual meeting of the association for computational linguistics, pp 7881\u20137892, held on-line","DOI":"10.18653\/v1\/2020.acl-main.704"},{"key":"1684_CR84","doi-asserted-by":"crossref","unstructured":"Sharma P, Ding N, Goodman S, Soricut R (2018) Conceptual captions: a cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Proceedings of the 56th annual meeting of the association for computational linguistics, Melbourne, Australia, pp 2556\u20132565","DOI":"10.18653\/v1\/P18-1238"},{"key":"1684_CR85","first-page":"3729","volume":"17","author":"H-C Shin","year":"2016","unstructured":"Shin H-C, Lu L, Kim L, Seff A, Yao J, Summers RM (2016a) Interleaved text\/image deep mining on a large-scale radiology database for automated image interpretation. JMLR 17:3729\u20133759","journal-title":"JMLR"},{"key":"1684_CR86","doi-asserted-by":"crossref","unstructured":"Shin H-C, Roberts K, Lu L, Demner-Fushman D, Yao J, Summers RM (2016b) Learning to read chest X-rays: Recurrent neural cascade model for automated image annotation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Las Vegas, NV, USA, pp 2497\u20132506","DOI":"10.1109\/CVPR.2016.274"},{"key":"1684_CR87","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556"},{"key":"1684_CR88","doi-asserted-by":"crossref","unstructured":"Singh S, Karimi S, Ho-Shon K, Hamey L (2021) Show, tell and summarise: learning to generate and summarise radiology findings from medical images. Neural Comput Appl pages 1\u201325","DOI":"10.1007\/s00521-021-05943-6"},{"key":"1684_CR89","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511596803","volume-title":"Fundamentals of medical imaging","author":"P Suetens","year":"2009","unstructured":"Suetens P (2009) Fundamentals of medical imaging. Cambridge University Press, Cambridge"},{"key":"1684_CR90","doi-asserted-by":"crossref","unstructured":"Sun S, Guzm\u00e1n F, Specia L (2020) Are we estimating or guesstimating translation quality? In: Proceedings of the 58th annual meeting of the association for computational linguistics, pp 6262\u20136267, held on-line","DOI":"10.18653\/v1\/2020.acl-main.558"},{"key":"1684_CR91","volume-title":"Reinforcement learning: an introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton RS, Barto AG (2018) Reinforcement learning: an introduction. MIT press, Cambridge"},{"key":"1684_CR92","doi-asserted-by":"publisher","first-page":"1285","DOI":"10.1126\/science.3287615","volume":"240","author":"JA Swets","year":"1988","unstructured":"Swets JA (1988) Measuring the accuracy of diagnostic systems. Science 240:1285\u20131293","journal-title":"Science"},{"key":"1684_CR93","doi-asserted-by":"crossref","unstructured":"Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2016) Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 2818\u20132826","DOI":"10.1109\/CVPR.2016.308"},{"key":"1684_CR94","doi-asserted-by":"crossref","unstructured":"Tsochantaridis I, Hofmann T, Joachims T, Altun Y (2004) Support vector machine learning for interdependent and structured output spaces. In: Proceedings of the international conference on machine learning, Banff, Alberta, Canada, pp 104\u2013114","DOI":"10.1145\/1015330.1015341"},{"key":"1684_CR95","doi-asserted-by":"publisher","first-page":"15","DOI":"10.1162\/0891201053630291","volume":"31","author":"K Van Deemter","year":"2005","unstructured":"Van Deemter K, Krahmer E, Theune M (2005) Real versus template-based natural language generation: a false opposition? Comput Linguist 31:15\u201324","journal-title":"Comput Linguist"},{"issue":"6","key":"1684_CR96","doi-asserted-by":"publisher","first-page":"3797","DOI":"10.1007\/s00330-021-07892-z","volume":"31","author":"KG van Leeuwen","year":"2021","unstructured":"van Leeuwen KG, Schalekamp S, Rutten MJ, van Ginneken B, de Rooij M (2021) Artificial intelligence in radiology: 100 commercially available products and their scientific evidence. Eur Radiol 31(6):3797\u20133804","journal-title":"Eur Radiol"},{"key":"1684_CR97","unstructured":"Varges S, Bieler H, Stede M, Faulstich LC, Irsig K, Atalla M (2012) SemScribe: natural language generation for medical reports. In: Proceedings of the eighth international conference on language resources and evaluation, Istanbul, Turkey, pp 2674\u20132681"},{"key":"1684_CR98","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017) Attention is all you need. In: NIPS, Red Hook, NY, USA. Curran Associates Inc, pp 6000\u20136010"},{"key":"1684_CR99","doi-asserted-by":"crossref","unstructured":"Vedantam R, Zitnick ZCL, Parikh D (2015) CIDEr: consensus-based image description evaluation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Boston, MA, USA, pp 4566\u20134575","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"1684_CR100","doi-asserted-by":"crossref","unstructured":"Vinyals O, Toshev A, Bengio S, Erhan D (2015) Show and tell: A neural image caption generator. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Boston, MA, USA, pp 3156\u20133164","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"1684_CR101","doi-asserted-by":"publisher","first-page":"652","DOI":"10.1109\/TPAMI.2016.2587640","volume":"39","author":"O Vinyals","year":"2017","unstructured":"Vinyals O, Toshev A, Bengio S, Erhan D (2017) Show and tell: lessons learned from the 2015 MSCOCO image captioning challenge. IEEE Trans Pattern Anal Mach Intell 39:652\u2013663","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"1684_CR102","doi-asserted-by":"crossref","unstructured":"Wang X, Peng Y, Lu L, Lu Z, Summers RM (2018) TieNet: text-image embedding network for common thorax disease classification and reporting in chest X-rays. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Quebec City, Canada, pp 9049\u20139058","DOI":"10.1109\/CVPR.2018.00943"},{"key":"1684_CR103","doi-asserted-by":"crossref","unstructured":"Wang Z, Zhou L, Wang L, Li X (2021) A self-boosting framework for automated radiographic report generation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 2433\u20132442, held on-line","DOI":"10.1109\/CVPR46437.2021.00246"},{"issue":"6","key":"1684_CR104","doi-asserted-by":"publisher","first-page":"e200057","DOI":"10.1148\/ryai.2020200057","volume":"2","author":"WF Wiggins","year":"2020","unstructured":"Wiggins WF, Caton MT, Magudia K, Glomski S-HA, George E, Rosenthal MH, Gaviola GC, Andriole KP (2020) Preparing radiologists to lead in the era of artificial intelligence: designing and implementing a focused data science pathway for senior radiology residents. Radiol Artif Intell 2(6):e200057","journal-title":"Radiol Artif Intell"},{"key":"1684_CR105","first-page":"229","volume":"8","author":"RJ Williams","year":"1992","unstructured":"Williams RJ (1992) Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach Learn 8:229\u2013256","journal-title":"Mach Learn"},{"key":"1684_CR106","doi-asserted-by":"crossref","unstructured":"Xenouleas S, Malakasiotis P, Apidianaki M, Androutsopoulos I (2019) Sumqe: a bert-based summary quality estimation model. In: Proceedings of the conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing, Hong Kong, China, pp 6005\u20136011","DOI":"10.18653\/v1\/D19-1618"},{"key":"1684_CR107","unstructured":"Xu K, Ba J, Kiros R, Cho K, Courville A, Salakhudinov R, Zemel R, Bengio Y (2015) Show, attend and tell: Neural image caption generation with visual attention. In: Proceedings of the international conference on machine learning, pp 2048\u20132057"},{"key":"1684_CR108","doi-asserted-by":"crossref","unstructured":"Xue Y, Xu T, Long LR, Xue Z, Antani S, Thoma GR, Huang X (2018) Multimodal recurrent model with attention for automated radiology report generation. In: Proceedings of the international conference on medical image computing and computer-assisted intervention, Granada, Spain, pp 457\u2013466","DOI":"10.1007\/978-3-030-00928-1_52"},{"key":"1684_CR109","doi-asserted-by":"crossref","unstructured":"Yin C, Qian B, Wei J, Li X, Zhang X, Li Y, Zheng Q (2019) Automatic generation of medical imaging diagnostic report with hierarchical recurrent neural network. In: Proceedings of the IEEE international conference on data mining, Beijing, China, pp 728\u2013737","DOI":"10.1109\/ICDM.2019.00083"},{"issue":"4","key":"1684_CR110","doi-asserted-by":"publisher","first-page":"e25759","DOI":"10.2196\/25759","volume":"23","author":"J Yin","year":"2021","unstructured":"Yin J, Ngiam KY, Teo HH (2021) Role of artificial intelligence applications in real-life clinical practice: Systematic review. J Med Internet Res 23(4):e25759","journal-title":"J Med Internet Res"},{"key":"1684_CR111","doi-asserted-by":"crossref","unstructured":"You Q, Jin H, Wang Z, Fang C, Luo J (2016) Image captioning with semantic attention. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Las Vegas, NV, USA, pp 4651\u20134659","DOI":"10.1109\/CVPR.2016.503"},{"key":"1684_CR112","doi-asserted-by":"crossref","unstructured":"Yuan J, Liao H, Luo R, Luo J (2019) Automatic radiology report generation based on multi-view image fusion and medical concept enrichment. In: Proceedings of the international conference on medical image computing and computer-assisted intervention, Shenzhen, China, pp 721\u2013729","DOI":"10.1007\/978-3-030-32226-7_80"},{"key":"1684_CR113","doi-asserted-by":"crossref","unstructured":"Zhang Y, Merck D, Tsai EB, Manning CD, Langlotz CP (2019) Optimizing the factual correctness of a summary: A study of summarizing radiology reports. arXiv:1911.02541","DOI":"10.18653\/v1\/2020.acl-main.458"},{"key":"1684_CR114","unstructured":"Zhang Y, Wang X, Guo Z, Li J (2018) ImageSem at ImageCLEF 2018 caption task: image retrieval and transfer learning. In: Proceedings of the CEUR workshop, CLEF2018 working notes, Avignon, France"},{"key":"1684_CR115","doi-asserted-by":"crossref","unstructured":"Zhang Z, Chen P, Sapkota M, Yang L (2017a) TandemNet: distilling knowledge from medical images using diagnostic reports as optional semantic references. In: Proceedings of the international conference on medical image computing and computer assisted intervention, Quebec City, Canada, pp 320\u2013328","DOI":"10.1007\/978-3-319-66179-7_37"},{"key":"1684_CR116","doi-asserted-by":"crossref","unstructured":"Zhang Z, Xie Y, Xing F, McGough M, Yang L (2017b) MDNet: a semantically and visually interpretable medical image diagnosis network. In: Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA, pp 6428\u20136436","DOI":"10.1109\/CVPR.2017.378"}],"container-title":["Knowledge and Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10115-022-01684-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10115-022-01684-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10115-022-01684-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,7,15]],"date-time":"2022-07-15T05:13:59Z","timestamp":1657862039000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10115-022-01684-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,16]]},"references-count":116,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2022,7]]}},"alternative-id":["1684"],"URL":"https:\/\/doi.org\/10.1007\/s10115-022-01684-7","relation":{},"ISSN":["0219-1377","0219-3116"],"issn-type":[{"value":"0219-1377","type":"print"},{"value":"0219-3116","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,16]]},"assertion":[{"value":"8 February 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 April 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 April 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 June 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}