{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T18:39:23Z","timestamp":1781375963026,"version":"3.54.1"},"reference-count":21,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2022,2,13]],"date-time":"2022-02-13T00:00:00Z","timestamp":1644710400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the National Program for Excellence in Software","award":["2016-0-00017"],"award-info":[{"award-number":["2016-0-00017"]}]},{"name":"Information Technology Research Center","award":["IITP-2022-2020-0-01789"],"award-info":[{"award-number":["IITP-2022-2020-0-01789"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Transformer-based approaches have shown good results in image captioning tasks. However, current approaches have a limitation in generating text from global features of an entire image. Therefore, we propose novel methods for generating better image captioning as follows: (1) The Global-Local Visual Extractor (GLVE) to capture both global features and local features. (2) The Cross Encoder-Decoder Transformer (CEDT) for injecting multiple-level encoder features into the decoding process. GLVE extracts not only global visual features that can be obtained from an entire image, such as size of organ or bone structure, but also local visual features that can be generated from a local region, such as lesion area. Given an image, CEDT can create a detailed description of the overall features by injecting both low-level and high-level encoder outputs into the decoder. Each method contributes to performance improvement and generates a description such as organ size and bone structure. The proposed model was evaluated on the IU X-ray dataset and achieved better performance than the transformer-based baseline results, by 5.6% in BLEU score, by 0.56% in METEOR, and by 1.98% in ROUGE-L.<\/jats:p>","DOI":"10.3390\/s22041429","type":"journal-article","created":{"date-parts":[[2022,2,13]],"date-time":"2022-02-13T20:34:45Z","timestamp":1644784485000},"page":"1429","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":22,"title":["Cross Encoder-Decoder Transformer with Global-Local Visual Extractor for Medical Image Captioning"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5082-425X","authenticated-orcid":false,"given":"Hojun","family":"Lee","sequence":"first","affiliation":[{"name":"Department of Computer Science and Engineering, Dongguk University, Seoul 04620, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1080-9394","authenticated-orcid":false,"given":"Hyunjun","family":"Cho","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Dongguk University, Seoul 04620, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3549-347X","authenticated-orcid":false,"given":"Jieun","family":"Park","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Dongguk University, Seoul 04620, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4769-9889","authenticated-orcid":false,"given":"Jinyeong","family":"Chae","sequence":"additional","affiliation":[{"name":"Department of Artificial Intelligence, Dongguk University, Seoul 04620, Korea"},{"name":"Okestro Ltd., Seoul 07326, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2358-4021","authenticated-orcid":false,"given":"Jihie","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Artificial Intelligence, Dongguk University, Seoul 04620, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Chen, Z., Song, Y., Chang, T., and Wan, X. (2020, January 16\u201320). Generating radiology reports via memory-driven transformer. Proceedings of the Empirical Methods in Natural Language Processing (EMNLP), Online.","DOI":"10.18653\/v1\/2020.emnlp-main.112"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Luo, Y., Ji, J., Sun, X., Cao, L., Wu, Y., Huang, F., Lin, C.W., and Ji, R. (2021). Dual-level collaborative transformer for image captioning. arXiv.","DOI":"10.1609\/aaai.v35i3.16328"},{"key":"ref_3","unstructured":"Liu, W., Chen, S., Guo, L., Zhu, X., and Liu, J. (2021). Cptr: Full transformer network for image captioning. arXiv."},{"key":"ref_4","unstructured":"Karan, D., and Justin, J. (2020). Virtex: Learning visual representations from textual annotations. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"4467","DOI":"10.1109\/TCSVT.2019.2947482","article-title":"Multimodal transformer with multi-view visual representation for image captioning","volume":"30","author":"Yu","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Cornia, M., Stefanini, M., Baraldi, L., and Cucchiara, R. (2020, January 13\u201319). Meshed-Memory Transformer for Image Captioning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Jiayi, J., Yunpeng, L., Xiaoshuai, S., Fuhai, C., Gen, L., Yongjian, W., Yue, G., and Rongrong, J. (2021, January 2\u20139). Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network. Proceedings of the AAAI Conference on Artificial Intelligence, Online.","DOI":"10.1609\/aaai.v35i2.16258"},{"key":"ref_8","unstructured":"Li, C.Y., Liang, X., Hu, Z., and Xing, E.P. (February, January 27). Knowledge-driven encode, retrieve, paraphrase for medical image report generation. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_9","unstructured":"Li, Y., Liang, X., Hu, Z., and Xing, E.P. (2018, January 3\u20138). Hybrid retrieval-generation reinforced agent for medical image report generation. Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada."},{"key":"ref_10","unstructured":"Liu, G., Hsu, T.-M.H., McDermott, M., Boag, W., Weng, W.-H., Szolovits, P., and Ghassemi, M. (2019). Clinically accurate chest X-ray report generation. arXiv."},{"key":"ref_11","unstructured":"Baoyu, J., Pengtao, X., and Eric, X. (2018, January 15\u201320). On the Automatic Generation of Medical Imaging Reports. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia."},{"key":"ref_12","unstructured":"Baoyu, J., Zeya, W., and Eric, X. (August, January 8). Show, describe and conclude: On exploiting the structure information of chest x-ray reports. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy."},{"key":"ref_13","unstructured":"Guan, Q., Huang, Y., Zhong, Z., Zheng, Z., Zheng, L., and Yang, Y. (2018). Diagnose like a radiologist: Attention guided convolutional neural network for thorax disease classification. arXiv."},{"key":"ref_14","unstructured":"Ashish, V., Noam, S., Niki, P., Jakob, U., Llion, J., Aidan, N.G., Lukasz, K., and Illia, P. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"304","DOI":"10.1093\/jamia\/ocv080","article-title":"Preparing a collection of radiology examinations for distribution and retrieval","volume":"23","author":"Kohli","year":"2016","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., and Summers, R.M. (2017, January 21\u201326). Chestx-ray8: Hospital scale chest X-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.369"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). ImageNet: A Large-Scale Hierarchical Image Database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_18","unstructured":"Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. (2014, January 8\u201311). How Transferable Are Features in Deep Neural Networks?. Proceedings of the Advances in Neural Information Processing Systems, NeurIPS Proceedings, Montreal, QC, Canada."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W. (2002, January 7\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_20","unstructured":"Denkowski, M., and Lavie, A. (2011). Meteor 1.3: Automatic Metric for Reliable Optimization and Evaluation of Machine Translation Systems. Sixth Workshop on Statistical Machine Translation, Association for Computational Linguistics."},{"key":"ref_21","unstructured":"Lin, C.Y. (2004). Rouge: A Package for Automatic Evaluation of Summaries; Text Summarization Branches Out, Association for Computational Linguistics."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/4\/1429\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:18:35Z","timestamp":1760134715000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/4\/1429"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,13]]},"references-count":21,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["s22041429"],"URL":"https:\/\/doi.org\/10.3390\/s22041429","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,13]]}}}