{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T15:35:14Z","timestamp":1778600114319,"version":"3.51.4"},"reference-count":49,"publisher":"Wiley","license":[{"start":{"date-parts":[[2024,2,13]],"date-time":"2024-02-13T00:00:00Z","timestamp":1707782400000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62002392"],"award-info":[{"award-number":["62002392"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62372478"],"award-info":[{"award-number":["62372478"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2019SK2022"],"award-info":[{"award-number":["2019SK2022"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2022JJ31019"],"award-info":[{"award-number":["2022JJ31019"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100019091","name":"Key Research and Development Program of Hunan Province of China","doi-asserted-by":"publisher","award":["62002392"],"award-info":[{"award-number":["62002392"]}],"id":[{"id":"10.13039\/501100019091","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100019091","name":"Key Research and Development Program of Hunan Province of China","doi-asserted-by":"publisher","award":["62372478"],"award-info":[{"award-number":["62372478"]}],"id":[{"id":"10.13039\/501100019091","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100019091","name":"Key Research and Development Program of Hunan Province of China","doi-asserted-by":"publisher","award":["2019SK2022"],"award-info":[{"award-number":["2019SK2022"]}],"id":[{"id":"10.13039\/501100019091","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100019091","name":"Key Research and Development Program of Hunan Province of China","doi-asserted-by":"publisher","award":["2022JJ31019"],"award-info":[{"award-number":["2022JJ31019"]}],"id":[{"id":"10.13039\/501100019091","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004735","name":"Natural Science Foundation of Hunan Province","doi-asserted-by":"publisher","award":["62002392"],"award-info":[{"award-number":["62002392"]}],"id":[{"id":"10.13039\/501100004735","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004735","name":"Natural Science Foundation of Hunan Province","doi-asserted-by":"publisher","award":["62372478"],"award-info":[{"award-number":["62372478"]}],"id":[{"id":"10.13039\/501100004735","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004735","name":"Natural Science Foundation of Hunan Province","doi-asserted-by":"publisher","award":["2019SK2022"],"award-info":[{"award-number":["2019SK2022"]}],"id":[{"id":"10.13039\/501100004735","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004735","name":"Natural Science Foundation of Hunan Province","doi-asserted-by":"publisher","award":["2022JJ31019"],"award-info":[{"award-number":["2022JJ31019"]}],"id":[{"id":"10.13039\/501100004735","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["International Journal of Intelligent Systems"],"published-print":{"date-parts":[[2024,2,13]]},"abstract":"<jats:p>Medical image description can be applied to clinical medical diagnosis, but the field still faces serious challenges. There is a serious problem of visual and textual data bias in medical datasets, which are the imbalanced distribution of health and disease data. This can greatly affect the learning performance of data-driven neural networks and finally lead to errors in the generated medical image descriptions. To address this problem, we propose a new medical image description network architecture named multimodal data-assisted knowledge fusion network (MDAKF), which introduces multimodal auxiliary signals to guide the Transformer network to generate more accurate medical reports. In detail, audio auxiliary signals provide clear abnormal visual regions to alleviate the visual data bias problem. However, the audio modality signals with similar pronunciation lack recognizability, which may lead to incorrect mapping of audio labels to medical image regions. Therefore, we further fuse the audio with text features as the auxiliary signal to improve the overall performance of the model. Through the experiments on two medical image description datasets, IU-X-ray and COV-CTR, it is found that the proposed model is superior to the previous models in terms of language generation evaluation indicators.<\/jats:p>","DOI":"10.1155\/2024\/6680546","type":"journal-article","created":{"date-parts":[[2024,2,13]],"date-time":"2024-02-13T21:21:39Z","timestamp":1707859299000},"page":"1-12","source":"Crossref","is-referenced-by-count":7,"title":["Medical Image Description Based on Multimodal Auxiliary Signals and Transformer"],"prefix":"10.1155","volume":"2024","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9855-8234","authenticated-orcid":true,"given":"Yun","family":"Tan","sequence":"first","affiliation":[{"name":"Central South University of Forestry and Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0695-2283","authenticated-orcid":true,"given":"Chunzhi","family":"Li","sequence":"additional","affiliation":[{"name":"Central South University of Forestry and Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7549-7731","authenticated-orcid":true,"given":"Jiaohua","family":"Qin","sequence":"additional","affiliation":[{"name":"Central South University of Forestry and Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Youyuan","family":"Xue","sequence":"additional","affiliation":[{"name":"Central South University of Forestry and Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuyu","family":"Xiang","sequence":"additional","affiliation":[{"name":"Central South University of Forestry and Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","reference":[{"key":"1","doi-asserted-by":"publisher","DOI":"10.1109\/tip.2018.2889922"},{"key":"2","doi-asserted-by":"publisher","DOI":"10.3390\/s21041270"},{"key":"3","doi-asserted-by":"publisher","DOI":"10.1109\/jas.2020.1003402"},{"key":"4","doi-asserted-by":"publisher","DOI":"10.1109\/access.2020.3013321"},{"key":"5","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.107823"},{"key":"6","doi-asserted-by":"publisher","DOI":"10.1016\/j.isprsjprs.2022.02.001"},{"key":"7","article-title":"On the automatic generation of medical imaging reports","author":"B. Jing","year":"2020"},{"key":"8","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-022-12291-7"},{"key":"9","doi-asserted-by":"publisher","DOI":"10.1007\/s13735-022-00228-7"},{"key":"10","article-title":"Transtrack: multiple object tracking with transformer","author":"J. Cao","year":"2020"},{"key":"11","article-title":"Mimic-cxr-jpg- chest radiographs with structured labels","author":"A. Johnson","year":"2023"},{"key":"12","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.107856"},{"key":"13","doi-asserted-by":"publisher","DOI":"10.1007\/s10278-021-00567-7"},{"issue":"1","key":"14","first-page":"13753","article-title":"Exploring and distilling posterior and prior knowledge for radiology report generation","volume":"3","author":"X. Wu","year":"2021","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"15","doi-asserted-by":"publisher","DOI":"10.1007\/s13042-022-01506-w"},{"key":"16","doi-asserted-by":"publisher","DOI":"10.1109\/tip.2020.2969330"},{"key":"17","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16328"},{"issue":"1","key":"18","first-page":"10819","article-title":"Metaformer is actually what you need for vision","volume":"3","author":"M. Luo","year":"2022","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"19","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2022.104570"},{"key":"20","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-13-9042-5_56"},{"key":"21","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16258"},{"issue":"1","key":"22","first-page":"10578","article-title":"Meshed-memory transformer for image captioning","volume":"3","author":"M. Stefanini","year":"2020","journal-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition"},{"key":"23","doi-asserted-by":"publisher","DOI":"10.1109\/access.2021.3124564"},{"issue":"5","key":"24","first-page":"1372","article-title":"Multi-level policy and reward-based deep reinforcement learning framework for image captioning","volume":"22","author":"N. X. Hanwang","year":"2019","journal-title":"IEEE Transactions on Multimedia"},{"key":"25","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-021-10632-6"},{"key":"26","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106730"},{"key":"27","article-title":"Addressing data bias problems for chest x-ray image report generation","author":"Y. Chen","year":"2019"},{"key":"28","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6989"},{"key":"29","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2019.102178"},{"key":"30","article-title":"Generating radiology reports via memory-driven transformer","author":"Y. Song","year":"2020"},{"key":"31","article-title":"Cross-modal memory networks for radiology report generation","author":"Y. Shen","year":"2022"},{"key":"32","doi-asserted-by":"publisher","DOI":"10.1016\/j.compbiomed.2022.105253"},{"key":"33","volume-title":"Neural Machine Translation with Universal Visual representation","author":"K. Chen","year":"2020"},{"key":"34","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2021.07.010"},{"issue":"4","key":"35","first-page":"8","article-title":"Multi-modal circulant fusion for video-to-language and backward","volume":"3","author":"A. Wu","year":"2018","journal-title":"International Joint Conference on Artificial Intelligence"},{"key":"36","first-page":"14200","article-title":"Attention bottlenecks for multimodal fusion","volume":"34","author":"S. Yang","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"1","key":"37","first-page":"4563","article-title":"Wav2clip: learning robust audio representations from clip","volume":"12","author":"P. Seetharaman","year":"2022","journal-title":"ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)"},{"key":"38","article-title":"Beit: bert pre-training of image transformers","author":"L. Dong","year":"2020"},{"key":"39","doi-asserted-by":"publisher","DOI":"10.1093\/jamia\/ocv080"},{"key":"40","article-title":"IU-xray dataset","author":"drive","year":"2023"},{"key":"41","doi-asserted-by":"publisher","DOI":"10.1007\/s11280-022-01013-6"},{"key":"42","article-title":"COV-CTR dataset","author":"github","year":"2020"},{"key":"43","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2017.10.018"},{"key":"44","article-title":"Attention is all you need","volume":"30","author":"N. Shazeer","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"45","article-title":"Hybrid retrieval-generation reinforced agent for medical image report generation","volume":"31","author":"X. Liang","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"46","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016666"},{"key":"47","first-page":"2048","article-title":"Show, attend and tell: neural image caption generation with visual attention","volume":"23","author":"K. Xu","year":"2015","journal-title":"International conference on machine learning"},{"key":"48","first-page":"375","article-title":"Knowing when to look: adaptive attention via a visual sentinel for image captioning","author":"J. Lu"},{"key":"49","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.121260"}],"container-title":["International Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2024\/6680546.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2024\/6680546.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2024\/6680546.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,2,13]],"date-time":"2024-02-13T21:21:52Z","timestamp":1707859312000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.hindawi.com\/journals\/ijis\/2024\/6680546\/"}},"subtitle":[],"editor":[{"given":"Costa","family":"Gianni","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2024,2,13]]},"references-count":49,"alternative-id":["6680546","6680546"],"URL":"https:\/\/doi.org\/10.1155\/2024\/6680546","relation":{},"ISSN":["1098-111X","0884-8173"],"issn-type":[{"value":"1098-111X","type":"electronic"},{"value":"0884-8173","type":"print"}],"subject":[],"published":{"date-parts":[[2024,2,13]]}}}