{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T05:10:39Z","timestamp":1778649039416,"version":"3.51.4"},"reference-count":42,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T00:00:00Z","timestamp":1776988800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MAKE"],"abstract":"<jats:p>Structured radiology reporting can improve clinical decision support by standardizing clinical findings into hierarchical formats. However, thousands of questions in structured report templates about clinical findings are prohibitively time-consuming, which can limit clinical adoption. Furthermore, early medical VQA datasets primarily focused on free-text and independent question\u2013answer pairs while a recent dataset, Rad-ReStruct, introduced a hierarchical VQA, but the accompanying model still relies heavily on flattened embedding representations and single-path text\u2013image fusion mechanisms that inadequately handle complex hierarchical dependencies in responses. In this paper, we propose DPA-HiVQA (Dual-Path Cross-Attention for Hierarchical VQA), addressing these limitations through two key contributions: (1) multi-scale image embedding representing global semantic embeddings with patch-level spatial features from domain-specific BioViL encoder; (2) dual-path cross-attention mechanism enabling simultaneous holistic semantic understanding and fine-grained spatial reasoning. Evaluated on the Rad-ReStruct benchmark, the model substantially outperforms the established benchmark baseline with an overall F1-score and Level 3 F1-score improvement by 21.2% and 31.9%, respectively. The proposed model demonstrates that dual-path cross-attention architectures can effectively connect holistic semantic understanding and fine-grained spatial detail, paving the way for practical AI-assisted structured reporting systems that reduce radiologist burden while maintaining diagnostic accuracy.<\/jats:p>","DOI":"10.3390\/make8050113","type":"journal-article","created":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T13:32:22Z","timestamp":1777469542000},"page":"113","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["DPA-HiVQA: Enhancing Structured Radiology Reporting with Dual-Path Cross-Attention"],"prefix":"10.3390","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-9507-5055","authenticated-orcid":false,"given":"Ngoc Tuyen","family":"Do","sequence":"first","affiliation":[{"name":"School of Information and Communication Technology, Hanoi University of Science and Technology, Hanoi 100000, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6796-191X","authenticated-orcid":false,"given":"Minh Nguyen","family":"Quang","sequence":"additional","affiliation":[{"name":"School of Electrical and Electronic Engineering, Hanoi University of Science and Technology, Hanoi 100000, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8325-1662","authenticated-orcid":false,"given":"Hai Van","family":"Pham","sequence":"additional","affiliation":[{"name":"School of Information and Communication Technology, Hanoi University of Science and Technology, Hanoi 100000, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,24]]},"reference":[{"key":"ref_1","unstructured":"Reichenpfader, D., Knupp, J., Sander, A., and Denecke, K. (2024). RadEx: A framework for structured information extraction from radiology reports based on large language models. arXiv."},{"key":"ref_2","unstructured":"Jain, S., Agrawal, A., Saporta, A., Truong, S.Q., Duong, D.N., Bui, T., Chambon, P., Zhang, Y., Lungren, M.P., and Ng, A.Y. (2021). Radgraph: Extracting clinical entities and relations from radiology reports. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2837","DOI":"10.1007\/s00330-021-08327-5","article-title":"Structured reporting in radiology: A systematic review to explore its potential","volume":"32","author":"Nobel","year":"2022","journal-title":"Eur. Radiol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"80","DOI":"10.1186\/s13244-024-01660-5","article-title":"A novel reporting workflow for automated integration of artificial intelligence results into structured radiology reports","volume":"15","author":"Jorg","year":"2024","journal-title":"Insights Imaging"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"194","DOI":"10.4103\/jmss.JMSS_21_20","article-title":"Automatic generation of structured radiology reports for volumetric computed tomography images using question-Specific deep feature extraction and learning","volume":"11","author":"Loveymi","year":"2021","journal-title":"J. Med. Signals Sens."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"368","DOI":"10.1109\/RBME.2024.3408456","article-title":"Automated radiology report generation: A review of recent advances","volume":"18","author":"Sloan","year":"2024","journal-title":"IEEE Rev. Biomed. Eng."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Bazi, Y., Rahhal, M.M.A., Bashmal, L., and Zuair, M. (2023). Vision\u2013language model for visual question answering in medical imagery. Bioengineering, 10.","DOI":"10.3390\/bioengineering10030380"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"180251","DOI":"10.1038\/sdata.2018.251","article-title":"A dataset of clinically generated visual questions and answers about radiology images","volume":"5","author":"Lau","year":"2018","journal-title":"Sci. Data"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"He, X., Zhang, Y., Mou, L., Xing, E., and Xie, P. (2020). Pathvqa: 30000+ questions for medical visual question answering. arXiv.","DOI":"10.36227\/techrxiv.13127537"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Liu, B., Zhan, L.M., Xu, L., Ma, L., Yang, Y., and Wu, X.M. (2021, January 13\u201316). Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. Proceedings of the 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), Nice, France.","DOI":"10.1109\/ISBI48211.2021.9434010"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Pellegrini, C., Keicher, M., \u00d6zsoy, E., and Navab, N. (2023). Rad-restruct: A novel vqa benchmark and method for structured radiology reporting. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer.","DOI":"10.1007\/978-3-031-43904-9_40"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Arsalane, W., Chikontwe, P., Luna, M., Kang, M., and Park, S.H. (2024). Context-Guided Medical Visual Question Answering. Proceedings of the Meets Africa Workshop, Springer.","DOI":"10.1007\/978-3-031-79103-1_25"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., and Parikh, D. (2015, January 7\u201313). Vqa: Visual question answering. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.279"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Yang, Z., He, X., Gao, J., Deng, L., and Smola, A. (2016, January 27\u201330). Stacked attention networks for image question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.10"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018, January 18\u201323). Bottom-up and top-down attention for image captioning and visual question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Tan, H., and Bansal, M. (2019). Lxmert: Learning cross-modality encoder representations from transformers. arXiv.","DOI":"10.18653\/v1\/D19-1514"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., and Wei, F. (2020). Oscar: Object-semantics aligned pre-training for vision-language tasks. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"ref_18","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, PMLR."},{"key":"ref_19","unstructured":"Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q., Sung, Y.H., Li, Z., and Duerig, T. (2021). Scaling up visual and vision-language representation learning with noisy text supervision. Proceedings of the International Conference on Machine Learning, PMLR."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1177\/1971400919845365","article-title":"The value of structured radiology reports to categorize intracranial metastases following radiation therapy","volume":"32","author":"Benson","year":"2019","journal-title":"Neuroradiol. J."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wetterauer, C., Winkel, D., Federer-Gsponer, J., Halla, A., Subotic, S., Deckart, A., Seifert, H., Boll, D., and Ebbing, J. (2019). Structured reporting of prostate magnetic resonance imaging has the potential to improve interdisciplinary communication. PLoS ONE, 14.","DOI":"10.1371\/journal.pone.0212444"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"437","DOI":"10.1007\/s00261-019-02287-7","article-title":"Impact of structured report on the quality of preoperative CT staging of pancreatic ductal adenocarcinoma: Assessment of intra-and inter-reader variability","volume":"45","author":"Dimarco","year":"2020","journal-title":"Abdom. Radiol."},{"key":"ref_23","first-page":"2644","article-title":"The impact of different radiology report formats on patient information processing: A systematic review","volume":"35","author":"Ottenheijm","year":"2025","journal-title":"Eur. Radiol."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"2589","DOI":"10.1007\/s00330-024-11107-6","article-title":"Large language models for structured reporting in radiology: Past, present, and future","volume":"35","author":"Busch","year":"2025","journal-title":"Eur. Radiol."},{"key":"ref_25","first-page":"2018","article-title":"Automatic structuring of radiology reports with on-premise open-source large language models","volume":"35","author":"Laqua","year":"2025","journal-title":"Eur. Radiol."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"2008","DOI":"10.1109\/TIP.2018.2882225","article-title":"Bi-directional spatial-semantic attention networks for image-text matching","volume":"28","author":"Huang","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., and Lazebnik, S. (2015, January 7\u201313). Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.303"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, C., Mao, Z., Liu, A.A., Zhang, T., Wang, B., and Zhang, Y. (2019). Focus your attention: A bidirectional focal attention network for image-text matching. Proceedings of the 27th ACM International Conference on Multimedia, ACM.","DOI":"10.1145\/3343031.3350869"},{"key":"ref_30","unstructured":"Zhang, Y., Jiang, H., Miura, Y., Manning, C.D., and Langlotz, C.P. (2022). Contrastive learning of medical visual representations from paired images and text. Proceedings of the Machine Learning for Healthcare Conference, PMLR."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Huang, S.C., Shen, L., Lungren, M.P., and Yeung, S. (2021, January 10\u201317). Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00391"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, Z., Wu, Z., Agarwal, D., and Sun, J. (2022). Medclip: Contrastive learning from unpaired medical images and text. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2022.emnlp-main.256"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1399","DOI":"10.1038\/s41551-022-00936-9","article-title":"Expert-level detection of pathologies from unannotated chest X-ray images via self-supervised learning","volume":"6","author":"Tiu","year":"2022","journal-title":"Nat. Biomed. Eng."},{"key":"ref_34","first-page":"28541","article-title":"Llava-med: Training a large language-and-vision assistant for biomedicine in one day","volume":"36","author":"Li","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","unstructured":"Moor, M., Huang, Q., Wu, S., Yasunaga, M., Dalmia, Y., Leskovec, J., Zakka, C., Reis, E.P., and Rajpurkar, P. (2023). Med-flamingo: A multimodal medical few-shot learner. Proceedings of the Machine Learning for Health (ML4H), PMLR."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Boecking, B., Usuyama, N., Bannur, S., Castro, D.C., Schwaighofer, A., Hyland, S., Wetscherek, M., Naumann, T., Nori, A., and Alvarez-Valle, J. (2022). Making the most of text semantics to improve biomedical vision\u2013language processing. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-031-20059-5_1"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"e210258","DOI":"10.1148\/ryai.210258","article-title":"RadBERT: Adapting transformer-based language models to radiology","volume":"4","author":"Yan","year":"2022","journal-title":"Radiol. Artif. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_39","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational Linguistics."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1038\/s41597-019-0322-0","article-title":"MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports","volume":"6","author":"Johnson","year":"2019","journal-title":"Sci. Data"},{"key":"ref_41","first-page":"304","article-title":"Preparing a collection of radiology examinations for distribution and retrieval","volume":"23","author":"Kohli","year":"2015","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Zhang, J., Li, B., and Zhou, S. (2025). Hierarchical modeling for medical visual question answering with cross-attention fusion. Appl. Sci., 15.","DOI":"10.3390\/app15094712"}],"container-title":["Machine Learning and Knowledge Extraction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/5\/113\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T04:24:17Z","timestamp":1778646257000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/5\/113"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,24]]},"references-count":42,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["make8050113"],"URL":"https:\/\/doi.org\/10.3390\/make8050113","relation":{},"ISSN":["2504-4990"],"issn-type":[{"value":"2504-4990","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,24]]}}}