{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T10:18:15Z","timestamp":1785406695896,"version":"3.56.0"},"reference-count":42,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,4,9]],"date-time":"2026-04-09T00:00:00Z","timestamp":1775692800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62262005"],"award-info":[{"award-number":["62262005"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"award":["62262005"],"award-info":[{"award-number":["62262005"]}],"id":[{"id":"https:\/\/ror.org\/01h0zpd94","id-type":"ROR","asserted-by":"publisher"}]},{"name":"High-level Innovative Talents in Guizhou Province","award":["GCC[2023]033"],"award-info":[{"award-number":["GCC[2023]033"]}]},{"name":"Natural Science Research Project of the Department of Education of Guizhou Province","award":["QJJ[2024]009"],"award-info":[{"award-number":["QJJ[2024]009"]}]},{"name":"Natural Science Research Project of the Department of Education of Guizhou Province","award":["QJJ[2023]011"],"award-info":[{"award-number":["QJJ[2023]011"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Difference visual question answering (Diff-VQA) aims to answer questions by identifying and reasoning about differences between medical images. Existing methods often rely on simple feature subtraction or fusion to model image differences, while overlooking the asymmetric descriptive requirements of changed and unchanged cases and providing limited task-specific guidance to pretrained language decoders. To address these limitations, we propose D2MNet (Difference-aware Decoupling and Multi-prompt Network), a framework for medical Diff-VQA that combines change-aware reasoning with prompt-guided answer generation. Specifically, a Change Analysis Module (CAM) predicts whether a change is present and produces a binary change-aware prompt; a Difference-Aware Module (DAM) uses dual attention to capture fine-grained difference features; and a multi-prompt learning mechanism (MLM) injects question-aware, change-aware, and learnable prompts into the language decoder to improve contextual alignment and response generation. Experiments on the MIMIC-DiffVQA benchmark show that D2MNet achieves a CIDEr score of 2.907 \u00b1 0.040, outperforming the strongest baseline, ReAl (2.409), under the same evaluation setting. These results demonstrate the effectiveness of the proposed design on benchmark medical Diff-VQA and suggest its potential for assisting difference-aware medical answer generation.<\/jats:p>","DOI":"10.3390\/jimaging12040162","type":"journal-article","created":{"date-parts":[[2026,4,9]],"date-time":"2026-04-09T07:48:07Z","timestamp":1775720887000},"page":"162","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["D2MNet: Difference-Aware Decoupling and Multi-Prompt Learning for Medical Difference Visual Question Answering"],"prefix":"10.3390","volume":"12","author":[{"given":"Lingge","family":"Lai","sequence":"first","affiliation":[{"name":"School of Big Data and Computer Science, Guizhou Normal University, Guiyang 550025, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weihua","family":"Ou","sequence":"additional","affiliation":[{"name":"School of Big Data and Computer Science, Guizhou Normal University, Guiyang 550025, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1413-0693","authenticated-orcid":false,"given":"Jianping","family":"Gou","sequence":"additional","affiliation":[{"name":"School of Computer and Information Sciences, Southwest University, Chongqing 400715, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9816-7471","authenticated-orcid":false,"given":"Zhonghua","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Zhejiang Ocean University, Zhoushan 316022, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Bazi, Y., Rahhal, M.M.A., Bashmal, L., and Zuair, M. (2023). Vision\u2013Language Model for Visual Question Answering in Medical Imagery. Bioengineering, 10.","DOI":"10.3390\/bioengineering10030380"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., and Parikh, D. (2015, January 7\u201313). VQA: Visual Question Answering. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.279"},{"key":"ref_3","unstructured":"Narayanan, A., Musthyala, R., Sankar, R., Nistala, A.P., Singh, P., and Cirrone, J. (2024). Free Form Medical Visual Question Answering in Radiology. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"102611","DOI":"10.1016\/j.artmed.2023.102611","article-title":"Medical Visual Question Answering: A Survey","volume":"143","author":"Lin","year":"2023","journal-title":"Artif. Intell. Med."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Do, T., Nguyen, B.X., Tjiputra, E., Tran, M., Tran, Q.D., and Nguyen, A. (October, January 27). Multiple Meta-Model Quantifying for Medical Visual Question Answering. Proceedings of the Medical Image Computing and Computer Assisted Intervention\u2014MICCAI 2021, Strasbourg, France.","DOI":"10.1007\/978-3-030-87240-3_7"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1146\/annurev-bioeng-071516-044442","article-title":"Deep Learning in Medical Image Analysis","volume":"19","author":"Shen","year":"2017","journal-title":"Annu. Rev. Biomed. Eng."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"20190043","DOI":"10.1259\/bjr.20190043","article-title":"Fatigue in radiology: A fertile area for future research","volume":"92","author":"Stinton","year":"2019","journal-title":"Br. J. Radiol."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Shen, B., Hou, W., Jiang, Z., Li, H., Singer, A.J., Hoshmand-Kochi, M., Abbasi, A., Glass, S., Thode, H.C., and Levsky, J. (2023). Longitudinal Chest X-ray Scores and their Relations with Clinical Variables and Outcomes in COVID-19 Patients. Diagnostics, 13.","DOI":"10.3390\/diagnostics13061107"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"101585","DOI":"10.1016\/j.imu.2024.101585","article-title":"Longitudinal Data and a Semantic Similarity Reward for Chest X-Ray Report Generation","volume":"50","author":"Nicolson","year":"2024","journal-title":"Inform. Med. Unlocked"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhu, Q., Mathai, T.S., Mukherjee, P., Peng, Y., Summers, R.M., and Lu, Z. (2023). Utilizing Longitudinal Chest X-Rays and Reports to Pre-Fill Radiology Reports. arXiv.","DOI":"10.1007\/978-3-031-43904-9_19"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Hu, X., Gu, L., An, Q., Zhang, M., Liu, L., Kobayashi, K., Harada, T., Summers, R.M., and Zhu, Y. (2023). Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering. Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD \u201923), Association for Computing Machinery. KDD \u201923.","DOI":"10.1145\/3580305.3599819"},{"key":"ref_12","unstructured":"Cho, Y., Kim, T., Shin, H., Cho, S., and Shin, D. (2024). Pretraining Vision\u2013Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Lu, Z., Xie, Y., Zeng, Q., Lu, M., Wu, Q., and Xia, Y. Spot the Difference: Difference Visual Question Answering with Residual Alignment. Proceedings of the Medical Image Computing and Computer Assisted Intervention\u2014MICCAI 2024, Springer. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-031-72086-4_61"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1958","DOI":"10.1097\/JTO.0b013e3181f2636e","article-title":"Follow-Up of Small (4 mm or Less) Incidentally Detected Nodules by Computed Tomography in Oncology Patients: A Retrospective Review","volume":"5","author":"Munden","year":"2010","journal-title":"J. Thorac. Oncol."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Huang, Z., and You, H. (2023). MFSFNet: Multi-Scale Feature Subtraction Fusion Network for Remote Sensing Image Change Detection. Remote. Sens., 15.","DOI":"10.3390\/rs15153740"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Woodland, M., and Castelo, A. (2024). Feature Extraction for Generative Medical Imaging Evaluation: New Evidence Against an Evolving Trend. Medical Image Understanding and Analysis, Springer.","DOI":"10.1007\/978-3-031-72390-2_9"},{"key":"ref_17","unstructured":"Chen, Z., Varma, M., Xu, J., Paschali, M., Veen, D.V., Johnston, A., Youssef, A., Blankemeier, L., Bluethgen, C., and Altmayer, S. (2024). A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"4374","DOI":"10.1007\/s00330-024-11339-6","article-title":"CXR-LLaVA: A multimodal large language model for interpreting chest X-ray images","volume":"35","author":"Lee","year":"2025","journal-title":"Eur. Radiol."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"7866","DOI":"10.1038\/s41467-025-62385-7","article-title":"Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data","volume":"16","author":"Wu","year":"2025","journal-title":"Nat. Commun."},{"key":"ref_20","unstructured":"Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G. (2021). Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. arXiv."},{"key":"ref_21","first-page":"1877","article-title":"Language Models are Few-Shot Learners","volume":"Volume 33","author":"Brown","year":"2020","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhou, K., Yang, J., Loy, C.C., and Liu, Z. (2022, January 19\u201320). Conditional Prompt Learning for Vision\u2013Language Models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01631"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Du, Y., Niu, T., and Zhao, R. (2024). Mixture of Prompt Learning for Vision\u2013Language Models. arXiv.","DOI":"10.3389\/frai.2025.1580973"},{"key":"ref_24","unstructured":"Kim, H., Jin, S., Sung, C., Kim, J., and Ok, J. (2024). Active Prompt Learning with Vision\u2013Language Model Priors. arXiv."},{"key":"ref_25","unstructured":"Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T.L., Cao, Y., and Narasimhan, K. (2023). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, B., Zeng, Z., Wang, Y., Liu, S., and Zhou, Y. (2023). Medical Visual Question Answering via Conditional Reasoning and Knowledge-Guided Learning. arXiv.","DOI":"10.1109\/TMI.2022.3232411"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Koleilat, T., Asgariandehkordi, H., Rivaz, H., and Xiao, Y. (2024). BiomedCoOp: Learning to Prompt for Biomedical Vision\u2013Language Models. arXiv.","DOI":"10.1109\/CVPR52734.2025.01376"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Chen, Q., Bian, M., and Xu, H. MMQL: Multi-Question Learning for Medical Visual Question Answering. Proceedings of the Medical Image Computing and Computer Assisted Intervention\u2014MICCAI 2024, Springer. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-031-72086-4_45"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Van Sonsbeek, T., Derakhshani, M.M., Najdenkoska, I., Snoek, C.G., and Worring, M. (2023). Open-Ended Medical Visual Question Answering Through Prefix Tuning of Language Models. arXiv.","DOI":"10.1007\/978-3-031-43904-9_70"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhu, H., Togo, R., Ogawa, T., and Haseyama, M. (2024). Prompt-based Personalized Federated Learning for Medical Visual Question Answering. arXiv.","DOI":"10.1109\/ICASSP48485.2024.10445933"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Guo, D., and Terzopoulos, D. (2024). Prompting Medical Large Vision\u2013Language Models to Diagnose Pathologies by Visual Question Answering. arXiv.","DOI":"10.59275\/j.melba.2025-1a8b"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Yan, Q., He, X., Yue, X., and Wang, X.E. (2024). Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA. arXiv.","DOI":"10.18653\/v1\/2025.findings-acl.981"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, J., Seng, K.P., Shen, Y., Ang, L.M., and Huang, D. (2024). Image to Label to Answer: An Efficient Framework for Enhanced Clinical Applications in Medical Visual Question Answering. Electronics, 13.","DOI":"10.3390\/electronics13122273"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1038\/s41597-019-0322-0","article-title":"MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports","volume":"6","author":"Johnson","year":"2019","journal-title":"Sci. Data"},{"key":"ref_35","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., and Antiga, L. (2019, January 8\u201314). PyTorch: An Imperative Style, High-Performance Deep Learning Library. Proceedings of the Advances in Neural Information Processing Systems, NeurIPS \u201919, Vancouver, BC, Canada."},{"key":"ref_36","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021). An Image is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002). BLEU: A Method for Automatic Evaluation of Machine Translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, PA, USA, Association for Computational Linguistics. ACL \u201902.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_38","unstructured":"Goldstein, J., Lavie, A., Lin, C.Y., and Voss, C.R. (2005). METEOR: An Automatic Metric for Machine Translation Evaluation with Improved Correlation with Human Judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, ACL Workshop."},{"key":"ref_39","unstructured":"Lin, C.Y. (2004). ROUGE: A Package for Automatic Evaluation of Summaries. Proceedings of the ACL-04 Workshop on Text Summarization Branches Out, Barcelona, Spain, Association for Computational Linguistics. W04-1013."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Zitnick, C.L., and Parikh, D. (2015, January 7\u201312). CIDEr: Consensus-based Image Description Evaluation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Qiu, Y., Yamamoto, S., Nakashima, K., Suzuki, R., Iwata, K., Kataoka, H., and Satoh, Y. (2021, January 10\u201317). Describing and Localizing Multiple Changes with Transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00198"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"3108","DOI":"10.1609\/aaai.v36i3.20218","article-title":"Image Difference Captioning with Pre-Training and Contrastive Learning","volume":"Volume 36","author":"Yao","year":"2022","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/4\/162\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,9]],"date-time":"2026-04-09T08:56:35Z","timestamp":1775724995000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/4\/162"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,9]]},"references-count":42,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["jimaging12040162"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12040162","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,9]]}}}