{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T22:58:09Z","timestamp":1778540289701,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":43,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Innovation and Transformation Fund of Peking University Third Hospital","award":["BYSYZHKC2021115"],"award-info":[{"award-number":["BYSYZHKC2021115"]}]},{"name":"Beijing Natural Science Foundation","award":["L192032"],"award-info":[{"award-number":["L192032"]}]},{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2019YFB1406500"],"award-info":[{"award-number":["2019YFB1406500"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548122","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:43:01Z","timestamp":1665416581000},"page":"3569-3577","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":31,"title":["Caption-Aware Medical VQA via Semantic Focusing and Progressive Cross-Modality Comprehension"],"prefix":"10.1145","author":[{"given":"Fuze","family":"Cong","sequence":"first","affiliation":[{"name":"Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shibiao","family":"Xu","sequence":"additional","affiliation":[{"name":"Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li","family":"Guo","sequence":"additional","affiliation":[{"name":"Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yinbing","family":"Tian","sequence":"additional","affiliation":[{"name":"Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"Asma Ben Abacha Vivek V Datla Sadid A Hasan Dina Demner-Fushman and Henning M\u00fcller. 2020. Overview of the vqa-med task at imageclef 2020: Visual question answering and generation in the medical domain. In CLEF (Working Notes). http:\/\/ceur-ws.org\/Vol-2696\/paper_106.pdf  Asma Ben Abacha Vivek V Datla Sadid A Hasan Dina Demner-Fushman and Henning M\u00fcller. 2020. Overview of the vqa-med task at imageclef 2020: Visual question answering and generation in the medical domain. In CLEF (Working Notes). http:\/\/ceur-ws.org\/Vol-2696\/paper_106.pdf"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007379606734"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_3_2_2_4_1","volume-title":"Anomaly Matters: An Anomaly-Oriented Model for Medical Visual Question Answering","author":"Cong Fuze","year":"2022","unstructured":"Fuze Cong , Shibiao Xu , Li Guo , and Yinbing Tian . 2022 . Anomaly Matters: An Anomaly-Oriented Model for Medical Visual Question Answering . IEEE Trans. Med. Imaging ( 2022). Fuze Cong, Shibiao Xu, Li Guo, and Yinbing Tian. 2022. Anomaly Matters: An Anomaly-Oriented Model for Medical Visual Question Answering. IEEE Trans. Med. Imaging (2022)."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_2_6_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). https:\/\/arxiv.org\/abs\/1810.04805 Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00048"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-87240-3_7"},{"key":"e_1_3_2_2_9_1","volume-title":"Article arXiv:2112.13906 (Dec.","author":"Eslami Sedigheh","year":"2021","unstructured":"Sedigheh Eslami , Gerard de Melo , and Christoph Meinel . 2021. Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain? arXiv e-prints , Article arXiv:2112.13906 (Dec. 2021 ). arxiv: 2112.13906 [cs.CV] Sedigheh Eslami, Gerard de Melo, and Christoph Meinel. 2021. Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain? arXiv e-prints, Article arXiv:2112.13906 (Dec. 2021). arxiv: 2112.13906 [cs.CV]"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240687"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1240"},{"key":"e_1_3_2_2_14_1","volume-title":"MMBERT: Multimodal BERT Pretraining for Improved Medical VQA. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI). IEEE","author":"Khare Yash","year":"2021","unstructured":"Yash Khare , Viraj Bagal , Minesh Mathew , Adithi Devi , U Deva Priyakumar , and CV Jawahar . 2021 . MMBERT: Multimodal BERT Pretraining for Improved Medical VQA. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI). IEEE , Nice, France, 1033--1036. https:\/\/doi.org\/10.1109\/ISBI48211. 2021.9434063 Yash Khare, Viraj Bagal, Minesh Mathew, Adithi Devi, U Deva Priyakumar, and CV Jawahar. 2021. MMBERT: Multimodal BERT Pretraining for Improved Medical VQA. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI). IEEE, Nice, France, 1033--1036. https:\/\/doi.org\/10.1109\/ISBI48211.2021.9434063"},{"key":"e_1_3_2_2_15_1","volume-title":"Advances in Neural Information Processing Systems","volume":"31","author":"Kim Jin-Hwa","year":"2018","unstructured":"Jin-Hwa Kim , Jaehyun Jun , and Byoung-Tak Zhang . 2018 . Bilinear attention networks . Advances in Neural Information Processing Systems , Vol. 31 (2018). Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. 2018. Bilinear attention networks. Advances in Neural Information Processing Systems, Vol. 31 (2018)."},{"key":"e_1_3_2_2_16_1","volume-title":"Proceedings 3rd International Conference on Learning Representations","author":"Kingma Diederik P","year":"2015","unstructured":"Diederik P Kingma and Jimmy Ba . 2015 . Adam: A method for stochastic optimization . In Proceedings 3rd International Conference on Learning Representations . San Diego, CA, USA. Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings 3rd International Conference on Learning Representations. San Diego, CA, USA."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1038\/sdata.2018.251"},{"key":"e_1_3_2_2_18_1","first-page":"4","article-title":"BioBERT: a pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee Jinhyuk","year":"2020","unstructured":"Jinhyuk Lee , Wonjin Yoon , Sungdong Kim , Donghyeon Kim , Sunkyu Kim , Chan Ho So , and Jaewoo Kang . 2020 . BioBERT: a pre-trained biomedical language representation model for biomedical text mining . Bioinformatics , Vol. 36 , 4 (Feb. 2020), 1234--1240. Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, Vol. 36, 4 (Feb. 2020), 1234--1240.","journal-title":"Bioinformatics"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"e_1_3_2_2_20_1","volume-title":"International Conference on Medical Image Computing and Computer-Assisted Intervention","author":"Liu Bo","unstructured":"Bo Liu , Li-Ming Zhan , and Xiao-Ming Wu. 2021b. Contrastive Pre-training and Representation Distillation for Medical Visual Question Answering Based on Radiology Images . In International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer , Strasbourg, France , 210--220. Bo Liu, Li-Ming Zhan, and Xiao-Ming Wu. 2021b. Contrastive Pre-training and Representation Distillation for Medical Visual Question Answering Based on Radiology Images. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, Strasbourg, France, 210--220."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISBI48211.2021.9434010"},{"key":"e_1_3_2_2_22_1","unstructured":"Shengyan Liu Haiyan Ding and Xiaobing Zhou. 2020. Shengyan at vqa-med 2020: An encoder-decoder model for medical domain visual question answering task. In CLEF (Working Notes). http:\/\/ceur-ws.org\/Vol-2696\/paper_73.pdf  Shengyan Liu Haiyan Ding and Xiaobing Zhou. 2020. Shengyan at vqa-med 2020: An encoder-decoder model for medical domain visual question answering task. In CLEF (Working Notes). http:\/\/ceur-ws.org\/Vol-2696\/paper_73.pdf"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00213"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-020-09832-7"},{"key":"e_1_3_2_2_25_1","first-page":"10","article-title":"The multimodal brain tumor image segmentation benchmark (BRATS)","volume":"34","author":"Menze Bjoern H","year":"2014","unstructured":"Bjoern H Menze , Andras Jakab , Stefan Bauer , Jayashree Kalpathy-Cramer , Keyvan Farahani , Justin Kirby , Yuliya Burren , Nicole Porz , Johannes Slotboom , Roland Wiest , 2014 . The multimodal brain tumor image segmentation benchmark (BRATS) . IEEE Trans. Med. Imaging , Vol. 34 , 10 (Dec. 2014), 1993--2024. Bjoern H Menze, Andras Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, et al. 2014. The multimodal brain tumor image segmentation benchmark (BRATS). IEEE Trans. Med. Imaging, Vol. 34, 10 (Dec. 2014), 1993--2024.","journal-title":"IEEE Trans. Med. Imaging"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-32251-9_57"},{"key":"e_1_3_2_2_27_1","first-page":"8026","article-title":"PyTorch: An Imperative Style, High-Performance Deep Learning Library","volume":"32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , 2019 . PyTorch: An Imperative Style, High-Performance Deep Learning Library . In Adv. Neural Inf. Process. Syst. , Vol. 32. 8026 -- 8037 . Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Adv. Neural Inf. Process. Syst., Vol. 32. 8026--8037.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_2_28_1","volume-title":"Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis","author":"Pelka Obioma","unstructured":"Obioma Pelka , Sven Koitka , Johannes R\u00fcckert , Felix Nensa , and Christoph M Friedrich . 2018. Radiology Objects in COntext (ROCO): a multimodal image dataset . In Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis . Springer , 180--189. Obioma Pelka, Sven Koitka, Johannes R\u00fcckert, Felix Nensa, and Christoph M Friedrich. 2018. Radiology Objects in COntext (ROCO): a multimodal image dataset. In Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis. Springer, 180--189."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_2_30_1","volume-title":"Visual Question Answering as a Multi-Task Problem. arXiv preprint arXiv:2007.01780","author":"Pollard Amelia Elizabeth","year":"2020","unstructured":"Amelia Elizabeth Pollard and Jonathan L Shapiro . 2020. Visual Question Answering as a Multi-Task Problem. arXiv preprint arXiv:2007.01780 ( 2020 ). https:\/\/arxiv.org\/abs\/2007.01780 Amelia Elizabeth Pollard and Jonathan L Shapiro. 2020. Visual Question Answering as a Multi-Task Problem. arXiv preprint arXiv:2007.01780 (2020). https:\/\/arxiv.org\/abs\/2007.01780"},{"key":"e_1_3_2_2_31_1","volume-title":"CGMVQA: a new classification and generative model for medical visual question answering","author":"Ren Fuji","year":"2020","unstructured":"Fuji Ren and Yangyang Zhou . 2020. CGMVQA: a new classification and generative model for medical visual question answering . IEEE Access , Vol . 8 ( Mar. 2020 ), 50626--50636. Fuji Ren and Yangyang Zhou. 2020. CGMVQA: a new classification and generative model for medical visual question answering. IEEE Access, Vol. 8 (Mar. 2020), 50626--50636."},{"key":"e_1_3_2_2_32_1","unstructured":"Mourad Sarrouti. 2020. Nlm at vqa-med 2020: Visual question answering and generation in the medical domain. In CLEF (Working Notes). http:\/\/ceur-ws.org\/Vol-2696\/paper_98.pdf  Mourad Sarrouti. 2020. Nlm at vqa-med 2020: Visual question answering and generation in the medical domain. In CLEF (Working Notes). http:\/\/ceur-ws.org\/Vol-2696\/paper_98.pdf"},{"key":"e_1_3_2_2_33_1","volume-title":"International Conference on Machine Learning. PMLR","author":"Shrikumar Avanti","year":"2017","unstructured":"Avanti Shrikumar , Peyton Greenside , and Anshul Kundaje . 2017 . Learning important features through propagating activation differences . In International Conference on Machine Learning. PMLR , Sydney, Australia, 3145--3153. Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017. Learning important features through propagating activation differences. In International Conference on Machine Learning. PMLR, Sydney, Australia, 3145--3153."},{"key":"e_1_3_2_2_34_1","unstructured":"Hideo Umada and Masaki Aono. 2020. kdevqa at vqa-med 2020: focusing on glu-based classification. In CLEF (Working Notes). http:\/\/ceur-ws.org\/Vol-2696\/paper_81.pdf  Hideo Umada and Masaki Aono. 2020. kdevqa at vqa-med 2020: focusing on glu-based classification. In CLEF (Working Notes). http:\/\/ceur-ws.org\/Vol-2696\/paper_81.pdf"},{"key":"e_1_3_2_2_35_1","volume-title":"Attention is all you need. Advances in neural information processing systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , \u0141ukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. Advances in neural information processing systems , Vol. 30 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, Vol. 30 (2017)."},{"key":"e_1_3_2_2_36_1","unstructured":"Harendra K Verma and S Sindhu Ramachandran. 2020. HARENDRAKV at VQA-Med 2020: Sequential VQA with Attention for Medical Visual Question Answering. (2020). http:\/\/ceur-ws.org\/Vol-2696\/paper_62.pdf  Harendra K Verma and S Sindhu Ramachandran. 2020. HARENDRAKV at VQA-Med 2020: Sequential VQA with Attention for Medical Visual Question Answering. (2020). http:\/\/ceur-ws.org\/Vol-2696\/paper_62.pdf"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMI.2020.2978284"},{"key":"e_1_3_2_2_38_1","volume-title":"Working Notes of CLEF","volume":"201","author":"Xiao Qian","year":"2021","unstructured":"Qian Xiao , Xiaobing Zhou , Y Xiao , and K Zhao . 2021 . Yunnan university at vqa-med 2021: Pretrained biobert for medical domain visual question answering . Working Notes of CLEF , Vol. 201 (2021). Qian Xiao, Xiaobing Zhou, Y Xiao, and K Zhao. 2021. Yunnan university at vqa-med 2021: Pretrained biobert for medical domain visual question answering. Working Notes of CLEF, Vol. 201 (2021)."},{"key":"e_1_3_2_2_39_1","volume-title":"International Conference on Machine Learning. PMLR","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu , Jimmy Ba , Ryan Kiros , Kyunghyun Cho , Aaron Courville , Ruslan Salakhudinov , Rich Zemel , and Yoshua Bengio . 2015 . Show, attend and tell: Neural image caption generation with visual attention . In International Conference on Machine Learning. PMLR , Lille, France , 2048--2057. Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In International Conference on Machine Learning. PMLR, Lille, France, 2048--2057."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.10"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.202"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413761"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.7005"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548122","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548122","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:19Z","timestamp":1750186819000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548122"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":43,"alternative-id":["10.1145\/3503161.3548122","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548122","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}