{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,31]],"date-time":"2025-12-31T00:31:32Z","timestamp":1767141092598,"version":"build-2238731810"},"reference-count":63,"publisher":"Springer Science and Business Media LLC","issue":"26","license":[{"start":{"date-parts":[[2024,5,20]],"date-time":"2024-05-20T00:00:00Z","timestamp":1716163200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,5,20]],"date-time":"2024-05-20T00:00:00Z","timestamp":1716163200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003500","name":"Universit\u00e0 degli Studi di Padova","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100003500","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2024,9]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Object detectors are used for searching all objects belonging to a pre-defined set of categories contained in a given picture. However, users are often not interested in finding all objects, but only those that pertain to a small set of categories or concepts. Nowadays, the standard approach to solve this task involves initially employing an object detector to identify all objects within the image, followed by refining the outcomes to retain only the ones of interest. Nevertheless, the object detector does not take advantage of the user\u2019s prior intent that, when used, can potentially improve the detection performance of the model. This work presents a method to condition an existing object detector with the user\u2019s intent, encoded as one or more concepts from the WordNet graph, to find just those objects of interest. The proposed approach takes advantage of existing datasets for object detection without the need for new annotations, and it allows to adapt the already existing object detector models with minor changes. The evaluation, performed on the COCO and the Visual Genome datasets considering several object detector architectures, shows that conditioning the search on concepts is actually beneficial. The code and the pre-trained model weights are released at:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/drigoni\/Concept-Conditioned-Object-Detector\" ext-link-type=\"uri\">https:\/\/github.com\/drigoni\/Concept-Conditioned-Object-Detector<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s00521-024-09914-5","type":"journal-article","created":{"date-parts":[[2024,5,20]],"date-time":"2024-05-20T15:01:43Z","timestamp":1716217303000},"page":"16001-16021","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Object search by a concept-conditioned object detector"],"prefix":"10.1007","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2092-3577","authenticated-orcid":false,"given":"Davide","family":"Rigoni","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Luciano","family":"Serafini","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alessandro","family":"Sperduti","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,5,20]]},"reference":[{"key":"9914_CR1","doi-asserted-by":"crossref","unstructured":"Antol S, Agrawal A, Lu J et\u00a0al (2015) VQA: Visual question answering. In: ICCV, pp 2425\u20132433","DOI":"10.1109\/ICCV.2015.279"},{"key":"9914_CR2","doi-asserted-by":"crossref","unstructured":"Bevilacqua M, Navigli R (2020) Breaking through the 80% glass ceiling: raising the state of the art in word sense disambiguation by incorporating knowledge graph information. In: Proceedings of the 58th annual meeting of the association for computational linguistics, pp 2854\u20132864","DOI":"10.18653\/v1\/2020.acl-main.255"},{"key":"9914_CR3","doi-asserted-by":"crossref","unstructured":"Chen K, Gao J, Nevatia R (2018) Knowledge aided consistency for weakly supervised phrase grounding. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4042\u20134050","DOI":"10.1109\/CVPR.2018.00425"},{"key":"9914_CR4","doi-asserted-by":"crossref","unstructured":"Cho J, Yoon Y, Kwak S (2022) Collaborative transformers for grounded situation recognition. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR), pp 19659\u201319668","DOI":"10.1109\/CVPR52688.2022.01904"},{"key":"9914_CR5","doi-asserted-by":"crossref","unstructured":"Dai X, Chen Y, Xiao B et\u00a0al (2021) Dynamic head: unifying object detection heads with attentions. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 7373\u20137382","DOI":"10.1109\/CVPR46437.2021.00729"},{"key":"9914_CR6","doi-asserted-by":"publisher","unstructured":"Deng J, Dong W, Socher R et\u00a0al (2009) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE computer society conference on computer vision and pattern recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA. IEEE computer society, pp 248\u2013255. https:\/\/doi.org\/10.1109\/CVPR.2009.5206848","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"9914_CR7","doi-asserted-by":"crossref","unstructured":"Dost S, Serafini L, Rospocher M et\u00a0al (2020a) Jointly linking visual and textual entity mentions with background knowledge. In: International conference on applications of natural language to information systems. Springer, Berlin, pp 264\u2013276","DOI":"10.1007\/978-3-030-51310-8_24"},{"key":"9914_CR8","doi-asserted-by":"crossref","unstructured":"Dost S, Serafini L, Rospocher M et\u00a0al (2020b) On visual-textual-knowledge entity linking. In: ICSC, IEEE, pp 190\u2013193","DOI":"10.1109\/ICSC.2020.00039"},{"key":"9914_CR9","doi-asserted-by":"crossref","unstructured":"Dost S, Serafini L, Rospocher M et\u00a0al (2020c) Vtkel: a resource for visual-textual-knowledge entity linking. In: ACM, pp 2021\u20132028","DOI":"10.1145\/3341105.3373958"},{"key":"9914_CR10","unstructured":"Fornoni M, Yan C, Luo L et\u00a0al (2021) Bridging the gap between object detection and user intent via query-modulation. arXiv preprint arXiv:2106.10258"},{"key":"9914_CR11","doi-asserted-by":"crossref","unstructured":"Frazzetto P, Pasa L, Navarin N et\u00a0al (2023) Topology preserving maps as aggregations for graph convolutional neural networks. In: Proceedings of the 38th ACM\/SIGAPP symposium on applied computing, pp 536\u2013543","DOI":"10.1145\/3555776.3577751"},{"key":"9914_CR12","unstructured":"Frome A, Corrado GS, Shlens J et\u00a0al (2013) Devise: a deep visual-semantic embedding model. In: Burges CJC, Bottou L, Ghahramani Z et\u00a0al (eds) NeurIPS, pp 2121\u20132129"},{"key":"9914_CR13","unstructured":"Gu X, Lin TY, Kuo W et\u00a0al (2021) Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv:2104.13921"},{"key":"9914_CR14","doi-asserted-by":"crossref","unstructured":"Gupta T, Vahdat A, Chechik G et\u00a0al (2020) Contrastive learning for weakly supervised phrase grounding. In: Computer vision\u2014ECCV 2020: 16th European conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part III 16. Springer, Berlin, pp 752\u2013768","DOI":"10.1007\/978-3-030-58580-8_44"},{"key":"9914_CR15","doi-asserted-by":"publisher","unstructured":"He K, Zhang X, Ren S et\u00a0al (2016) Deep residual learning for image recognition. In: 2016 IEEE conference on computer vision and pattern recognition. CVPR 2016, Las Vegas, NV, USA, June 27\u201330, 2016. IEEE computer society, pp 770\u2013778. https:\/\/doi.org\/10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"9914_CR16","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1016\/j.artint.2012.06.001","volume":"194","author":"J Hoffart","year":"2013","unstructured":"Hoffart J, Suchanek FM, Berberich K et al (2013) Yago2: a spatially and temporally enhanced knowledge base from Wikipedia. Artif. Intell. 194:28\u201361","journal-title":"Artif. Intell."},{"key":"9914_CR17","doi-asserted-by":"publisher","unstructured":"Kamath A, Singh M, LeCun Y et\u00a0al (2021) MDETR\u2014modulated detection for end-to-end multi-modal understanding. In: 2021 IEEE\/CVF international conference on computer vision, ICCV 2021, Montreal, QC, Canada, October 10\u201317, 2021. IEEE, pp 1760\u20131770. https:\/\/doi.org\/10.1109\/ICCV48922.2021.00180","DOI":"10.1109\/ICCV48922.2021.00180"},{"key":"9914_CR18","doi-asserted-by":"crossref","unstructured":"Kim D, Angelova A, Kuo W (2023) Region-aware pretraining for open-vocabulary object detection with vision transformers. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 11144\u201311154","DOI":"10.1109\/CVPR52729.2023.01072"},{"key":"9914_CR19","unstructured":"Kipf TN, Welling M (2016) Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907"},{"key":"9914_CR20","unstructured":"Kiros R, Salakhutdinov R, Zemel RS (2014) Unifying visual-semantic embeddings with multimodal neural language models. arXiv preprint arXiv:1411.2539"},{"key":"9914_CR21","unstructured":"Klein B, Lev G, Sadeh G et\u00a0al (2014) Fisher vectors derived from hybrid Gaussian\u2013Laplacian mixture models for image annotation. arXiv preprint arXiv:1411.7399"},{"key":"9914_CR22","doi-asserted-by":"crossref","unstructured":"Kokane CD, Babar SD, Mahalle PN et\u00a0al (2023) Word sense disambiguation: adaptive word embedding with adaptive-lexical resource. In: International conference on data analytics and insights. Springer, Berlin, pp 421\u2013429","DOI":"10.1007\/978-981-99-3878-0_36"},{"issue":"1","key":"9914_CR23","doi-asserted-by":"publisher","first-page":"32","DOI":"10.1007\/s11263-016-0981-7","volume":"123","author":"R Krishna","year":"2017","unstructured":"Krishna R, Zhu Y, Groth O et al (2017) Visual genome: connecting language and vision using crowdsourced dense image annotations. Int J Comput Vis 123(1):32\u201373. https:\/\/doi.org\/10.1007\/s11263-016-0981-7","journal-title":"Int J Comput Vis"},{"key":"9914_CR24","doi-asserted-by":"publisher","unstructured":"Kumar S, Jat S, Saxena K et\u00a0al (2019) Zero-shot word sense disambiguation using sense definition embeddings. In: Korhonen A, Traum DR, M\u00e0rquez L (eds) Proceedings of the 57th conference of the association for computational linguistics, ACL 2019, Florence, Italy, July 28\u2013August 2, 2019, Volume 1: long papers. Association for computational linguistics, pp 5670\u20135681. https:\/\/doi.org\/10.18653\/V1\/P19-1568","DOI":"10.18653\/V1\/P19-1568"},{"key":"9914_CR25","doi-asserted-by":"crossref","unstructured":"Lerner P, Ferret O, Guinaudeau C (2023) Multimodal inverse cloze task for knowledge-based visual question answering. In: European conference on information retrieval. Springer, Berlin, pp 569\u2013587","DOI":"10.1007\/978-3-031-28244-7_36"},{"key":"9914_CR26","doi-asserted-by":"crossref","unstructured":"Lin TY, Maire M, Belongie S et\u00a0al (2014) Microsoft coco: common objects in context. In: European conference on computer vision. Springer, Berlin, pp 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"9914_CR27","doi-asserted-by":"crossref","unstructured":"Lin TY, Doll\u00e1r P, Girshick R et\u00a0al (2017a) Feature pyramid networks for object detection. In: 2017 IEEE conference on computer vision and pattern recognition (CVPR), IEEE, pp 936\u2013944","DOI":"10.1109\/CVPR.2017.106"},{"key":"9914_CR28","doi-asserted-by":"crossref","unstructured":"Lin TY, Goyal P, Girshick R et\u00a0al (2017b) Focal loss for dense object detection. In: Proceedings of the IEEE international conference on computer vision, pp 2980\u20132988","DOI":"10.1109\/ICCV.2017.324"},{"key":"9914_CR29","doi-asserted-by":"crossref","unstructured":"Liu W, Anguelov D, Erhan D et\u00a0al (2016) SSD: single shot multibox detector. In: European conference on computer vision. Springer, Berlin, pp 21\u201337","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"9914_CR30","doi-asserted-by":"crossref","unstructured":"Liu Z, Lin Y, Cao Y et\u00a0al (2021) Swin transformer: hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 10012\u201310022","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"9914_CR31","doi-asserted-by":"crossref","unstructured":"Luo Z, Zhao P, Xu C et\u00a0al (2023) Lexlip: lexicon-bottlenecked language-image pre-training for large-scale image-text sparse retrieval. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 11206\u201311217","DOI":"10.1109\/ICCV51070.2023.01029"},{"key":"9914_CR32","unstructured":"Mahdisoltani F, Biega J, Suchanek F (2014) YAGO3: a knowledge base from multilingual wikipedias. In: 7th biennial conference on innovative data systems research, CIDR conference"},{"key":"9914_CR33","unstructured":"Mao J, Xu W, Yang Y et\u00a0al (2015) Deep captioning with multimodal recurrent neural networks (m-RNN). In: Bengio Y, LeCun Y (eds) ICLR"},{"key":"9914_CR34","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.101988","volume":"101","author":"R Mao","year":"2024","unstructured":"Mao R, He K, Zhang X et al (2024) A survey on semantic processing techniques. Inf Fusion 101:101988","journal-title":"Inf Fusion"},{"key":"9914_CR35","volume-title":"WordNet: an electronic lexical database","author":"GA Miller","year":"1998","unstructured":"Miller GA (1998) WordNet: an electronic lexical database. MIT Press, Cambridge"},{"key":"9914_CR36","doi-asserted-by":"crossref","unstructured":"Minderer M, Gritsenko A, Stone A et\u00a0al (2022) Simple open-vocabulary object detection with vision transformers. arXiv preprint arXiv:2205.06230","DOI":"10.1007\/978-3-031-20080-9_42"},{"issue":"2","key":"9914_CR37","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/1459352.1459355","volume":"41","author":"R Navigli","year":"2009","unstructured":"Navigli R (2009) Word sense disambiguation: a survey. ACM Comput Surv (CSUR) 41(2):1\u201369","journal-title":"ACM Comput Surv (CSUR)"},{"key":"9914_CR38","doi-asserted-by":"crossref","unstructured":"Nickel M, Rosasco L, Poggio TA (2016) Holographic embeddings of knowledge graphs. In: Schuurmans D, Wellman MP (eds) Proceedings of the thirtieth AAAI conference on artificial intelligence, Febr 12\u201317, 2016, Phoenix, Arizona, USA. AAAI Press, Washington, pp 1955\u20131961. http:\/\/www.aaai.org\/ocs\/index.php\/AAAI\/AAAI16\/paper\/view\/12484","DOI":"10.1609\/aaai.v30i1.10314"},{"key":"9914_CR39","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s00521-020-05484-4","volume":"34","author":"L Pasa","year":"2022","unstructured":"Pasa L, Navarin N, Sperduti A (2022) SOM-based aggregation for graph convolutional neural networks. Neural Comput Appl 34:1\u201320","journal-title":"Neural Comput Appl"},{"key":"9914_CR40","doi-asserted-by":"crossref","unstructured":"Pellissier\u00a0Tanon T, Weikum G, Suchanek F (2020) YAGO 4: a reason-able knowledge base. In: European semantic web conference. Springer, pp 583\u2013596","DOI":"10.1007\/978-3-030-49461-2_34"},{"issue":"1","key":"9914_CR41","first-page":"43","volume":"18","author":"V Raj","year":"2024","unstructured":"Raj V, Abbas N (2024) Contextual sense model: word sense disambiguation using sense and sense value of context surrounding the target. Int J Cognit Lang Sci 18(1):43\u201350","journal-title":"Int J Cognit Lang Sci"},{"key":"9914_CR42","doi-asserted-by":"publisher","unstructured":"Rigoni D, Serafini L, Sperduti A (2022) A better loss for visual-textual grounding. In: Hong J, Bures M, Park JW et\u00a0al (eds) SAC\u201922: the 37th ACM\/SIGAPP symposium on applied computing, virtual event, April 25\u201329, 2022. ACM, pp 49\u201357. https:\/\/doi.org\/10.1145\/3477314.3507047","DOI":"10.1145\/3477314.3507047"},{"key":"9914_CR43","doi-asserted-by":"crossref","unstructured":"Rigoni D, Elliott D, Frank S (2023a) Cleaner categories improve object detection and visual-textual grounding. In: Scandinavian conference on image analysis. Springer, Berlin, pp 412\u2013442","DOI":"10.1007\/978-3-031-31435-3_28"},{"key":"9914_CR44","unstructured":"Rigoni D, Parolari L, Serafini L et al (2023b) Weakly-supervised visual-textual grounding with semantic prior refinement. In: 34th British machine vision conference 2023. BMVA Press, Aberdeen, UK. http:\/\/proceedings.bmvc2023.org\/229\/"},{"key":"9914_CR45","doi-asserted-by":"crossref","unstructured":"Rohrbach A, Rohrbach M, Hu R et\u00a0al (2016) Grounding of textual phrases in images by reconstruction. In: European conference on computer vision. Springer, Berlin, pp 817\u2013834","DOI":"10.1007\/978-3-319-46448-0_49"},{"key":"9914_CR46","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.118669","volume":"212","author":"A Salaberria","year":"2023","unstructured":"Salaberria A, Azkune G, de Lacalle OL et al (2023) Image captioning for effective use of language models in knowledge-based visual question answering. Expert Syst Appl 212:118669","journal-title":"Expert Syst Appl"},{"key":"9914_CR47","doi-asserted-by":"crossref","unstructured":"Shi C, Yang S (2023) EDADET: open-vocabulary object detection using early dense alignment. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 15724\u201315734","DOI":"10.1109\/ICCV51070.2023.01441"},{"key":"9914_CR48","doi-asserted-by":"crossref","unstructured":"Shih KJ, Singh S, Hoiem D (2016) Where to look: focus regions for visual question answering. In: CVPR, pp 4613\u20134621","DOI":"10.1109\/CVPR.2016.499"},{"key":"9914_CR49","first-page":"249","volume":"249","author":"M Stevenson","year":"2003","unstructured":"Stevenson M, Wilks Y (2003) Word sense disambiguation. Oxf Handb Comput Linguist 249:249","journal-title":"Oxf Handb Comput Linguist"},{"key":"9914_CR50","doi-asserted-by":"crossref","unstructured":"Su W, Miao P, Dou H et\u00a0al (2023) Language adaptive weight generation for multi-task visual grounding. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10857\u201310866","DOI":"10.1109\/CVPR52729.2023.01045"},{"key":"9914_CR51","unstructured":"Suchanek F, Alam M, Bonald T et\u00a0al (2023) Integrating the wikidata taxonomy into YAGO. arXiv preprint arXiv:2308.11884"},{"key":"9914_CR52","unstructured":"Veli\u010dkovi\u0107 P, Cucurull G, Casanova A et\u00a0al (2017) Graph attention networks. arXiv preprint arXiv:1710.10903"},{"key":"9914_CR53","doi-asserted-by":"crossref","unstructured":"Wang J, Zhang H, Hong H et\u00a0al (2023) Open-vocabulary object detection with an open corpus. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 6759\u20136769","DOI":"10.1109\/ICCV51070.2023.00622"},{"issue":"7","key":"9914_CR54","doi-asserted-by":"publisher","first-page":"5397","DOI":"10.1007\/S00521-021-06696-Y","volume":"34","author":"J Wu","year":"2022","unstructured":"Wu J, Weng W, Fu J et al (2022) Deep semantic hashing with dual attention for cross-modal retrieval. Neural Comput Appl 34(7):5397\u20135416. https:\/\/doi.org\/10.1007\/S00521-021-06696-Y","journal-title":"Neural Comput Appl"},{"key":"9914_CR55","doi-asserted-by":"crossref","unstructured":"Wu S, Zhang W, Jin S et\u00a0al (2023) Aligning bag of regions for open-vocabulary object detection. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 15254\u201315264","DOI":"10.1109\/CVPR52729.2023.01464"},{"issue":"4","key":"9914_CR56","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3572844","volume":"19","author":"S Yang","year":"2023","unstructured":"Yang S, Li Q, Li W et al (2023) Semantic completion and filtration for image-text retrieval. ACM Trans Multimedia Comput Commun Appl 19(4):1\u201320","journal-title":"ACM Trans Multimedia Comput Commun Appl"},{"key":"9914_CR57","doi-asserted-by":"crossref","unstructured":"Yang Z, Gong B, Wang L et\u00a0al (2019) A fast and accurate one-stage approach to visual grounding. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 4683\u20134693","DOI":"10.1109\/ICCV.2019.00478"},{"key":"9914_CR58","unstructured":"Zaheer M, Kottur S, Ravanbakhsh S et\u00a0al (2017) Deep sets. Advances in neural information processing systems 30"},{"key":"9914_CR59","doi-asserted-by":"crossref","unstructured":"Zhang H, Niu Y, Chang SF (2018) Grounding referring expressions in images by variational context. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4158\u20134166","DOI":"10.1109\/CVPR.2018.00437"},{"key":"9914_CR60","first-page":"8772","volume":"2023","author":"X Zhang","year":"2023","unstructured":"Zhang X, Mao R, He K et al (2023) Neuro-symbolic sentiment analysis with dynamic word sense disambiguation. Find Assoc Comput Linguist: EMNLP 2023:8772\u20138783","journal-title":"Find Assoc Comput Linguist: EMNLP"},{"key":"9914_CR61","doi-asserted-by":"crossref","unstructured":"Zhang X, Zhen T, Zhang J et\u00a0al (2023b) SRCB at semeval-2023 task 1: prompt based and cross-modal retrieval enhanced visual word sense disambiguation. In: Proceedings of the 17th international workshop on semantic evaluation (SemEval-2023), pp 439\u2013446","DOI":"10.18653\/v1\/2023.semeval-1.60"},{"issue":"11","key":"9914_CR62","doi-asserted-by":"publisher","first-page":"9015","DOI":"10.1007\/S00521-022-06923-0","volume":"34","author":"J Zhao","year":"2022","unstructured":"Zhao J, Zhang X, Wang X et al (2022) Overcoming language priors in VQA via adding visual module. Neural Comput Appl 34(11):9015\u20139023. https:\/\/doi.org\/10.1007\/S00521-022-06923-0","journal-title":"Neural Comput Appl"},{"key":"9914_CR63","unstructured":"Zhou B, Tian Y, Sukhbaatar S et\u00a0al (2015) Simple baseline for visual question answering. arXiv preprint arXiv:1512.02167"}],"updated-by":[{"DOI":"10.1007\/s00521-025-11232-3","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2025,4,15]],"date-time":"2025-04-15T00:00:00Z","timestamp":1744675200000}}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-024-09914-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-024-09914-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-024-09914-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,15]],"date-time":"2025-04-15T07:15:03Z","timestamp":1744701303000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-024-09914-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,20]]},"references-count":63,"journal-issue":{"issue":"26","published-print":{"date-parts":[[2024,9]]}},"alternative-id":["9914"],"URL":"https:\/\/doi.org\/10.1007\/s00521-024-09914-5","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,20]]},"assertion":[{"value":"30 October 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 April 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 May 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 April 2025","order":4,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Correction","order":5,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"A Correction to this paper has been published:","order":6,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"https:\/\/doi.org\/10.1007\/s00521-025-11232-3","URL":"https:\/\/doi.org\/10.1007\/s00521-025-11232-3","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that there are no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval"}},{"value":"The authors have consented to the submission of this manuscript to the journal.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"The authors have consented to the publication of this manuscript to the journal.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}