{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T03:27:10Z","timestamp":1773372430560,"version":"3.50.1"},"reference-count":40,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2022,2,17]],"date-time":"2022-02-17T00:00:00Z","timestamp":1645056000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"The National Pre-Research Foundation of China","award":["No.104040402"],"award-info":[{"award-number":["No.104040402"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Collaborative reasoning for knowledge-based visual question answering is challenging but vital and efficient in understanding the features of the images and questions. While previous methods jointly fuse all kinds of features by attention mechanism or use handcrafted rules to generate a layout for performing compositional reasoning, which lacks the process of visual reasoning and introduces a large number of parameters for predicting the correct answer. For conducting visual reasoning on all kinds of image\u2013question pairs, in this paper, we propose a novel reasoning model of a question-guided tree structure with a knowledge base (QGTSKB) for addressing these problems. In addition, our model consists of four neural module networks: the attention model that locates attended regions based on the image features and question embeddings by attention mechanism, the gated reasoning model that forgets and updates the fused features, the fusion reasoning model that mines high-level semantics of the attended visual features and knowledge base and knowledge-based fact model that makes up for the lack of visual and textual information with external knowledge. Therefore, our model performs visual analysis and reasoning based on tree structures, knowledge base and four neural module networks. Experimental results show that our model achieves superior performance over existing methods on the VQA v2.0 and CLVER dataset, and visual reasoning experiments prove the interpretability of the model.<\/jats:p>","DOI":"10.3390\/s22041575","type":"journal-article","created":{"date-parts":[[2022,2,17]],"date-time":"2022-02-17T20:26:41Z","timestamp":1645129601000},"page":"1575","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Learning to Reason on Tree Structures for Knowledge-Based Visual Question Answering"],"prefix":"10.3390","volume":"22","author":[{"given":"Qifeng","family":"Li","sequence":"first","affiliation":[{"name":"Shanghai Institute of Technical Physics of the Chinese Academy of Sciences, Shanghai 200083, China"},{"name":"School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"Key Laboratory of Infrared System Detection and Imaging Technology, Chinese Academy of Sciences, Shanghai 200083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyi","family":"Tang","sequence":"additional","affiliation":[{"name":"Shanghai Institute of Technical Physics of the Chinese Academy of Sciences, Shanghai 200083, China"},{"name":"Key Laboratory of Infrared System Detection and Imaging Technology, Chinese Academy of Sciences, Shanghai 200083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Jian","sequence":"additional","affiliation":[{"name":"Shanghai Institute of Technical Physics of the Chinese Academy of Sciences, Shanghai 200083, China"},{"name":"Key Laboratory of Infrared System Detection and Imaging Technology, Chinese Academy of Sciences, Shanghai 200083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,17]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Shih, K.J., Singh, S., and Hoiem, D. (2016, January 27\u201330). Where To Look: Focus Regions for Visual Question Answering. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR.2016.499"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M. (2016, January 1\u20135). Multimodal compact bilinear pooling for visual question answering and visual grounding. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, TX, USA.","DOI":"10.18653\/v1\/D16-1044"},{"key":"ref_3","unstructured":"Lu, J., Yang, J., Batra, D., and Parikh, D. (2016, January 5\u201310). Hierarchical Question-Image Co-Attention for Visual Question Answering. Proceedings of the 30th Conference on Neural Information Processing Systems (NIPS), Barcelona, Spain."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Shi, Y., Furlanello, T., Zha, S., and Anandkumar, A. (2018, January 8\u201314). Question Type Guided Attention in Visual Question Answering. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_10"},{"key":"ref_5","first-page":"2156","article-title":"Dual attention networks for multimodal reasoning and matching","volume":"Volume 2017-January","author":"Nam","year":"2017","journal-title":"Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Manjunatha, V., Saini, N., Davis, L.S., and Soc, I.C. (2019, January 16\u201320). Explicit Bias Discovery in Visual Question Answering Models. Proceedings of the 32nd IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00979"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Johnson, J., Hariharan, B., Van Der Maaten, L., Hoffman, J., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R. (2017, January 22\u201329). Inferring and Executing Programs for Visual Reasoning. Proceedings of the 16th IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.325"},{"key":"ref_8","first-page":"804","article-title":"Learning to Reason: End-to-End Module Networks for Visual Question Answering","volume":"Volume 2017-October","author":"Hu","year":"2017","journal-title":"Proceedings of the 16th IEEE International Conference on Computer Vision, ICCV 2017"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1016\/j.inffus.2019.08.009","article-title":"Multimodal feature fusion by relational reasoning and attention for visual question answering","volume":"55","author":"Zhang","year":"2020","journal-title":"Inf. Fusion"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"887","DOI":"10.1109\/TPAMI.2019.2943456","article-title":"Interpretable Visual Question Answering by Reasoning on Dependency Trees","volume":"43","author":"Cao","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Cao, Q., Liang, X., Li, B., Li, G., and Lin, L. (2018, January 18\u201322). Visual Question Reasoning on General Dependency Tree. Proceedings of the 31st Meeting of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00757"},{"key":"ref_12","first-page":"722","article-title":"DBpedia: A nucleus for a Web of open data","volume":"Volume 4825 LNCS","author":"Auer","year":"2007","journal-title":"Proceedings of the 6th International Semantic Web Conference, ISWC 2007 and 2nd Asian Semantic Web Conference, ASWC 2007"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tandon, N., De Melo, G., Suchanek, F., and Weikum, G. (2014, January 24\u201328). WebChild: Harvesting and organizing commonsense knowledge from the web. Proceedings of the 7th ACM International Conference on Web Search and Data Mining, WSDM 2014, New York, NY, USA.","DOI":"10.1145\/2556195.2556245"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Speer, R., Chin, J., and Havasi, C. (2017, January 4\u20139). ConceptNet 5.5: An Open Multilingual Graph of General Knowledge. Proceedings of the 31st AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11164"},{"key":"ref_15","first-page":"2953","article-title":"Exploring models and data for image question answering","volume":"Volume 2015-January","author":"Ren","year":"2015","journal-title":"Proceedings of the 29th Annual Conference on Neural Information Processing Systems, NIPS 2015"},{"key":"ref_16","unstructured":"Zhang, Y., Hare, J., and Prugel-Bennett, A. (May, January 30). Learning to count objects in natural images for visual question answering. Proceedings of the 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yu, Z., Yu, J., Fan, J., and Tao, D. (2017, January 22\u201329). Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering. Proceedings of the 16th IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.202"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ben-Younes, H., Cadene, R., Thome, N., and Cord, M. (February, January 27). BLOCK: Bilinear superdiagonal fusion for visual question answering and visual relationship detection. Proceedings of the 33rd AAAI Conference on Artificial Intelligence, AAAI 2019, 31st Annual Conference on Innovative Applications of Artificial Intelligence, IAAI 2019 and the 9th AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, HI, USA.","DOI":"10.1609\/aaai.v33i01.33018102"},{"key":"ref_19","first-page":"451","article-title":"Ask, attend and answer: Exploring question-guided spatial attention for visual question answering","volume":"Volume 9911 LNCS","author":"Xu","year":"2016","journal-title":"Proceedings of the 21st ACM Conference on Computer and Communications Security, CCS 2014"},{"key":"ref_20","first-page":"4995","article-title":"Visual7W: Grounded question answering in images","volume":"Volume 2016-December","author":"Zhu","year":"2016","journal-title":"Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Nguyen, D.-K., and Okatani, T. (2018, January 18\u201322). Improved Fusion of Visual and Language Representations by Dense Symmetric Co-attention for Visual Question Answering. Proceedings of the 31st Meeting of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00637"},{"key":"ref_22","first-page":"21","article-title":"Stacked attention networks for image question answering","volume":"Volume 2016-December","author":"Yang","year":"2016","journal-title":"Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018, January 18\u201322). Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. Proceedings of the 31st Meeting of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_24","first-page":"4968","article-title":"A simple neural network module for relational reasoning","volume":"Volume 2017-December","author":"Santoro","year":"2017","journal-title":"Proceedings of the 31st Annual Conference on Neural Information Processing Systems, NIPS 2017"},{"key":"ref_25","first-page":"275","article-title":"Chain of reasoning for visual question answering","volume":"Volume 2018-December","author":"Wu","year":"2018","journal-title":"Proceedings of the 32nd Conference on Neural Information Processing Systems, NeurIPS 2018"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"2413","DOI":"10.1109\/TPAMI.2017.2754246","article-title":"FVQA: Fact-Based Visual Question Answering","volume":"40","author":"Wang","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","first-page":"4622","article-title":"Ask me anything: Free-form visual question answering based on knowledge from external sources","volume":"Volume 2016-December","author":"Wu","year":"2016","journal-title":"Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wang, P., Wu, Q., Shen, C., Dick, A., and van den Hengel, A. (2017, January 19\u201325). Explicit knowledge-based reasoning for visual question answering. Proceedings of the 26th International Joint Conference on Artificial Intelligence, IJCAI 2017, Melbourne, VIC, Australia.","DOI":"10.24963\/ijcai.2017\/179"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"107563","DOI":"10.1016\/j.patcog.2020.107563","article-title":"Cross-modal knowledge reasoning for knowledge-based visual question answering","volume":"108","author":"Yu","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. (2019, January 15\u201320). OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00331"},{"key":"ref_31","first-page":"39","article-title":"Neural module networks","volume":"Volume 2016-December","author":"Andreas","year":"2016","journal-title":"Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Andreas, J., Rohrbach, M., Darrell, T., and Klein, D. (2016, January 12\u201317). Learning to compose neural networks for question answering. Proceedings of the 15th Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL HLT 2016, San Diego, CA, USA.","DOI":"10.18653\/v1\/N16-1181"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Chen, D., and Manning, C.D. (2014, January 25\u201329). A fast and accurate dependency parser using neural networks. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, Doha, Qatar.","DOI":"10.3115\/v1\/D14-1082"},{"key":"ref_34","first-page":"770","article-title":"Deep residual learning for image recognition","volume":"Volume 2016-December","author":"He","year":"2016","journal-title":"Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"5393","DOI":"10.30534\/ijatcse\/2020\/175942020","article-title":"Binary cross entropy with deep learning technique for Image classification","volume":"9","author":"Ruby","year":"2020","journal-title":"Int. J. Adv. Trends Comput. Sci. Eng."},{"key":"ref_36","first-page":"1988","article-title":"CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning","volume":"Volume 2017-January","author":"Johnson","year":"2017","journal-title":"Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"398","DOI":"10.1007\/s11263-018-1116-0","article-title":"Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering","volume":"127","author":"Goyal","year":"2019","journal-title":"Int. J. Comput. Vis."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"740","DOI":"10.1007\/978-3-319-10602-1_48","article-title":"Microsoft COCO: Common objects in context","volume":"Volume 8693 LNCS","author":"Lin","year":"2014","journal-title":"Proceedings of the 13th European Conference on Computer Vision, ECCV 2014"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). GloVe: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_40","unstructured":"Kingma, D.P., and Ba, J.L. (2015, January 7\u20139). Adam: A method for stochastic optimization. Proceedings of the 3rd International Conference on Learning Representations, ICLR 2015, International Conference on Learning Representations, ICLR, San Diego, CA, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/4\/1575\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:21:50Z","timestamp":1760134910000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/4\/1575"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,17]]},"references-count":40,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["s22041575"],"URL":"https:\/\/doi.org\/10.3390\/s22041575","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,17]]}}}