{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T13:22:06Z","timestamp":1782220926700,"version":"3.54.5"},"reference-count":56,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2021,6,30]],"date-time":"2021-06-30T00:00:00Z","timestamp":1625011200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Object detection, visual relationship detection, and image captioning, which are the three main visual tasks in scene understanding, are highly correlated and correspond to different semantic levels of scene image. However, the existing captioning methods convert the extracted image features into description text, and the obtained results are not satisfactory. In this work, we propose a Multi-level Semantic Context Information (MSCI) network with an overall symmetrical structure to leverage the mutual connections across the three different semantic layers and extract the context information between them, to solve jointly the three vision tasks for achieving the accurate and comprehensive description of the scene image. The model uses a feature refining structure to mutual connections and iteratively updates the different semantic features of the image. Then a context information extraction network is used to extract the context information between the three different semantic layers, and an attention mechanism is introduced to improve the accuracy of image captioning while using the context information between the different semantic layers to improve the accuracy of object detection and relationship detection. Experiments on the VRD and COCO datasets demonstrate that our proposed model can leverage the context information between semantic layers to improve the accuracy of those visual tasks generation.<\/jats:p>","DOI":"10.3390\/sym13071184","type":"journal-article","created":{"date-parts":[[2021,7,1]],"date-time":"2021-07-01T02:44:39Z","timestamp":1625107479000},"page":"1184","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["Image Caption Generation Using Multi-Level Semantic Context Information"],"prefix":"10.3390","volume":"13","author":[{"given":"Peng","family":"Tian","sequence":"first","affiliation":[{"name":"College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongwei","family":"Mo","sequence":"additional","affiliation":[{"name":"College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Laihao","family":"Jiang","sequence":"additional","affiliation":[{"name":"College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,6,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., and Ren, S. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_2","unstructured":"Karpathy, A., and Li, F.-F. (2016, January 7\u201312). Deep Visual-Semantic Alignments for Generating Image Descriptions. Proceedings of the IEEE Transactions on Pattern Analysis and Machine Intelligence, Boston, MA, USA."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1016\/j.neucom.2015.09.116","article-title":"Deep learning for visual understanding: A review","volume":"187","author":"Yan","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1041","DOI":"10.1049\/el.2017.0326","article-title":"Multimodal object description network for dense captioning","volume":"53","author":"Wang","year":"2017","journal-title":"IEEE Electron. Lett."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Johnson, J., Karpathy, A., and Li, F.-F. (2016, January 27\u201330). DenseCap: Fully Convolutional Localization Networks for Dense Captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.494"},{"key":"ref_6","unstructured":"Xu, K., Ba, J., and Kiros, R. (2015, January 6\u20137). Show, attend and tell: Neural image caption generation with visual attention. Proceedings of the 32nd International Conference on International Conference on Machine Learning (ICML), Lille, France."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Gu, J., Wang, G., and Cai, J. (2017, January 22\u201329). An Empirical Study of Language CNN for Image Captioning. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.138"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1016\/j.patrec.2020.12.020","article-title":"Image Captioning with Transformer and Knowledge Graph","volume":"143","author":"Zhang","year":"2021","journal-title":"Pattern Recognit. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"104146","DOI":"10.1016\/j.imavis.2021.104146","article-title":"Exploring Region Relationships Implicitly: Image Captioning with Visual Relationship Attention","volume":"109","author":"Zhang","year":"2021","journal-title":"Image Vis. Comput."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Sun, Y., and Honavar, V. (2019, January 8\u201310). Improving Image Captioning by Leveraging Knowledge Graphs. Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa Village, HI, USA.","DOI":"10.1109\/WACV.2019.00036"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"You, Q., Jin, H., and Wang, Z. (2016, January 27\u201330). Image captioning with semantic attention. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.503"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Lu, J., Xiong, C., and Parikh, D. (2017, January 21\u201326). Knowing when to look: Adaptive attention via a visual sentinel for image captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.345"},{"key":"ref_13","unstructured":"Gao, L., Fan, K., and Song, J. (2019, January 27\u201331). Deliberate Attention Networks for Image Captioning. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), Honolulu, HI, USA."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Yang, X., Tang, K., and Zhang, H. (2019, January 15\u201321). Auto-Encoding Scene Graphs for Image Captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01094"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhong, Y., Wang, L., and Chen, J. (2020, January 23\u201328). Comprehensive Image Captioning via Scene Graph Decomposition. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-58568-6_13"},{"key":"ref_16","unstructured":"Li, Y., Tarlow, D., and Brockschmidt, M. (2016, January 2\u20134). Gated Graph Sequence Neural Networks. Proceedings of the IEEE International Conference on Learning Representations (ICLR), San Juan, PR, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., and Darrell, T. (2014, January 23\u201328). Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","unstructured":"Bochkovskiy, A., Wang, C.-Y., and Liao, H. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., and Erhan, D. (2016, January 11\u201314). SSD: Single Shot MultiBox Detector. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Li, Y., Ouyang, W., and Zhou, B. (2018, January 8\u201314). Factorizable net: An efficient subgraph-based framework for scene graph generation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01246-5_21"},{"key":"ref_22","unstructured":"Lv, J., Xiao, Q., and Zhong, J. (2020). AVR: Attention based Salient Visual Relationship Detection. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liang, X., Lee, L., and Xing, E.P. (2017, January 21\u201326). Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.469"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Lu, C., Krishna, R., and Bernstein, M. (2016, January 11\u201314). Visual Relationship Detection with Language Priors. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_51"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Dai, B., Zhang, Y., and Lin, D. (2017, January 21\u201326). Detecting visual relationships with deep relational networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.352"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Chen, T., Yu, W., and Chen, R. (2019, January 16\u201320). Knowledge-Embedded Routing Network for Scene Graph Generation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00632"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"2891","DOI":"10.1109\/TPAMI.2012.162","article-title":"Babytalk: Understanding and generating simple image descriptions","volume":"35","author":"Kulkarni","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","unstructured":"Elliott, D., and Keller, F. (2013, January 18\u201321). Image description using visual dependency representations. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP), Seattle, WA, USA."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Verma, Y., Gupta, A., and Mannem, P. (2013, January 23\u201328). Generating image descriptions using semantic similarities in the output space. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPRW.2013.50"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Devlin, J., Cheng, H., and Fang, H. (2015, January 26\u201331). Language models for image captioning: The quirks and what works. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), Beijing, China.","DOI":"10.3115\/v1\/P15-2017"},{"key":"ref_31","unstructured":"Zhu, Z., Wang, T., and Qu, H. (2021). Macroscopic Control of Text Generation for Image Captioning. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"7615","DOI":"10.1109\/TIP.2020.3004729","article-title":"Spatio-Temporal Memory Attention for Image Captioning","volume":"29","author":"Ji","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., and Buehler, C. (2018, January 18\u201323). Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake, UT, USA.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_34","unstructured":"Wang, W., Chen, Z., and Hu, H. (2019, January 27\u201331). Hierarchical Attention Network for Image Captioning. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), Honolulu, HI, USA."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"43","DOI":"10.3389\/fnbot.2020.00043","article-title":"Interactive Natural Language Grounding via Referring Expression Comprehension and Scene Graph Parsing","volume":"14","author":"Mi","year":"2020","journal-title":"Front. Neurorobot."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"2117","DOI":"10.1109\/TMM.2019.2896516","article-title":"Know More Say Less: Image Captioning Based on Scene Graphs","volume":"21","author":"Li","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Mottaghi, R., Chen, X., and Liu, X. (2014, January 23\u201328). The Role of Context for Object Detection and Semantic Segmentation in the Wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.119"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Zeng, X., Ouyang, W., and Yang, B. (2016, January 11\u201314). Gated Bi-directional CNN for Object Detection. Proceedings of the IEEE European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46478-7_22"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Ma, Y., Guo, Y., and Liu, H. (2020, January 2\u20135). Global Context Reasoning for Semantic Segmentation of 3D Point Clouds. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Snowmass Village, CO, USA.","DOI":"10.1109\/WACV45572.2020.9093411"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Lin, C.-Y., Chiu, Y.-C., and Ng, H.-F. (2020). Global-and-Local Context Network for Semantic Segmentation of Street View Images. Sensors, 20.","DOI":"10.3390\/s20102907"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"2014","DOI":"10.1109\/TPAMI.2019.2961896","article-title":"On the Importance of Visual Context for Data Augmentation in Scene Understanding","volume":"43","author":"Dvornik","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Zhuang, B., Liu, L., and Shen, C. (2017, January 22\u201329). Towards Context-Aware Interaction Recognition for Visual Relationship Detection. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.71"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zellers, R., Yatskar, M., and Thomson, S. (2018, January 18\u201323). Neural motifs: Scene graph parsing with global context. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake, UT, USA.","DOI":"10.1109\/CVPR.2018.00611"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Qi, X., Liao, R., and Jia, J. (2017, January 22\u201329). 3D Graph Neural Networks for RGBD Semantic Segmentation. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.556"},{"key":"ref_45","unstructured":"Kenneth, M., Ruslan, S., and Abhinav, G. (2017). The More You Know: Using Knowledge Graphs for Image Classification. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Maire, M., and Belongie, S. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the IEEE European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Plummer, B.-A., Wang, L., and Cervantes, C.-M. (2015, January 13\u201316). Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.303"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., and Ward, T. (2002, January 6\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (ACL), Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_50","unstructured":"Banerjee, S., and Lavie, A. (2005, January 29). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, Ann Arbor, MI, USA."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Lin, C.-Y., and Hovy, E. (2003, January 1\u201311). Automatic evaluation of summaries using n-gram co-occurrence statistics. Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), Baltimore, MD, USA.","DOI":"10.3115\/1073445.1073465"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Zitnick, C., and Parikh, D. (2015, January 7\u201312). Cider: Consensus-based image description evaluation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Anderson, P., Fernando, B., and Johnson, M. (2016, January 11\u201314). Spice: Semantic propositional image caption evaluation. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"ref_54","unstructured":"Kingma, D., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1016\/j.patrec.2017.10.018","article-title":"Image caption generation with part of speech guidance","volume":"119","author":"He","year":"2019","journal-title":"Pattern Recognit. Lett."},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"30615","DOI":"10.1007\/s11042-020-09539-5","article-title":"Reference-based model using multimodal gated recurrent units for image captioning","volume":"79","author":"Nogueira","year":"2020","journal-title":"Multimed. Tools Appl."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/7\/1184\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:24:05Z","timestamp":1760163845000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/7\/1184"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,30]]},"references-count":56,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2021,7]]}},"alternative-id":["sym13071184"],"URL":"https:\/\/doi.org\/10.3390\/sym13071184","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,6,30]]}}}