{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:06:19Z","timestamp":1760241979172,"version":"build-2065373602"},"reference-count":49,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2018,11,12]],"date-time":"2018-11-12T00:00:00Z","timestamp":1541980800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the National Key R&amp;G Program of China","award":["2018YFB1004600"],"award-info":[{"award-number":["2018YFB1004600"]}]},{"name":"the National Key Research and Development Program of China","award":["2016YFB10005000"],"award-info":[{"award-number":["2016YFB10005000"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Image caption generation is a fundamental task to build a bridge between image and its description in text, which is drawing increasing interest in artificial intelligence. Images and textual sentences are viewed as two different carriers of information, which are symmetric and unified in the same content of visual scene. The existing image captioning methods rarely consider generating a final description sentence in a coarse-grained to fine-grained way, which is how humans understand the surrounding scenes; and the generated sentence sometimes only describes coarse-grained image content. Therefore, we propose a coarse-to-fine-grained hierarchical generation method for image captioning, named SDA-CFGHG, to address the two problems above. The core of our SDA-CFGHG method is a sequential dual attention that is used to fuse different grained visual information with sequential means. The advantage of our SDA-CFGHG method is that it can achieve image captioning in a coarse-to-fine-grained way and the generated textual sentence can capture details of the raw image to some degree. Moreover, we validate the impressive performance of our method on benchmark datasets\u2014MS COCO, Flickr\u2014with several popular evaluation metrics\u2014CIDEr, SPICE, METEOR, ROUGE-L, and BLEU.<\/jats:p>","DOI":"10.3390\/sym10110626","type":"journal-article","created":{"date-parts":[[2018,11,14]],"date-time":"2018-11-14T10:58:22Z","timestamp":1542193102000},"page":"626","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Sequential Dual Attention: Coarse-to-Fine-Grained Hierarchical Generation for Image Captioning"],"prefix":"10.3390","volume":"10","author":[{"given":"Zhibin","family":"Guan","sequence":"first","affiliation":[{"name":"School of Mechanical Electronic &amp; Information Engineering, China University of Mining &amp; Technology (Beijing), Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8413-123X","authenticated-orcid":false,"given":"Kang","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Mechanical Electronic &amp; Information Engineering, China University of Mining &amp; Technology (Beijing), Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9243-8322","authenticated-orcid":false,"given":"Yan","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Mechanical Electronic &amp; Information Engineering, China University of Mining &amp; Technology (Beijing), Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xu","family":"Qian","sequence":"additional","affiliation":[{"name":"School of Mechanical Electronic &amp; Information Engineering, China University of Mining &amp; Technology (Beijing), Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tongkai","family":"Ji","sequence":"additional","affiliation":[{"name":"School of Mechanical Electronic &amp; Information Engineering, China University of Mining &amp; Technology (Beijing), Beijing 100083, China"},{"name":"G-Cloud Technology Corporation, Cloud Computing Center, Chinese Academy of Sciences, Dongguan 523808, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,11,12]]},"reference":[{"key":"ref_1","unstructured":"Wei, Y., Xia, W., Huang, J., Ni, B., Dong, J., Zhao, Y., and Yan, S. (arXiv, 2014). CNN: Single-label to Multi-label, arXiv."},{"key":"ref_2","unstructured":"Simonyan, K., and Zisserman, A. (arXiv, 2015). Very Deep Convolutional Networks for Large-Scale Image Recognition, arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (arXiv, 2015). Deep Residual Learning for Image Recognition, arXiv.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1016\/j.neunet.2017.10.009","article-title":"Adaptive neuro-heuristic hybrid model for fruit peel defects detection","volume":"98","year":"2018","journal-title":"Neural Netw."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1016\/j.neunet.2017.12.005","article-title":"STDP-based spiking deep convolutional neural networks for object recognition","volume":"99","author":"Kheradpisheh","year":"2018","journal-title":"Neural Netw."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1016\/j.cmpb.2018.04.025","article-title":"Small lung nodules detection based on local variance analysis and probabilistic neural network","volume":"161","author":"Capizzi","year":"2018","journal-title":"Comput. Methods Programs Biomed."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"610","DOI":"10.1016\/j.patcog.2016.07.026","article-title":"Facial expression recognition with Convolutional Neural Networks: Coping with few data and the training sample order","volume":"61","author":"Lopes","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"76","DOI":"10.1016\/j.neucom.2018.09.003","article-title":"Object detection and recognition via clustered features","volume":"320","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Po\u0142ap, D., Winnicka, A., Serwata, K., K\u0119sik, K., and Wo\u017aniak, M. (2018). An Intelligent System for Monitoring Skin Diseases. Sensors, 18.","DOI":"10.3390\/s18082552"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Mikolov, T., Karafi\u00e1t, M., Burget, L., \u010cernock\u1ef3, J., and Khudanpur, S. (2010, January 26\u201330). Recurrent neural network based language model. Proceedings of the Eleventh Annual Conference of the International Speech Communication Association, Chiba, Japan.","DOI":"10.21437\/Interspeech.2010-343"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2673","DOI":"10.1109\/78.650093","article-title":"Bidirectional recurrent neural networks","volume":"45","author":"Schuster","year":"1997","journal-title":"IEEE Trans. Signal Process."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Hochreiter, S., and Schmidhuber, J. (1997). Long Short-term Memory. Neural Computation, MIT Press.","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merrienboer, B., G\u00fcl\u00e7ehre, \u00c7., Bougares, F., Schwenk, H., and Bengio, Y. (arXiv, 2014). Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation, arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_14","unstructured":"Ren, S., He, K., Girshick, R.B., and Sun, J. (arXiv, 2015). Faster R-CNN: Towards Real-Time Object Detection with Region Proposal, arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Zitnick, C.L., and Parikh, D. (arXiv, 2014). CIDEr: Consensus-based Image Description Evaluation, arXiv.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Anderson, P., Fernando, B., Johnson, M., and Gould, S. (2016). SPICE: Semantic Propositional Image Caption Evaluation. ECCV, Springer.","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Denkowski, M., and Lavie, A. (2014, January 26\u201327). Meteor Universal: Language Specific Translation Evaluation for Any Target Language. Proceedings of the Ninth Workshop on Statistical Machine Translation, Association for Computational Linguistics, Baltimore, MD, USA.","DOI":"10.3115\/v1\/W14-3348"},{"key":"ref_18","unstructured":"Lin, C.Y. (2004, January 25\u201326). ROUGE: A Package for Automatic Evaluation of Summaries. Proceedings of the ACL-04 Workshop Text Summarization Branches Out, Barcelona, Spain."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., and Ward, T. (2002, January 7\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Farhadi, A., Hejrati, M., Sadeghi, M.A., Young, P., Rashtchian, C., Hockenmaier, J., and Forsyth, D. (2010). Every Picture Tells a Story: Generating Sentences from Images. ECCV 2010, Springer.","DOI":"10.1007\/978-3-642-15561-1_2"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Kulkarni, G., Premraj, V., Dhar, S., Li, S., Choi, Y., Berg, A.C., and Berg, T.L. (2011, January 20\u201325). Baby talk: Understanding and generating simple image descriptions. Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995466"},{"key":"ref_22","unstructured":"Yang, Y., Teo, C.L., Daum\u00e9, H., and Aloimonos, Y. (2011, January 27\u201331). Corpus-Guided Sentence Generation of Natural Images. Proceedings of the EMNLP \u201911 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, Stroudsburg, PA, USA."},{"key":"ref_23","unstructured":"Kuznetsova, P., Ordonez, V., Berg, A.C., Berg, T.L., and Choi, Y. (2012, January 8\u201314). Collective Generation of Natural Image Descriptions. Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Mason, R., and Charniak, E. (2014, January 22\u201327). Nonparametric Method for Data-driven Image Captioning. Proceedings of the 52th Annual Meeting of the Association for Computational Linguistics, Baltimore, MD, USA.","DOI":"10.3115\/v1\/P14-2097"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"46","DOI":"10.1007\/s11263-015-0840-y","article-title":"Large Scale Retrieval and Generation of Image Descriptions","volume":"119","author":"Ordonez","year":"2016","journal-title":"Int. J. Comput. Vis. (IJCV)"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Jia, X., Gavves, E., Fernando, B., and Tuytelaars, T. (arXiv, 2015). Guiding Long-Short Term Memory for Image Caption Generation, arXiv.","DOI":"10.1109\/ICCV.2015.277"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Luo, R., Price, B.L., Cohen, S., and Shakhnarovich, G. (arXiv, 2018). Discriminability objective for training descriptive captions, arXiv.","DOI":"10.1109\/CVPR.2018.00728"},{"key":"ref_28","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A.C., Salakhutdinov, R., Zemel, R.S., and Bengio, Y. (arXiv, 2015). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ren, Z., Wang, X., Zhang, N., Lv, X., and Li, L. (2017, January 21\u201326). Deep Reinforcement Learning-based Image Captioning with Embedding Reward. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.128"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"726","DOI":"10.1109\/TMM.2017.2751140","article-title":"GLA: Global-Local Attention for Image Description","volume":"20","author":"Li","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"5514","DOI":"10.1109\/TIP.2018.2855406","article-title":"Attentive Linear Transformation for Image Captioning","volume":"27","author":"Ye","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1016\/j.neucom.2018.08.069","article-title":"Image captioning with triple-attention and stack parallel LSTM","volume":"319","author":"Zhu","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhou, L., Xu, C., Koch, P.A., and Corso, J.J. (arXiv, 2016). Watch What You Just Said: Image Caption Generation with Text-Conditional Semantic Attention, arXiv.","DOI":"10.1145\/3126686.3126717"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Pedersoli, M., Lucas, T., Schmid, C., and Verbeek, J. (arXiv, 2016). Areas of Attention for Image Captioning, arXiv.","DOI":"10.1109\/ICCV.2017.140"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Liu, C., Sun, F., Wang, C., Wang, F., and Yuille, A.L. (arXiv, 2017). MAT: A Multimodal Attentive Translator for Image Captioning, arXiv.","DOI":"10.24963\/ijcai.2017\/563"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (arXiv, 2017). Bottom-Up and Top-Down Attention for Image Captioning and VQA, arXiv.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1016\/j.image.2018.06.002","article-title":"Modeling visual and word-conditional semantic attention for image captioning","volume":"67","author":"Wu","year":"2018","journal-title":"Signal Process. Image Commun."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lu, J., Xiong, C., Parikh, D., and Socher, R. (2017, January 21\u201326). Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.345"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Wu, Q., Shen, C., Liu, L., Dick, A., and Hengel, A. (2016, January 27\u201330). What Value Do Explicit High Level Concepts Have in Vision to Language Problems?. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.29"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Wang, L., Chu, X., Zhang, W., Wei, Y., Sun, W., and Wu, C. (2018). Social Image Captioning: Exploring Visual Attention and User Attention. Sensors, 18.","DOI":"10.3390\/s18020646"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Yao, T., Pan, Y., Li, Y., Qiu, Z., and Mei, T. (2017, January 22\u201329). Boosting Image Captioning with Attributes. Proceedings of the IEEE International Conference on Computer Vision ICCV, Venice, Italy.","DOI":"10.1109\/ICCV.2017.524"},{"key":"ref_42","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012). ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems 25, Curran Associates, Inc."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Rennie, S.J., Marcheret, E., Mroueh, Y., Ross, J., and Goel, V. (arXiv, 2016). Self-critical Sequence Training for Image Captioning, arXiv.","DOI":"10.1109\/CVPR.2017.131"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Lin, T., Maire, M., Belongie, S.J., Bourdev, L.D., Girshick, R.B., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (arXiv, 2014). Microsoft COCO: Common Objects in Context, arXiv.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_45","first-page":"853","article-title":"Framing Image Description As a Ranking Task: Data, Models and Evaluation Metrics","volume":"47","author":"Hodosh","year":"2013","journal-title":"J. Abbr."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1162\/tacl_a_00166","article-title":"From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions","volume":"2","author":"Young","year":"2014","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_47","unstructured":"Kiros, R., Salakhutdinov, R., and Zemel, R. (2014, January 21\u201326). Multimodal neural language models. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"You, Q., Jin, H., Wang, Z., Fang, C., and Luo, J. (2016, January 27\u201330). Image Captioning with Semantic Attention. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.503"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"2959","DOI":"10.1007\/s11042-017-4593-1","article-title":"Fine-grained attention for image caption generation","volume":"77","author":"Chang","year":"2018","journal-title":"Multimed. Tools Appl."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/10\/11\/626\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:29:10Z","timestamp":1760196550000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/10\/11\/626"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,11,12]]},"references-count":49,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2018,11]]}},"alternative-id":["sym10110626"],"URL":"https:\/\/doi.org\/10.3390\/sym10110626","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2018,11,12]]}}}