{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,20]],"date-time":"2026-08-20T16:04:41Z","timestamp":1787241881766,"version":"build-2736575974"},"reference-count":46,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2023,8,16]],"date-time":"2023-08-16T00:00:00Z","timestamp":1692144000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"The Nature Science Foundation of China","doi-asserted-by":"publisher","award":["62061042"],"award-info":[{"award-number":["62061042"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"The Nature Science Foundation of China","doi-asserted-by":"publisher","award":["2018[98]"],"award-info":[{"award-number":["2018[98]"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"The Nature Science Foundation of China","doi-asserted-by":"publisher","award":["31920220037"],"award-info":[{"award-number":["31920220037"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Innovative Research for the innovative research team of SEAC","award":["62061042"],"award-info":[{"award-number":["62061042"]}]},{"name":"Innovative Research for the innovative research team of SEAC","award":["2018[98]"],"award-info":[{"award-number":["2018[98]"]}]},{"name":"Innovative Research for the innovative research team of SEAC","award":["31920220037"],"award-info":[{"award-number":["31920220037"]}]},{"DOI":"10.13039\/501100012226","name":"Key Laboratory of China\u2019s Ethnic Languages and Information Technology of Ministry of Education","doi-asserted-by":"publisher","award":["62061042"],"award-info":[{"award-number":["62061042"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Key Laboratory of China\u2019s Ethnic Languages and Information Technology of Ministry of Education","doi-asserted-by":"publisher","award":["2018[98]"],"award-info":[{"award-number":["2018[98]"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Key Laboratory of China\u2019s Ethnic Languages and Information Technology of Ministry of Education","doi-asserted-by":"publisher","award":["31920220037"],"award-info":[{"award-number":["31920220037"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Thangka images exhibit a high level of diversity and richness, and the existing deep learning-based image captioning methods generate poor accuracy and richness of Chinese captions for Thangka images. To address this issue, this paper proposes a Semantic Concept Prompt and Multimodal Feature Optimization network (SCAMF-Net). The Semantic Concept Prompt (SCP) module is introduced in the text encoding stage to obtain more semantic information about the Thangka by introducing contextual prompts, thus enhancing the richness of the description content. The Multimodal Feature Optimization (MFO) module is proposed to optimize the correlation between Thangka images and text. This module enhances the correlation between the image features and text features of the Thangka through the Captioner and Filter to more accurately describe the visual concept features of the Thangka. The experimental results demonstrate that our proposed method outperforms baseline models on the Thangka dataset in terms of BLEU-4, METEOR, ROUGE, CIDEr, and SPICE by 8.7%, 7.9%, 8.2%, 76.6%, and 5.7%, respectively. Furthermore, this method also exhibits superior performance compared to the state-of-the-art methods on the public MSCOCO dataset.<\/jats:p>","DOI":"10.3390\/jimaging9080162","type":"journal-article","created":{"date-parts":[[2023,8,16]],"date-time":"2023-08-16T09:57:02Z","timestamp":1692179822000},"page":"162","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Thangka Image Captioning Based on Semantic Concept Prompt and Multimodal Feature Optimization"],"prefix":"10.3390","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3120-5231","authenticated-orcid":false,"given":"Wenjin","family":"Hu","sequence":"first","affiliation":[{"name":"School of Mathematics and Computer Science, Northwest Minzu Univsersity, Lanzhou 730030, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lang","family":"Qiao","sequence":"additional","affiliation":[{"name":"Key Laboratory of China\u2019s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wendong","family":"Kang","sequence":"additional","affiliation":[{"name":"Key Laboratory of China\u2019s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinyue","family":"Shi","sequence":"additional","affiliation":[{"name":"Key Laboratory of China\u2019s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,8,16]]},"reference":[{"key":"ref_1","unstructured":"(2021, January 07). Thanka Introduction. Available online: https:\/\/en.wikipedia.org\/wiki."},{"key":"ref_2","first-page":"67","article-title":"Research outline and progress of digital protection on thangka","volume":"2","author":"Wang","year":"2012","journal-title":"Adv. Top. Multimed. Res."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2891","DOI":"10.1109\/TPAMI.2012.162","article-title":"Babytalk: Understanding and generating simple image descriptions","volume":"35","author":"Kulkarni","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Farhadi, A., Hejrati, M., Sadeghi, M.A., Young, P., Rashtchian, C., Hockenmaier, J., and Forsyth, D. (2010, January 5\u201311). Every picture tells a story: Generating sentences from images. Proceedings of the European Conference on Computer Vision, Heraklion, Greece.","DOI":"10.1007\/978-3-642-15561-1_2"},{"key":"ref_5","unstructured":"Ordonez, V., Kulkarni, G., and Berg, T.L. (2011, January 16\u201317). Im2text: Describing images using 1 million captioned photographs. Proceedings of the Advances in Neural Information Processing Systems, Sierra Nevada, Spain."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1162\/tacl_a_00177","article-title":"Grounded compositional semantics for finding and describing images with sentences","volume":"2","author":"Socher","year":"2014","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Mason, R., and Charniak, E. (2014, January 22\u201327). Nonparametric Method for Data-driven Image Captioning. Proceedings of the Meeting of the Association for Computational Linguistics, Baltimore, MD, USA.","DOI":"10.3115\/v1\/P14-2097"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Fang, H., Gupta, S., Iandola, F.N., Srivastava, R.K., Deng, L., Doll\u00e1r, P., Gao, J., He, X., Mitchell, M., and Platt, J.C. (2015, January 7\u201312). From Captions to Visual Concepts and Back. Proceedings of the IEEE Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298754"},{"key":"ref_9","unstructured":"Herdade, S., Kappeler, A., Boakye, K., and Soares, J. (2019, January 8\u201314). Image captioning: Transforming objects into words. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lu, J.S., Xiong, C.M., Parikh, D., and Socher, R. (2017, January 21\u201326). Knowing when to look: Adaptive attention via a visual sentinel for image captioning. Proceedings of the 2017 International Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.345"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Pan, Y., Yao, T., Li, Y., and Mei, T. (2020, January 13\u201319). X-linear attention networks for image captioning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01098"},{"key":"ref_12","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, attend and tell: Neural image caption generation with visual attention. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018, January 18\u201323). Bottom-up and top-down attention for image captioning and visual question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"ref_14","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. arXiv.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Khare, Y., Bagal, V., Mathew, M., Devi, A., Priyakumar, U.D., and Jawahar, C.V. (2021, January 13\u201316). Mmbert: Multimodal bert pretraining for improved medical vqa. Proceedings of the 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), Nice, France.","DOI":"10.1109\/ISBI48211.2021.9434063"},{"key":"ref_17","unstructured":"Lu, J., Batra, D., Parikh, D., and Lee, S. (2019, January 8\u201314). Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Tan, H., and Bansal, M. (2019). Lxmert: Learning cross-modality encoder representations from transformers. arXiv.","DOI":"10.18653\/v1\/D19-1514"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chen, Y.C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J. (2020, January 23\u201328). Uniter: Universal image-text representation learning. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhou, L., Palangi, H., Zhang, L., Hu, H., Corso, J., and Gao, J. (2020, January 7\u201312). Unified vision language pre-training for image captioning and vqa. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.7005"},{"key":"ref_21","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021). Learn-ing transferable visual models from natural language super-vision. arXiv."},{"key":"ref_22","unstructured":"Mokady, R., Hertz, A., and Bermano, A.H. (2021). Clipcap: Clip prefix for image captioning. arXiv."},{"key":"ref_23","unstructured":"Zeng, A., Attarian, M., Ichter, B., Choromanski, K., Wong, A., Welker, S., Tombari, F., Purohit, A., Ryoo, M., and Sindhwani, V. (2022). Socratic models: Composing zero-shot multimodal reasoning with language. arXiv."},{"key":"ref_24","unstructured":"Furlanello, T., Lipton, Z., Tschannen, M., Itti, L., and An Kumar, A. (2018, January 10\u201315). Born-again neural networks. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_25","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2020). Training data-efficient image transformers & distillation through attention. arXiv."},{"key":"ref_26","unstructured":"Tsimpoukelli, M., Menick, J.L., Cabi, S., Eslami, S.M., Vinyals, O., and Hill, F. (2021, January 6\u201314). Multimodal few-shot learning with frozen language models. Proceedings of the Advances in Neural Information Processing Systems, Virtual."},{"key":"ref_27","unstructured":"Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., and Reynolds, M. (2022). Flamingo: A visual language model for few-shot learning. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1162","DOI":"10.1109\/TPAMI.2022.3144984","article-title":"Switchable novel object captioner","volume":"45","author":"Wu","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Petroni, F., Rockt\u00e4schel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A.H., and Riedel, S. (2019). Language models as knowledge bases?. arXiv.","DOI":"10.18653\/v1\/D19-1250"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, X., Ye, Y., and Gupta, A. (2018, January 18\u201322). Zero-shot recognition via semantic embeddings and knowledge graphs. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00717"},{"key":"ref_31","unstructured":"Sanh, V., Debut, L., Chaumond, J., and Wolf, T. (2019). Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Xiang, T., Hospedales, T.M., and Lu, H. (2018, January 18\u201323). Deep mutual learning. Proceedings of the Computer Vision and Pattern Recognition CVPR, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00454"},{"key":"ref_33","unstructured":"Anil, R., Pereyra, G., Passos, A., Ormandi, R., Dahl, G.E., and Hinton, G.E. (May, January 30). Large scale distributed neural network training through online distillation. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1016\/j.neucom.2018.03.030","article-title":"Face detection using deep learning: An improved faster RCNN approach","volume":"299","author":"Sun","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_35","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"423","DOI":"10.1162\/tacl_a_00324","article-title":"How can we know what language models know?","volume":"8","author":"Jiang","year":"2020","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_37","unstructured":"Andrej, K., and Li, F. (2015, January 7\u201312). Deep visual-semantic alignments for generating image descriptions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 7\u201312). Bleu: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_39","unstructured":"Satanjeev, B. (2005, January 25\u201330). METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. Proceedings of the ACL-2005, Ann Arbor, MI, USA."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Lin, C.Y., and Och, F.J. (2004, January 21\u201326). Automatic Evaluation of Machine Translation Quality Using Longest Common Subsequence and Skip-Bigram Statistics. Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics, Barcelona, Spain.","DOI":"10.3115\/1218955.1219032"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Lawrence Zitnick, C., and Parikh, D. (2015, January 7\u201312). Cider: Consensus-based image description evaluation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_42","first-page":"382","article-title":"SPICE: Semantic Propositional Image Caption Evaluation","volume":"11","author":"Anderson","year":"2016","journal-title":"Adapt. Behav."},{"key":"ref_43","unstructured":"Huang, L., Wang, W., Chen, J., and Wei, X.Y. (November, January 27). Attention on attention for image captioning. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., and Wei, F. (2020, January 23\u201328). Oscar: Object semantics aligned pre-training for vision-language tasks. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"ref_45","unstructured":"Jia, X., Gavves, E., Fernando, B., and Tuytelaars, T. (2019). Guiding long-short term memory for image caption generation. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Hu, X., Gan, Z., Wang, J., Yang, Z., Liu, Z., Lu, Y., and Wang, L. (2021). Scaling up vision-language pre-training for image captioning. arXiv.","DOI":"10.1109\/CVPR52688.2022.01745"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/9\/8\/162\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:34:54Z","timestamp":1760128494000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/9\/8\/162"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,16]]},"references-count":46,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2023,8]]}},"alternative-id":["jimaging9080162"],"URL":"https:\/\/doi.org\/10.3390\/jimaging9080162","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,16]]}}}