{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T09:59:12Z","timestamp":1775901552072,"version":"3.50.1"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62276278"],"award-info":[{"award-number":["62276278"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100021171","name":"GuangDong Basic and Applied Basic Research Foundation","doi-asserted-by":"crossref","award":["2022A1515110006 and 2024A1515011259"],"award-info":[{"award-number":["2022A1515110006 and 2024A1515011259"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>\n                    Image captioning is a cross-modal text generation task aimed at understanding the relationships among various objects in an image. Therefore, accurately expressing object\u2013object relations remains a key bottleneck for transformer-based image captioning. Prior methods usually inject semantic and geometric relations once and keep them fixed while only updating visual features, creating a mismatch\u2014evolving visuals vs. frozen relations\u2014that weakens relational guidance and leads to feature entanglement. We propose the Relationship-Experts Transformer (RET), which treats semantic and geometric relations as learnable experts that guide object visual features (students) and co-evolve with them. In RET, we first design the Relationship-Guided Feature Aggregation (RGFA) module, which is analogous to experts-guided student learning, specifically utilizing the relationship kernel (the expert\u2019s knowledge brain) to guide the learning of the object visual features (students). Secondly, we develop the Experts Knowledge Updating (EKU) module, which continuously iterates expert knowledge during training to enhance the expert\u2019s guiding ability over the student. Finally, we design the Student Knowledge Selector (SKS) module to adaptively select object visual features enhanced with different relations under the guidance of semantic and geometric experts to generate descriptive texts embodying semantic and geometric knowledge. Experiments on the MSCOCO dataset demonstrate that our model achieves state-of-the-art performance. All codes are available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/songchuanle-1\/RET\">https:\/\/github.com\/songchuanle-1\/RET<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3796710","type":"journal-article","created":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T20:19:13Z","timestamp":1773778753000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Relationship-Experts Transformer for Image Captioning"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-9893-6195","authenticated-orcid":false,"given":"Chuanle","family":"Song","sequence":"first","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-Sen University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8861-4263","authenticated-orcid":false,"given":"Wenjin","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-Sen University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2021-8015","authenticated-orcid":false,"given":"Han","family":"Jiao","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-Sen University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-8773-7765","authenticated-orcid":false,"given":"Junfeng","family":"Li","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-Sen University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6736-7913","authenticated-orcid":false,"given":"Yihua","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-Sen University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,11]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3671000"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3243725"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3638558"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106560"},{"issue":"3","key":"e_1_3_1_6_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3618301","article-title":"Cross-modality multiple relations learning for knowledge-based visual question answering","volume":"20","author":"Wang Y.","year":"2023","unstructured":"Y. Wang, P. Li, Q. Si, H. Zhang, W. Zang, Z. Lin, and P. Fu. 2023. Cross-modality multiple relations learning for knowledge-based visual question answering. ACM Transactions on Multimedia Computing, Communications, and Applications 20, 3 (2023), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3715141"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458281"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3631356"},{"key":"e_1_3_1_10_2","first-page":"107028","article-title":"Text-guided image restoration and semantic enhancement for text-to-image person retrieval","volume":"180","author":"Liu D.","year":"2024","unstructured":"D. Liu, H. Li, Z. Zhao, Y. Dong, and N. V. Boulgouris. 2024. Text-guided image restoration and semantic enhancement for text-to-image person retrieval. Neural Networks: The Official Journal of the International Neural Network Society 180 (2024), 107028.","journal-title":"Neural Networks: The Official Journal of the International Neural Network Society"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16328"},{"key":"e_1_3_1_14_2","first-page":"1081","volume-title":"Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI \u201922)","author":"Li J.","year":"2022","unstructured":"J. Li, Z. Mao, S. Fang, and H. Li. 2022. ER-SAN: Enhanced-adaptive relation self-attention network for image captioning. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI \u201922), 1081\u20131087."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106710"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106813"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2025.113127"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3769866"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","first-page":"128597","DOI":"10.1016\/j.eswa.2025.128597","article-title":"Dual dynamic transformer for image captioning","volume":"292","author":"Shan C.","year":"2025","unstructured":"C. Shan, C. Song, T. Zou, J. Li, and S. Liu. 2025. Dual dynamic transformer for image captioning. Expert Systems with Applications 292 (2025), 128597.","journal-title":"Expert Systems with Applications"},{"key":"e_1_3_1_20_2","unstructured":"A. Dosovitskiy L. Beyer A. Kolesnikov D. Weissenborn X. Zhai T. Unterthiner M. Dehghani M. Minderer G. Heigold S. Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_21_2","first-page":"1","article-title":"Attention is all you need","volume":"30","author":"Vaswani A.","year":"2017","unstructured":"A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, \u0141. Kaiser, and I. Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017), 1\u201311.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_22_2","first-page":"2048","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xu K.","year":"2015","unstructured":"K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the International Conference on Machine Learning. PMLR, 2048\u20132057."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.503"},{"key":"e_1_3_1_24_2","first-page":"11137","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"Herdade S.","year":"2019","unstructured":"S. Herdade, A. Kappeler, K. Boakye, and J. Soares. 2019. Image captioning: Transforming objects into words. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, 11137\u201311147."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2947482"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00902"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"e_1_3_1_28_2","first-page":"1608","volume-title":"Proceedings of the 31st International Joint Conferences on Artificial Intelligence","volume":"5","author":"Zeng P.","year":"2022","unstructured":"P. Zeng, H. Zhang, J. Song, and L. Gao. 2022. S2 transformer for image captioning. In Proceedings of the 31st International Joint Conferences on Artificial Intelligence, Vol. 5, 1608\u20131614."},{"key":"e_1_3_1_29_2","unstructured":"X. Yang J. Peng Z. Wang H. Xu Q. Ye C. Li M. Yan F. Huang Z. Li and Y. Zhang. 2023. Transforming visual scene graphs to image captions. arXiv:2305.02177. Retrieved from https:\/\/arxiv.org\/abs\/2305.02177"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3181490"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3215861"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3169061"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01034"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.131"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_1_36_2","first-page":"740","volume-title":"Proceedings of the Computer Vision: 13th European Conference (ECCV \u201914)","volume":"13","author":"Lin T.-Y.","year":"2014","unstructured":"T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll\u00e1r, and C. L. Zitnick. 2014. Microsoft COCO: Common objects in context. In Proceedings of the Computer Vision: 13th European Conference (ECCV \u201914), Vol. 13. Springer, 740\u2013755."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_1_38_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni K.","year":"2002","unstructured":"K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311\u2013318."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-3348"},{"key":"e_1_3_1_40_2","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin C.-Y.","year":"2004","unstructured":"C.-Y. Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out. ACL, pp. 74\u201381."},{"key":"e_1_3_1_41_2","first-page":"382","volume-title":"Proceedings of the Computer Vision: 14th European Conference (ECCV \u201916)","volume":"14","author":"Anderson P.","year":"2016","unstructured":"P. Anderson, B. Fernando, M. Johnson, and S. Gould. 2016. SPICE: Semantic propositional image caption evaluation. In Proceedings of the Computer Vision: 14th European Conference (ECCV \u201916), Vol. 14. Springer, 382\u2013398."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00473"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01098"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.109420"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01521"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01749"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3549555.3549585"},{"issue":"2","key":"e_1_3_1_48_2","doi-asserted-by":"crossref","first-page":"1785","DOI":"10.1109\/TNNLS.2022.3185320","article-title":"Adaptive semantic-enhanced transformer for image captioning","volume":"35","author":"Zhang J.","year":"2022","unstructured":"J. Zhang, Z. Fang, H. Sun, and Z. Wang. 2022. Adaptive semantic-enhanced transformer for image captioning. IEEE Transactions on Neural Networks and Learning Systems 35, 2 (2022), 1785\u20131796.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01748"},{"key":"e_1_3_1_50_2","unstructured":"Y. Zhou Z. Hu D. Liu H. Ben and M. Wang. 2022. Compact bidirectional transformer for image captioning. arXiv:2201.01984. Retrieved from https:\/\/arxiv.org\/abs\/2201.01984"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01744"},{"key":"e_1_3_1_52_2","first-page":"19730","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Li J.","year":"2023","unstructured":"J. Li, D. Li, S. Savarese, and S. Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning. PMLR, 19730\u201319742."},{"issue":"3","key":"e_1_3_1_53_2","doi-asserted-by":"crossref","first-page":"2585","DOI":"10.1609\/aaai.v36i3.20160","article-title":"End-to-end transformer based model for image captioning","volume":"36","author":"Wang Y.","year":"2022","unstructured":"Y. Wang, J. Xu, and Y. Sun. End-to-end transformer based model for image captioning. Proceedings of the AAAI Conference on Artificial Intelligence 36, 3 (2022), 2585\u20132594.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01746"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3796710","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T09:19:34Z","timestamp":1775899174000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3796710"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,11]]},"references-count":53,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3796710"],"URL":"https:\/\/doi.org\/10.1145\/3796710","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,11]]},"assertion":[{"value":"2025-04-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-04","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}