{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:53:51Z","timestamp":1782312831603,"version":"3.54.5"},"reference-count":93,"publisher":"Association for Computing Machinery (ACM)","issue":"7","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U21B2024, 62472303"],"award-info":[{"award-number":["U21B2024, 62472303"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>\n                    News captioning aims to describe an image with its news article body as input. It greatly relies on a set of detected named entities, including real-world people, organizations, and places. This article exploits commonsense knowledge to understand named entities for news captioning. By \u201cunderstand,\u201d we mean correlating the news content with commonsense in the wild, which helps an agent to (1) distinguish semantically similar named entities and (2) describe named entities using words outside of training corpora. Our approach consists of three modules: (a)\n                    <jats:italic toggle=\"yes\">Filter Module<\/jats:italic>\n                    aims to clarify the commonsense concerning a named entity from two aspects:\n                    <jats:italic toggle=\"yes\">what does it mean<\/jats:italic>\n                    ? and\n                    <jats:italic toggle=\"yes\">what is it related to<\/jats:italic>\n                    ?, which divide the commonsense into\n                    <jats:italic toggle=\"yes\">explanatory knowledge<\/jats:italic>\n                    and\n                    <jats:italic toggle=\"yes\">relevant knowledge<\/jats:italic>\n                    , respectively. (b)\n                    <jats:italic toggle=\"yes\">Distinguish Module<\/jats:italic>\n                    aggregates\n                    <jats:italic toggle=\"yes\">explanatory knowledge<\/jats:italic>\n                    from\n                    <jats:italic toggle=\"yes\">node-degree<\/jats:italic>\n                    ,\n                    <jats:italic toggle=\"yes\">dependency<\/jats:italic>\n                    , and\n                    <jats:italic toggle=\"yes\">distinguish<\/jats:italic>\n                    three aspects to distinguish semantically similar named entities. (c)\u00a0\n                    <jats:italic toggle=\"yes\">Enrich Module<\/jats:italic>\n                    attaches\n                    <jats:italic toggle=\"yes\">relevant knowledge<\/jats:italic>\n                    to named entities to enrich the entity description by commonsense information (e.g., identity and social position). Finally, all of information is integrated into the large multimodal model to generate the news caption. Extensive experiments on two challenging datasets (i.e., GoodNews and NYTimes) demonstrate the superiority of our method. Ablation studies and visualization further validate its effectiveness in understanding named entities.\n                  <\/jats:p>","DOI":"10.1145\/3769085","type":"journal-article","created":{"date-parts":[[2025,10,3]],"date-time":"2025-10-03T16:09:24Z","timestamp":1759507764000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["How to Understand Named Entities: Using Commonsense for News Captioning"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8613-7651","authenticated-orcid":false,"given":"Shenyuan","family":"Zhang","sequence":"first","affiliation":[{"name":"Tianjin University, Tianjin, China and People\u2019s Daily, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7526-4356","authenticated-orcid":false,"given":"Ning","family":"Xu","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3027-2114","authenticated-orcid":false,"given":"Yanhui","family":"Wang","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-7218-6302","authenticated-orcid":false,"given":"Tongle","family":"Ma","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1633-7575","authenticated-orcid":false,"given":"Wu","family":"Liu","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3047-6520","authenticated-orcid":false,"given":"Jinlin","family":"Guo","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2792-6588","authenticated-orcid":false,"given":"Chao","family":"Xue","sequence":"additional","affiliation":[{"name":"Tiandy Technologies Co., Ltd., Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5755-9145","authenticated-orcid":false,"given":"An-An","family":"Liu","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","first-page":"6077","DOI":"10.1109\/CVPR.2018.00636","volume-title":"Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201918)","author":"Anderson Peter","year":"2018","unstructured":"Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201918), 6077\u20136086."},{"key":"e_1_3_2_3_2","first-page":"313","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) (NAACL-HLT \u201918)","author":"Annervaz K. M.","year":"2018","unstructured":"K. M. Annervaz, Somnath Basu Roy Chowdhury, and Ambedkar Dukkipati. 2018. Learning beyond datasets: Knowledge graph augmented neural networks for natural language processing. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) (NAACL-HLT \u201918). Marilyn A. Walker, Heng Ji, and Amanda Stent (Eds.), 313\u2013322."},{"key":"e_1_3_2_4_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)","author":"Bahdanau Dzmitry","year":"2015","unstructured":"Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)."},{"key":"e_1_3_2_5_2","first-page":"65","volume-title":"Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization, 65\u201372."},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","unstructured":"Huixia Ben Yingwei Pan Yehao Li Ting Yao Richang Hong Meng Wang and Tao Mei. 2022. Unpaired image captioning with semantic-constrained self-learning. IEEE Transactions on Multimedia 24 (2022) 904\u2013916.","DOI":"10.1109\/TMM.2021.3060948"},{"key":"e_1_3_2_7_2","first-page":"12466","volume-title":"Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201919)","author":"Biten Ali Furkan","year":"2019","unstructured":"Ali Furkan Biten, Llu\u00eds G\u00f3mez, Mar\u00e7al Rusi\u00f1ol, and Dimosthenis Karatzas. 2019. Good news, everyone! context driven entity-aware captioning for news images. In Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201919), 12466\u201312475."},{"key":"e_1_3_2_8_2","first-page":"2787","volume-title":"Proceedings of the 26th International Conference on Neural Information Processing Systems (NeurIPS \u201913)","author":"Bordes Antoine","year":"2013","unstructured":"Antoine Bordes, Nicolas Usunier, Alberto Garc\u00eda-Dur\u00e1n, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. In Proceedings of the 26th International Conference on Neural Information Processing Systems (NeurIPS \u201913), 2787\u20132795."},{"key":"e_1_3_2_9_2","first-page":"35","volume-title":"Proceedings of the 3rd ACM International Conference on Multimedia (ACM MM)","author":"Brown Martin G.","year":"1995","unstructured":"Martin G. Brown, J. T. Foote, Gareth J. F. Jones, Karen Sparck Jones, and Steve J. Young. 1995. Automatic content-based retrieval of broadcast news. In Proceedings of the 3rd ACM International Conference on Multimedia (ACM MM), 35\u201343."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3178844"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2016.11.004"},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"6298","DOI":"10.1109\/CVPR.2017.667","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Chen Long","year":"2017","unstructured":"Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua. 2017. SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 6298\u20136306."},{"key":"e_1_3_2_13_2","unstructured":"Zhe Chen Weiyun Wang Yue Cao Yangzhou Liu Zhangwei Gao Erfei Cui Jinguo Zhu Shenglong Ye Hao Tian Zhaoyang Liu et al. 2024. Expanding performance boundaries of open-source multimodal models with model data and test-time scaling. arXiv:2412.05271. Retrieved from https:\/\/arxiv.org\/abs\/2412.05271"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-024-4231-5"},{"key":"e_1_3_2_15_2","first-page":"24185","volume-title":"Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201924)","author":"Chen Zhe","year":"2024","unstructured":"Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201924), 24185\u201324198."},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","first-page":"10173","DOI":"10.18653\/v1\/2023.emnlp-main.629","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923)","author":"Chowdhury Sanjoy","year":"2023","unstructured":"Sanjoy Chowdhury, Sayan Nag, and Dinesh Manocha. 2023. APoLLo : Unified adapter and prompt learning for vision language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923), 10173\u201310187."},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Chaorui Deng Ning Ding Mingkui Tan and Qi Wu. 2020. Length-controllable image captioning. In Proceedings of the European Conference on Computer Vision (ECCV \u201920) Vol. 12358 712\u2013729.","DOI":"10.1007\/978-3-030-58601-0_42"},{"key":"e_1_3_2_18_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201919)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201919), 4171\u20134186."},{"key":"e_1_3_2_19_2","first-page":"3113","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201923)","author":"Fei Junjie","year":"2023","unstructured":"Junjie Fei, Teng Wang, Jinrui Zhang, Zhenyu He, Chengjie Wang, and Feng Zheng. 2023. Transferable decoding with visual entities for zero-shot image captioning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201923), 3113\u20133123."},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"Yansong Feng and Mirella Lapata. 2013. Automatic caption generation for news images. IEEE Transactions on Pattern Analysis and Machine Intelligence 35 4 (2013) 797\u2013812.","DOI":"10.1109\/TPAMI.2012.118"},{"key":"e_1_3_2_21_2","volume-title":"Proceedings of the ACL Workshop on Text Summarization (WAS \u201918)","author":"Flick Carlos","year":"2018","unstructured":"Carlos Flick. 2018. Rouge: A package for automatic evaluation of summaries. In Proceedings of the ACL Workshop on Text Summarization (WAS \u201918)."},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","first-page":"1815","DOI":"10.1609\/aaai.v38i3.27950","article-title":"LAMM: Label alignment for multi-modal prompt learning","volume":"38","author":"Gao Jingsheng","year":"2024","unstructured":"Jingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu, Ke Ji, Mingye Xie, Ting Liu, and Yuzhuo Fu. 2024. LAMM: Label alignment for multi-modal prompt learning. Proceedings of the AAAI Conference on Artificial Intelligence 38 (2024), 1815\u20131823.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3120867"},{"issue":"5","key":"e_1_3_2_24_2","first-page":"1112","article-title":"Hierarchical LSTMs with adaptive attention for visual captioning","volume":"42","author":"Gao Lianli","year":"2020","unstructured":"Lianli Gao, Xiangpeng Li, Jingkuan Song, and Heng Tao Shen. 2020. Hierarchical LSTMs with adaptive attention for visual captioning. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 5 (2020), 1112\u20131131.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/s44267-024-00067-6"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"Matt Gardner Joel Grus Mark Neumann Oyvind Tafjord Pradeep Dasigi Nelson F. Liu Matthew E. Peters Michael Schmitz and Luke Zettlemoyer. 2018. AllenNLP: A deep semantic natural language processing platform. arXiv:1803.07640. Retrieved from https:\/\/arxiv.org\/abs\/1803.07640","DOI":"10.18653\/v1\/W18-2501"},{"key":"e_1_3_2_27_2","first-page":"10324","volume-title":"Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920)","author":"Guo Longteng","year":"2020","unstructured":"Longteng Guo, Jing Liu, Xinxin Zhu, Peng Yao, Shichen Lu, and Hanqing Lu. 2020. Normalized and geometry-aware self-attention network for image captioning. In Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920), 10324\u201310333."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"issue":"1","key":"e_1_3_2_29_2","first-page":"411","article-title":"spaCy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing","volume":"7","author":"Honnibal Matthew","year":"2017","unstructured":"Matthew Honnibal and Ines Montani. 2017. spaCy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing. To Appear 7, 1 (2017), 411\u2013420. Manuscript submitted for review.","journal-title":"To Appear"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413576"},{"key":"e_1_3_2_31_2","first-page":"4633","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201919)","author":"Huang Lun","year":"2019","unstructured":"Lun Huang, Wenmin Wang, Jie Chen, and Xiaoyong Wei. 2019. Attention on attention for image captioning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201919), 4633\u20134642."},{"key":"e_1_3_2_32_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)."},{"key":"e_1_3_2_33_2","first-page":"15259","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201923)","author":"Kornblith Simon","year":"2023","unstructured":"Simon Kornblith, Lala Li, Zirui Wang, and Thao Nguyen. 2023. Guiding image captioning models toward more specific captions. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201923), 15259\u201315269."},{"key":"e_1_3_2_34_2","first-page":"11039","volume-title":"Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201923)","author":"Kuo Chia-Wen","year":"2023","unstructured":"Chia-Wen Kuo and Zsolt Kira. 2023. HAAV: Hierarchical aggregation of augmented views for image captioning. In Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201923), 11039\u201311049."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0957-4174(02)00015-5"},{"key":"e_1_3_2_36_2","unstructured":"Bo Li Yuanhan Zhang Dong Guo Renrui Zhang Feng Li Hao Zhang Kaichen Zhang Yanwei Li Ziwei Liu and Chunyuan Li. 2024. LLaVA-OneVision: Easy visual task transfer. arXiv:2408.03326. Retrieved from https:\/\/arxiv.org\/abs\/2408.03326"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772758"},{"key":"e_1_3_2_38_2","first-page":"2829","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Lin Bill Yuchen","year":"2019","unstructured":"Bill Yuchen Lin, Xinyue Chen, Jamin Chen, and Xiang Ren. 2019. KagNet: Knowledge-aware graph networks for commonsense reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2829\u20132839."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_40_2","first-page":"6761","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201921)","author":"Liu Fuxiao","year":"2021","unstructured":"Fuxiao Liu, Yinghan Wang, Tianlu Wang, and Vicente Ordonez. 2021. Visual news: Benchmark and challenges in news image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201921), 6761\u20136771."},{"key":"e_1_3_2_41_2","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923)","author":"Liu Haotian","year":"2023","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923)."},{"key":"e_1_3_2_42_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_43_2","first-page":"3242","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Lu Jiasen","year":"2017","unstructured":"Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2017. Knowing when to look: Adaptive attention via a visual sentinel for image captioning. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 3242\u20133250."},{"key":"e_1_3_2_44_2","first-page":"8449","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201920)","author":"Lv Shangwen","year":"2020","unstructured":"Shangwen Lv, Daya Guo, Jingjing Xu, Duyu Tang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, and Songlin Hu. 2020. Graph-based reasoning over heterogeneous external knowledge for commonsense question answering. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI \u201920), 8449\u20138456."},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","first-page":"821","DOI":"10.18653\/v1\/P18-1076","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (ACL \u201918)","author":"Mihaylov Todor","year":"2018","unstructured":"Todor Mihaylov and Anette Frank. 2018. Knowledgeable reader: Enhancing cloze-style reading comprehension with external commonsense knowledge. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (ACL \u201918). Iryna Gurevych and Yusuke Miyao (Eds.), 821\u2013832."},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","unstructured":"Yujie Mo Yuhuan Chen Yajie Lei Liang Peng Xiaoshuang Shi Changan Yuan and Xiaofeng Zhu. 2023. Multiplex graph representation learning via dual correlation reduction. IEEE Transactions on Knowledge and Data Engineering 35 12 (2023) 12814\u201312827.","DOI":"10.1109\/TKDE.2023.3268069"},{"issue":"4","key":"e_1_3_2_47_2","first-page":"279","article-title":"An approach based on combination of features for automatic news retrieval","volume":"14","author":"Moradi Mohammad","year":"2019","unstructured":"Mohammad Moradi, Elham Ghanbari, Mehrdad Maeen, and Sasan Harifi. 2019. An approach based on combination of features for automatic news retrieval. Journal of Information and Computing Science 14, 4 (2019), 279\u2013290.","journal-title":"Journal of Information and Computing Science"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548180"},{"key":"e_1_3_2_49_2","first-page":"10968","volume-title":"Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920)","author":"Pan Yingwei","year":"2020","unstructured":"Yingwei Pan, Ting Yao, Yehao Li, and Tao Mei. 2020. X-linear attention networks for image captioning. In Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920), 10968\u201310977."},{"key":"e_1_3_2_50_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL \u201902)","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL \u201902), 311\u2013318."},{"key":"e_1_3_2_51_2","volume-title":"Proceedings of the 31st Conference on Neural Information Processing Systems (NeurIPS \u201917)","author":"Paszke Adam","year":"2017","unstructured":"Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017. Automatic differentiation in PyTorch. In Proceedings of the 31st Conference on Neural Information Processing Systems (NeurIPS \u201917)."},{"key":"e_1_3_2_52_2","first-page":"1194","volume-title":"Proceedings of the 29th International World Wide Web Conference (WWW \u201920)","author":"Patro Gourab K.","year":"2020","unstructured":"Gourab K. Patro, Arpita Biswas, Niloy Ganguly, Krishna P. Gummadi, and Abhijnan Chakraborty. 2020. FairRec: Two-sided fairness for personalized recommendations in two-sided platforms. In Proceedings of the 29th International World Wide Web Conference (WWW \u201920), 1194\u20131204."},{"issue":"12","key":"e_1_3_2_53_2","first-page":"8609","article-title":"GRLC: Graph representation learning with constraints","volume":"34","author":"Peng Liang","year":"2023","unstructured":"Liang Peng, Yujie Mo, Jie Xu, Jialie Shen, Xiaoshuang Shi, Xiaoxiao Li, Heng Tao Shen, and Xiaofeng Zhu. 2023. GRLC: Graph representation learning with constraints. IEEE Transactions on Neural Networks and Learning Systems 34, 12 (2023), 8609\u20138622.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_2_54_2","unstructured":"Tingyu Qu Tinne Tuytelaars and Marie-Francine Moens. 2023. Visually-Aware Context Modeling for News Image Captioning. arXiv:2308.08325. Retrieved form https:\/\/arxiv.org\/abs\/2308.08325"},{"key":"e_1_3_2_55_2","doi-asserted-by":"crossref","unstructured":"Arnau Ramisa Fei Yan Francesc Moreno-Noguer and Krystian Mikolajczyk. 2018. BreakingNews: Article annotation by image and text processing. IEEE Transactions on Pattern Analysis and Machine Intelligence 40 5 (2018) 1072\u20131085.","DOI":"10.1109\/TPAMI.2017.2721945"},{"key":"e_1_3_2_56_2","unstructured":"Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An incremental improvement. arXiv:1804.02767. Retrieved from https:\/\/arxiv.org\/abs\/1804.02767"},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","first-page":"1179","DOI":"10.1109\/CVPR.2017.131","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Rennie Steven J.","year":"2017","unstructured":"Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel. 2017. Self-critical sequence training for image captioning. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 1179\u20131195."},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1162"},{"key":"e_1_3_2_60_2","doi-asserted-by":"crossref","unstructured":"Zhuang Shao Jungong Han Demetris Marnerides and Kurt Debattista. 2022. Region-object relation-aware dense captioning via transformer. IEEE Transactions on Neural Networks and Learning Systems 36 3 (2022) 4184\u20134195.","DOI":"10.1109\/TNNLS.2022.3152990"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1238"},{"key":"e_1_3_2_62_2","doi-asserted-by":"crossref","first-page":"4615","DOI":"10.18653\/v1\/2020.emnlp-main.373","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920)","author":"Shwartz Vered","year":"2020","unstructured":"Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Unsupervised commonsense question answering with self-talk. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920), 4615\u20134629."},{"key":"e_1_3_2_63_2","doi-asserted-by":"crossref","unstructured":"Jingkuan Song Yuyu Guo Lianli Gao Xuelong Li Alan Hanjalic and Heng Tao Shen. 2019. From deterministic to generative: Multimodal stochastic RNNs for video captioning. IEEE Transactions on Neural Networks and Learning Systems 30 10 (2019) 3047\u20133058.","DOI":"10.1109\/TNNLS.2018.2851077"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11164"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2628585"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","first-page":"4593","DOI":"10.18653\/v1\/P19-1452","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL \u201919)","author":"Tenney Ian","year":"2019","unstructured":"Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL \u201919), 4593\u20134601."},{"key":"e_1_3_2_67_2","first-page":"13032","volume-title":"Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920)","author":"Tran Alasdair","year":"2020","unstructured":"Alasdair Tran, Alexander Patrick Mathews, and Lexing Xie. 2020. Transform and tell: Entity-aware news image captioning. In Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920), 13032\u201313042."},{"key":"e_1_3_2_68_2","first-page":"5998","volume-title":"Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS \u201917)","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS \u201917), 5998\u20136008."},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","first-page":"4566","DOI":"10.1109\/CVPR.2015.7299087","volume-title":"Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201915)","author":"Vedantam Ramakrishna","year":"2015","unstructured":"Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. CIDEr: Consensus-based image description evaluation. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201915), 4566\u20134575."},{"key":"e_1_3_2_70_2","doi-asserted-by":"crossref","first-page":"3156","DOI":"10.1109\/CVPR.2015.7298935","volume-title":"Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201915)","author":"Vinyals Oriol","year":"2015","unstructured":"Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201915), 3156\u20133164."},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612245"},{"key":"e_1_3_2_72_2","doi-asserted-by":"crossref","unstructured":"Ting Wang Weidong Chen Yuanhe Tian Yan Song and Zhendong Mao. 2023. Improving image captioning via predicting structured concepts. arXiv:2311.08223. Retrieved from https:\/\/arxiv.org\/abs\/2311.08223","DOI":"10.18653\/v1\/2023.emnlp-main.25"},{"key":"e_1_3_2_73_2","unstructured":"Weiyun Wang Zhe Chen Wenhai Wang Yue Cao Yangzhou Liu Zhangwei Gao Jinguo Zhu Xizhou Zhu Lewei Lu Yu Qiao and Jifeng Dai. 2024. Enhancing the reasoning ability of multimodal large language models via mixed preference optimization. arXiv:2411.10442. Retrieved from https:\/\/arxiv.org\/abs\/2411.10442"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3121062"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/2806416.2806533"},{"key":"e_1_3_2_76_2","first-page":"481","volume-title":"Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data (SIGMOD \u201912)","author":"Wu Wentao","year":"2012","unstructured":"Wentao Wu, Hongsong Li, Haixun Wang, and Kenny Qili Zhu. 2012. Probase: A probabilistic taxonomy for text understanding. In Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data (SIGMOD \u201912), 481\u2013492."},{"key":"e_1_3_2_77_2","unstructured":"Zhiyu Wu Xiaokang Chen Zizheng Pan Xingchao Liu Wen Liu Damai Dai Huazuo Gao Yiyang Ma Chengyue Wu Bingxuan Wang et al. 2024. DeepSeek-VL2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv:2412.10302. Retrieved from https:\/\/arxiv.org\/abs\/2412.10302"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3155795"},{"key":"e_1_3_2_79_2","first-page":"2048","volume-title":"Proceedings of the 32nd International Conference on Machine Learning (ICML \u201915)","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the 32nd International Conference on Machine Learning (ICML \u201915), 2048\u20132057."},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3165934"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3067449"},{"key":"e_1_3_2_82_2","doi-asserted-by":"crossref","first-page":"1436","DOI":"10.18653\/v1\/P17-1132","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (ACL \u201917)","author":"Yang Bishan","year":"2017","unstructured":"Bishan Yang and Tom M. Mitchell. 2017. Leveraging knowledge bases in LSTMs for improving machine reading. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (ACL \u201917). Regina Barzilay and Min-Yen Kan (Eds.), 1436\u20131446."},{"key":"e_1_3_2_83_2","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201921)","author":"Yang Xuewen","year":"2021","unstructured":"Xuewen Yang, Svebor Karaman, Joel R. Tetreault, and Alex Jaimes. 2021. Journalistic guidelines aware news image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201921)."},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.45"},{"key":"e_1_3_2_85_2","first-page":"4651","volume-title":"Proceedings of the 2016 Conference on Computer Vision and Pattern Recognition (CVPR \u201916)","author":"You Quanzeng","year":"2016","unstructured":"Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. 2016. Image captioning with semantic attention. In Proceedings of the 2016 Conference on Computer Vision and Pattern Recognition (CVPR \u201916), 4651\u20134659."},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611891"},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2947482"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547883"},{"key":"e_1_3_2_89_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612480"},{"key":"e_1_3_2_90_2","doi-asserted-by":"crossref","unstructured":"Kaipeng Zhang Zhanpeng Zhang Zhifeng Li and Yu Qiao. 2016. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Processing Letters 23 10 (2016) 1499\u20131503.","DOI":"10.1109\/LSP.2016.2603342"},{"key":"e_1_3_2_91_2","first-page":"15465","volume-title":"Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201921)","author":"Zhang Xuying","year":"2021","unstructured":"Xuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji, Yiyi Zhou, Yongjian Wu, Feiyue Huang, and Rongrong Ji. 2021. RSTNet: Captioning with adaptive attention on visual and non-visual words. In Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201921), 15465\u201315474."},{"key":"e_1_3_2_92_2","unstructured":"Wentian Zhao Yao Hu Heda Wang Xinxiao Wu and Jiebo Luo. 2021. Boosting entity-aware image captioning with multi-modal knowledge graph. arXiv:2107.11970. Retrieved from https:\/\/arxiv.org\/abs\/2107.11970"},{"key":"e_1_3_2_93_2","doi-asserted-by":"crossref","unstructured":"Mingyang Zhou Grace Luo Anna Rohrbach and Zhou Yu. 2022. Focus! Relevant and sufficient context selection for news image captioning. arXiv:2212.00843. Retrieved from https:\/\/arxiv.org\/abs\/2212.00843","DOI":"10.18653\/v1\/2022.findings-emnlp.450"},{"key":"e_1_3_2_94_2","unstructured":"Jinguo Zhu Weiyun Wang Zhe Chen Zhaoyang Liu Shenglong Ye Lixin Gu Hao Tian Yuchen Duan Weijie Su Jie Shao et al. 2025. InternVL3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv:2504.10479. Retrieved. Retrieved from https:\/\/arxiv.org\/abs\/2504.10479"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3769085","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:41:48Z","timestamp":1782312108000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3769085"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":93,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3769085"],"URL":"https:\/\/doi.org\/10.1145\/3769085","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-01-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-16","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}