{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T19:34:11Z","timestamp":1777577651901,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":33,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,10,21]],"date-time":"2023-10-21T00:00:00Z","timestamp":1697846400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key R&D Program of China","award":["2022YFB3102600"],"award-info":[{"award-number":["2022YFB3102600"]}]},{"name":"NSFC","award":["62272469 and U19B2024"],"award-info":[{"award-number":["62272469 and U19B2024"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,10,21]]},"DOI":"10.1145\/3583780.3614967","type":"proceedings-article","created":{"date-parts":[[2023,10,21]],"date-time":"2023-10-21T07:45:26Z","timestamp":1697874326000},"page":"639-648","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["MGICL: Multi-Grained Interaction Contrastive Learning for Multimodal Named Entity Recognition"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3896-7236","authenticated-orcid":false,"given":"Aibo","family":"Guo","sequence":"first","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6339-0219","authenticated-orcid":false,"given":"Xiang","family":"Zhao","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8643-4683","authenticated-orcid":false,"given":"Zhen","family":"Tan","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6957-3769","authenticated-orcid":false,"given":"Weidong","family":"Xiao","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,10,21]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"A multimodal deep learning approach for named entity recognition from social media. CoRR","author":"Asgari-Chenaghlu Meysam","year":"2020","unstructured":"Meysam Asgari-Chenaghlu , Mohammad-Reza Feizi-Derakhshi , Leili Farzinvash , and Cina Motamed . 2020. A multimodal deep learning approach for named entity recognition from social media. CoRR , Vol. abs\/ 2001 .06888 ( 2020 ). showeprint[arXiv]2001.06888 https:\/\/arxiv.org\/abs\/2001.06888 Meysam Asgari-Chenaghlu, Mohammad-Reza Feizi-Derakhshi, Leili Farzinvash, and Cina Motamed. 2020. A multimodal deep learning approach for named entity recognition from social media. CoRR, Vol. abs\/2001.06888 (2020). showeprint[arXiv]2001.06888 https:\/\/arxiv.org\/abs\/2001.06888"},{"key":"e_1_3_2_2_2_1","volume-title":"DASFAA 2021, Taipei, Taiwan, April 11--14, 2021, Proceedings, Part II (Lecture Notes in Computer Science","volume":"201","author":"Chen Dawei","year":"2021","unstructured":"Dawei Chen , Zhixu Li , Binbin Gu , and Zhigang Chen . 2021 . Multimodal Named Entity Recognition with Image Attributes and Image Knowledge. In Database Systems for Advanced Applications - 26th International Conference , DASFAA 2021, Taipei, Taiwan, April 11--14, 2021, Proceedings, Part II (Lecture Notes in Computer Science , Vol. 12682), Christian S. Jensen, Ee-Peng Lim, De-Nian Yang, Wang-Chien Lee, Vincent S. Tseng, Vana Kalogeraki, Jen-Wei Huang, and Chih-Ya Shen (Eds.). Springer, 186-- 201 . https:\/\/doi.org\/10.1007\/978--3-030--73197--7_12 10.1007\/978--3-030--73197--7_12 Dawei Chen, Zhixu Li, Binbin Gu, and Zhigang Chen. 2021. Multimodal Named Entity Recognition with Image Attributes and Image Knowledge. In Database Systems for Advanced Applications - 26th International Conference, DASFAA 2021, Taipei, Taiwan, April 11--14, 2021, Proceedings, Part II (Lecture Notes in Computer Science, Vol. 12682), Christian S. Jensen, Ee-Peng Lim, De-Nian Yang, Wang-Chien Lee, Vincent S. Tseng, Vana Kalogeraki, Jen-Wei Huang, and Chih-Ya Shen (Eds.). Springer, 186--201. https:\/\/doi.org\/10.1007\/978--3-030--73197--7_12"},{"key":"e_1_3_2_2_3_1","volume-title":"Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph Completion. In SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Chen Xiang","year":"2022","unstructured":"Xiang Chen , Ningyu Zhang , Lei Li , Shumin Deng , Chuanqi Tan , Changliang Xu , Fei Huang , Luo Si , and Huajun Chen . 2022 a. Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph Completion. In SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , Madrid, Spain, July 11 - 15 , 2022, Enrique Amig\u00f3, Pablo Castells, Julio Gonzalo, Ben Carterette, J. Shane Culpepper, and Gabriella Kazai (Eds.). ACM, 904--915. https:\/\/doi.org\/10.1145\/3477495.3531992 10.1145\/3477495.3531992 Xiang Chen, Ningyu Zhang, Lei Li, Shumin Deng, Chuanqi Tan, Changliang Xu, Fei Huang, Luo Si, and Huajun Chen. 2022a. Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph Completion. In SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022, Enrique Amig\u00f3, Pablo Castells, Julio Gonzalo, Ben Carterette, J. Shane Culpepper, and Gabriella Kazai (Eds.). ACM, 904--915. https:\/\/doi.org\/10.1145\/3477495.3531992"},{"key":"e_1_3_2_2_4_1","volume-title":"Good Visual Guidance Makes A Better Extractor: Hierarchical Visual Prefix for Multimodal Entity and Relation Extraction. CoRR","author":"Chen Xiang","year":"2022","unstructured":"Xiang Chen , Ningyu Zhang , Lei Li , Yunzhi Yao , Shumin Deng , Chuanqi Tan , Fei Huang , Luo Si , and Huajun Chen . 2022b. Good Visual Guidance Makes A Better Extractor: Hierarchical Visual Prefix for Multimodal Entity and Relation Extraction. CoRR , Vol. abs\/ 2205 .03521 ( 2022 ). https:\/\/doi.org\/10.48550\/arXiv.2205.03521 showeprint[arXiv]2205.03521 10.48550\/arXiv.2205.03521 Xiang Chen, Ningyu Zhang, Lei Li, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2022b. Good Visual Guidance Makes A Better Extractor: Hierarchical Visual Prefix for Multimodal Entity and Relation Extraction. CoRR, Vol. abs\/2205.03521 (2022). https:\/\/doi.org\/10.48550\/arXiv.2205.03521 showeprint[arXiv]2205.03521"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"e_1_3_2_2_6_1","volume-title":"MNER-QG: An End-to-End MRC framework for Multimodal Named Entity Recognition with Query Grounding. CoRR","author":"Jia Meihuizi","year":"2022","unstructured":"Meihuizi Jia , Lei Shen , Xin Shen , Lejian Liao , Meng Chen , Xiaodong He , Zhendong Chen , and Jiaqi Li. 2022a. MNER-QG: An End-to-End MRC framework for Multimodal Named Entity Recognition with Query Grounding. CoRR , Vol. abs\/ 2211 .14739 ( 2022 ). https:\/\/doi.org\/10.48550\/arXiv.2211.14739 showeprint[arXiv]2211.14739 10.48550\/arXiv.2211.14739 Meihuizi Jia, Lei Shen, Xin Shen, Lejian Liao, Meng Chen, Xiaodong He, Zhendong Chen, and Jiaqi Li. 2022a. MNER-QG: An End-to-End MRC framework for Multimodal Named Entity Recognition with Query Grounding. CoRR, Vol. abs\/2211.14739 (2022). https:\/\/doi.org\/10.48550\/arXiv.2211.14739 showeprint[arXiv]2211.14739"},{"key":"e_1_3_2_2_7_1","volume-title":"Query Prior Matters: A MRC Framework for Multimodal Named Entity Recognition. In MM '22: The 30th ACM International Conference on Multimedia","author":"Jia Meihuizi","year":"2022","unstructured":"Meihuizi Jia , Xin Shen , Lei Shen , Jinhui Pang , Lejian Liao , Yang Song , Meng Chen , and Xiaodong He . 2022 b. Query Prior Matters: A MRC Framework for Multimodal Named Entity Recognition. In MM '22: The 30th ACM International Conference on Multimedia , Lisboa, Portugal, October 10 - 14 , 2022, Jo a o Magalh a es, Alberto Del Bimbo, Shin'ichi Satoh, Nicu Sebe, Xavier Alameda-Pineda, Qin Jin, Vincent Oria, and Laura Toni (Eds.). ACM, 3549--3558. https:\/\/doi.org\/10.1145\/3503161.3548427 10.1145\/3503161.3548427 Meihuizi Jia, Xin Shen, Lei Shen, Jinhui Pang, Lejian Liao, Yang Song, Meng Chen, and Xiaodong He. 2022b. Query Prior Matters: A MRC Framework for Multimodal Named Entity Recognition. In MM '22: The 30th ACM International Conference on Multimedia, Lisboa, Portugal, October 10 - 14, 2022, Jo a o Magalh a es, Alberto Del Bimbo, Shin'ichi Satoh, Nicu Sebe, Xavier Alameda-Pineda, Qin Jin, Vincent Oria, and Laura Toni (Eds.). ACM, 3549--3558. https:\/\/doi.org\/10.1145\/3503161.3548427"},{"key":"e_1_3_2_2_8_1","volume-title":"Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021","author":"Li Junnan","year":"2021","unstructured":"Junnan Li , Ramprasaath R. Selvaraju , Akhilesh Gotmare , Shafiq R. Joty , Caiming Xiong , and Steven Chu-Hong Hoi . 2021 . Align before Fuse: Vision and Language Representation Learning with Momentum Distillation . In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021 , NeurIPS 2021, December 6--14, 2021, virtual, Marc'Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (Eds.). 9694--9705. https:\/\/proceedings.neurips.cc\/paper\/2021\/hash\/505259756244493872b7709a8a01b536-Abstract.html Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. 2021. Align before Fuse: Vision and Language Representation Learning with Momentum Distillation. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6--14, 2021, virtual, Marc'Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (Eds.). 9694--9705. https:\/\/proceedings.neurips.cc\/paper\/2021\/hash\/505259756244493872b7709a8a01b536-Abstract.html"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-021-02546-5"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1185"},{"key":"e_1_3_2_2_11_1","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea, October 12--17","author":"Lu Junyu","year":"2022","unstructured":"Junyu Lu , Dixiang Zhang , Jiaxing Zhang , and Pingjian Zhang . 2022 . Flat Multi-modal Interaction Transformer for Named Entity Recognition . In Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea, October 12--17 , 2022, Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wanner, Key-Sun Choi, Pum-Mo Ryu, Hsin-Hsi Chen, Lucia Donatelli, Heng Ji, Sadao Kurohashi, Patrizia Paggio, Nianwen Xue, Seokhwan Kim, Younggyun Hahm, Zhong He, Tony Kyungil Lee, Enrico Santus, Francis Bond, and Seung-Hoon Na (Eds.). International Committee on Computational Linguistics , 2055--2064. https:\/\/aclanthology.org\/2022.coling-1.179 Junyu Lu, Dixiang Zhang, Jiaxing Zhang, and Pingjian Zhang. 2022. Flat Multi-modal Interaction Transformer for Named Entity Recognition. In Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea, October 12--17, 2022, Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wanner, Key-Sun Choi, Pum-Mo Ryu, Hsin-Hsi Chen, Lucia Donatelli, Heng Ji, Sadao Kurohashi, Patrizia Paggio, Nianwen Xue, Seokhwan Kim, Younggyun Hahm, Zhong He, Tony Kyungil Lee, Enrico Santus, Francis Bond, and Seung-Hoon Na (Eds.). International Committee on Computational Linguistics, 2055--2064. https:\/\/aclanthology.org\/2022.coling-1.179"},{"key":"e_1_3_2_2_12_1","volume-title":"X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval. In MM '22: The 30th ACM International Conference on Multimedia","author":"Ma Yiwei","year":"2022","unstructured":"Yiwei Ma , Guohai Xu , Xiaoshuai Sun , Ming Yan , Ji Zhang , and Rongrong Ji . 2022 . X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval. In MM '22: The 30th ACM International Conference on Multimedia , Lisboa, Portugal, October 10 - 14 , 2022, Jo a o Magalh a es, Alberto Del Bimbo, Shin'ichi Satoh, Nicu Sebe, Xavier Alameda-Pineda, Qin Jin, Vincent Oria, and Laura Toni (Eds.). ACM, 638--647. https:\/\/doi.org\/10.1145\/3503161.3547910 10.1145\/3503161.3547910 Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, and Rongrong Ji. 2022. X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval. In MM '22: The 30th ACM International Conference on Multimedia, Lisboa, Portugal, October 10 - 14, 2022, Jo a o Magalh a es, Alberto Del Bimbo, Shin'ichi Satoh, Nicu Sebe, Xavier Alameda-Pineda, Qin Jin, Vincent Oria, and Laura Toni (Eds.). ACM, 638--647. https:\/\/doi.org\/10.1145\/3503161.3547910"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1078"},{"key":"e_1_3_2_2_14_1","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18--24","volume":"8763","author":"Radford Alec","year":"2021","unstructured":"Alec Radford , Jong Wook Kim , Chris Hallacy , Aditya Ramesh , Gabriel Goh , Sandhini Agarwal , Girish Sastry , Amanda Askell , Pamela Mishkin , Jack Clark , Gretchen Krueger , and Ilya Sutskever . 2021 . Learning Transferable Visual Models From Natural Language Supervision . In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18--24 July 2021, Virtual Event (Proceedings of Machine Learning Research , Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, 8748-- 8763 . http:\/\/proceedings.mlr.press\/v139\/radford21a.html Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18--24 July 2021, Virtual Event (Proceedings of Machine Learning Research, Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, 8748--8763. http:\/\/proceedings.mlr.press\/v139\/radford21a.html"},{"key":"e_1_3_2_2_15_1","volume-title":"An Overview of the Tesseract OCR Engine. In 9th International Conference on Document Analysis and Recognition (ICDAR 2007","author":"Smith R.","year":"2007","unstructured":"R. Smith . 2007 . An Overview of the Tesseract OCR Engine. In 9th International Conference on Document Analysis and Recognition (ICDAR 2007 ), 23--26 September, Curitiba, Paran\u00e1, Brazil. IEEE Computer Society, 629--633. https:\/\/doi.org\/10.1109\/ICDAR. 2007.4376991 10.1109\/ICDAR.2007.4376991 R. Smith. 2007. An Overview of the Tesseract OCR Engine. In 9th International Conference on Document Analysis and Recognition (ICDAR 2007), 23--26 September, Curitiba, Paran\u00e1, Brazil. IEEE Computer Society, 629--633. https:\/\/doi.org\/10.1109\/ICDAR.2007.4376991"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.168"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i15.17633"},{"key":"e_1_3_2_2_18_1","volume-title":"Named Entity and Relation Extraction with Multi-Modal Retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Wang Xinyu","year":"2022","unstructured":"Xinyu Wang , Jiong Cai , Yong Jiang , Pengjun Xie , Kewei Tu , and Wei Lu . 2022 a. Named Entity and Relation Extraction with Multi-Modal Retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2022 , Abu Dhabi, United Arab Emirates, December 7--11 , 2022, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, 5925--5936. https:\/\/aclanthology.org\/2022.findings-emnlp.437 Xinyu Wang, Jiong Cai, Yong Jiang, Pengjun Xie, Kewei Tu, and Wei Lu. 2022a. Named Entity and Relation Extraction with Multi-Modal Retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7--11, 2022, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, 5925--5936. https:\/\/aclanthology.org\/2022.findings-emnlp.437"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.232"},{"key":"e_1_3_2_2_20_1","volume-title":"DASFAA 2022, Virtual Event, April 11--14, 2022, Proceedings, Part III (Lecture Notes in Computer Science","volume":"305","author":"Wang Xuwu","year":"2022","unstructured":"Xuwu Wang , Junfeng Tian , Min Gui , Zhixu Li , Jiabo Ye , Ming Yan , and Yanghua Xiao . 2022 c. PromptMNER: Prompt-Based Entity-Related Visual Clue Extraction and Integration for Multimodal Named Entity Recognition. In Database Systems for Advanced Applications - 27th International Conference , DASFAA 2022, Virtual Event, April 11--14, 2022, Proceedings, Part III (Lecture Notes in Computer Science , Vol. 13247), Arnab Bhattacharya, Janice Lee, Mong Li, Divyakant Agrawal, P. Krishna Reddy, Mukesh K. Mohania, Anirban Mondal, Vikram Goyal, and Rage Uday Kiran (Eds.). Springer, 297-- 305 . https:\/\/doi.org\/10.1007\/978--3-031-00129--1_24 10.1007\/978--3-031-00129--1_24 Xuwu Wang, Junfeng Tian, Min Gui, Zhixu Li, Jiabo Ye, Ming Yan, and Yanghua Xiao. 2022c. PromptMNER: Prompt-Based Entity-Related Visual Clue Extraction and Integration for Multimodal Named Entity Recognition. In Database Systems for Advanced Applications - 27th International Conference, DASFAA 2022, Virtual Event, April 11--14, 2022, Proceedings, Part III (Lecture Notes in Computer Science, Vol. 13247), Arnab Bhattacharya, Janice Lee, Mong Li, Divyakant Agrawal, P. Krishna Reddy, Mukesh K. Mohania, Anirban Mondal, Vikram Goyal, and Rage Uday Kiran (Eds.). Springer, 297--305. https:\/\/doi.org\/10.1007\/978--3-031-00129--1_24"},{"key":"e_1_3_2_2_21_1","volume-title":"CAT-MNER: Multimodal Named Entity Recognition with Knowledge-Refined Cross-Modal Attention. In IEEE International Conference on Multimedia and Expo, ICME 2022","author":"Wang Xuwu","year":"2022","unstructured":"Xuwu Wang , Jiabo Ye , Zhixu Li , Junfeng Tian , Yong Jiang , Ming Yan , Ji Zhang , and Yanghua Xiao . 2022 d. CAT-MNER: Multimodal Named Entity Recognition with Knowledge-Refined Cross-Modal Attention. In IEEE International Conference on Multimedia and Expo, ICME 2022 , Taipei, Taiwan, July 18--22 , 2022. IEEE, 1--6. https:\/\/doi.org\/10.1109\/ICME52920.2022.9859972 10.1109\/ICME52920.2022.9859972 Xuwu Wang, Jiabo Ye, Zhixu Li, Junfeng Tian, Yong Jiang, Ming Yan, Ji Zhang, and Yanghua Xiao. 2022d. CAT-MNER: Multimodal Named Entity Recognition with Knowledge-Refined Cross-Modal Attention. In IEEE International Conference on Multimedia and Expo, ICME 2022, Taipei, Taiwan, July 18--22, 2022. IEEE, 1--6. https:\/\/doi.org\/10.1109\/ICME52920.2022.9859972"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413650"},{"key":"#cr-split#-e_1_3_2_2_23_1.1","doi-asserted-by":"crossref","unstructured":"Bo Xu Shizhou Huang Chaofeng Sha and Hongya Wang. 2022. MAF: A General Matching and Alignment Framework for Multimodal Named Entity Recognition. In WSDM '22: The Fifteenth ACM International Conference on Web Search and Data Mining Virtual Event \/ Tempe AZ USA February 21 - 25 2022 K. Selcuk Candan Huan Liu Leman Akoglu Xin Luna Dong and Jiliang Tang (Eds.). ACM 1215--1223. https:\/\/doi.org\/10.1145\/3488560.3498475 10.1145\/3488560.3498475","DOI":"10.1145\/3488560.3498475"},{"key":"#cr-split#-e_1_3_2_2_23_1.2","doi-asserted-by":"crossref","unstructured":"Bo Xu Shizhou Huang Chaofeng Sha and Hongya Wang. 2022. MAF: A General Matching and Alignment Framework for Multimodal Named Entity Recognition. In WSDM '22: The Fifteenth ACM International Conference on Web Search and Data Mining Virtual Event \/ Tempe AZ USA February 21 - 25 2022 K. Selcuk Candan Huan Liu Leman Akoglu Xin Luna Dong and Jiliang Tang (Eds.). ACM 1215--1223. https:\/\/doi.org\/10.1145\/3488560.3498475","DOI":"10.1145\/3488560.3498475"},{"key":"e_1_3_2_2_24_1","volume-title":"TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment. In 2021 IEEE\/CVF International Conference on Computer Vision, ICCV 2021","author":"Yang Jianwei","year":"2021","unstructured":"Jianwei Yang , Yonatan Bisk , and Jianfeng Gao . 2021 . TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment. In 2021 IEEE\/CVF International Conference on Computer Vision, ICCV 2021 , Montreal, QC, Canada, October 10--17 , 2021. IEEE, 11542--11552. https:\/\/doi.org\/10.1109\/ICCV48922.2021.01136 10.1109\/ICCV48922.2021.01136 Jianwei Yang, Yonatan Bisk, and Jianfeng Gao. 2021. TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment. In 2021 IEEE\/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10--17, 2021. IEEE, 11542--11552. https:\/\/doi.org\/10.1109\/ICCV48922.2021.01136"},{"key":"e_1_3_2_2_25_1","volume-title":"FILIP: Fine-grained Interactive Language-Image Pre-Training. In The Tenth International Conference on Learning Representations, ICLR 2022","author":"Yao Lewei","year":"2022","unstructured":"Lewei Yao , Runhui Huang , Lu Hou , Guansong Lu , Minzhe Niu , Hang Xu , Xiaodan Liang , Zhenguo Li , Xin Jiang , and Chunjing Xu . 2022 . FILIP: Fine-grained Interactive Language-Image Pre-Training. In The Tenth International Conference on Learning Representations, ICLR 2022 , Virtual Event, April 25--29 , 2022. OpenReview.net. https:\/\/openreview.net\/forum?id=cpDhcsEDC2 Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2022. FILIP: Fine-grained Interactive Language-Image Pre-Training. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25--29, 2022. OpenReview.net. https:\/\/openreview.net\/forum?id=cpDhcsEDC2"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.306"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i16.17687"},{"key":"e_1_3_2_2_28_1","volume-title":"VinVL: Revisiting Visual Representations in Vision-Language Models. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021","author":"Zhang Pengchuan","year":"2021","unstructured":"Pengchuan Zhang , Xiujun Li , Xiaowei Hu , Jianwei Yang , Lei Zhang , Lijuan Wang , Yejin Choi , and Jianfeng Gao . 2021 a. VinVL: Revisiting Visual Representations in Vision-Language Models. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021 , virtual, June 19 --25 , 2021. Computer Vision Foundation \/ IEEE, 5579--5588. https:\/\/doi.org\/10.1109\/CVPR46437.2021.00553 10.1109\/CVPR46437.2021.00553 Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021a. VinVL: Revisiting Visual Representations in Vision-Language Models. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19--25, 2021. Computer Vision Foundation \/ IEEE, 5579--5588. https:\/\/doi.org\/10.1109\/CVPR46437.2021.00553"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11962"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548228"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3013398"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2022.3224228"}],"event":{"name":"CIKM '23: The 32nd ACM International Conference on Information and Knowledge Management","location":"Birmingham United Kingdom","acronym":"CIKM '23","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGIR ACM Special Interest Group on Information Retrieval"]},"container-title":["Proceedings of the 32nd ACM International Conference on Information and Knowledge Management"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3583780.3614967","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3583780.3614967","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:44Z","timestamp":1750178204000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3583780.3614967"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,21]]},"references-count":33,"alternative-id":["10.1145\/3583780.3614967","10.1145\/3583780"],"URL":"https:\/\/doi.org\/10.1145\/3583780.3614967","relation":{},"subject":[],"published":{"date-parts":[[2023,10,21]]},"assertion":[{"value":"2023-10-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}