{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T17:57:51Z","timestamp":1770227871383,"version":"3.49.0"},"publisher-location":"New York, NY, USA","reference-count":59,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,7,6]],"date-time":"2022-07-06T00:00:00Z","timestamp":1657065600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Major Scientific and Technological Projects of CNPC","award":["ZD2019-183-008"],"award-info":[{"award-number":["ZD2019-183-008"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61671480"],"award-info":[{"award-number":["61671480"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Open Program of the National Laboratory of Pattern Recognition","award":["202000009"],"award-info":[{"award-number":["202000009"]}]},{"name":"National Natural Science Foundation","award":["61902093"],"award-info":[{"award-number":["61902093"]}]},{"name":"Shenzhen Foundational Research Funding","award":["20200805173048001"],"award-info":[{"award-number":["20200805173048001"]}]},{"name":"Natural Science Foundation of Guangdong","award":["2020A1515010652"],"award-info":[{"award-number":["2020A1515010652"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,7,6]]},"DOI":"10.1145\/3477495.3531715","type":"proceedings-article","created":{"date-parts":[[2022,7,7]],"date-time":"2022-07-07T15:12:13Z","timestamp":1657206733000},"page":"2727-2737","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":19,"title":["Where Does the Performance Improvement Come From?"],"prefix":"10.1145","author":[{"given":"Jun","family":"Rao","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fei","family":"Wang","sequence":"additional","affiliation":[{"name":"China University of Petroleum (East China), qingdao, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liang","family":"Ding","sequence":"additional","affiliation":[{"name":"JD Explore Academy, beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuhan","family":"Qi","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yibing","family":"Zhan","sequence":"additional","affiliation":[{"name":"JD Explore Academy, beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weifeng","family":"Liu","sequence":"additional","affiliation":[{"name":"China University of Petroleum (East China), qingdao, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dacheng","family":"Tao","sequence":"additional","affiliation":[{"name":"JD Explore Academy, beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,7,7]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Peter Anderson Xiaodong He Chris Buehler Damien Teney Mark Johnson Stephen Gould and Lei Zhang. 2018. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In CVPR.  Peter Anderson Xiaodong He Chris Buehler Damien Teney Mark Johnson Stephen Gould and Lei Zhang. 2018. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In CVPR.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_1_2_1","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2015. Neural Machine Translation by Jointly Learning to Align and Translate. In ICLR.  Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2015. Neural Machine Translation by Jointly Learning to Align and Translate. In ICLR."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Federico Bianchi and Dirk Hovy. 2021. On the Gap between Adoption and Understanding in NLP. In Findings of ACL.  Federico Bianchi and Dirk Hovy. 2021. On the Gap between Adoption and Understanding in NLP. In Findings of ACL.","DOI":"10.18653\/v1\/2021.findings-acl.340"},{"key":"e_1_3_2_1_4_1","volume-title":"Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu.","author":"Chen Yen-Chun","year":"2020","unstructured":"Yen-Chun Chen , Linjie Li , Licheng Yu , Ahmed El Kholy , Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 . UNITER : UNiversal Image-TExt Representation Learning. In ECCV. Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. UNITER: UNiversal Image-TExt Representation Learning. In ECCV."},{"key":"e_1_3_2_1_5_1","volume-title":"Cross-lingual language model pretraining. NeurIPS","author":"Conneau Alexis","year":"2019","unstructured":"Alexis Conneau and Guillaume Lample . 2019. Cross-lingual language model pretraining. NeurIPS ( 2019 ). Alexis Conneau and Guillaume Lample. 2019. Cross-lingual language model pretraining. NeurIPS (2019)."},{"key":"e_1_3_2_1_6_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . In NAACL-HLT. Association for Computational Linguistics . Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT. Association for Computational Linguistics."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Haiwen Diao Ying Zhang Lin Ma and Huchuan Lu. 2021. Similarity Reasoning and Filtration for Image-Text Matching. In AAAI.  Haiwen Diao Ying Zhang Lin Ma and Huchuan Lu. 2021. Similarity Reasoning and Filtration for Image-Text Matching. In AAAI.","DOI":"10.1609\/aaai.v35i2.16209"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Liang Ding Longyue Wang Xuebo Liu Derek F Wong Dacheng Tao and Zhaopeng Tu. 2021. Progressive Multi-Granularity Training for NonAutoregressive Translation. In Fingdings of ACL.  Liang Ding Longyue Wang Xuebo Liu Derek F Wong Dacheng Tao and Zhaopeng Tu. 2021. Progressive Multi-Granularity Training for NonAutoregressive Translation. In Fingdings of ACL.","DOI":"10.18653\/v1\/2021.findings-acl.247"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Liang Ding Longyue Wang and Dacheng Tao. 2020. Self-attention with crosslingual position representation. In ACL.  Liang Ding Longyue Wang and Dacheng Tao. 2020. Self-attention with crosslingual position representation. In ACL.","DOI":"10.18653\/v1\/2020.acl-main.153"},{"key":"e_1_3_2_1_10_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly etal 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR.  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR."},{"key":"e_1_3_2_1_11_1","volume-title":"Jamie Ryan Kiros, and Sanja Fidler","author":"Faghri Fartash","year":"2018","unstructured":"Fartash Faghri , David J. Fleet , Jamie Ryan Kiros, and Sanja Fidler . 2018 . VSE++: Improving Visual-Semantic Embeddings with Hard Negatives. In BMVC. Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018. VSE++: Improving Visual-Semantic Embeddings with Hard Negatives. In BMVC."},{"key":"e_1_3_2_1_12_1","unstructured":"Dehong Gao Linbo Jin Ben Chen Minghui Qiu Peng Li Yi Wei Yi Hu and Hao Wang. 2020. FashionBERT: Text and Image Matching with Adaptive Loss for Cross-modal Retrieval. In SIGIR.  Dehong Gao Linbo Jin Ben Chen Minghui Qiu Peng Li Yi Wei Yi Hu and Hao Wang. 2020. FashionBERT: Text and Image Matching with Adaptive Loss for Cross-modal Retrieval. In SIGIR."},{"key":"e_1_3_2_1_13_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR.  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR."},{"key":"e_1_3_2_1_14_1","unstructured":"Peng Hu Liangli Zhen Dezhong Peng and Pei Liu. 2019. Scalable Deep Multimodal Learning for Cross-Modal Retrieval. In SIGIR.  Peng Hu Liangli Zhen Dezhong Peng and Pei Liu. 2019. Scalable Deep Multimodal Learning for Cross-Modal Retrieval. In SIGIR."},{"key":"e_1_3_2_1_15_1","unstructured":"Zhibin Hu Yongsheng Luo Jiong Lin Yan Yan and Jian Chen. 2019. Multi-Level Visual-Semantic Alignments with Relation-Wise Dual Attention Network for Image and Text Matching. In IJCAI.  Zhibin Hu Yongsheng Luo Jiong Lin Yan Yan and Jian Chen. 2019. Multi-Level Visual-Semantic Alignments with Relation-Wise Dual Attention Network for Image and Text Matching. In IJCAI."},{"key":"e_1_3_2_1_16_1","volume-title":"Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers. CoRR abs\/2004.00849","author":"Huang Zhicheng","year":"2020","unstructured":"Zhicheng Huang , Zhaoyang Zeng , Bei Liu , Dongmei Fu , and Jianlong Fu. 2020. Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers. CoRR abs\/2004.00849 ( 2020 ). Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020. Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers. CoRR abs\/2004.00849 (2020)."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Kalervo J\u00e4rvelin and Jaana Kek\u00e4l\u00e4inen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Trans. Inf. Syst. (2002).  Kalervo J\u00e4rvelin and Jaana Kek\u00e4l\u00e4inen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Trans. Inf. Syst. (2002).","DOI":"10.1145\/582415.582418"},{"key":"e_1_3_2_1_18_1","unstructured":"Chao Jia Yinfei Yang Ye Xia Yi-Ting Chen Zarana Parekh Hieu Pham Quoc V. Le Yun-Hsuan Sung Zhen Li and Tom Duerig. 2021. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. In ICML.  Chao Jia Yinfei Yang Ye Xia Yi-Ting Chen Zarana Parekh Hieu Pham Quoc V. Le Yun-Hsuan Sung Zhen Li and Tom Duerig. 2021. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. In ICML."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Huaizu Jiang Ishan Misra Marcus Rohrbach Erik G. Learned-Miller and Xinlei Chen. 2020. In Defense of Grid Features for Visual Question Answering. In CVPR.  Huaizu Jiang Ishan Misra Marcus Rohrbach Erik G. Learned-Miller and Xinlei Chen. 2020. In Defense of Grid Features for Visual Question Answering. In CVPR.","DOI":"10.1109\/CVPR42600.2020.01028"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2598339"},{"key":"e_1_3_2_1_21_1","unstructured":"Wonjae Kim Bokyung Son and Ildoo Kim. 2021. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. In ICML.  Wonjae Kim Bokyung Son and Ildoo Kim. 2021. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. In ICML."},{"key":"e_1_3_2_1_22_1","volume-title":"Almost Surely Stable Deep Dynamics. NeurIPS","author":"Lawrence Nathan","year":"2020","unstructured":"Nathan Lawrence , Philip Loewen , Michael Forbes , Johan Backstrom , and Bhushan Gopaluni . 2020. Almost Surely Stable Deep Dynamics. NeurIPS ( 2020 ). Nathan Lawrence, Philip Loewen, Michael Forbes, Johan Backstrom, and Bhushan Gopaluni. 2020. Almost Surely Stable Deep Dynamics. NeurIPS (2020)."},{"key":"e_1_3_2_1_23_1","unstructured":"Kuang-Huei Lee Xi Chen Gang Hua Houdong Hu and Xiaodong He. 2018. Stacked cross attention for image-text matching. In ECCV.  Kuang-Huei Lee Xi Chen Gang Hua Houdong Hu and Xiaodong He. 2018. Stacked cross attention for image-text matching. In ECCV."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Gen Li Nan Duan Yuejian Fang Ming Gong and Daxin Jiang. 2020. Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training. In AAAI.  Gen Li Nan Duan Yuejian Fang Ming Gong and Daxin Jiang. 2020. Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training. In AAAI.","DOI":"10.1609\/aaai.v34i07.6795"},{"key":"e_1_3_2_1_25_1","unstructured":"Kunpeng Li Yulun Zhang Kai Li Yuanyuan Li and Yun Fu. 2019. Visual Semantic Reasoning for Image-Text Matching. In ICCV.  Kunpeng Li Yulun Zhang Kai Li Yuanyuan Li and Yun Fu. 2019. Visual Semantic Reasoning for Image-Text Matching. In ICCV."},{"key":"e_1_3_2_1_26_1","volume-title":"VisualBERT: A Simple and Performant Baseline for Vision and Language. CoRR abs\/1908.03557","author":"Li Liunian Harold","year":"2019","unstructured":"Liunian Harold Li , Mark Yatskar , Da Yin , Cho-Jui Hsieh , and Kai-Wei Chang . 2019. VisualBERT: A Simple and Performant Baseline for Vision and Language. CoRR abs\/1908.03557 ( 2019 ). Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. VisualBERT: A Simple and Performant Baseline for Vision and Language. CoRR abs\/1908.03557 (2019)."},{"key":"e_1_3_2_1_27_1","volume-title":"UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning. In ACL\/IJCNLP.","author":"Li Wei","year":"2021","unstructured":"Wei Li , Can Gao , Guocheng Niu , Xinyan Xiao , Hao Liu , Jiachen Liu , Hua Wu , and Haifeng Wang . 2021 . UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning. In ACL\/IJCNLP. Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. 2021. UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning. In ACL\/IJCNLP."},{"key":"e_1_3_2_1_28_1","volume-title":"Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks. In ECCV.","author":"Li Xiujun","year":"2020","unstructured":"Xiujun Li , Xi Yin , Chunyuan Li , Pengchuan Zhang , Xiaowei Hu , Lei Zhang , Lijuan Wang , Houdong Hu , Li Dong , Furu Wei , Yejin Choi , and Jianfeng Gao . 2020 . Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks. In ECCV. Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. 2020. Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks. In ECCV."},{"key":"e_1_3_2_1_29_1","volume-title":"Belongie","author":"Lin Tsung-Yi","year":"2017","unstructured":"Tsung-Yi Lin , Piotr Doll\u00e1r , Ross B. Girshick , Kaiming He , Bharath Hariharan , and Serge J . Belongie . 2017 . Feature Pyramid Networks for Object Detection. In CVPR. 936--944. Tsung-Yi Lin, Piotr Doll\u00e1r, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. 2017. Feature Pyramid Networks for Object Detection. In CVPR. 936--944."},{"key":"e_1_3_2_1_30_1","unstructured":"Tsung-Yi Lin Michael Maire Serge J. Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In ECCV.  Tsung-Yi Lin Michael Maire Serge J. Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In ECCV."},{"key":"e_1_3_2_1_31_1","unstructured":"Jiasen Lu Dhruv Batra Devi Parikh and Stefan Lee. 2019. ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In NeurIPS.  Jiasen Lu Dhruv Batra Devi Parikh and Stefan Lee. 2019. ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In NeurIPS."},{"key":"e_1_3_2_1_32_1","unstructured":"Xiaopeng Lu Tiancheng Zhao and Kyusong Lee. 2021. VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words. In ACL.  Xiaopeng Lu Tiancheng Zhao and Kyusong Lee. 2021. VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words. In ACL."},{"key":"e_1_3_2_1_33_1","volume-title":"Yi Tay, Liam Fedus, Thibault F\u00e9vry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel.","author":"Narang Sharan","year":"2021","unstructured":"Sharan Narang , Hyung Won Chung , Yi Tay, Liam Fedus, Thibault F\u00e9vry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel. 2021 . Do Transformer Modifications Transfer Across Implementations and Applications?. In EMNLP. Sharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault F\u00e9vry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel. 2021. Do Transformer Modifications Transfer Across Implementations and Applications?. In EMNLP."},{"key":"e_1_3_2_1_34_1","volume-title":"Berg","author":"Ordonez Vicente","year":"2011","unstructured":"Vicente Ordonez , Girish Kulkarni , and Tamara L . Berg . 2011 . Im2Text: Describing Images Using 1 Million Captioned Photographs. In NeurIPS. Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. 2011. Im2Text: Describing Images Using 1 Million Captioned Photographs. In NeurIPS."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"crossref","unstructured":"James Philbin Ondrej Chum Michael Isard Josef Sivic and Andrew Zisserman. 2007. Object retrieval with large vocabularies and fast spatial matching. In CVPR.  James Philbin Ondrej Chum Michael Isard Josef Sivic and Andrew Zisserman. 2007. Object retrieval with large vocabularies and fast spatial matching. In CVPR.","DOI":"10.1109\/CVPR.2007.383172"},{"key":"e_1_3_2_1_36_1","unstructured":"Leigang Qu Meng Liu Da Cao Liqiang Nie and Qi Tian. 2020. Context-Aware Multi-View Summarization Network for Image-Text Matching. In ACM Multimedia.  Leigang Qu Meng Liu Da Cao Liqiang Nie and Qi Tian. 2020. Context-Aware Multi-View Summarization Network for Image-Text Matching. In ACM Multimedia."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"crossref","unstructured":"Jun Rao Tao Qian Shuhan Qi Yulin Wu Qing Liao and Xuan Wang. 2021. Student Can Also be a Good Teacher: Extracting Knowledge from Vision-andLanguage Model for Cross-Modal Retrieval. In CIKM.  Jun Rao Tao Qian Shuhan Qi Yulin Wu Qing Liao and Xuan Wang. 2021. Student Can Also be a Good Teacher: Extracting Knowledge from Vision-andLanguage Model for Cross-Modal Retrieval. In CIKM.","DOI":"10.1145\/3459637.3482194"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000019"},{"key":"e_1_3_2_1_40_1","unstructured":"Joshua David Robinson Ching-Yao Chuang Suvrit Sra and Stefanie Jegelka. 2021. Contrastive Learning with Hard Negative Samples. In ICLR.  Joshua David Robinson Ching-Yao Chuang Suvrit Sra and Stefanie Jegelka. 2021. Contrastive Learning with Hard Negative Samples. In ICLR."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/78.650093"},{"key":"e_1_3_2_1_42_1","volume-title":"Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In ACL.","author":"Sharma Piyush","year":"2018","unstructured":"Piyush Sharma , Nan Ding , Sebastian Goodman , and Radu Soricut . 2018 . Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In ACL. Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In ACL."},{"key":"e_1_3_2_1_43_1","unstructured":"Weijie Su Xizhou Zhu Yue Cao Bin Li Lewei Lu Furu Wei and Jifeng Dai. 2020. VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In ICLR.  Weijie Su Xizhou Zhu Yue Cao Bin Li Lewei Lu Furu Wei and Jifeng Dai. 2020. VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In ICLR."},{"key":"e_1_3_2_1_44_1","volume-title":"Representation Learning with Contrastive Predictive Coding. CoRR abs\/1807.03748","author":"van den Oord A\u00e4ron","year":"2018","unstructured":"A\u00e4ron van den Oord , Yazhe Li , and Oriol Vinyals . 2018. Representation Learning with Contrastive Predictive Coding. CoRR abs\/1807.03748 ( 2018 ). A\u00e4ron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation Learning with Contrastive Predictive Coding. CoRR abs\/1807.03748 (2018)."},{"key":"e_1_3_2_1_45_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In NeurIPS.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In NeurIPS."},{"key":"e_1_3_2_1_46_1","unstructured":"Petar Velickovic Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Li\u00f2 and Yoshua Bengio. 2018. Graph Attention Networks. In ICLR.  Petar Velickovic Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Li\u00f2 and Yoshua Bengio. 2018. Graph Attention Networks. In ICLR."},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"crossref","unstructured":"Jian Wang Feng Zhou Shilei Wen Xiao Liu and Yuanqing Lin. 2017. Deep Metric Learning with Angular Loss. In ICCV.  Jian Wang Feng Zhou Shilei Wen Xiao Liu and Yuanqing Lin. 2017. Deep Metric Learning with Angular Loss. In ICCV.","DOI":"10.1109\/ICCV.2017.283"},{"key":"e_1_3_2_1_48_1","volume-title":"CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval. In ICCV.","author":"Wang Zihao","year":"2019","unstructured":"Zihao Wang , Xihui Liu , Hongsheng Li , Lu Sheng , Junjie Yan , Xiaogang Wang , and Jing Shao . 2019 . CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval. In ICCV. Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao. 2019. CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval. In ICCV."},{"key":"e_1_3_2_1_49_1","volume-title":"Slua: A super lightweight unsupervised word alignment model via cross-lingual contrastive learning. arXiv preprint","author":"Wu Di","year":"2021","unstructured":"Di Wu , Liang Ding , Shuo Yang , and Dacheng Tao . 2021 . Slua: A super lightweight unsupervised word alignment model via cross-lingual contrastive learning. arXiv preprint (2021). Di Wu, Liang Ding, Shuo Yang, and Dacheng Tao. 2021. Slua: A super lightweight unsupervised word alignment model via cross-lingual contrastive learning. arXiv preprint (2021)."},{"key":"e_1_3_2_1_50_1","unstructured":"Yiling Wu Shuhui Wang Guoli Song and Qingming Huang. 2019. Learning Fragment Self-Attention Embeddings for Image-Text Matching. In ACM Multimedia.  Yiling Wu Shuhui Wang Guoli Song and Qingming Huang. 2019. Learning Fragment Self-Attention Embeddings for Image-Text Matching. In ACM Multimedia."},{"key":"e_1_3_2_1_51_1","volume-title":"Vitae: Vision transformer advanced by exploring intrinsic inductive bias. NeurIPS","author":"Xu Yufei","year":"2021","unstructured":"Yufei Xu , Qiming Zhang , Jing Zhang , and Dacheng Tao . 2021 . Vitae: Vision transformer advanced by exploring intrinsic inductive bias. NeurIPS (2021). Yufei Xu, Qiming Zhang, Jing Zhang, and Dacheng Tao. 2021. Vitae: Vision transformer advanced by exploring intrinsic inductive bias. NeurIPS (2021)."},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00166"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"crossref","unstructured":"Jun Yu Hao Zhou Yibing Zhan and Dacheng Tao. 2021. Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal Hashing. In AAAI.  Jun Yu Hao Zhou Yibing Zhan and Dacheng Tao. 2021. Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal Hashing. In AAAI.","DOI":"10.1609\/aaai.v35i5.16592"},{"key":"e_1_3_2_1_54_1","unstructured":"Changtong Zan Liang Ding Li Shen Yu Cao Weifeng Liu and Dacheng Tao. 2022. Bridging Cross-Lingual Gaps During Leveraging the Multilingual Sequenceto-Sequence Pretraining for Text Generation. In arXiv preprint.  Changtong Zan Liang Ding Li Shen Yu Cao Weifeng Liu and Dacheng Tao. 2022. Bridging Cross-Lingual Gaps During Leveraging the Multilingual Sequenceto-Sequence Pretraining for Text Generation. In arXiv preprint."},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"crossref","unstructured":"Yibing Zhan Jun Yu Zhou Yu Rong Zhang Dacheng Tao and Qi Tian. 2018. Comprehensive distance-preserving autoencoders for cross-modal retrieval. In ACM Multimedia.  Yibing Zhan Jun Yu Zhou Yu Rong Zhang Dacheng Tao and Qi Tian. 2018. Comprehensive distance-preserving autoencoders for cross-modal retrieval. In ACM Multimedia.","DOI":"10.1145\/3240508.3240607"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"crossref","unstructured":"Bowen Zhang Hexiang Hu Vihan Jain Eugene Ie and Fei Sha. 2020. Learning to Represent Image and Text with Denotation Graph. In EMNLP.  Bowen Zhang Hexiang Hu Vihan Jain Eugene Ie and Fei Sha. 2020. Learning to Represent Image and Text with Denotation Graph. In EMNLP.","DOI":"10.18653\/v1\/2020.emnlp-main.60"},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"crossref","unstructured":"Pengchuan Zhang Xiujun Li Xiaowei Hu Jianwei Yang Lei Zhang Lijuan Wang Yejin Choi and Jianfeng Gao. 2021. VinVL: Revisiting Visual Representations in Vision-Language Models. In CVPR.  Pengchuan Zhang Xiujun Li Xiaowei Hu Jianwei Yang Lei Zhang Lijuan Wang Yejin Choi and Jianfeng Gao. 2021. VinVL: Revisiting Visual Representations in Vision-Language Models. In CVPR.","DOI":"10.1109\/CVPR46437.2021.00553"},{"key":"e_1_3_2_1_58_1","volume-title":"Li","author":"Zhang Qi","year":"2020","unstructured":"Qi Zhang , Zhen Lei , Zhaoxiang Zhang , and Stan Z . Li . 2020 . Context-Aware Attention Network for Image-Text Retrieval. In CVPR. Qi Zhang, Zhen Lei, Zhaoxiang Zhang, and Stan Z. Li. 2020. Context-Aware Attention Network for Image-Text Retrieval. In CVPR."},{"key":"e_1_3_2_1_59_1","volume-title":"ViTAEv2: Vision Transformer Advanced by Exploring Inductive Bias for Image Recognition and Beyond. arXiv preprint","author":"Zhang Qiming","year":"2022","unstructured":"Qiming Zhang , Yufei Xu , Jing Zhang , and Dacheng Tao . 2022. ViTAEv2: Vision Transformer Advanced by Exploring Inductive Bias for Image Recognition and Beyond. arXiv preprint ( 2022 ). Qiming Zhang, Yufei Xu, Jing Zhang, and Dacheng Tao. 2022. ViTAEv2: Vision Transformer Advanced by Exploring Inductive Bias for Image Recognition and Beyond. arXiv preprint (2022)."}],"event":{"name":"SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","location":"Madrid Spain","acronym":"SIGIR '22","sponsor":["SIGIR ACM Special Interest Group on Information Retrieval"]},"container-title":["Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3477495.3531715","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3477495.3531715","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:07Z","timestamp":1750186927000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3477495.3531715"}},"subtitle":["- A Reproducibility Concern about Image-Text Retrieval"],"short-title":[],"issued":{"date-parts":[[2022,7,6]]},"references-count":59,"alternative-id":["10.1145\/3477495.3531715","10.1145\/3477495"],"URL":"https:\/\/doi.org\/10.1145\/3477495.3531715","relation":{},"subject":[],"published":{"date-parts":[[2022,7,6]]},"assertion":[{"value":"2022-07-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}