{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T15:22:27Z","timestamp":1759332147009,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":42,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"the National Natural Science Foundation of China","award":["91838303"],"award-info":[{"award-number":["91838303"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,17]]},"DOI":"10.1145\/3474085.3481539","type":"proceedings-article","created":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T08:03:42Z","timestamp":1634544222000},"page":"1129-1137","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["AsyNCE"],"prefix":"10.1145","author":[{"given":"Cheng","family":"Da","sequence":"first","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanhao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yun","family":"Zheng","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pan","family":"Pan","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yinghui","family":"Xu","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chunhong","family":"Pan","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413614"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Jingyuan Chen Xinpeng Chen Lin Ma Zequn Jie and Tat-Seng Chua. 2018. Temporally Grounding Natural Sentence in Video. In EMNLP. 162--171.  Jingyuan Chen Xinpeng Chen Lin Ma Zequn Jie and Tat-Seng Chua. 2018. Temporally Grounding Natural Sentence in Video. In EMNLP. 162--171.","DOI":"10.18653\/v1\/D18-1015"},{"key":"e_1_3_2_1_3_1","first-page":"1597","article-title":"c. A Simple Framework for Contrastive Learning of Visual Representations","volume":"119","author":"Chen Ting","year":"2020","journal-title":"ICML"},{"volume-title":"2020 b. Improved Baselines with Momentum Contrastive Learning. CoRR","year":"2020","author":"Chen Xinlei","key":"e_1_3_2_1_4_1"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Zhenfang Chen Lin Ma Wenhan Luo and Kwan-Yee Kenneth Wong. 2019. Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video. In ACL.  Zhenfang Chen Lin Ma Wenhan Luo and Kwan-Yee Kenneth Wong. 2019. Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video. In ACL.","DOI":"10.18653\/v1\/P19-1183"},{"volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT. 4171--4186.","year":"2019","author":"Devlin Jacob","key":"e_1_3_2_1_6_1"},{"key":"e_1_3_2_1_7_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR. 770--778.  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR. 770--778."},{"volume-title":"Russell","year":"2017","author":"Hendricks Lisa Anne","key":"e_1_3_2_1_8_1"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"crossref","unstructured":"De-An Huang Shyamal Buch Lucio M. Dery Animesh Garg Li Fei-Fei and Juan Carlos Niebles. 2018. Finding \"It\": Weakly-Supervised Reference-Aware Visual Grounding in Instructional Videos. In CVPR. 5948--5957.  De-An Huang Shyamal Buch Lucio M. Dery Animesh Garg Li Fei-Fei and Juan Carlos Niebles. 2018. Finding \"It\": Weakly-Supervised Reference-Aware Visual Grounding in Instructional Videos. In CVPR. 5948--5957.","DOI":"10.1109\/CVPR.2018.00623"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413902"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2598339"},{"key":"e_1_3_2_1_12_1","unstructured":"Jingyu Liu Liang Wang and Ming-Hsuan Yang. 2017. Referring Expression Generation and Comprehension via Attributes. In ICCV. 4866--4874.  Jingyu Liu Liang Wang and Ming-Hsuan Yang. 2017. Referring Expression Generation and Comprehension via Attributes. In ICCV. 4866--4874."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454289"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Antoine Miech Jean-Baptiste Alayrac Lucas Smaira Ivan Laptev Josef Sivic and Andrew Zisserman. 2020. End-to-End Learning of Visual Representations From Uncurated Instructional Videos. In CVPR. 9876--9886.  Antoine Miech Jean-Baptiste Alayrac Lucas Smaira Ivan Laptev Josef Sivic and Andrew Zisserman. 2020. End-to-End Learning of Visual Representations From Uncurated Instructional Videos. In CVPR. 9876--9886.","DOI":"10.1109\/CVPR42600.2020.00990"},{"volume-title":"Manning","year":"2014","author":"Pennington Jeffrey","key":"e_1_3_2_1_15_1"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0965-7"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414053"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Anna Rohrbach Marcus Rohrbach Ronghang Hu Trevor Darrell and Bernt Schiele. 2016. Grounding of Textual Phrases in Images by Reconstruction. In ECCV Bastian Leibe Jiri Matas Nicu Sebe and Max Welling (Eds.). 817--834.  Anna Rohrbach Marcus Rohrbach Ronghang Hu Trevor Darrell and Bernt Schiele. 2016. Grounding of Textual Phrases in Images by Reconstruction. In ECCV Bastian Leibe Jiri Matas Nicu Sebe and Max Welling (Eds.). 817--834.","DOI":"10.1007\/978-3-319-46448-0_49"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Arka Sadhu Kan Chen and Ram Nevatia. 2020. Video Object Grounding Using Semantic Roles in Language Description. In CVPR. 10414--10424.  Arka Sadhu Kan Chen and Ram Nevatia. 2020. Video Object Grounding Using Semantic Roles in Language Description. In CVPR. 10414--10424.","DOI":"10.1109\/CVPR42600.2020.01043"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Jing Shi Jia Xu Boqing Gong and Chenliang Xu. 2019. Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering Losses. In CVPR. 10444--10452.  Jing Shi Jia Xu Boqing Gong and Chenliang Xu. 2019. Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering Losses. In CVPR. 10444--10452.","DOI":"10.1109\/CVPR.2019.01069"},{"key":"e_1_3_2_1_22_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR.  Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR."},{"volume-title":"Vl-bert: Pre-training of generic visual-linguistic representations. In ICLR. 1--12.","year":"2020","author":"Su Weijie","key":"e_1_3_2_1_23_1"},{"volume-title":"2019 a. Learning Video Representations using Contrastive Bidirectional Transformer. CoRR","year":"2019","author":"Sun Chen","key":"e_1_3_2_1_24_1"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Chen Sun Austin Myers Carl Vondrick Kevin Murphy and Cordelia Schmid. 2019 b. VideoBERT: A Joint Model for Video and Language Representation Learning. In ICCV. 7463--7472.  Chen Sun Austin Myers Carl Vondrick Kevin Murphy and Cordelia Schmid. 2019 b. VideoBERT: A Joint Model for Video and Language Representation Learning. In ICCV. 7463--7472.","DOI":"10.1109\/ICCV.2019.00756"},{"volume-title":"LXMERT: Learning Cross-Modality Encoder Representations from Transformers. In EMNLP. 5099--5110.","year":"2019","author":"Tan Hao","key":"e_1_3_2_1_26_1"},{"volume-title":"J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov.","year":"2019","author":"Hubert Tsai Yao-Hung","key":"e_1_3_2_1_27_1"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"volume-title":"Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation. CoRR","year":"1951","author":"Wang Liwei","key":"e_1_3_2_1_29_1"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413862"},{"key":"e_1_3_2_1_31_1","first-page":"447","article-title":"Visual Relation Grounding in Videos","volume":"12351","author":"Xiao Junbin","year":"2020","journal-title":"ECCV"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401151"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413610"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-017-1018-6"},{"volume-title":"Berg","year":"2018","author":"Yu Licheng","key":"e_1_3_2_1_35_1"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Songyang Zhang Houwen Peng Jianlong Fu and Jiebo Luo. 2020 a. Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language. In AAAI. 12870--12877.  Songyang Zhang Houwen Peng Jianlong Fu and Jiebo Luo. 2020 a. Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language. In AAAI. 12870--12877.","DOI":"10.1609\/aaai.v34i07.6984"},{"key":"e_1_3_2_1_37_1","unstructured":"Zhu Zhang Zhou Zhao Zhijie Lin Jieming Zhu and Xiuqiang He. 2020 b. Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language Grounding. In NeurIPS.  Zhu Zhang Zhou Zhao Zhijie Lin Jieming Zhu and Xiuqiang He. 2020 b. Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language Grounding. In NeurIPS."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"crossref","unstructured":"Zhu Zhang Zhou Zhao Yang Zhao Qi Wang Huasheng Liu and Lianli Gao. 2020 c. Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences. In CVPR. 10665--10674.  Zhu Zhang Zhou Zhao Yang Zhao Qi Wang Huasheng Liu and Lianli Gao. 2020 c. Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences. In CVPR. 10665--10674.","DOI":"10.1109\/CVPR42600.2020.01068"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Chen Zheng Quan Guo and Parisa Kordjamshidi. 2020. Cross-Modality Relevance for Reasoning on Language and Vision. In ACL. 7642--7651.  Chen Zheng Quan Guo and Parisa Kordjamshidi. 2020. Cross-Modality Relevance for Reasoning on Language and Vision. In ACL. 7642--7651.","DOI":"10.18653\/v1\/2020.acl-main.683"},{"volume-title":"Corso","year":"2018","author":"Zhou Luowei","key":"e_1_3_2_1_40_1"},{"volume-title":"Corso","year":"2018","author":"Zhou Luowei","key":"e_1_3_2_1_41_1"},{"key":"e_1_3_2_1_42_1","unstructured":"Linchao Zhu and Yi Yang. 2020. ActBERT: Learning Global-Local Video-Text Representations. In CVPR. 8743--8752.  Linchao Zhu and Yi Yang. 2020. ActBERT: Learning Global-Local Video-Text Representations. In CVPR. 8743--8752."}],"event":{"name":"MM '21: ACM Multimedia Conference","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Virtual Event China","acronym":"MM '21"},"container-title":["Proceedings of the 29th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3481539","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3481539","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:17:35Z","timestamp":1750191455000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3481539"}},"subtitle":["Disentangling False-Positives for Weakly-Supervised Video Grounding"],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":42,"alternative-id":["10.1145\/3474085.3481539","10.1145\/3474085"],"URL":"https:\/\/doi.org\/10.1145\/3474085.3481539","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}