{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:17:08Z","timestamp":1750220228698,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":48,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,6,27]],"date-time":"2022-06-27T00:00:00Z","timestamp":1656288000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Beijing Natural Science Foundation","award":["Z190001"],"award-info":[{"award-number":["Z190001"]}]},{"name":"Science and Technology Innovation 2030 - New Generation Artificial Intelligence of China","award":["2020AAA0104401"],"award-info":[{"award-number":["2020AAA0104401"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,6,27]]},"DOI":"10.1145\/3512527.3531371","type":"proceedings-article","created":{"date-parts":[[2022,6,23]],"date-time":"2022-06-23T22:23:32Z","timestamp":1656023012000},"page":"342-350","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Video2Subtitle: Matching Weakly-Synchronized Sequences via Dynamic Temporal Alignment"],"prefix":"10.1145","author":[{"given":"Ben","family":"Xue","sequence":"first","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chenchen","family":"Liu","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yadong","family":"Mu","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,6,27]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Youtube-8m: A large-scale video classification benchmark. CoRR","author":"Abu-El-Haija Sami","year":"2016","unstructured":"Sami Abu-El-Haija , Nisarg Kothari , Joonseok Lee , Paul Natsev , George Toderici , Balakrishnan Varadarajan , and Sudheendra Vijayanarasimhan . 2016. Youtube-8m: A large-scale video classification benchmark. CoRR ( 2016 ), 1609.08675. Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. 2016. Youtube-8m: A large-scale video classification benchmark. CoRR (2016), 1609.08675."},{"key":"e_1_3_2_2_2_1","volume-title":"International Conference on Machine Learning","author":"Andrew Galen","year":"2013","unstructured":"Galen Andrew , Raman Arora , Jeff Bilmes , and Karen Livescu . 2013 . Deep canonical correlation analysis . International Conference on Machine Learning (2013), 1247--1255. Galen Andrew, Raman Arora, Jeff Bilmes, and Karen Livescu. 2013. Deep canonical correlation analysis. International Conference on Machine Learning (2013), 1247--1255."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.14778\/2350229.2350266"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01065"},{"key":"e_1_3_2_2_6_1","volume-title":"Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. Advances in Neural Information Processing Systems Workshop","author":"Chung Junyoung","year":"2014","unstructured":"Junyoung Chung , \u00c7aglar G\u00fcl\u00e7ehre , KyungHyun Cho , and Yoshua Bengio . 2014 . Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. Advances in Neural Information Processing Systems Workshop (2014). Junyoung Chung, \u00c7aglar G\u00fcl\u00e7ehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. Advances in Neural Information Processing Systems Workshop (2014)."},{"key":"e_1_3_2_2_7_1","volume-title":"Asian Conference on Computer Vision","author":"Chung Joon Son","year":"2016","unstructured":"Joon Son Chung and Andrew Zisserman . 2016 . Lip reading in the wild . Asian Conference on Computer Vision (2016), 87--103. Joon Son Chung and Andrew Zisserman. 2016. Lip reading in the wild. Asian Conference on Computer Vision (2016), 87--103."},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Joon Son Chung and AP Zisserman. 2017. Lip reading in profile. (2017).  Joon Son Chung and AP Zisserman. 2017. Lip reading in profile. (2017).","DOI":"10.5244\/C.31.155"},{"key":"#cr-split#-e_1_3_2_2_9_1.1","unstructured":"Martin Cooke Jon Barker Stuart Cunningham and Xu Shao. 2006. The Grid Audio-Visual Speech Corpus. https:\/\/doi.org\/10.5281\/zenodo.3625687 10.5281\/zenodo.3625687"},{"key":"#cr-split#-e_1_3_2_2_9_1.2","unstructured":"Martin Cooke Jon Barker Stuart Cunningham and Xu Shao. 2006. The Grid Audio-Visual Speech Corpus. https:\/\/doi.org\/10.5281\/zenodo.3625687"},{"key":"e_1_3_2_2_10_1","volume-title":"International Conference on Machine Learning","author":"Cuturi Marco","year":"2017","unstructured":"Marco Cuturi and Mathieu Blondel . 2017 . Soft-dtw: a differentiable loss function for time-series . International Conference on Machine Learning (2017), 894--903. Marco Cuturi and Mathieu Blondel. 2017. Soft-dtw: a differentiable loss function for time-series. International Conference on Machine Learning (2017), 894--903."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_2_12_1","volume-title":"CoRR (2018)","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding . CoRR (2018) , 1810.04805. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. CoRR (2018), 1810.04805."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00957"},{"key":"e_1_3_2_2_14_1","volume-title":"Jamie Ryan Kiros, and Sanja Fidler","author":"Faghri Fartash","year":"2017","unstructured":"Fartash Faghri , David J Fleet , Jamie Ryan Kiros, and Sanja Fidler . 2017 . Vse++: Improving visual-semantic embeddings with hard negatives. CoRR ( 2017), 1707.05612. Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017. Vse++: Improving visual-semantic embeddings with hard negatives. CoRR (2017), 1707.05612."},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"crossref","unstructured":"Felix A Gers J\u00fcrgen Schmidhuber and Fred Cummins. 1999. Learning to forget: Continual prediction with LSTM. (1999).  Felix A Gers J\u00fcrgen Schmidhuber and Fred Cummins. 1999. Learning to forget: Continual prediction with LSTM. (1999).","DOI":"10.1049\/cp:19991218"},{"key":"e_1_3_2_2_16_1","volume-title":"Segmented Pairwise Distance for Time Series with Large Discontinuities. International Joint Conference on Neural Networks","author":"He Jiabo","year":"2020","unstructured":"Jiabo He , Sarah Erfani , Sudanthi Wijewickrema , Stephen O'Leary , and Kotagiri Ramamohanarao . 2020 . Segmented Pairwise Distance for Time Series with Large Discontinuities. International Joint Conference on Neural Networks (2020), 1--8. Jiabo He, Sarah Erfani, Sudanthi Wijewickrema, Stephen O'Leary, and Kotagiri Ramamohanarao. 2020. Segmented Pairwise Distance for Time Series with Large Discontinuities. International Joint Conference on Neural Networks (2020), 1--8."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_2_2_18_1","volume-title":"MovieNet: A Holistic Dataset for Movie Understanding. European Conference on Computer Vision","author":"Huang Qingqiu","year":"2020","unstructured":"Qingqiu Huang , Yu Xiong , Anyi Rao , Jiaze Wang , and Dahua Lin . 2020 . MovieNet: A Holistic Dataset for Movie Understanding. European Conference on Computer Vision (2020). Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, and Dahua Lin. 2020. MovieNet: A Holistic Dataset for Movie Understanding. European Conference on Computer Vision (2020)."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00133"},{"key":"e_1_3_2_2_20_1","volume-title":"et al","author":"Kay Will","year":"2017","unstructured":"Will Kay , Joao Carreira , Karen Simonyan , Brian Zhang , Chloe Hillier , Sudheendra Vijayanarasimhan , Fabio Viola , Tim Green , Trevor Back , Paul Natsev , et al . 2017 . The kinetics human action video dataset. CoRR ( 2017), 1705.06950. Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al . 2017. The kinetics human action video dataset. CoRR (2017), 1705.06950."},{"key":"e_1_3_2_2_21_1","volume-title":"IEEE International Conference on Computer Vision","author":"Kuehne H.","year":"2011","unstructured":"H. Kuehne , H. Jhuang , E. Garrote , T. Poggio , and T. Serre . 2011. HMDB: A large video database for human motion recognition . IEEE International Conference on Computer Vision ( 2011 ), 2556--2563. H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. 2011. HMDB: A large video database for human motion recognition. IEEE International Conference on Computer Vision (2011), 2556--2563."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2833157.2833162"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1167"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58589-1_27"},{"key":"e_1_3_2_2_25_1","volume-title":"ACM International Conference on Multimedia","author":"Li Xirong","year":"2019","unstructured":"Xirong Li , Chaoxi Xu , Gang Yang , Zhineng Chen , and Jianfeng Dong . 2019 . W2VV++ Fully Deep Learning for Ad-hoc Video Search . ACM International Conference on Multimedia (2019), 1786--1794. Xirong Li, Chaoxi Xu, Gang Yang, Zhineng Chen, and Jianfeng Dong. 2019. W2VV++ Fully Deep Learning for Ad-hoc Video Search. ACM International Conference on Multimedia (2019), 1786--1794."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.502"},{"key":"e_1_3_2_2_27_1","volume-title":"CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval. CoRR","author":"Luo Huaishao","year":"2021","unstructured":"Huaishao Luo , Lei Ji , Ming Zhong , Yang Chen , Wen Lei , Nan Duan , and Tianrui Li. 2021. CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval. CoRR ( 2021 ), 2104.08860. Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2021. CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval. CoRR (2021), 2104.08860."},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00272"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3206025.3206064"},{"key":"e_1_3_2_2_30_1","volume-title":"International Conference on Acoustics, Speech, and Signal Processing","author":"Nagrani Arsha","year":"2020","unstructured":"Arsha Nagrani , Joon Son Chung , Samuel Albanie , and Andrew Zisserman . 2020 . Disentangled Speech Embeddings using Cross-Modal Self-Supervision . International Conference on Acoustics, Speech, and Signal Processing (2020). Arsha Nagrani, Joon Son Chung, Samuel Albanie, and Andrew Zisserman. 2020. Disentangled Speech Embeddings using Cross-Modal Self-Supervision. International Conference on Acoustics, Speech, and Signal Processing (2020)."},{"key":"e_1_3_2_2_31_1","volume-title":"Conference of the International Speech Communication Association","author":"Nagrani A.","year":"2017","unstructured":"A. Nagrani , J. S. Chung , and A. Zisserman . 2017. VoxCeleb: a large-scale speaker identification dataset . Conference of the International Speech Communication Association ( 2017 ). A. Nagrani, J. S. Chung, and A. Zisserman. 2017. VoxCeleb: a large-scale speaker identification dataset. Conference of the International Speech Communication Association (2017)."},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00772"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_2_34_1","volume-title":"Generative adversarial text to image synthesis. CoRR","author":"Reed Scott","year":"2016","unstructured":"Scott Reed , Zeynep Akata , Xinchen Yan , Lajanugen Logeswaran , Bernt Schiele , and Honglak Lee . 2016. Generative adversarial text to image synthesis. CoRR ( 2016 ), 1605.05396. Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016. Generative adversarial text to image synthesis. CoRR (2016), 1605.05396."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0987-1"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00756"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.501"},{"key":"e_1_3_2_2_38_1","volume-title":"Canonical correlation analysis. Encyclopedia of statistics in behavioral science","author":"Thompson Bruce","year":"2005","unstructured":"Bruce Thompson . 2005. Canonical correlation analysis. Encyclopedia of statistics in behavioral science ( 2005 ). Bruce Thompson. 2005. Canonical correlation analysis. Encyclopedia of statistics in behavioral science (2005)."},{"key":"e_1_3_2_2_39_1","volume-title":"Attention is all you need. Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , Lukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. Advances in Neural Information Processing Systems ( 2017 ), 5998--6008. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017), 5998--6008."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-1452"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2964328"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v29i1.9512"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.347"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2017.02.018"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.09.020"},{"key":"e_1_3_2_2_47_1","volume-title":"Places: A 10 million Image Database for Scene Recognition","author":"Zhou Bolei","year":"2017","unstructured":"Bolei Zhou , Agata Lapedriza , Aditya Khosla , Aude Oliva , and Antonio Torralba . 2017 . Places: A 10 million Image Database for Scene Recognition . IEEE Transactions on Pattern Analysis and Machine Intelligence ( 2017). Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017. Places: A 10 million Image Database for Scene Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence (2017)."}],"event":{"name":"ICMR '22: International Conference on Multimedia Retrieval","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Newark NJ USA","acronym":"ICMR '22"},"container-title":["Proceedings of the 2022 International Conference on Multimedia Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512527.3531371","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3512527.3531371","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:12Z","timestamp":1750188612000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512527.3531371"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,27]]},"references-count":48,"alternative-id":["10.1145\/3512527.3531371","10.1145\/3512527"],"URL":"https:\/\/doi.org\/10.1145\/3512527.3531371","relation":{},"subject":[],"published":{"date-parts":[[2022,6,27]]},"assertion":[{"value":"2022-06-27","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}