{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T03:38:01Z","timestamp":1782358681544,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":40,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"he National Natural Science Foundation of China","award":["No.:U1936203"],"award-info":[{"award-number":["No.:U1936203"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548180","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:43:12Z","timestamp":1665416592000},"page":"3234-3243","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":29,"title":["Search-oriented Micro-video Captioning"],"prefix":"10.1145","author":[{"given":"Liqiang","family":"Nie","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Leigang","family":"Qu","sequence":"additional","affiliation":[{"name":"Shandong University, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dai","family":"Meng","sequence":"additional","affiliation":[{"name":"Kuaishou, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Min","family":"Zhang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qi","family":"Tian","sequence":"additional","affiliation":[{"name":"Huawei Cloud &amp; AI, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alberto Del","family":"Bimbo","sequence":"additional","affiliation":[{"name":"University of Florence, Florence, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Proceedings of the Neural Information Processing Systems Conference. 25--37","author":"Alayrac Jean-Baptiste","year":"2020","unstructured":"Jean-Baptiste Alayrac , Adria Recasens , Rosalia Schneider , Relja Arandjelovi?, Jason Ramapuram , Jeffrey De Fauw , Lucas Smaira , Sander Dieleman , and Andrew Zisserman . 2020 . Self-supervised Multimodal Versatile Networks . In Proceedings of the Neural Information Processing Systems Conference. 25--37 . Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelovi?, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. 2020. Self-supervised Multimodal Versatile Networks. In Proceedings of the Neural Information Processing Systems Conference. 25--37."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00436"},{"key":"e_1_3_2_2_3_1","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics. ACL, 190--200","author":"Chen David","year":"2011","unstructured":"David Chen and William B Dolan . 2011 . Collecting Highly Parallel Data for Paraphrase Evaluation . In Proceedings of the Annual Meeting of the Association for Computational Linguistics. ACL, 190--200 . David Chen and William B Dolan. 2011. Collecting Highly Parallel Data for Paraphrase Evaluation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. ACL, 190--200."},{"key":"e_1_3_2_2_4_1","volume-title":"Focus-Constrained Attention Mechanism for CVAE-based Response Generation. In Findings of the Conference on Empirical Methods in Natural Language Processing. ACL","author":"Cui Zhi","year":"2020","unstructured":"Zhi Cui , Yanran Li , Jiayi Zhang , Jianwei Cui , Chen Wei , and Bin Wang . 2020 . Focus-Constrained Attention Mechanism for CVAE-based Response Generation. In Findings of the Conference on Empirical Methods in Natural Language Processing. ACL , 2021--2030. Zhi Cui, Yanran Li, Jiayi Zhang, Jianwei Cui, Chen Wei, and Bin Wang. 2020. Focus-Constrained Attention Mechanism for CVAE-based Response Generation. In Findings of the Conference on Empirical Methods in Natural Language Processing. ACL, 2021--2030."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.323"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01095"},{"key":"e_1_3_2_2_7_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing. ACL, 1100--1111","author":"Gimpel Kevin","year":"2013","unstructured":"Kevin Gimpel , Dhruv Batra , Chris Dyer , and Gregory Shakhnarovich . 2013 . A Systematic Exploration of Diversity in Machine Translation . In Proceedings of the Conference on Empirical Methods in Natural Language Processing. ACL, 1100--1111 . Kevin Gimpel, Dhruv Batra, Chris Dyer, and Gregory Shakhnarovich. 2013. A Systematic Exploration of Diversity in Machine Translation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. ACL, 1100--1111."},{"key":"e_1_3_2_2_8_1","volume-title":"Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 289--297","author":"Ging Simon","year":"2020","unstructured":"Simon Ging , Mohammadreza Zolfaghari , Hamed Pirsiavash , and Thomas Brox . 2020 . Coot: Cooperative Hierarchical Transformer for Video-text Representation Learning . In Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 289--297 . Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox. 2020. Coot: Cooperative Hierarchical Transformer for Video-text Representation Learning. In Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 289--297."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.337"},{"key":"e_1_3_2_2_10_1","volume-title":"Proceedings of the International Conference on Machine Learning. PMLR, 448--456","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy . 2015 . Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift . In Proceedings of the International Conference on Machine Learning. PMLR, 448--456 . Sergey Ioffe and Christian Szegedy. 2015. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In Proceedings of the International Conference on Machine Learning. PMLR, 448--456."},{"key":"e_1_3_2_2_11_1","volume-title":"Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 10236--10245","author":"Kingma Durk P","year":"2018","unstructured":"Durk P Kingma and Prafulla Dhariwal . 2018 . Glow: Generative Flow with Invertible 1x1 Convolutions . In Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 10236--10245 . Durk P Kingma and Prafulla Dhariwal. 2018. Glow: Generative Flow with Invertible 1x1 Convolutions. In Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 10236--10245."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1020346032608"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v27i1.8679"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.233"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.161"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240549"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502081.2502084"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01088"},{"key":"e_1_3_2_2_22_1","volume-title":"Proceedings of the International Conference on Machine Learning. PMLR, 1530--1538","author":"Rezende Danilo","year":"2015","unstructured":"Danilo Rezende and Shakir Mohamed . 2015 . Variational Inference with Normalizing Flows . In Proceedings of the International Conference on Machine Learning. PMLR, 1530--1538 . Danilo Rezende and Shakir Mohamed. 2015. Variational Inference with Normalizing Flows. In Proceedings of the International Conference on Machine Learning. PMLR, 1530--1538."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.61"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.445"},{"key":"e_1_3_2_2_25_1","volume-title":"Proceedings of the Neural Information Processing Systems Conference","author":"Sohn Kihyuk","year":"2015","unstructured":"Kihyuk Sohn , Honglak Lee , and Xinchen Yan . 2015 . Learning Structured Output Representation using Deep Conditional Generative Models . Proceedings of the Neural Information Processing Systems Conference (2015), 1--9. Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015. Learning Structured Output Representation using Deep Conditional Generative Models. Proceedings of the Neural Information Processing Systems Conference (2015), 1--9."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00756"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967199"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/104"},{"key":"e_1_3_2_2_29_1","volume-title":"Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 5998--6008","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , Lukasz Kaiser , and Illia Polosukhin . 2017 . Attention is All you Need . In Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 5998--6008 . Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 5998--6008."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.515"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1173"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12340"},{"key":"e_1_3_2_2_33_1","volume-title":"Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 5756--5766","author":"Wang Liwei","year":"2017","unstructured":"Liwei Wang , Alexander Schwing , and Svetlana Lazebnik . 2017 . Diverse and Accurate Image Description using a Variational Auto-encoder with an Additive Gaussian Encoding Space . In Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 5756--5766 . Liwei Wang, Alexander Schwing, and Svetlana Lazebnik. 2017. Diverse and Accurate Image Description using a Variational Auto-encoder with an Additive Gaussian Encoding Space. In Proceedings of the Neural Information Processing Systems Conference. Curran Associates, Inc., 5756--5766."},{"key":"e_1_3_2_2_34_1","volume-title":"Diverse Video Captioning through Latent Variable Expansion. arXiv preprint arXiv:1910.12019","author":"Xiao Huanhou","year":"2019","unstructured":"Huanhou Xiao and Jinglun Shi . 2019. Diverse Video Captioning through Latent Variable Expansion. arXiv preprint arXiv:1910.12019 ( 2019 ), 1--11. Huanhou Xiao and Jinglun Shi. 2019. Diverse Video Captioning through Latent Variable Expansion. arXiv preprint arXiv:1910.12019 (2019), 1--11."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i4.16421"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.512"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1061"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00911"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00877"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548180","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548180","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:20Z","timestamp":1750186820000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548180"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":40,"alternative-id":["10.1145\/3503161.3548180","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548180","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}