{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:09:49Z","timestamp":1750219789813,"version":"3.41.0"},"reference-count":73,"publisher":"Association for Computing Machinery (ACM)","issue":"5s","license":[{"start":{"date-parts":[[2023,6,7]],"date-time":"2023-06-07T00:00:00Z","timestamp":1686096000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61832001, 62276047"],"award-info":[{"award-number":["61832001, 62276047"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Open Fund of Intelligent Terminal Key Laboratory of Sichuan Province","award":["SCITLAB-20008"],"award-info":[{"award-number":["SCITLAB-20008"]}]},{"name":"Sichuan Science and Technology Program","award":["2022JDRC0064"],"award-info":[{"award-number":["2022JDRC0064"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>\n            Pre-training model on large-scale unlabeled web videos followed by task-specific fine-tuning is a canonical approach to learning video and language representations. However, the accompanying Automatic Speech Recognition (ASR) transcripts in these videos are directly transcribed from audio, which may be inconsistent with visual information and would impair the language modeling ability of the model. Meanwhile, previous V-L models fuse visual and language modality features using single- or dual-stream architectures, which are not suitable for the current situation. Besides, traditional V-L research focuses mainly on the interaction between vision and language modalities and leaves the modeling of relationships within modalities untouched. To address these issues and maintain a small manual labor cost, we add automatically extracted dense captions as a supplementary text and propose a new trilinear video-language interaction framework TEVL (Trilinear Encoder for Video-Language representation learning). TEVL contains three unimodal encoders, a TRIlinear encOder (TRIO) block, and a temporal Transformer. TRIO is specially designed to support effective text-vision-text interaction, which encourages inter-modal cooperation while maintaining intra-modal dependencies. We pre-train TEVL on the HowTo100M and TV datasets with four task objectives. Experimental results demonstrate that TEVL can learn powerful video-text representation and achieve competitive performance on three downstream tasks, including multimodal video captioning, video Question Answering (QA), as well as video and language inference. Implementation code is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/Gufrannn\/TEVL\">https:\/\/github.com\/Gufrannn\/TEVL<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3585388","type":"journal-article","created":{"date-parts":[[2023,2,24]],"date-time":"2023-02-24T11:05:52Z","timestamp":1677236752000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["TEVL: Trilinear Encoder for Video-language Representation Learning"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9684-4424","authenticated-orcid":false,"given":"Xin","family":"Man","sequence":"first","affiliation":[{"name":"University of Electronic Science and Technology of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2615-1555","authenticated-orcid":false,"given":"Jie","family":"Shao","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0928-6899","authenticated-orcid":false,"given":"Feiyu","family":"Chen","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, China Intelligent Terminal Key Laboratory of Sichuan Province, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9939-158X","authenticated-orcid":false,"given":"Mingxing","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, China Sichuan Artificial Intelligence Research Institute, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2999-2088","authenticated-orcid":false,"given":"Heng Tao","family":"Shen","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, China Sichuan Artificial Intelligence Research Institute, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,6,7]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"65","volume-title":"Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization@ACL 2005","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization@ACL 2005. 65\u201372."},{"key":"e_1_3_2_3_2","article-title":"iPerceive: Applying common-sense reasoning to multi-modal dense video captioning and video question answering","volume":"2011","author":"Chadha Aman","year":"2020","unstructured":"Aman Chadha, Gurneet Arora, and Navpreet Kaloty. 2020. iPerceive: Applying common-sense reasoning to multi-modal dense video captioning and video question answering. CoRR abs\/2011.07735 (2020).","journal-title":"CoRR"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00203"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_7_2","first-page":"4171","volume-title":"Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171\u20134186."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00048"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3150959"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3120867"},{"key":"e_1_3_2_12_2","volume-title":"Annual Conference on Neural Information Processing Systems","author":"Ging Simon","year":"2020","unstructured":"Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox. 2020. COOT: Cooperative hierarchical transformer for video-text representation learning. In Annual Conference on Neural Information Processing Systems."},{"issue":"2","key":"e_1_3_2_13_2","first-page":"53:1","article-title":"Visual semantic-based representation learning using deep CNNs for scene recognition","volume":"17","author":"Gupta Shikha","year":"2021","unstructured":"Shikha Gupta, Krishan Sharma, Dileep Aroor Dinesh, and Veena Thenkanidiyoor. 2021. Visual semantic-based representation learning using deep CNNs for scene recognition. ACM Trans. Multim. Comput. Commun. Appl. 17, 2 (2021), 53:1\u201353:24.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1252"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.494"},{"key":"e_1_3_2_17_2","article-title":"Exploring the limits of language modeling","volume":"1602","author":"J\u00f3zefowicz Rafal","year":"2016","unstructured":"Rafal J\u00f3zefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016. Exploring the limits of language modeling. CoRR abs\/1602.02410 (2016).","journal-title":"CoRR"},{"key":"e_1_3_2_18_2","article-title":"The kinetics human action video dataset","volume":"1705","author":"Kay Will","year":"2017","unstructured":"Will Kay, Jo\u00e3o Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman. 2017. The kinetics human action video dataset. CoRR abs\/1705.06950 (2017).","journal-title":"CoRR"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1351"},{"key":"e_1_3_2_20_2","first-page":"4812","volume-title":"58th Annual Meeting of the Association for Computational Linguistics","author":"Kim Hyounghun","year":"2020","unstructured":"Hyounghun Kim, Zineng Tang, and Mohit Bansal. 2020. Dense-caption matching and frame-selection gating for temporal localization in VideoQA. In 58th Annual Meeting of the Association for Computational Linguistics. 4812\u20134822."},{"key":"e_1_3_2_21_2","first-page":"1","volume-title":"International Joint Conference on Neural Networks","author":"Kim Junyeong","year":"2019","unstructured":"Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, and Chang D. Yoo. 2019. Gaining extra supervision via multi-task learning for multi-modal video question answering. In International Joint Conference on Neural Networks. 1\u20138."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00853"},{"key":"e_1_3_2_23_2","article-title":"Video understanding as machine translation","volume":"2006","author":"Korbar Bruno","year":"2020","unstructured":"Bruno Korbar, Fabio Petroni, Rohit Girdhar, and Lorenzo Torresani. 2020. Video understanding as machine translation. CoRR abs\/2006.07203 (2020).","journal-title":"CoRR"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1167"},{"key":"e_1_3_2_26_2","first-page":"8211","volume-title":"58th Annual Meeting of the Association for Computational Linguistics","author":"Lei Jie","year":"2020","unstructured":"Jie Lei, Licheng Yu, Tamara L. Berg, and Mohit Bansal. 2020. TVQA+: Spatio-temporal grounding for video question answering. In 58th Annual Meeting of the Association for Computational Linguistics. 8211\u20138225."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58589-1_27"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6795"},{"key":"e_1_3_2_29_2","article-title":"A CLIP-enhanced method for video-language understanding","volume":"2110","author":"Li Guohao","year":"2021","unstructured":"Guohao Li, Feng He, and Zhifan Feng. 2021. A CLIP-enhanced method for video-language understanding. CoRR abs\/2110.07137 (2021).","journal-title":"CoRR"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.161"},{"key":"e_1_3_2_31_2","volume-title":"Annual Conference on Neural Information Processing Systems","author":"Li Linjie","year":"2021","unstructured":"Linjie Li, Jie Lei, Zhe Gan, Licheng Yu, Yen-Chun Chen, Rohit Pillai, Yu Cheng, Luowei Zhou, Xin Eric Wang, William Yang Wang, Tamara Lee Berg, Mohit Bansal, Jingjing Liu, Lijuan Wang, and Zicheng Liu. 2021. VALUE: A multi-task benchmark for video-and-language understanding evaluation. In Annual Conference on Neural Information Processing Systems."},{"key":"e_1_3_2_32_2","article-title":"VisualBERT: A simple and performant baseline for vision and language","volume":"1908","author":"Li Liunian Harold","year":"2019","unstructured":"Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. VisualBERT: A simple and performant baseline for vision and language. CoRR abs\/1908.03557 (2019).","journal-title":"CoRR"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3149329"},{"issue":"2","key":"e_1_3_2_34_2","first-page":"48:1","article-title":"Uni-EDEN: Universal encoder-decoder network by multi-granular vision-language pre-training","volume":"18","author":"Li Yehao","year":"2022","unstructured":"Yehao Li, Jiahao Fan, Yingwei Pan, Ting Yao, Weiyao Lin, and Tao Mei. 2022. Uni-EDEN: Universal encoder-decoder network by multi-granular vision-language pre-training. ACM Trans. Multim. Comput. Commun. Appl. 18, 2 (2022), 48:1\u201348:16.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3103782"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2624140"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2852750"},{"key":"e_1_3_2_38_2","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, Association for Computational Linguistics, 74\u201381."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01742"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01091"},{"key":"e_1_3_2_41_2","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","volume":"1907","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. CoRR abs\/1907.11692 (2019).","journal-title":"CoRR"},{"key":"e_1_3_2_42_2","volume-title":"7th International Conference on Learning Representations","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations."},{"key":"e_1_3_2_43_2","first-page":"13","volume-title":"Annual Conference on Neural Information Processing Systems","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Annual Conference on Neural Information Processing Systems. 13\u201323."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475703"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.778"},{"issue":"4","key":"e_1_3_2_46_2","first-page":"128:1","article-title":"Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders","volume":"17","author":"Messina Nicola","year":"2021","unstructured":"Nicola Messina, Giuseppe Amato, Andrea Esuli, Fabrizio Falchi, Claudio Gennaro, and St\u00e9phane Marchand-Maillet. 2021. Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders. ACM Trans. Multim. Comput. Commun. Appl. 17, 4 (2021), 128:1\u2013128:23.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00272"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3551581"},{"key":"e_1_3_2_49_2","first-page":"311","volume-title":"40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In 40th Annual Meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_2_50_2","first-page":"8748","volume-title":"38th International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In 38th International Conference on Machine Learning. 8748\u20138763."},{"key":"e_1_3_2_51_2","article-title":"Winning the ICCV\u20192021 VALUE challenge: Task-aware ensemble and transfer learning with visual concepts","volume":"2110","author":"Shin Minchul","year":"2021","unstructured":"Minchul Shin, Jonghwan Mun, Kyoung-Woon On, Woo-Young Kang, Gunsoo Han, and Eun-Sol Kim. 2021. Winning the ICCV\u20192021 VALUE challenge: Task-aware ensemble and transfer learning with visual concepts. CoRR abs\/2110.06476 (2021).","journal-title":"CoRR"},{"key":"e_1_3_2_52_2","article-title":"Contrastive bidirectional transformer for temporal representation learning","volume":"1906","author":"Sun Chen","year":"2019","unstructured":"Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid. 2019. Contrastive bidirectional transformer for temporal representation learning. CoRR abs\/1906.05743 (2019).","journal-title":"CoRR"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00756"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1514"},{"issue":"2","key":"e_1_3_2_55_2","first-page":"54:1","article-title":"Show, reward, and tell: Adversarial visual story generation","volume":"15","author":"Tang Jinhui","year":"2019","unstructured":"Jinhui Tang, Jing Wang, Zechao Li, Jianlong Fu, and Tao Mei. 2019. Show, reward, and tell: Adversarial visual story generation. ACM Trans. Multim. Comput. Commun. Appl. 15, 2s (2019), 54:1\u201354:20.","journal-title":"ACM Trans. Multim. Comput. Commun. Appl."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.193"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3080928"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-240"},{"key":"e_1_3_2_59_2","first-page":"5998","volume-title":"Annual Conference on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Annual Conference on Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_62_2","article-title":"Google\u2019s neural machine translation system: Bridging the gap between human and machine translation","volume":"1609","author":"Wu Yonghui","year":"2016","unstructured":"Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016. Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. CoRR abs\/1609.08144 (2016).","journal-title":"CoRR"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.214"},{"key":"e_1_3_2_64_2","first-page":"1545","volume-title":"IEEE Winter Conference on Applications of Computer Vision","author":"Yang Zekun","year":"2020","unstructured":"Zekun Yang, Noa Garcia, Chenhui Chu, Mayu Otani, Yuta Nakashima, and Haruo Takemura. 2020. BERT representations for video question answering. In IEEE Winter Conference on Applications of Computer Vision. 1545\u20131554."},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.10"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00142"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3002667"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3205212"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548024"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00553"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00911"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3235523"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00877"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00148"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3585388","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3585388","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:56Z","timestamp":1750178276000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3585388"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,7]]},"references-count":73,"journal-issue":{"issue":"5s","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3585388"],"URL":"https:\/\/doi.org\/10.1145\/3585388","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2023,6,7]]},"assertion":[{"value":"2022-09-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}