{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T17:08:20Z","timestamp":1781284100169,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":60,"publisher":"ACM","funder":[{"name":"Beijing Natural Science Foundation","award":["4242028,L251032,L231012"],"award-info":[{"award-number":["4242028,L251032,L231012"]}]},{"name":"Aviation Science Foundation of China","award":["202400550M5002"],"award-info":[{"award-number":["202400550M5002"]}]},{"name":"Open Project of Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, Anhui University","award":["MMC202401"],"award-info":[{"award-number":["MMC202401"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62376021,62006015"],"award-info":[{"award-number":["62376021,62006015"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003787","name":"Natural Science Foundation of Hebei Province","doi-asserted-by":"publisher","award":["F2025105036"],"award-info":[{"award-number":["F2025105036"]}],"id":[{"id":"10.13039\/501100003787","id-type":"DOI","asserted-by":"publisher"}]},{"name":"CEFLA Audio-Video Restoration and Evaluation Key Lab of Ministry of Culture and Tourism"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746027.3754999","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T07:38:54Z","timestamp":1761377934000},"page":"3163-3172","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["InstructStep: Fine-Grained Localization of Step Content and Relation in Instructional Video"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-9845-7888","authenticated-orcid":false,"given":"Wangsheng","family":"He","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China and Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, Anhui University, Anhui, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2206-5051","authenticated-orcid":false,"given":"Wanru","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China and Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, Anhui University, Anhui, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1166-4042","authenticated-orcid":false,"given":"Ping","family":"Guo","sequence":"additional","affiliation":[{"name":"Intel Labs China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8032-5769","authenticated-orcid":false,"given":"Zhenjiang","family":"Miao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6054-7970","authenticated-orcid":false,"given":"Yi","family":"Tian","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Jiaotong University, Beijing, Beijing Jiaotong University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Michael Ahn Anthony Brohan Noah Brown Yevgen Chebotar Omar Cortes Byron David Chelsea Finn Chuyuan Fu Keerthana Gopalakrishnan Karol Hausman Alex Herzog Daniel Ho Jasmine Hsu Julian Ibarz Brian Ichter Alex Irpan Eric Jang Rosario Jauregui Ruano Kyle Jeffrey Sally Jesmonth Nikhil J Joshi Ryan Julian Dmitry Kalashnikov Yuheng Kuang Kuang-Huei Lee Sergey Levine Yao Lu Linda Luu Carolina Parada Peter Pastor Jornell Quiambao Kanishka Rao Jarek Rettinghouse Diego Reyes Pierre Sermanet Nicolas Sievers Clayton Tan Alexander Toshev Vincent Vanhoucke Fei Xia Ted Xiao Peng Xu Sichun Xu Mengyuan Yan and Andy Zeng. 2022. Do As I Can Not As I Say: Grounding Language in Robotic Affordances. arXiv:2204.01691 [cs.RO] https:\/\/arxiv.org\/abs\/2204.01691"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3495883"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Joao Carreira and Andrew Zisserman. 2018. Quo Vadis Action Recognition? A New Model and the Kinetics Dataset. arXiv:1705.07750 [cs.CV] https:\/\/arxiv.org\/abs\/1705.07750","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","unstructured":"Shaoxiang Chen and Yu-Gang Jiang. 2019. Semantic proposal for activity localization in videos via sentence query. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence (Honolulu Hawaii USA) (AAAI'19\/IAAI'19\/EAAI'19). AAAI Press Article 1005 8 pages. doi:10.1609\/aaai.v33i01.33018199","DOI":"10.1609\/aaai.v33i01.33018199"},{"key":"e_1_3_2_1_5_1","unstructured":"Zhe Chen Weiyun Wang Yue Cao Yangzhou Liu Zhangwei Gao Erfei Cui Jinguo Zhu Shenglong Ye Hao Tian Zhaoyang Liu Lixin Gu Xuehui Wang Qingyun Li Yimin Ren Zixuan Chen Jiapeng Luo Jiahao Wang Tan Jiang Bo Wang Conghui He Botian Shi Xingcheng Zhang Han Lv Yi Wang Wenqi Shao Pei Chu Zhongying Tu Tong He Zhiyong Wu Huipeng Deng Jiaye Ge Kai Chen Kaipeng Zhang Limin Wang Min Dou Lewei Lu Xizhou Zhu Tong Lu Dahua Lin Yu Qiao Jifeng Dai and Wenhai Wang. 2025. Expanding Performance Boundaries of Open-Source Multimodal Models with Model Data and Test-Time Scaling. arXiv:2412.05271 [cs.CV] https:\/\/arxiv.org\/abs\/2412.05271"},{"key":"e_1_3_2_1_6_1","volume-title":"Daisy Zhe Wang, and Doo Soon Kim","author":"Colas Anthony","year":"2020","unstructured":"Anthony Colas, Seokhwan Kim, Franck Dernoncourt, Siddhesh Gupte, Daisy Zhe Wang, and Doo Soon Kim. 2020a. TutorialVQA: Question Answering Dataset for Tutorial Videos. arXiv:1912.01046 [cs.CL] https:\/\/arxiv.org\/abs\/1912.01046"},{"key":"e_1_3_2_1_7_1","unstructured":"Anthony Colas Seokhwan Kim Franck Dernoncourt Siddhesh Gupte Zhe Wang and Doo Soon Kim. 2020b. TutorialVQA: Question Answering Dataset for Tutorial Videos. In Proceedings of the Twelfth Language Resources and Evaluation Conference Nicoletta Calzolari Fr\u00e9d\u00e9ric B\u00e9chet Philippe Blache Khalid Choukri Christopher Cieri Thierry Declerck Sara Goggi Hitoshi Isahara Bente Maegaard Joseph Mariani H\u00e9l\u00e8ne Mazo Asuncion Moreno Jan Odijk and Stelios Piperidis (Eds.). European Language Resources Association Marseille France 5450-5455. https:\/\/aclanthology.org\/2020.lrec-1.670\/"},{"key":"e_1_3_2_1_8_1","unstructured":"Danny Driess Fei Xia Mehdi S. M. Sajjadi Corey Lynch Aakanksha Chowdhery Brian Ichter Ayzaan Wahid Jonathan Tompson Quan Vuong Tianhe Yu Wenlong Huang Yevgen Chebotar Pierre Sermanet Daniel Duckworth Sergey Levine Vincent Vanhoucke Karol Hausman Marc Toussaint Klaus Greff Andy Zeng Igor Mordatch and Pete Florence. 2023. PaLM-E: An Embodied Multimodal Language Model. arXiv:2303.03378 [cs.LG] https:\/\/arxiv.org\/abs\/2303.03378"},{"key":"e_1_3_2_1_9_1","volume-title":"StepFormer: Self-Supervised Step Discovery and Localization in Instructional Videos. 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023)","author":"Dvornik Nikita","unstructured":"Nikita Dvornik, Isma Hadji, Ran Zhang, Konstantinos G. Derpanis, Animesh Garg, Richard P. Wildes, and Allan D. Jepson. 2023. StepFormer: Self-Supervised Step Discovery and Localization in Instructional Videos. 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023), 18952-18961. https:\/\/api.semanticscholar.org\/CorpusID:258331832"},{"key":"e_1_3_2_1_10_1","volume-title":"TALL: Temporal Activity Localization via Language Query. arXiv:1705.02101 [cs.CV] https:\/\/arxiv.org\/abs\/1705.02101","author":"Gao Jiyang","year":"2017","unstructured":"Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017. TALL: Temporal Activity Localization via Language Query. arXiv:1705.02101 [cs.CV] https:\/\/arxiv.org\/abs\/1705.02101"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00020"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1198"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3372833"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01453-z"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2024.108170"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Deepak Gupta Kush Attal and Dina Demner-Fushman. 2022. A Dataset for Medical Instructional Video Classification and Question Answering. arXiv:2201.12888 [cs.CV] https:\/\/arxiv.org\/abs\/2201.12888","DOI":"10.1038\/s41597-023-02036-y"},{"key":"e_1_3_2_1_17_1","unstructured":"Pengcheng He Jianfeng Gao and Weizhu Chen. 2023. DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing. arXiv:2111.09543 [cs.CL] https:\/\/arxiv.org\/abs\/2111.09543"},{"key":"e_1_3_2_1_18_1","unstructured":"Pengcheng He Xiaodong Liu Jianfeng Gao and Weizhu Chen. 2021. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. arXiv:2006.03654 [cs.CL] https:\/\/arxiv.org\/abs\/2006.03654"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.618"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475281"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01353"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.bionlp-1.43"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3532626"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3411045"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210003"},{"key":"e_1_3_2_1_26_1","unstructured":"Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. arXiv:1711.05101 [cs.LG] https:\/\/arxiv.org\/abs\/1711.05101"},{"key":"e_1_3_2_1_27_1","unstructured":"Hongyin Luo Mitra Mohtarami James R. Glass Karthik Krishnamurthy and Brigitte Richardson. 2019. Integrating Video Retrieval and Moment Detection in a Unified Corpus for Video Question Answering. In Interspeech. https:\/\/api.semanticscholar.org\/CorpusID:203141820"},{"key":"e_1_3_2_1_28_1","volume-title":"Makarand Tapaswi, Ivan Laptev, and Josef Sivic.","author":"Miech Antoine","year":"2019","unstructured":"Antoine Miech, Dimitri Zhukov, Jean Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. IEEE (2019)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3570955"},{"key":"e_1_3_2_1_30_1","volume-title":"Spoken Moments: Learning Joint Audio-Visual Representations from Video Descriptions. arXiv:2105.04489 [cs.CV] https:\/\/arxiv.org\/abs\/2105.04489","author":"Monfort Mathew","year":"2021","unstructured":"Mathew Monfort, SouYoung Jin, Alexander Liu, David Harwath, Rogerio Feris, James Glass, and Aude Oliva. 2021. Spoken Moments: Learning Joint Audio-Visual Representations from Video Descriptions. arXiv:2105.04489 [cs.CV] https:\/\/arxiv.org\/abs\/2105.04489"},{"key":"e_1_3_2_1_31_1","unstructured":"OpenAI. 2022. Introducing ChatGPT. https:\/\/openai.com\/blog\/chatgpt\/."},{"key":"e_1_3_2_1_32_1","unstructured":"OpenAI and Josh Achiam. 2024. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_1_33_1","unstructured":"Wonpyo Park Dongju Kim Yan Lu and Minsu Cho. 2019. Relational Knowledge Distillation. arXiv:1904.05068 [cs.CV] https:\/\/arxiv.org\/abs\/1904.05068"},{"key":"e_1_3_2_1_34_1","unstructured":"Xiao Pu Mingqi Gao and Xiaojun Wan. 2023. Summarization is (Almost) Dead. arXiv:2309.09558 [cs.CL] https:\/\/arxiv.org\/abs\/2309.09558"},{"key":"e_1_3_2_1_35_1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., Vol. 21, 1, Article 140 (Jan. 2020), 67 pages.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_1_36_1","unstructured":"Shuhuai Ren Linli Yao Shicheng Li Xu Sun and Lu Hou. 2024. TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding. arXiv:2312.02051 [cs.CV] https:\/\/arxiv.org\/abs\/2312.02051"},{"key":"e_1_3_2_1_37_1","volume-title":"Hongdong Li, and Stephen Gould.","author":"Rodriguez-Opazo Cristian","year":"2020","unstructured":"Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Fatemeh Sadat Saleh, Hongdong Li, and Stephen Gould. 2020a. Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention. arXiv:1908.07236 [cs.CV] https:\/\/arxiv.org\/abs\/1908.07236"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093328"},{"key":"e_1_3_2_1_39_1","first-page":"1","article-title":"Dropout: a simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava Nitish","year":"2014","unstructured":"Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: a simple way to prevent neural networks from overfitting. J. Mach. Learn. Res., Vol. 15, 1 (Jan. 2014), 1929-1958.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3063631"},{"key":"e_1_3_2_1_41_1","volume-title":"COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis. arXiv:1903.02874 [cs.CV] https:\/\/arxiv.org\/abs\/1903.02874","author":"Tang Yansong","year":"2019","unstructured":"Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. 2019. COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis. arXiv:1903.02874 [cs.CV] https:\/\/arxiv.org\/abs\/1903.02874"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/tpami.2020.2980824"},{"key":"e_1_3_2_1_43_1","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar Aurelien Rodriguez Armand Joulin Edouard Grave and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971 [cs.CL] https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123354"},{"key":"e_1_3_2_1_45_1","unstructured":"Aaron van den Oord Yazhe Li and Oriol Vinyals. 2019. Representation Learning with Contrastive Predictive Coding. arXiv:1807.03748 [cs.LG] https:\/\/arxiv.org\/abs\/1807.03748"},{"key":"e_1_3_2_1_46_1","first-page":"2579","article-title":"Visualizing Data using t-SNE","volume":"9","author":"van der Maaten Laurens","year":"2008","unstructured":"Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing Data using t-SNE. Journal of Machine Learning Research, Vol. 9, 86 (2008), 2579-2605. http:\/\/jmlr.org\/papers\/v9\/vandermaaten08a.html","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-024-02855-5"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"crossref","unstructured":"Xiao Wang Qingyi Si Jianlong Wu Shiyu Zhu Li Cao and Liqiang Nie. 2025. AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding. arXiv:2503.12559 [cs.CV] https:\/\/arxiv.org\/abs\/2503.12559","DOI":"10.18653\/v1\/2025.findings-acl.283"},{"key":"e_1_3_2_1_49_1","volume-title":"Visual Answer Localization with Cross-modal Mutual Knowledge Transfer. In 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1-5.","author":"Weng Yixuan","year":"2023","unstructured":"Yixuan Weng and Bin Li. 2023. Visual Answer Localization with Cross-modal Mutual Knowledge Transfer. In 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1-5."},{"key":"e_1_3_2_1_50_1","unstructured":"Yongliang Wu Xinting Hu Yuyang Sun Yizhou Zhou Wenbo Zhu Fengyun Rao Bernt Schiele and Xu Yang. 2025. Number it: Temporal Grounding Videos like Flipping Manga. arXiv:2411.10332 [cs.CV] https:\/\/arxiv.org\/abs\/2411.10332"},{"key":"e_1_3_2_1_51_1","unstructured":"Huijuan Xu Kun He Bryan A. Plummer Leonid Sigal Stan Sclaroff and Kate Saenko. 2018. Multilevel Language and Vision Integration for Text-to-Clip Retrieval. arXiv:1804.05113 [cs.CV] https:\/\/arxiv.org\/abs\/1804.05113"},{"key":"e_1_3_2_1_52_1","unstructured":"An Yang Anfeng Li Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chang Gao Chengen Huang Chenxu Lv Chujie Zheng Dayiheng Liu Fan Zhou Fei Huang Feng Hu Hao Ge Haoran Wei Huan Lin Jialong Tang Jian Yang Jianhong Tu Jianwei Zhang Jianxin Yang Jiaxi Yang Jing Zhou Jingren Zhou Junyang Lin Kai Dang Keqin Bao Kexin Yang Le Yu Lianghao Deng Mei Li Mingfeng Xue Mingze Li Pei Zhang Peng Wang Qin Zhu Rui Men Ruize Gao Shixuan Liu Shuang Luo Tianhao Li Tianyi Tang Wenbiao Yin Xingzhang Ren Xinyu Wang Xinyu Zhang Xuancheng Ren Yang Fan Yang Su Yichang Zhang Yinger Zhang Yu Wan Yuqiong Liu Zekun Wang Zeyu Cui Zhenru Zhang Zhipeng Zhou and Zihan Qiu. 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] https:\/\/arxiv.org\/abs\/2505.09388"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","unstructured":"Jihun Yoon and Min-Kook Choi. 2023. Exploring Video Frame Redundancies for Efficient Data Sampling and Annotation in Instance Segmentation. In 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 3308-3317. doi:10.1109\/CVPRW59228.2023.00333","DOI":"10.1109\/CVPRW59228.2023.00333"},{"key":"e_1_3_2_1_54_1","unstructured":"Yitian Yuan Tao Mei and Wenwu Zhu. 2018. To Find Where You Talk: Temporal Sentence Localization in Video with Attention Based Location Regression. arXiv:1804.07014 [cs.CV] https:\/\/arxiv.org\/abs\/1804.07014"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.585"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3258628"},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"crossref","unstructured":"Songyang Zhang Houwen Peng Jianlong Fu and Jiebo Luo. 2020a. Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language. arXiv:1912.03590 [cs.CV] https:\/\/arxiv.org\/abs\/1912.03590","DOI":"10.1609\/aaai.v34i07.6984"},{"key":"e_1_3_2_1_58_1","volume-title":"The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Bb21JPnhhr","author":"Zhao Qi","year":"2024","unstructured":"Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu, Minh Quan Do, Nakul Agarwal, Kwonjoon Lee, and Chen Sun. 2024. AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Bb21JPnhhr"},{"key":"e_1_3_2_1_59_1","volume-title":"Corso","author":"Zhou Luowei","year":"2017","unstructured":"Luowei Zhou, Chenliang Xu, and Jason J. Corso. 2017. Towards Automatic Learning of Procedures from Web Instructional Videos. arXiv:1703.09788 [cs.CV] https:\/\/arxiv.org\/abs\/1703.09788"},{"key":"e_1_3_2_1_60_1","volume-title":"David Fouhey, Ivan Laptev, and Josef Sivic.","author":"Zhukov Dimitri","year":"2019","unstructured":"Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. 2019. Cross-task weakly supervised learning from instructional videos. arXiv:1903.08225 [cs.CV] https:\/\/arxiv.org\/abs\/1903.08225"}],"event":{"name":"MM '25: The 33rd ACM International Conference on Multimedia","location":"Dublin Ireland","acronym":"MM '25","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 33rd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3754999","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T04:14:09Z","timestamp":1765340049000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746027.3754999"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":60,"alternative-id":["10.1145\/3746027.3754999","10.1145\/3746027"],"URL":"https:\/\/doi.org\/10.1145\/3746027.3754999","relation":{},"subject":[],"published":{"date-parts":[[2025,10,27]]},"assertion":[{"value":"2025-10-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}