{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T15:44:07Z","timestamp":1781279047070,"version":"3.54.1"},"reference-count":85,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T00:00:00Z","timestamp":1770854400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T00:00:00Z","timestamp":1770854400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,3]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    The pursuit of artificial general intelligence (AGI) has been accelerated by Multimodal Large Language Models (MLLMs), which exhibit superior reasoning, generalization capabilities, and proficiency in processing multimodal inputs. A crucial milestone in the evolution of AGI is the attainment of human-level planning, a fundamental ability for making informed decisions in complex environments, and solving a wide range of real-world problems. Despite the impressive advancements in MLLMs, a question remains:\n                    <jats:bold>How far are current MLLMs from achieving human-level planning?<\/jats:bold>\n                    To shed light on this question, we introduce EgoPlan-Bench, a comprehensive benchmark to evaluate the planning abilities of MLLMs in real-world scenarios from an egocentric perspective, mirroring human perception. EgoPlan-Bench emphasizes the evaluation of planning capabilities of MLLMs, featuring realistic tasks, diverse action plans, and intricate visual observations. Our rigorous evaluation of a wide range of MLLMs reveals that EgoPlan-Bench poses significant challenges, highlighting a substantial scope for improvement in MLLMs to achieve human-level task planning. To facilitate this advancement, we further present EgoPlan-IT, a specialized instruction-tuning dataset that effectively enhances model performance on EgoPlan-Bench. We have made all the codes, data, and a maintained benchmark leaderboard available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/chenyi99.github.io\/ego_plan\/\" ext-link-type=\"uri\">https:\/\/chenyi99.github.io\/ego_plan\/<\/jats:ext-link>\n                    to advance future research.\n                  <\/jats:p>","DOI":"10.1007\/s11263-025-02676-0","type":"journal-article","created":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T04:13:56Z","timestamp":1770869636000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning"],"prefix":"10.1007","volume":"134","author":[{"given":"Yi","family":"Chen","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuying","family":"Ge","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yixiao","family":"Ge","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mingyu","family":"Ding","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bohao","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rui","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruifeng","family":"Xu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ying","family":"Shan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1831-9952","authenticated-orcid":false,"given":"Xihui","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,2,12]]},"reference":[{"key":"2676_CR1","unstructured":"Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A. and others (2020). Language models are few-shot learners. Advances in neural information processing systems 33, 1877\u20131901."},{"key":"2676_CR2","unstructured":"Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., and others (2022). Training language models to follow instructions with human feedback. NeurIPS."},{"key":"2676_CR3","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., Azhar, F., and others (2023). Llama: Open and efficient foundation language models. arXiv:2302.13971."},{"key":"2676_CR4","unstructured":"Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J.E., and others (2023). Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https:\/\/vicuna.lmsys.org (accessed 14 April 2023)."},{"key":"2676_CR5","unstructured":"Arcas, B.A., & Norvig, P. (2023). Artificial general intelligence is already here. Noema, October."},{"key":"2676_CR6","unstructured":"Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y.T., Li, Y., Lundberg, S., and others (2023). Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv:2303.12712."},{"key":"2676_CR7","unstructured":"Morris, M.R., Sohl-dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., & Legg, S. (2023). Levels of agi: Operationalizing progress on the path to agi. arXiv:2311.02462."},{"key":"2676_CR8","doi-asserted-by":"crossref","unstructured":"Fan, C. (2019). Egovqa-an egocentric video question answering benchmark dataset. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops, pp. 0\u20130.","DOI":"10.1109\/ICCVW.2019.00536"},{"key":"2676_CR9","first-page":"3343","volume":"35","author":"B Jia","year":"2022","unstructured":"Jia, B., Lei, T., Zhu, S.-C., & Huang, S. (2022). Egotaskqa: Understanding human tasks in egocentric videos. Advances in Neural Information Processing Systems,35, 3343\u20133360.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2676_CR10","doi-asserted-by":"crossref","unstructured":"Li, B., Ge, Y., Ge, Y., Wang, G., Wang, R., Zhang, R., & Shan, Y. (2023). Seed-bench-2: Benchmarking multimodal large language models. arXiv:2311.17092.","DOI":"10.1109\/CVPR52733.2024.01263"},{"key":"2676_CR11","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1007\/s11263-021-01531-2","volume":"130","author":"D Damen","year":"2022","unstructured":"Damen, D., Doughty, H., Farinella, G. M., Furnari, A., Ma, J., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., & Wray, M. (2022). Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision (IJCV),130, 33\u201355.","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2676_CR12","doi-asserted-by":"crossref","unstructured":"Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., and others (2022). Ego4d: Around the world in 3,000 hours of egocentric video. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 18995\u201319012.","DOI":"10.1109\/CVPR52688.2022.01842"},{"key":"2676_CR13","unstructured":"Chung, H.W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., and others (2022). Scaling instruction-finetuned language models. arXiv:2210.11416."},{"key":"2676_CR14","unstructured":"Zhang, A., Fei, H., Yao, Y., Ji, W., Li, L., Liu, Z., & Chua, T.-S. (2023). Transfer visual prompt generator across llms. arXiv:abs\/2304.50127"},{"key":"2676_CR15","unstructured":"Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., and others (2023). mplug-owl: Modularization empowers large language models with multimodality. arXiv:2304.14178."},{"key":"2676_CR16","unstructured":"Li, B., Zhang, Y., Chen, L., Wang, J., Yang, J., & Liu, Z. (2023). Otter: A multi-modal model with in-context instruction tuning. arXiv:2305.03726."},{"key":"2676_CR17","unstructured":"Gong, T., Lyu, C., Zhang, S., Wang, Y., Zheng, M., Zhao, Q., Liu, K., Zhang, W., Luo, P., & Chen, K. (2023). MultiModal-GPT: A Vision and Language Model for Dialogue with Humans."},{"key":"2676_CR18","unstructured":"Awadalla, A., Gao, I., Gardner, J., Hessel, J., Hanafy, Y., Zhu, W., Marathe, K., Bitton, Y., Gadre, S., Sagawa, S., Jitsev, J., Kornblith, S., Koh, P.W., Ilharco, G., Wortsman, M., & Schmidt, L. (2023). Openflamingo: An open-source framework for training large autoregressive vision-language models. arXiv:2308.01390."},{"key":"2676_CR19","unstructured":"Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., & Qiao, Y. (2023). Llama-adapter v2: Parameter-efficient visual instruction model. arXiv:2304.15010."},{"key":"2676_CR20","unstructured":"Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., & Wei, F. (2023). Kosmos-2: Grounding multimodal large language models to the world. arXiv:2306.14824."},{"key":"2676_CR21","unstructured":"Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., & Zhou, J. (2023). Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv:2308.12966"},{"key":"2676_CR22","unstructured":"Zhang, P., Wang, X.D.B., Cao, Y., Xu, C., Ouyang, L., Zhao, Z., Ding, S., Zhang, S., Duan, H., Yan, H. and others (2023). Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv:2309.15112."},{"key":"2676_CR23","unstructured":"Wang, G., Ge, Y., Ding, X., Kankanhalli, M., & Shan, Y. (2023). What makes for good visual tokenizers for large language models? arXiv:2305.12223."},{"key":"2676_CR24","unstructured":"Li, J., Li, D., Savarese, S., & Hoi, S. (2023). Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. ICML."},{"key":"2676_CR25","doi-asserted-by":"crossref","unstructured":"Dai, W., Li, J., Li, D., Tiong, A.M.H., Zhao, J., Wang, W., Li, B., Fung, P., & Hoi, S. (2023). Instructblip: Towards general-purpose vision-language models with instruction tuning. arXiv:2305.06500.","DOI":"10.52202\/075280-2142"},{"key":"2676_CR26","unstructured":"Zhu, D., Chen, J., Shen, X., Li, X., & Elhoseiny, M. (2023). Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592."},{"key":"2676_CR27","unstructured":"Liu, H., Li, C., Wu, Q., & Lee, Y.J. (2023). Visual instruction tuning. arXiv:2304.08485."},{"key":"2676_CR28","doi-asserted-by":"crossref","unstructured":"Liu, H., Li, C., Li, Y., & Lee, Y.J.(2023). Improved baselines with visual instruction tuning. arXiv:2310.03744.","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"2676_CR29","doi-asserted-by":"crossref","unstructured":"Ye, Q., Xu, H., Ye, J., Yan, M., Hu, A., Liu, H., Qian, Q., Zhang, J., Huang, F., & Zhou, J. (2023). mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.","DOI":"10.1109\/CVPR52733.2024.01239"},{"key":"2676_CR30","unstructured":"Lu, H., Liu, W., Zhang, B., Wang, B., Dong, K., Liu, B., Sun, J., Ren, T., Li, Z., Sun, Y. and others (2024). Deepseek-vl: towards real-world vision-language understanding. arXiv:2403.05525."},{"key":"2676_CR31","unstructured":"Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J. and others (2024). Yi: Open foundation models by 01. ai. arXiv:2403.04652."},{"key":"2676_CR32","doi-asserted-by":"crossref","unstructured":"Zhang, H., Li, X., & Bing, L. (2023). Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv:2306.02858.","DOI":"10.18653\/v1\/2023.emnlp-demo.49"},{"key":"2676_CR33","unstructured":"Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., & Qiao, Y. (2023). Videochat: Chat-centric video understanding. arXiv:2305.06355."},{"key":"2676_CR34","doi-asserted-by":"crossref","unstructured":"Maaz, M., Rasheed, H., Khan, S., & Khan, F.S. (2023). Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv:2306.05424.","DOI":"10.18653\/v1\/2024.acl-long.679"},{"key":"2676_CR35","unstructured":"Ge, Y., Ge, Y., Zeng, Z., Wang, X., & Shan, Y. (2023). Planting a seed of vision in large language model. arXiv:2307.08041."},{"key":"2676_CR36","unstructured":"Ge, Y., Zhao, S., Zeng, Z., Ge, Y., Li, C., Wang, X., & Shan, Y. (2023). Making llama see and draw with seed tokenizer. arXiv:2310.01218."},{"key":"2676_CR37","unstructured":"Ge, Y., Zhao, S., Zhu, J., Ge, Y., Yi, K., Song, L., Li, C., Ding, X., & Shan, Y. (2024). Seed-x: Multimodal models with unified multi-granularity comprehension and generation. arXiv:2404.14396."},{"key":"2676_CR38","unstructured":"Gpt-4v(ision) system card. (2023). https:\/\/api.semanticscholar.org\/CorpusID:263218031"},{"key":"2676_CR39","unstructured":"Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., Hauth, A. and others (2023). Gemini: a family of highly capable multimodal models. arXiv:2312.11805."},{"key":"2676_CR40","unstructured":"Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., and others (2023). Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv:2306.13394."},{"key":"2676_CR41","first-page":"26650","volume":"36","author":"Z Yin","year":"2024","unstructured":"Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Huang, X., Wang, Z., Sheng, L., Bai, L., et al. (2024). Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark. Advances in Neural Information Processing Systems,36, 26650\u201326685.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2676_CR42","doi-asserted-by":"crossref","unstructured":"Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z. and others (2023). Mmbench: Is your multi-modal model an all-around player? arXiv:2307.06281.","DOI":"10.1007\/978-3-031-72658-3_13"},{"key":"2676_CR43","unstructured":"Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., & Wang, L. (2023). Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv:2308.02490."},{"key":"2676_CR44","doi-asserted-by":"crossref","unstructured":"Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y. and others (2023). Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv:2311.16502.","DOI":"10.1109\/CVPR52733.2024.00913"},{"key":"2676_CR45","unstructured":"Xu, P., Shao, W., Zhang, K., Gao, P., Liu, S., Lei, M., Meng, F., Huang, S., Qiao, Y.,& Luo, P. (2023). Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models. arXiv:2306.09265."},{"key":"2676_CR46","unstructured":"Ying, K., Meng, F., Wang, J., Li, Z., Lin, H., Yang, Y., Zhang, H., Zhang, W., Lin, Y., Liu, S. and others (2024). Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi. arXiv:2404.16006."},{"key":"2676_CR47","unstructured":"Cheng, S., Guo, Z., Wu, J., Fang, K., Li, P., Liu, H., & Liu, Y. (2023). Can vision-language models think from a first-person perspective? arXiv:2311.15596."},{"key":"2676_CR48","unstructured":"Chen, L., Zhang, Y., Ren, S., Zhao, H., Cai, Z., Wang, Y., Wang, P., Liu, T., & Chang, B. (2023). Towards end-to-end embodied decision making via multi-modal large language model: Explorations with gpt4-vision and beyond. arXiv:2310.02071."},{"key":"2676_CR49","doi-asserted-by":"crossref","unstructured":"Sener, F., & Yao, A. (2019). Zero-shot anticipation for instructional activities. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 862\u2013871.","DOI":"10.1109\/ICCV.2019.00095"},{"key":"2676_CR50","doi-asserted-by":"crossref","unstructured":"Abu\u00a0Farha, Y., Ke, Q., Schiele, B., & Gall, J. (2021). Long-term anticipation of activities with cycle consistency. In: Pattern Recognition: 42nd DAGM German Conference, DAGM GCPR 2020, T\u00fcbingen, Germany, September 28\u2013October 1, 2020, Proceedings 42, pp. 159\u2013173 . Springer.","DOI":"10.1007\/978-3-030-71278-5_12"},{"key":"2676_CR51","doi-asserted-by":"publisher","first-page":"401","DOI":"10.1016\/j.jvcir.2017.10.004","volume":"49","author":"A Furnari","year":"2017","unstructured":"Furnari, A., Battiato, S., Grauman, K., & Farinella, G. M. (2017). Next-active-object prediction from egocentric videos. Journal of Visual Communication and Image Representation,49, 401\u2013411.","journal-title":"Journal of Visual Communication and Image Representation"},{"key":"2676_CR52","doi-asserted-by":"crossref","unstructured":"Liu, S., Tripathi, S., Majumdar, S., & Wang, X. (2022). Joint hand motion and interaction hotspots prediction from egocentric videos. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 3282\u20133292.","DOI":"10.1109\/CVPR52688.2022.00328"},{"key":"2676_CR53","unstructured":"Lotter, W., Kreiman, G., & Cox, D. (2016). Deep predictive coding networks for video prediction and unsupervised learning. In: International Conference on Learning Representations."},{"key":"2676_CR54","unstructured":"Villegas, R., Yang, J., Hong, S., Lin, X., & Lee, H. (2017). Decomposing motion and content for natural video sequence prediction. arXiv:1706.08033."},{"key":"2676_CR55","doi-asserted-by":"crossref","unstructured":"Mendonca, R., Bahl, S., & Pathak, D. (2023). Structured world models from human videos. arXiv:2308.10901.","DOI":"10.15607\/RSS.2023.XIX.012"},{"key":"2676_CR56","doi-asserted-by":"crossref","unstructured":"Lin, X., Petroni, F., Bertasius, G., Rohrbach, M., Chang, S.-F., & Torresani, L. (2022). Learning to recognize procedural activities with distant supervision. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 13853\u201313863.","DOI":"10.1109\/CVPR52688.2022.01348"},{"key":"2676_CR57","doi-asserted-by":"crossref","unstructured":"Zhong, Y., Yu, L., Bai, Y., Li, S., Yan, X., & Li, Y. (2023). Learning procedure-aware video representation from instructional videos and their narrations. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 14825\u201314835.","DOI":"10.1109\/CVPR52729.2023.01424"},{"key":"2676_CR58","doi-asserted-by":"crossref","unstructured":"Chang, C.-Y., Huang, D.-A., Xu, D., Adeli, E., Fei-Fei, L., & Niebles, J.C. (2020). Procedure planning in instructional videos. In: European Conference on Computer Vision, pp. 334\u2013350. Springer.","DOI":"10.1007\/978-3-030-58621-8_20"},{"key":"2676_CR59","doi-asserted-by":"crossref","unstructured":"Bi, J., Luo, J., & Xu, C. (2021). Procedure planning in instructional videos via contextual modeling and model-based policy learning. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 15611\u201315620.","DOI":"10.1109\/ICCV48922.2021.01532"},{"issue":"2","key":"2676_CR60","doi-asserted-by":"publisher","first-page":"4924","DOI":"10.1109\/LRA.2022.3150855","volume":"7","author":"J Sun","year":"2022","unstructured":"Sun, J., Huang, D.-A., Lu, B., Liu, Y.-H., Zhou, B., & Garg, A. (2022). Plate: Visually-grounded planning with transformers in procedural tasks. IEEE Robotics and Automation Letters,7(2), 4924\u20134930.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2676_CR61","doi-asserted-by":"crossref","unstructured":"Zhao, H., Hadji, I., Dvornik, N., Derpanis, K.G., Wildes, R.P., & Jepson, A.D. (2022). P3iv: Probabilistic procedure planning from instructional videos with weak supervision. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 2938\u20132948.","DOI":"10.1109\/CVPR52688.2022.00295"},{"key":"2676_CR62","doi-asserted-by":"crossref","unstructured":"Patel, D., Eghbalzadeh, H., Kamra, N., Iuzzolino, M.L., Jain, U., & Desai, R. (2023). Pretrained language models as visual planners for human assistance. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 15302\u201315314.","DOI":"10.1109\/ICCV51070.2023.01404"},{"key":"2676_CR63","doi-asserted-by":"crossref","unstructured":"Zhukov, D., Alayrac, J.-B., Cinbis, R.G., Fouhey, D., Laptev, I., & Sivic, J. (2019). Cross-task weakly supervised learning from instructional videos. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 3537\u20133545.","DOI":"10.1109\/CVPR.2019.00365"},{"key":"2676_CR64","doi-asserted-by":"crossref","unstructured":"Tang, Y., Ding, D., Rao, Y., Zheng, Y., Zhang, D., Zhao, L., Lu, J., & Zhou, J. (2019). Coin: A large-scale dataset for comprehensive instructional video analysis. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 1207\u20131216.","DOI":"10.1109\/CVPR.2019.00130"},{"key":"2676_CR65","doi-asserted-by":"crossref","unstructured":"Lai, B., Dai, X., Chen, L., Pang, G., Rehg, J.M., & Liu, M. (2023). Lego: Learning egocentric action frame generation via visual instruction tuning. arXiv:2312.03849.","DOI":"10.1007\/978-3-031-72673-6_8"},{"key":"2676_CR66","doi-asserted-by":"crossref","unstructured":"Pirsiavash, H., & Ramanan, D. (2012). Detecting activities of daily living in first-person camera views. In: 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 2847\u20132854. IEEE.","DOI":"10.1109\/CVPR.2012.6248010"},{"key":"2676_CR67","doi-asserted-by":"crossref","unstructured":"Li, Y., Ye, Z., & Rehg, J.M. (2015). Delving into egocentric actions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 287\u2013295.","DOI":"10.1109\/CVPR.2015.7298625"},{"key":"2676_CR68","unstructured":"Sigurdsson, G.A., Gupta, A., Schmid, C., Farhadi, A., & Alahari, K. (2018). Charades-ego: A large-scale dataset of paired third and first person videos. arXiv:1804.09626."},{"key":"2676_CR69","first-page":"38863","volume":"36","author":"Y Song","year":"2024","unstructured":"Song, Y., Byrne, E., Nagarajan, T., Wang, H., Martin, M., & Torresani, L. (2024). Ego4d goal-step: Toward hierarchical understanding of procedural activities. Advances in Neural Information Processing Systems,36, 38863\u201338886.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2676_CR70","unstructured":"OpenAI: GPT-4 Technical Report (2023)"},{"key":"2676_CR71","doi-asserted-by":"crossref","unstructured":"Lin, K.Q., Wang, J., Soldan, M., Wray, M., Yan, R., XU, E.Z., Gao, D., Tu, R.-C., Zhao, W., Kong, W. and others (2022). Egocentric video-language pretraining. Advances in Neural Information Processing Systems 35, 7575\u20137586.","DOI":"10.52202\/068431-0550"},{"key":"2676_CR72","unstructured":"Brohan, A., Chebotar, Y., Finn, C., Hausman, K., Herzog, A., Ho, D., Ibarz, J., Irpan, A., Jang, E., Julian, R. and others (2023). Do as i can, not as i say: Grounding language in robotic affordances. In: Conference on Robot Learning, pp. 287\u2013318. PMLR"},{"key":"2676_CR73","doi-asserted-by":"crossref","unstructured":"Lin, S., Hilton, J., & Evans, O. (2021). Truthfulqa: Measuring how models mimic human falsehoods. arXiv:2109.07958.","DOI":"10.18653\/v1\/2022.acl-long.229"},{"key":"2676_CR74","unstructured":"Luo, R., Zhao, Z., Yang, M., Dong, J., Qiu, M., Lu, P., Wang, T., & Wei, Z. (2023). Valley: Video assistant with large language model enhanced ability. arXiv:2306.07207."},{"key":"2676_CR75","unstructured":"Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X. and others (2023). Cogvlm: Visual expert for pretrained language models. arXiv:2311.03079."},{"key":"2676_CR76","doi-asserted-by":"crossref","unstructured":"Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., Luo, P. and others (2024). Mvbench: A comprehensive multi-modal video understanding benchmark. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 22195\u201322206.","DOI":"10.1109\/CVPR52733.2024.02095"},{"key":"2676_CR77","doi-asserted-by":"crossref","unstructured":"Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., & Cao, Y. (2023). Eva: Exploring the limits of masked visual representation learning at scale. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 19358\u201319369.","DOI":"10.1109\/CVPR52729.2023.01855"},{"key":"2676_CR78","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J. and others (2021). Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748\u20138763. PmLR."},{"key":"2676_CR79","doi-asserted-by":"publisher","unstructured":"Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., & Schmidt, L. OpenCLIP. If you use this software, please cite it as below. https:\/\/doi.org\/10.5281\/zenodo.5143773","DOI":"10.5281\/zenodo.5143773"},{"key":"2676_CR80","doi-asserted-by":"crossref","unstructured":"Zhai, X., Mustafa, B., Kolesnikov, A., & Beyer, L. (2023). Sigmoid loss for language image pre-training. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 11975\u201311986.","DOI":"10.1109\/ICCV51070.2023.01100"},{"key":"2676_CR81","doi-asserted-by":"crossref","unstructured":"Li, Y., Mao, H., Girshick, R., & He, K. (2022). Exploring plain vision transformer backbones for object detection. In: European Conference on Computer Vision, pp. 280\u2013296. Springer.","DOI":"10.1007\/978-3-031-20077-9_17"},{"key":"2676_CR82","unstructured":"Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., & Artzi, Y. (2019). Bertscore: Evaluating text generation with bert. In: International Conference on Learning Representations."},{"key":"2676_CR83","unstructured":"Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). Lora: Low-rank adaptation of large language models. arXiv:2106.09685."},{"key":"2676_CR84","first-page":"10935","volume":"36","author":"H Yuan","year":"2023","unstructured":"Yuan, H., Yuan, Z., Tan, C., Wang, W., Huang, S., & Huang, F. (2023). Rrhf: Rank responses to align language models with human feedback. Thirty-seventh Conference on Neural Information Processing Systems,36, 10935\u201310950.","journal-title":"Thirty-seventh Conference on Neural Information Processing Systems"},{"key":"2676_CR85","unstructured":"Hurst, A., Lerer, A., Goucher, A.P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A. and others (2024). Gpt-4o system card. arXiv:2410.21276."}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02676-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-025-02676-0","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02676-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T08:40:32Z","timestamp":1774600832000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-025-02676-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,12]]},"references-count":85,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,3]]}},"alternative-id":["2676"],"URL":"https:\/\/doi.org\/10.1007\/s11263-025-02676-0","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,12]]},"assertion":[{"value":"30 September 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 September 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 February 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 March 2026","order":5,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Update","order":6,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The original version is revised due to update in affiliation.","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"118"}}