{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,23]],"date-time":"2026-05-23T09:09:28Z","timestamp":1779527368134,"version":"3.53.1"},"reference-count":99,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T00:00:00Z","timestamp":1775606400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T00:00:00Z","timestamp":1775606400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Multimodal Large Language Models (MLLMs) has recently demonstrated superior multimodal comprehension abilities, heralding a new era for artificial general intelligence (AGI). However, achieving AGI necessitates more than just comprehension. A crucial capability required is effective planning in diverse scenarios, which involves making reasonable decisions based on complex environments to solve real-world problems. Despite its importance, the planning abilities of current MLLMs in varied scenarios remain underexplored, leaving a significant gap in our understanding of their full potential. In this paper, we introduce EgoPlan-Bench2, a rigorous and comprehensive benchmark designed to\n                    <jats:italic>assess the planning capabilities of MLLMs across a wide range of real-world scenarios<\/jats:italic>\n                    . EgoPlan-Bench2 encompasses everyday tasks spanning 4 major domains and 24 detailed scenarios, closely aligned with human daily life. It is constructed through a semi-automatic process utilizing egocentric videos, complemented by manual verification. Grounded in a first-person perspective, it mirrors the way humans approach problem-solving in everyday life. We evaluate 25 competitive MLLMs and provide an in-depth analysis of their limitations, revealing that they face significant challenges in real-world planning. To diagnose the underlying bottlenecks, we investigate the effectiveness of various prompts via a training-free multimodal prompting method. We find that MLLMs\u2019 planning performance on EgoPlan-Bench2 is critically dependent on temporally structured action sequences in historical task progress and interactions between objects and humans in current observation state. This dependency also underscores the necessity for strong reasoning abilities to integrate diverse multimodal cues and analysis before making final decision. Building on this insight, we demonstrate that EgoPlan-Bench2 is also an effective video reasoning benchmark. Experiments with Gemini-2.5-Flash and a post-trained Qwen-2.5-VL confirm its ability to distinguish between models with and without explicit deliberate reasoning mechanisms, showcasing the tangible impact of DeepSeek-R1 paradigm reasoning on planning tasks. We have made data and code available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/qiulu66.github.io\/egoplanbench2\/\" ext-link-type=\"uri\">https:\/\/qiulu66.github.io\/egoplanbench2\/<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s11263-026-02826-y","type":"journal-article","created":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T08:38:30Z","timestamp":1775637510000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios"],"prefix":"10.1007","volume":"134","author":[{"given":"Lu","family":"Qiu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuying","family":"Ge","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yixiao","family":"Ge","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ying","family":"Shan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1831-9952","authenticated-orcid":false,"given":"Xihui","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,4,8]]},"reference":[{"key":"2826_CR1","doi-asserted-by":"publisher","first-page":"23716","DOI":"10.52202\/068431-1723","volume":"35","author":"J-B Alayrac","year":"2022","unstructured":"Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. (2022). Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35, 23716\u201323736.","journal-title":"Advances in neural information processing systems"},{"key":"2826_CR2","unstructured":"Arcas, B.A., & Norvig, P. (2023) Artificial general intelligence is already here. Noema"},{"key":"2826_CR3","unstructured":"Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., & Zhou, J. (2023) Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966"},{"key":"2826_CR4","unstructured":"Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., ... Lin, J. (2025) Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923"},{"key":"2826_CR5","unstructured":"Bird, S., Klein, E., & Loper, E. (2009) Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit. \" O\u2019Reilly Media, Inc.\""},{"key":"2826_CR6","unstructured":"Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Doll\u00e1r, P., & Zitnick, C.L. (2015) Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325"},{"key":"2826_CR7","unstructured":"Chen, Y., Ge, Y., Ge, Y., Ding, M., Li, B., Wang, R., Xu, R., Shan, Y., & Liu, X. (2023) Egoplan-bench: Benchmarking egocentric embodied planning with multimodal large language models. arXiv preprint arXiv:2312.06722"},{"key":"2826_CR8","unstructured":"Chen, Y., Ge, Y., Wang, R., Ge, Y., Cheng, J., Shan, Y., & Liu, X. (2025) Grpo-care: Consistency-aware reinforcement learning for multimodal reasoning. arXiv preprint arXiv:2506.16141"},{"key":"2826_CR9","unstructured":"Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Cui, E., Zhu, J., Ye, S., Tian, H., & Liu, Z., et al. (2024) Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271"},{"key":"2826_CR10","doi-asserted-by":"crossref","unstructured":"Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., & Ma, Z., et al. (2024) How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821","DOI":"10.1007\/s11432-024-4231-5"},{"key":"2826_CR11","doi-asserted-by":"crossref","unstructured":"Chen, L., Wei, X., Li, J., Dong, X., Zhang, P., Zang, Y., Chen, Z., Duan, H., Lin, B., & Tang, Z., et al. (2024) Sharegpt4video: Improving video understanding and generation with better captions. arXiv preprint arXiv:2406.04325","DOI":"10.52202\/079017-0614"},{"key":"2826_CR12","doi-asserted-by":"crossref","unstructured":"Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., & Lu, L., et al. (2024) Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 24185\u201324198","DOI":"10.1109\/CVPR52733.2024.02283"},{"key":"2826_CR13","unstructured":"Cheng, S., Fang, K., Yu, Y., Zhou, S., Li, B., Tian, Y., Li, T., Han, L., & Liu, Y. (2024) Videgothink: Assessing egocentric video understanding capabilities for embodied ai. arXiv preprint arXiv:2410.11623"},{"key":"2826_CR14","unstructured":"Cheng, J., Ge, Y., Wang, T., Ge, Y., Liao, J., & Shan, Y. (2025) Video-holmes: Can mllm think like holmes for complex video reasoning? arXiv preprint arXiv:2505.21374"},{"key":"2826_CR15","unstructured":"Cheng, S., Guo, Z., Wu, J., Fang, K., Li, P., Liu, H., & Liu, Y. (2023) Can vision-language models think from a first-person perspective? arXiv preprint arXiv:2311.15596"},{"key":"2826_CR16","unstructured":"Cheng, Z., Leng, S., Zhang, H., Xin, Y., Li, X., Chen, G., Zhu, Y., Zhang, W., Luo, Z., & Zhao, D., et al. (2024) Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476"},{"key":"2826_CR17","unstructured":"Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., & Gonzalez, J.E., et al. (2023) Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https:\/\/vicuna.lmsys.org (accessed 14 April 2023) 2(3), 6"},{"issue":"240","key":"2826_CR18","first-page":"1","volume":"24","author":"A Chowdhery","year":"2023","unstructured":"Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. (2023). Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240), 1\u2013113.","journal-title":"Journal of Machine Learning Research"},{"key":"2826_CR19","doi-asserted-by":"crossref","unstructured":"Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P., & Hoi, S. (2023) Instructblip: Towards general-purpose vision-language models with instruction tuning. 2 arxiv 2023. arXiv preprint arXiv:2305.06500","DOI":"10.52202\/075280-2142"},{"issue":"11","key":"2826_CR20","doi-asserted-by":"publisher","first-page":"4125","DOI":"10.1109\/TPAMI.2020.2991965","volume":"43","author":"D Damen","year":"2021","unstructured":"Damen, D., Doughty, H., Farinella, G. M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., & Wray, M. (2021). The epic-kitchens dataset: Collection, challenges and baselines. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11), 4125\u20134141. https:\/\/doi.org\/10.1109\/TPAMI.2020.2991965","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2826_CR21","unstructured":"DeepMind, G. (2023) Gemini. https:\/\/deepmind.google\/technologies\/gemini\/"},{"key":"2826_CR22","unstructured":"Dosovitskiy, A. (2020) An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929"},{"key":"2826_CR23","unstructured":"Driess, D., Xia, F., Sajjadi, M.S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., & Yu, T., et al. (2023) Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378"},{"key":"2826_CR24","doi-asserted-by":"crossref","unstructured":"Fan, C. (2019) Egovqa-an egocentric video question answering benchmark dataset. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops, pp. 0\u20130","DOI":"10.1109\/ICCVW.2019.00536"},{"key":"2826_CR25","doi-asserted-by":"crossref","unstructured":"Fathi, A., Hodgins, J.K., & Rehg, J.M. (2012) Social interactions: A first-person perspective. In: 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1226\u20131233. IEEE","DOI":"10.1109\/CVPR.2012.6247805"},{"key":"2826_CR26","unstructured":"Feng, K., Gong, K., Li, B., Guo, Z., Wang, Y., Peng, T., Wu, J., Zhang, X., Wang, B., & Yue, X. (2025) Video-r1: Reinforcing video reasoning in mllms. arXiv preprint arXiv:2503.21776"},{"key":"2826_CR27","doi-asserted-by":"crossref","unstructured":"Fu, C., Dai, Y., Luo, Y., Li, L., Ren, S., Zhang, R., Wang, Z., Zhou, C., Shen, Y., & Zhang, M., et al. (2024) Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075","DOI":"10.1109\/CVPR52734.2025.02245"},{"key":"2826_CR28","unstructured":"Ge, Y., Ge, Y., Zeng, Z., Wang, X., & Shan, Y. (2023) Planting a seed of vision in large language model. arXiv preprint arXiv:2307.08041"},{"key":"2826_CR29","unstructured":"Ge, Y., Zhao, S., Zeng, Z., Ge, Y., Li, C., Wang, X., & Shan, Y. (2023) Making llama see and draw with seed tokenizer. arXiv preprint arXiv:2310.01218"},{"key":"2826_CR30","unstructured":"Ge, Y., Zhao, S., Zhu, J., Ge, Y., Yi, K., Song, L., Li, C., Ding, X., & Shan, Y. (2024) Seed-x: Multimodal models with unified multi-granularity comprehension and generation. arXiv preprint arXiv:2404.14396"},{"key":"2826_CR31","unstructured":"Gong, T., Lyu, C., Zhang, S., Wang, Y., Zheng, M., Zhao, Q., Liu, K., Zhang, W., Luo, P., & Chen, K. (2023) Multimodal-gpt: A vision and language model for dialogue with humans. arXiv preprint arXiv:2305.04790"},{"key":"2826_CR32","doi-asserted-by":"publisher","unstructured":"Grauman, K., Westbury, A., Byrne, E., Cartillier, V., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Kukreja, D., Liu, M., Liu, X., Martin, M., Nagarajan, T., Radosavovic, I., Ramakrishnan, S. K., Ryan, F., Sharma, J., Wray, M., \u2026 J. (2024). Ego4d: Around the world in 3,000 hours of egocentric video. IEEE Transactions on Pattern Analysis and Machine Intelligence,1\u201332,. https:\/\/doi.org\/10.1109\/TPAMI.2024.3381075","DOI":"10.1109\/TPAMI.2024.3381075"},{"key":"2826_CR33","unstructured":"Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., & Bi, X., et al. (2025) Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948"},{"issue":"12","key":"2826_CR34","doi-asserted-by":"publisher","first-page":"10284","DOI":"10.1109\/TPAMI.2024.3437288","volume":"46","author":"Y Guo","year":"2024","unstructured":"Guo, Y., Jiao, F., Shen, Z., Nie, L., & Kankanhalli, M. (2024). Unk-vqa: A dataset and a probe into the abstention ability of multi-modal large models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 10284\u201310296. https:\/\/doi.org\/10.1109\/TPAMI.2024.3437288","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2826_CR35","unstructured":"He, X., Feng, W., Zheng, K., Lu, Y., Zhu, W., Li, J., Fan, Y., Wang, J., Li, L., & Yang, Z., et al. (2024) Mmworld: Towards multi-discipline multi-faceted world model evaluation in videos. arXiv preprint arXiv:2406.08407"},{"key":"2826_CR36","doi-asserted-by":"crossref","unstructured":"Huang, J., Zhang, J., Jiang, K., Qiu, H., Zhang, X., Shao, L., Lu, S., & Tao, D. (2025) Visual instruction tuning towards general-purpose multimodal large language model: A survey. International Journal of Computer Vision, 1\u201339","DOI":"10.1007\/s11263-025-02572-7"},{"key":"2826_CR37","first-page":"3343","volume":"35","author":"B Jia","year":"2022","unstructured":"Jia, B., Lei, T., Zhu, S.-C., & Huang, S. (2022). Egotaskqa: Understanding human tasks in egocentric videos. Advances in Neural Information Processing Systems, 35, 3343\u20133360.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2826_CR38","doi-asserted-by":"crossref","unstructured":"Jiang, S., Zhang, Y., Zhou, C., Jin, Y., Feng, Y., Wu, J., & Liu, Z. (2024) Joint visual and text prompting for improved object-centric perception with multimodal large language models. arXiv preprint arXiv:2404.04514","DOI":"10.1007\/978-3-032-00274-7_17"},{"key":"2826_CR39","unstructured":"Kenton, J.D.M.-W.C., & Toutanova, L.K. (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of naacL-HLT, vol. 1, p. 2. Minneapolis, Minnesota"},{"key":"2826_CR40","doi-asserted-by":"crossref","unstructured":"Lee, Y.J., Ghosh, J., & Grauman, K. (2012) Discovering important people and objects for egocentric video summarization. In: 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1346\u20131353. IEEE","DOI":"10.1109\/CVPR.2012.6247820"},{"key":"2826_CR41","doi-asserted-by":"crossref","unstructured":"Li, B., Ge, Y., Chen, Y., Ge, Y., Zhang, R., & Shan, Y. (2024) Seed-bench-2-plus: Benchmarking multimodal large language models with text-rich visual comprehension. arXiv preprint arXiv:2404.16790","DOI":"10.1109\/CVPR52733.2024.01263"},{"key":"2826_CR42","doi-asserted-by":"crossref","unstructured":"Li, B., Ge, Y., Ge, Y., Wang, G., Wang, R., Zhang, R., & Shan, Y. (2024) Seed-bench: Benchmarking multimodal large language models. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 13299\u201313308","DOI":"10.1109\/CVPR52733.2024.01263"},{"key":"2826_CR43","unstructured":"Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., & Qiao, Y. (2023) Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355"},{"key":"2826_CR44","unstructured":"Li, J., Li, D., Savarese, S., & Hoi, S. (2023) Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In: International Conference on Machine Learning, pp. 19730\u201319742. PMLR"},{"key":"2826_CR45","doi-asserted-by":"crossref","unstructured":"Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., & Luo, P., et al. (2024) Mvbench: A comprehensive multi-modal video understanding benchmark. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 22195\u201322206","DOI":"10.1109\/CVPR52733.2024.02095"},{"key":"2826_CR46","doi-asserted-by":"crossref","unstructured":"Li, B., Wang, R., Wang, G., Ge, Y., Ge, Y., & Shan, Y. (2023) Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125","DOI":"10.1109\/CVPR52733.2024.01263"},{"key":"2826_CR47","doi-asserted-by":"crossref","unstructured":"Li, L., Wang, Y., Xu, R., Wang, P., Feng, X., Kong, L., & Liu, Q. (2024) Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models. arXiv preprint arXiv:2403.00231","DOI":"10.18653\/v1\/2024.acl-long.775"},{"key":"2826_CR48","unstructured":"Li, X., Yan, Z., Meng, D., Dong, L., Zeng, X., He, Y., Wang, Y., Qiao, Y., Wang, Y., & Wang, L. (2025) Videochat-r1: Enhancing spatio-temporal perception via reinforcement fine-tuning. arXiv preprint arXiv:2504.06958"},{"key":"2826_CR49","doi-asserted-by":"crossref","unstructured":"Li, W., Yuan, Y., Liu, J., Tang, D., Wang, S., Qin, J., Zhu, J., & Zhang, L. (2025) Tokenpacker: Efficient visual projector for multimodal llm. International Journal of Computer Vision, 1\u201319","DOI":"10.1007\/s11263-025-02491-7"},{"key":"2826_CR50","doi-asserted-by":"crossref","unstructured":"Lin, J., Yin, H., Ping, W., Molchanov, P., Shoeybi, M., & Han, S. (2024) Vila: On pre-training for visual language models. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 26689\u201326699","DOI":"10.1109\/CVPR52733.2024.02520"},{"key":"2826_CR51","first-page":"7575","volume":"35","author":"KQ Lin","year":"2022","unstructured":"Lin, K. Q., Wang, J., Soldan, M., Wray, M., Yan, R., Xu, E. Z., Gao, D., Tu, R.-C., Zhao, W., Kong, W., et al. (2022). Egocentric video-language pretraining. Advances in Neural Information Processing Systems, 35, 7575\u20137586.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2826_CR52","doi-asserted-by":"crossref","unstructured":"Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., & Liu, Z., et al. (2023) Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281","DOI":"10.1007\/978-3-031-72658-3_13"},{"key":"2826_CR53","doi-asserted-by":"crossref","unstructured":"Liu, H., Li, C., Li, Y., & Lee, Y.J. (2024) Improved baselines with visual instruction tuning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 26296\u201326306","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"2826_CR54","unstructured":"Liu, H., Li, C., Wu, Q., & Lee, Y.J. (2024) Visual instruction tuning. Advances in neural information processing systems36"},{"key":"2826_CR55","doi-asserted-by":"crossref","unstructured":"Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., & Zhu, J., et al. (2023) Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499","DOI":"10.1007\/978-3-031-72970-6_3"},{"key":"2826_CR56","unstructured":"Lu, H., Liu, W., Zhang, B., Wang, B., Dong, K., Liu, B., Sun, J., Ren, T., Li, Z., & Sun, Y., et al. (2024) Deepseek-vl: towards real-world vision-language understanding. arXiv preprint arXiv:2403.05525"},{"key":"2826_CR57","unstructured":"Luo, R., Zhao, Z., Yang, M., Dong, J., Li, D., Lu, P., Wang, T., Hu, L., Qiu, M., & Wei, Z. (2023) Valley: Video assistant with large language model enhanced ability. arXiv preprint arXiv:2306.07207"},{"key":"2826_CR58","doi-asserted-by":"crossref","unstructured":"Maaz, M., Rasheed, H., Khan, S., & Khan, F.S. (2023) Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424","DOI":"10.18653\/v1\/2024.acl-long.679"},{"key":"2826_CR59","doi-asserted-by":"crossref","unstructured":"Mitra, C., Huang, B., Darrell, T., & Herzig, R. (2024) Compositional chain-of-thought prompting for large multimodal models. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 14420\u201314431","DOI":"10.1109\/CVPR52733.2024.01367"},{"key":"2826_CR60","unstructured":"Morris, M.R., Sohl-Dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., & Legg, S. (2023) Levels of agi: Operationalizing progress on the path to agi. arXiv preprint arXiv:2311.02462"},{"key":"2826_CR61","unstructured":"OpenAI (2023) Gpt-4v(ision) system card. https:\/\/openai.com\/index\/gpt-4v-system-card\/"},{"key":"2826_CR62","unstructured":"Patraucean, V., Smaira, L., Gupta, A., Recasens, A., Markeeva, L., Banarse, D., Koppula, S., Malinowski, M., Yang, Y., & Doersch, C., et al. (2024) Perception test: A diagnostic benchmark for multimodal video models. Advances in Neural Information Processing Systems 36"},{"key":"2826_CR63","unstructured":"Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., & Wei, F. (2023) Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824"},{"key":"2826_CR64","doi-asserted-by":"crossref","unstructured":"Pirsiavash, H., & Ramanan, D. (2012) Detecting activities of daily living in first-person camera views. In: 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 2847\u20132854. IEEE","DOI":"10.1109\/CVPR.2012.6248010"},{"key":"2826_CR65","unstructured":"Qin, M., Liu, X., Liang, Z., Shu, Y., Yuan, H., Zhou, J., Xiao, S., Zhao, B., & Liu, Z. (2025) Video-xl-2: Towards very long-video understanding through task-aware kv sparsification. arXiv preprint arXiv:2506.19225"},{"key":"2826_CR66","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., & Clark, J., et al. (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748\u20138763. PMLR"},{"key":"2826_CR67","unstructured":"Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., Khedr, H., R\u00e4dle, R., Rolland, C., & Gustafson, L., et al. (2024) Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714"},{"key":"2826_CR68","unstructured":"Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., & Wu, Y., et al. (2024) Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300"},{"key":"2826_CR69","unstructured":"Shi, L., Lv, Q., Deng, X., & Nie, L. (2024) Epd: Long-term memory extraction, context-awared planning and multi-iteration decision@ egoplan challenge icml 2024. arXiv preprint arXiv:2407.19510"},{"key":"2826_CR70","unstructured":"Sigurdsson, G.A., Gupta, A., Schmid, C., Farhadi, A., & Alahari, K. (2018) Charades-ego: A large-scale dataset of paired third and first person videos. arXiv preprint arXiv:1804.09626"},{"key":"2826_CR71","doi-asserted-by":"crossref","unstructured":"Su, Y.-C., & Grauman, K. (2016) Detecting engagement in egocentric video. In: Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14, pp. 454\u2013471. Springer","DOI":"10.1007\/978-3-319-46454-1_28"},{"key":"2826_CR72","unstructured":"Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., & Hauth, A., et al. (2023) Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805"},{"key":"2826_CR73","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., & Azhar, F., et al. (2023) Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971"},{"key":"2826_CR74","unstructured":"Wang, W., Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Zhu, J., Zhu, X., Lu, L., Qiao, Y., & Dai, J. (2024) Enhancing the reasoning ability of multimodal large language models via mixed preference optimization. arXiv preprint arXiv:2411.10442"},{"key":"2826_CR75","unstructured":"Wang, Y., Li, X., Yan, Z., He, Y., Yu, J., Zeng, X., Wang, C., Ma, C., Huang, H., Gao, J., Dou, M., Chen, K., Wang, W., Qiao, Y., Wang, Y., & Wang, L. (2025) Internvideo2.5: Empowering video mllms with long and rich context modeling. arXiv preprint arXiv:2501.12386"},{"key":"2826_CR76","doi-asserted-by":"crossref","unstructured":"Wang, B., Zhang, J., Dong, S., Fang, I., & Feng, C. (2024) Vlm see, robot do: Human demo video to robot action plan via vision language model. arXiv preprint arXiv:2410.08792","DOI":"10.1109\/IROS60139.2025.11246682"},{"key":"2826_CR77","first-page":"8483","volume":"35","author":"Z Wang","year":"2022","unstructured":"Wang, Z., Li, M., Xu, R., Zhou, L., Lei, J., Lin, X., Wang, S., Yang, Z., Zhu, C., Hoiem, D., et al. (2022). Language models with image descriptors are strong few-shot video-language learners. Advances in Neural Information Processing Systems, 35, 8483\u20138497.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2826_CR78","unstructured":"Xu, J., Guo, Z., He, J., Hu, H., He, T., Bai, S., Chen, K., Wang, J., Fan, Y., & Dang, K., et al. (2025) Qwen2.5-omni technical report. arXiv preprint arXiv:2503.20215"},{"key":"2826_CR79","doi-asserted-by":"crossref","unstructured":"Xu, G., Jin, P., Wu, Z., Li, H., Song, Y., Sun, L., & Yuan, L. (2024) Llava-cot: Let vision language models reason step-by-step. arXiv preprint arXiv:2411.10440","DOI":"10.1109\/ICCV51701.2025.00202"},{"key":"2826_CR80","unstructured":"Xu, P., Shao, W., Zhang, K., Gao, P., Liu, S., Lei, M., Meng, F., Huang, S., Qiao, Y., & Luo, P. (2023) Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models. arXiv preprint arXiv:2306.09265"},{"key":"2826_CR81","unstructured":"Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., & Shi, Y., et al. (2023) mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178"},{"key":"2826_CR82","doi-asserted-by":"crossref","unstructured":"Ye, Q., Xu, H., Ye, J., Yan, M., Hu, A., Liu, H., Qian, Q., Zhang, J., & Huang, F. (2024) mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 13040\u201313051","DOI":"10.1109\/CVPR52733.2024.01239"},{"key":"2826_CR83","unstructured":"Ye, H., Zhang, H., Daxberger, E., Chen, L., Lin, Z., Li, Y., Zhang, B., You, H., Xu, D., & Gan, Z., et al. (2024) Mm-ego: Towards building egocentric multimodal llms for video qa. arXiv preprint arXiv:2410.07177"},{"key":"2826_CR84","unstructured":"Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Huang, X., Wang, Z., Sheng, L., & Bai, L., et al. (2024) Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark. Advances in Neural Information Processing Systems 36"},{"key":"2826_CR85","unstructured":"Ying, K., Meng, F., Wang, J., Li, Z., Lin, H., Yang, Y., Zhang, H., Zhang, W., Lin, Y., & Liu, S., et al. (2024) Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi. arXiv preprint arXiv:2404.16006"},{"key":"2826_CR86","unstructured":"Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., & Chang, J., et al. (2024) Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652"},{"key":"2826_CR87","unstructured":"Yuan, Y., Dang, R., Li, L., Li, W., Jiao, D., Li, X., Zhao, D., Wang, F., Zhang, W., & Xiao, J., et al. (2025) Eoc-bench: Can mllms identify, recall, and forecast objects in an egocentric world? arXiv preprint arXiv:2506.05287"},{"key":"2826_CR88","doi-asserted-by":"crossref","unstructured":"Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., & Sun, Y., et al. (2024) Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 9556\u20139567","DOI":"10.1109\/CVPR52733.2024.00913"},{"issue":"5","key":"2826_CR89","doi-asserted-by":"publisher","first-page":"3156","DOI":"10.1109\/TPAMI.2023.3339661","volume":"46","author":"Y Zeng","year":"2024","unstructured":"Zeng, Y., Zhang, X., Li, H., Wang, J., Zhang, J., & Zhou, W. (2024). $$\\text{ X}^{2}$$2-vlm: All-in-one pre-trained model for vision-language tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(5), 3156\u20133168. https:\/\/doi.org\/10.1109\/TPAMI.2023.3339661","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2826_CR90","doi-asserted-by":"crossref","unstructured":"Zhang, H., Dong, L., Liu, Y., Huang, Y., Wang, Y., Wang, L., & Qiao, Y. (2025) Lvbench: A benchmark for long-form video understanding with versatile multi-modal question answering. International Journal of Computer Vision, 1\u201322","DOI":"10.1007\/s11263-025-02555-8"},{"key":"2826_CR91","doi-asserted-by":"crossref","unstructured":"Zhang, H., Li, X., & Bing, L. (2023) Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858","DOI":"10.18653\/v1\/2023.emnlp-demo.49"},{"key":"2826_CR92","unstructured":"Zhang, Y., Li, B., Liu, h., Lee, Y.j., Gui, L., Fu, D., Feng, J., Liu, Z., & Li, C. (2024) LLaVA-NeXT: A Strong Zero-shot Video Understanding Model. https:\/\/llava-vl.github.io\/blog\/2024-04-30-llava-next-video\/"},{"key":"2826_CR93","unstructured":"Zhang, Y., Zhang, K., Li, B., Pu, F., Setiadharma, C.A., Yang, J., & Liu, Z. (2024) Worldqa: Multimodal world knowledge in videos through long-chain reasoning. arXiv preprint arXiv:2405.03272"},{"key":"2826_CR94","unstructured":"Zhang, P., Zhang, K., Li, B., Zeng, G., Yang, J., Zhang, Y., Wang, Z., Tan, H., Li, C., & Liu, Z. (2024) Long context transfer from language to vision. arXiv preprint arXiv:2406.16852"},{"issue":"8","key":"2826_CR95","doi-asserted-by":"publisher","first-page":"5625","DOI":"10.1109\/TPAMI.2024.3369699","volume":"46","author":"J Zhang","year":"2024","unstructured":"Zhang, J., Huang, J., Jin, S., & Lu, S. (2024). Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8), 5625\u20135644. https:\/\/doi.org\/10.1109\/TPAMI.2024.3369699","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"12","key":"2826_CR96","doi-asserted-by":"publisher","first-page":"10404","DOI":"10.1109\/TPAMI.2024.3445770","volume":"46","author":"Z Zhang","year":"2024","unstructured":"Zhang, Z., Wu, H., Zhang, E., Zhai, G., & Lin, W. (2024). Q-$$\\text{ bench}^+$$+: A benchmark for multi-modal foundation models on low-level vision from single images to pairs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 10404\u201310418. https:\/\/doi.org\/10.1109\/TPAMI.2024.3445770","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2826_CR97","doi-asserted-by":"crossref","unstructured":"Zhou, J., Shu, Y., Zhao, B., Wu, B., Xiao, S., Yang, X., Xiong, Y., Zhang, B., Huang, T., & Liu, Z. (2024) Mlvu: A comprehensive benchmark for multi-task long video understanding. arXiv preprint arXiv:2406.04264","DOI":"10.1109\/CVPR52734.2025.01278"},{"key":"2826_CR98","unstructured":"Zhou, Q., Zhou, R., Hu, Z., Lu, P., Gao, S., & Zhang, Y.(2024) Image-of-thought prompting for visual reasoning refinement in multimodal large language models. arXiv preprint arXiv:2405.13872"},{"key":"2826_CR99","unstructured":"Zhu, D., Chen, J., Shen, X., Li, X., & Elhoseiny, M. (2023) Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02826-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-026-02826-y","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02826-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,23]],"date-time":"2026-05-23T08:43:40Z","timestamp":1779525820000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-026-02826-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,8]]},"references-count":99,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5]]}},"alternative-id":["2826"],"URL":"https:\/\/doi.org\/10.1007\/s11263-026-02826-y","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,8]]},"assertion":[{"value":"10 September 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 March 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 April 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"222"}}