{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T15:53:09Z","timestamp":1776441189215,"version":"3.51.2"},"reference-count":75,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T00:00:00Z","timestamp":1764547200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T00:00:00Z","timestamp":1765152000000},"content-version":"vor","delay-in-days":7,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Beijing Natural Science Foundation","award":["L242019"],"award-info":[{"award-number":["L242019"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"published-print":{"date-parts":[[2025,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In the gaze behavior understanding task, existing vision-based models demonstrate inherent limitations in high-dimensional semantic understanding, while vision-language models (VLMs) encounter challenges in precise object localization. To address this issue, we propose GazeLLM, the first zero-shot large language model (LLM) boosted framework for gaze target reasoning. Our key innovations include three aspects. First, we have structured object extraction. Using off-the-shelf detectors (e.g., MM-GroundingDINO and Depth Anything V2), we convert images into 3D object representations, including head and gaze direction, object categories, and metric depth. Second, we implemented an autonomous chain-of-thought (CoT) reasoning system. We designed self-generated CoT prompts to guide pretrained LLMs, such as ChatGPT o3-mini-high, to predict gaze targets via spatial-semantic analysis. Third, we proposed a plug-and-play module. We employed a novel cross-modal fusion mechanism that combines the LLM\u2019s probability dictionaries with vision-based gaze heatmaps via Gaussian-weighted multi-hot mapping. Extensive experiments show that GazeLLM significantly improves state-of-the-art models, increasing their performance from 17% to 34% on challenging cases, such as long-range targets or rare categories, without the need for retraining. It also extends seamlessly to multi-person social gaze tasks (e.g., a 42% LAEO AP gain on the AVA-LAEO benchmark). Our framework demonstrates superior generalizability and interpretability compared to VLMs, validating the efficacy of LLMs in understanding gaze behavior by mining semantic cues.<\/jats:p>","DOI":"10.1007\/s44267-025-00101-1","type":"journal-article","created":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T02:29:59Z","timestamp":1765160999000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["GazeLLM: a plug-and-play zero-shot LLM reasoning framework for boosting gaze target detection"],"prefix":"10.1007","volume":"3","author":[{"given":"Yaokun","family":"Yang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9064-7964","authenticated-orcid":false,"given":"Feng","family":"Lu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,12,8]]},"reference":[{"key":"101_CR1","doi-asserted-by":"publisher","first-page":"2033","DOI":"10.1109\/TPAMI.2014.2313123","volume":"36","author":"F. Lu","year":"2014","unstructured":"Lu, F., Sugano, Y., Okabe, T., & Sato, Y. (2014). Adaptive linear regression for appearance-based gaze estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36, 2033\u20132046.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"101_CR2","first-page":"4511","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"X. Zhang","year":"2015","unstructured":"Zhang, X., Sugano, Y., Fritz, M., & Bulling, A. (2015). Appearance-based gaze estimation in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4511\u20134520). Piscataway: IEEE."},{"key":"101_CR3","first-page":"365","volume-title":"Proceedings of the 16th European conference on computer vision","author":"X. Zhang","year":"2020","unstructured":"Zhang, X., Park, S., Beeler, T., Bradley, D., Tang, S., & Hilliges, O. (2020). ETH-XGaze: a large scale dataset for gaze estimation under extreme head pose and gaze variation. In A. Vedaldi, H. Bischof, T. Brox, & J. Frahm (Eds.), Proceedings of the 16th European conference on computer vision (pp. 365\u2013381). Cham: Springer."},{"key":"101_CR4","doi-asserted-by":"publisher","first-page":"7509","DOI":"10.1109\/TPAMI.2024.3393571","volume":"46","author":"Y. Cheng","year":"2024","unstructured":"Cheng, Y., Wang, H., Bao, Y., & Lu, F. (2024). Appearance-based gaze estimation with deep learning: a review and benchmark. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46, 7509\u20137528.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"101_CR5","first-page":"6333","volume-title":"Proceedings of the 38th AAAI conference on artificial intelligence","author":"M. Xu","year":"2024","unstructured":"Xu, M., & Lu, F. (2024). Gaze from origin: learning for generalized gaze estimation by embedding the gaze frontalization process. In M. J. Wooldridge, J. G. Dy, & S. Natarajan (Eds.), Proceedings of the 38th AAAI conference on artificial intelligence (pp. 6333\u20136341). Palo Alto: AAAI Press."},{"key":"101_CR6","first-page":"1409","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Bao","year":"2024","unstructured":"Bao, Y., & Lu, F. (2024). From feature to gaze: a generalizable replacement of linear layer for gaze estimation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 1409\u20131418). Piscataway: IEEE."},{"key":"101_CR7","first-page":"1419","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Bao","year":"2024","unstructured":"Bao, Y., & Lu, F. (2024). Unsupervised gaze representation learning from multi-view face images. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 1419\u20131428). Piscataway: IEEE."},{"key":"101_CR8","first-page":"3693","volume-title":"Proceedings of the 38th AAAI conference on artificial intelligence","author":"R. Liu","year":"2024","unstructured":"Liu, R., & Lu, F. (2024). Uvagaze: unsupervised 1-to-2 views adaptation for gaze estimation. In M. J. Wooldridge, J. G. Dy, & S. Natarajan (Eds.), Proceedings of the 38th AAAI conference on artificial intelligence (pp. 3693\u20133701). Palo Alto: AAAI Press."},{"key":"101_CR9","doi-asserted-by":"publisher","first-page":"1290","DOI":"10.1007\/s11263-024-02233-1","volume":"133","author":"R. Liu","year":"2025","unstructured":"Liu, R., Wang, H., & Lu, F. (2025). From gaze jitter to domain adaptation: generalizing gaze estimation by manipulating high-frequency components. International Journal of Computer Vision, 133, 1290\u20131305.","journal-title":"International Journal of Computer Vision"},{"key":"101_CR10","first-page":"199","volume-title":"Proceedings of the 29th international conference on neural information processing systems","author":"A. Recasens","year":"2015","unstructured":"Recasens, A., Khosla, A., Vondrick, C., & Torralba, A. (2015). Where are they looking? In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, & R. Garnett (Eds.), Proceedings of the 29th international conference on neural information processing systems (pp. 199\u2013207). Red Hook: Curran Associates."},{"key":"101_CR11","first-page":"35","volume-title":"Proceedings of the 14th Asian conference on computer vision","author":"D. Lian","year":"2018","unstructured":"Lian, D., Yu, Z., & Gao, S. (2018). Believe it or not, we know what you are looking at! In C. V. Jawahar, H. Li, G. Mori, & K. Schindler (Eds.), Proceedings of the 14th Asian conference on computer vision (pp. 35\u201350). Cham: Springer."},{"key":"101_CR12","first-page":"383","volume-title":"Proceedings of the 15th European conference on computer vision","author":"E. Chong","year":"2018","unstructured":"Chong, E., Ruiz, N., Wang, Y., Zhang, Y., Rozga, A., & Rehg, J. M. (2018). Connecting gaze, scene, and attention: generalized attention estimation via joint modeling of gaze and scene saliency. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Proceedings of the 15th European conference on computer vision (pp. 383\u2013398). Cham: Springer."},{"key":"101_CR13","first-page":"5396","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"E. Chong","year":"2020","unstructured":"Chong, E., Wang, Y., Ruiz, N., & Rehg, J. M. (2020). Detecting attended visual targets in video. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 5396\u20135406). Piscataway: IEEE."},{"key":"101_CR14","first-page":"6460","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"L. Fan","year":"2018","unstructured":"Fan, L., Chen, Y., Wei, P., Wang, W., & Zhu, S.-C. (2018). Inferring shared attention in social scene videos. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6460\u20136468). Piscataway: IEEE."},{"key":"101_CR15","first-page":"3477","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"M. J. Marin-Jimenez","year":"2019","unstructured":"Marin-Jimenez, M. J., Kalogeiton, V., Medina-Suarez, P., & Zisserman, A. (2019). LAEO-Net: revisiting people looking at each other in videos. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3477\u20133485). Piscataway: IEEE."},{"key":"101_CR16","first-page":"11390","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Fang","year":"2021","unstructured":"Fang, Y., Tang, J., Shen, W., Shen, W., Gu, X., Song, L., & Zhai, G. (2021). Dual attention guided gaze target detection in the wild. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11390\u201311399). Piscataway: IEEE."},{"key":"101_CR17","first-page":"14126","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. Bao","year":"2022","unstructured":"Bao, J., Liu, B., & Yu, J. (2022). ESCNet: gaze target detection with the understanding of 3D scenes. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 14126\u201314135). Piscataway: IEEE."},{"key":"101_CR18","first-page":"5041","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"A. Gupta","year":"2022","unstructured":"Gupta, A., Tafasca, S., & Odobez, J.-M. (2022). A modular multimodal architecture for gaze target prediction: application to privacy-sensitive settings. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 5041\u20135050). Piscataway: IEEE."},{"key":"101_CR19","first-page":"880","volume-title":"Proceedings of the IEEE\/CVF winter conference on applications of computer vision","author":"Q. Miao","year":"2023","unstructured":"Miao, Q., Hoai, M., & Samaras, D. (2023). Patch-level gaze distribution prediction for gaze following. In Proceedings of the IEEE\/CVF winter conference on applications of computer vision (pp. 880\u2013889). Piscataway: IEEE."},{"key":"101_CR20","first-page":"20935","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"S. Tafasca","year":"2023","unstructured":"Tafasca, S., Gupta, A., & Odobez, J.-M. (2023). Childplay: a new benchmark for understanding children\u2019s gaze behaviour. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 20935\u201320946). Piscataway: IEEE."},{"key":"101_CR21","first-page":"2008","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Tafasca","year":"2024","unstructured":"Tafasca, S., Gupta, A., & Odobez, J.-M. (2024). Sharingan: a transformer architecture for multi-person gaze following. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 2008\u20132017). Piscataway: IEEE."},{"key":"101_CR22","first-page":"1","volume-title":"Proceedings of the 38th international conference on neural information processing systems","author":"S. Tafasca","year":"2024","unstructured":"Tafasca, S., Gupta, A., Bros, V., & Odobez, J.-M. (2024). Toward semantic gaze target detection. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, & C. Zhang (Eds.), Proceedings of the 38th international conference on neural information processing systems. (pp. 1\u201327). Red Hook: Curran Associates."},{"key":"101_CR23","first-page":"15646","volume-title":"Proceedings of the 38th international conference on neural information processing systems","author":"A. Gupta","year":"2024","unstructured":"Gupta, A., Tafasca, S., Farkhondeh, A., Vuillecard, P., & Odobez, J. M. (2024). MTGS: a novel framework for multi-person temporal gaze following and social gaze prediction. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, & C. Zhang (Eds.), Proceedings of the 38th international conference on neural information processing systems. (pp. 15646\u201315673). Red Hook: Curran Associates."},{"key":"101_CR24","first-page":"28874","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"F. Ryan","year":"2024","unstructured":"Ryan, F., Bati, A., Lee, S., Bolya, D., Hoffman, J., & Rehg, J. M. (2024). Gaze-LLE: gaze target estimation via large-scale learned encoders. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 28874\u201328884). Piscataway: IEEE."},{"key":"101_CR25","first-page":"26296","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"H. Liu","year":"2024","unstructured":"Liu, H., Li, C., Li, Y., & Lee, Y. J. (2024). Improved baselines with visual instruction tuning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 26296\u201326306). Piscataway: IEEE."},{"key":"101_CR26","first-page":"1","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"H. Liu","year":"2023","unstructured":"Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2023). Visual instruction tuning. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Proceedings of the 37th international conference on neural information processing systems. (pp. 1\u201325). Red Hook: Curran Associates."},{"key":"101_CR27","first-page":"19730","volume-title":"Proceedings of the international conference on machine learning","author":"J. Li","year":"2023","unstructured":"Li, J., Hu, D., Yang, J., Shen, X., Keutzer, K., Yuille, A., Gao, J., & Hariharan, B. (2023). Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, & J. Scarlett (Eds.), Proceedings of the international conference on machine learning (pp. 19730\u201319742). Retrieved November 19, 2025, from https:\/\/proceedings.mlr.press\/v202\/li23q.html."},{"key":"101_CR28","unstructured":"Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., & Wei, F. (2023). Kosmos-2: grounding multimodal large language models to the world. arXiv preprint. arXiv:2306.14824."},{"key":"101_CR29","first-page":"1","volume-title":"Proceedings of the 12th international conference on learning representations","author":"D. Zhu","year":"2024","unstructured":"Zhu, D., Chen, J., Shen, X., Li, X., & Elhoseiny, M. (2024). MiniGPT-4: enhancing vision-language understanding with advanced large language models. In Proceedings of the 12th international conference on learning representations (pp. 1\u201317). Retrieved November 19, 2025, from https:\/\/openreview.net\/forum?id=1tZbq88f27."},{"key":"101_CR30","unstructured":"Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., & Zhao, R. (2023). Shikra: unleashing multimodal LLM\u2019s referential dialogue magic. arXiv preprint. arXiv:2306.15195."},{"key":"101_CR31","unstructured":"Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et\u00a0al. (2024). Qwen2-VL: enhancing vision-language model\u2019s perception of the world at any resolution. arXiv preprint. arXiv:2409.12191."},{"key":"101_CR32","unstructured":"Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et\u00a0al. (2025). Qwen2.5-VL technical report. arXiv preprint. arXiv:2502.13923."},{"key":"101_CR33","unstructured":"Wu, Z., Chen, X., Pan, Z., Liu, X., Liu, W., Dai, D., Gao, H., Ma, Y., Wu, C., Wang, B., et\u00a0al. (2024). Deepseek-VL2: mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint. arXiv:2412.10302."},{"key":"101_CR34","doi-asserted-by":"publisher","DOI":"10.1007\/s44267-024-00067-6","volume":"2","author":"Z. Gao","year":"2024","unstructured":"Gao, Z., Chen, Z., Cui, E., Ren, Y., Wang, W., Zhu, J., Tian, H., Ye, S., He, J., Zhu, X., et al. (2024). Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance. Visual Intelligence, 2, 32.","journal-title":"Visual Intelligence"},{"key":"101_CR35","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et\u00a0al. (2023). GPT-4 technical report. arXiv preprint. arXiv:2303.08774."},{"key":"101_CR36","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., Azhar, F., et\u00a0al. (2023). LLaMA: open and efficient foundation language models. arXiv preprint. arXiv:2302.13971."},{"key":"101_CR37","unstructured":"Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et\u00a0al. (2023). Qwen technical report. arXiv preprint. arXiv:2309.16609."},{"key":"101_CR38","unstructured":"Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et\u00a0al. (2024). Qwen2.5 technical report. arXiv preprint. arXiv:2412.15115."},{"key":"101_CR39","unstructured":"Liu, A., Feng, B., Wang, B., Wang, B., Liu, B., Zhao, C., Deng, C., Ruan, C., Dai, D., Guo, D., et\u00a0al. (2024). DeepSeek-V2: a strong, economical, and efficient mixture-of-experts language model. arXiv preprint. arXiv:2405.04434."},{"key":"101_CR40","unstructured":"Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et\u00a0al. (2024). DeepSeek-V3 technical report. arXiv preprint. arXiv:2412.19437."},{"key":"101_CR41","unstructured":"Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et\u00a0al. (2025). DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint. arXiv:2501.12948."},{"key":"101_CR42","unstructured":"OpenAI (2025). ChatGPT o3-mini. Retrieved November 19, 2025, from https:\/\/openai.com\/index\/openai-o3-mini."},{"key":"101_CR43","first-page":"10867","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. Guo","year":"2023","unstructured":"Guo, J., Li, J., Li, D., Tiong, A. M. H., Li, B., Tao, D., & Hoi, S. (2023). From images to textual prompts: zero-shot visual question answering with frozen large language models. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10867\u201310877). Piscataway: IEEE."},{"key":"101_CR44","unstructured":"Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., & Duan, N. (2023). Visual ChatGPT: talking, drawing and editing with visual foundation models. arXiv preprint. arXiv:2303.04671."},{"key":"101_CR45","unstructured":"Mao, J., Qian, Y., Ye, J., Zhao, H., & Wang, Y. (2023). GPT-Driver: learning to drive with GPT. arXiv preprint. arXiv:2310.01415."},{"key":"101_CR46","first-page":"1","volume-title":"Proceedings of the 38th international conference on neural information processing systems","author":"X. Wu","year":"2024","unstructured":"Wu, X., Li, Y.-L., Sun, J., & Lu, C. (2024). Symbol-LLM: leverage language models for symbolic system in visual human activity reasoning. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Proceedings of the 38th international conference on neural information processing systems (pp. 1\u201312). Red Hook: Curran Associates."},{"key":"101_CR47","doi-asserted-by":"publisher","first-page":"3707","DOI":"10.1109\/TPAMI.2023.3348528","volume":"46","author":"R. Liu","year":"2024","unstructured":"Liu, R., Liu, Y., Wang, H., & Lu, F. (2024). PnP-GA+: plug-and-play domain adaptation for gaze estimation using model variants. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46, 3707\u20133721.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"101_CR48","first-page":"6585","volume-title":"Proceedings of the 38th AAAI conference on artificial intelligence","author":"Y. Yang","year":"2024","unstructured":"Yang, Y., Yin, Y., & Lu, F. (2024). Gaze target detection by merging human attention and activity cues. In M. J. Wooldridge, J. G. Dy, & S. Natarajan (Eds.), Proceedings of the 38th AAAI conference on artificial intelligence (pp. 6585\u20136593). Palo Alto: AAAI Press."},{"key":"101_CR49","first-page":"305","volume-title":"Proceedings of the 18th European conference on computer vision","author":"Y. Yang","year":"2024","unstructured":"Yang, Y., & Lu, F. (2024). Gaze target detection based on head-local-global coordination. In A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, & G. Varol (Eds.), Proceedings of the 18th European conference on computer vision (pp. 305\u2013322). Cham: Springer."},{"key":"101_CR50","doi-asserted-by":"publisher","DOI":"10.1007\/s44267-024-00064-9","volume":"2","author":"Y. Song","year":"2024","unstructured":"Song, Y., Wang, X., Yao, J., Liu, W., Zhang, J., & Xu, X. (2024). ViTGaze: gaze following with interaction features in vision transformers. Visual Intelligence, 2, 31.","journal-title":"Visual Intelligence"},{"key":"101_CR51","first-page":"2192","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"D. Tu","year":"2022","unstructured":"Tu, D., Min, X., Duan, H., Guo, G., Zhai, G., & Shen, W. (2022). End-to-end human-gaze-target detection with transformers. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 2192\u20132200). Piscataway: IEEE."},{"key":"101_CR52","first-page":"21860","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"F. Tonini","year":"2023","unstructured":"Tonini, F., Dall\u2019Asen, N., Beyan, C., & Ricci, E. (2023). Object-aware gaze target detection. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 21860\u201321869). Piscataway: IEEE."},{"key":"101_CR53","doi-asserted-by":"publisher","first-page":"3271","DOI":"10.1109\/TCSVT.2023.3318839","volume":"34","author":"D. Tu","year":"2024","unstructured":"Tu, D., Shen, W., Sun, W., Min, X., Zhai, G., & Chen, C. (2024). Un-gaze: a unified transformer for joint gaze-location and gaze-object detection. IEEE Transactions on Circuits and Systems for Video Technology, 34, 3271\u20133285.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"101_CR54","first-page":"4171","volume-title":"Proceedings of the 2019 conference of the North American chapter of the Association for Computational Linguistics: human language technologies","author":"J. Devlin","year":"2019","unstructured":"Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 conference of the North American chapter of the Association for Computational Linguistics: human language technologies (pp. 4171\u20134186). Stroudsburg: ACL."},{"key":"101_CR55","first-page":"1","volume":"21","author":"C. Raffel","year":"2020","unstructured":"Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21, 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"101_CR56","first-page":"1877","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"T. Brown","year":"2020","unstructured":"Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, & H. Lin (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 1877\u20131901). Red Hook: Curran Associates."},{"key":"101_CR57","unstructured":"Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et\u00a0al. (2023). PaLM2 technical report. arXiv preprint. arXiv:2305.10403."},{"key":"101_CR58","doi-asserted-by":"publisher","DOI":"10.1007\/s44267-024-00050-1","volume":"2","author":"Y. Jiang","year":"2024","unstructured":"Jiang, Y., Yan, X., Ji, G.-P., Fu, K., Sun, M., Xiong, H., Fan, D.-P., & Khan, F. S. (2024). Effectiveness assessment of recent large vision-language models. Visual Intelligence, 2, 17.","journal-title":"Visual Intelligence"},{"key":"101_CR59","doi-asserted-by":"publisher","DOI":"10.1007\/s44267-025-00074-1","volume":"3","author":"H. Zhang","year":"2025","unstructured":"Zhang, H., Zhang, W., Qu, H., & Liu, J. (2025). Enhancing human-centered dynamic scene understanding via multiple LLMs collaborated reasoning. Visual Intelligence, 3, 3.","journal-title":"Visual Intelligence"},{"key":"101_CR60","doi-asserted-by":"publisher","DOI":"10.1007\/s44267-024-00065-8","volume":"2","author":"X. Tu","year":"2024","unstructured":"Tu, X., He, Z., Huang, Y., Zhang, Z.-H., Yang, M., & Zhao, J. (2024). An overview of large AI models and their applications. Visual Intelligence, 2, 34.","journal-title":"Visual Intelligence"},{"key":"101_CR61","doi-asserted-by":"publisher","first-page":"2337","DOI":"10.1007\/s11263-022-01653-1","volume":"130","author":"K. Zhou","year":"2022","unstructured":"Zhou, K., Yang, J., Loy, C. C., & Liu, Z. (2022). Learning to prompt for vision-language models. International Journal of Computer Vision, 130, 2337\u20132348.","journal-title":"International Journal of Computer Vision"},{"key":"101_CR62","unstructured":"Zhao, X., Chen, Y., Xu, S., Li, X., Wang, X., Li, Y., & Huang, H. (2024). An open and comprehensive pipeline for unified object grounding and detection. arXiv preprint. arXiv:2401.02361."},{"key":"101_CR63","first-page":"5356","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"A. Gupta","year":"2019","unstructured":"Gupta, A., Dollar, P., & Girshick, R. (2019). LVIS: a dataset for large vocabulary instance segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 5356\u20135364). Piscataway: IEEE."},{"key":"101_CR64","first-page":"4197","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"M. Zhang","year":"2022","unstructured":"Zhang, M., Liu, Y., & Lu, F. (2022). Gazeonce: real-time multi-person gaze estimation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 4197\u20134206). Piscataway: IEEE."},{"key":"101_CR65","first-page":"1","volume-title":"Proceedings of the 38th international conference on neural information processing systems","author":"L. Yang","year":"2024","unstructured":"Yang, L., Kang, B., Huang, Z., Zhao, Z., Xu, X., Feng, J., & Zhao, H. (2024). Depth anything v2. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, & C. Zhang (Eds.), Proceedings of the 38th international conference on neural information processing systems. (pp. 1\u201337). Red Hook: Curran Associates."},{"key":"101_CR66","first-page":"24824","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"J. Wei","year":"2022","unstructured":"Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Proceedings of the 36th international conference on neural information processing systems (pp. 24824\u201324837). Red Hook: Curran Associates."},{"key":"101_CR67","first-page":"1","volume-title":"Proceedings of the 11th international conference on learning representations","author":"X. Wang","year":"2023","unstructured":"Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., & Zhou, D. (2023). Self-consistency improves chain of thought reasoning in language models. In Proceedings of the 11th international conference on learning representations (pp. 1\u201324). Retrieved November 19, 2025, from https:\/\/openreview.net\/forum?id=1PL1NIMMrw."},{"key":"101_CR68","unstructured":"Zhang, Z., Zhang, A., Li, M., & Smola, A. (2022). Automatic chain of thought prompting in large language models. arXiv preprint. arXiv:2210.03493."},{"key":"101_CR69","first-page":"13003","volume-title":"Findings of the Association for Computational Linguistics","author":"M. Suzgun","year":"2023","unstructured":"Suzgun, M., Scales, N., Sch\u00e4rli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al. (2023). Challenging big-bench tasks and whether chain-of-thought can solve them. In A. Rogers, J. L. Boyd-Graber, & N. Okazaki (Eds.), Findings of the Association for Computational Linguistics (pp. 13003\u201313051). Stroudsburg: ACL."},{"key":"101_CR70","unstructured":"Zhang, Z., Zhang, A., Li, M., Zhao, H., Karypis, G., & Smola, A. (2023). Multimodal chain-of-thought reasoning in language models. arXiv preprint. arXiv:2302.00923."},{"key":"101_CR71","doi-asserted-by":"publisher","first-page":"3069","DOI":"10.1109\/TPAMI.2020.3048482","volume":"44","author":"M. J. Marin-Jimenez","year":"2022","unstructured":"Marin-Jimenez, M. J., Kalogeiton, V., Medina-Suarez, P., & Zisserman, A. (2022). LAEO-Net++: revisiting people looking at each other in videos. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44, 3069\u20133081.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"101_CR72","first-page":"2106","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"T. Judd","year":"2009","unstructured":"Judd, T., Ehinger, K., Durand, F., & Torralba, A. (2009). Learning to predict where humans look. In Proceedings of the IEEE international conference on computer vision (pp. 2106\u20132113). Piscataway: IEEE."},{"key":"101_CR73","unstructured":"Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et\u00a0al. (2025). Qwen3 technical report. arXiv preprint. arXiv:2505.09388."},{"key":"101_CR74","unstructured":"Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Cui, E., Zhu, J., Ye, S., Tian, H., Liu, Z., et\u00a0al. (2024). Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint. arXiv:2412.05271."},{"issue":"1","key":"101_CR75","volume":"69","author":"F. Lu","year":"2026","unstructured":"Lu, F., & Zhao, Q. (2026). Towards cobodied\/symbodied AI: concept and eight scientific and technical problems. Science China. Information Sciences, 69(1), 116101.","journal-title":"Science China. Information Sciences"}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00101-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-025-00101-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00101-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T04:15:02Z","timestamp":1765167302000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-025-00101-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12]]},"references-count":75,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["101"],"URL":"https:\/\/doi.org\/10.1007\/s44267-025-00101-1","relation":{},"ISSN":["2097-3330","2731-9008"],"issn-type":[{"value":"2097-3330","type":"print"},{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12]]},"assertion":[{"value":"29 August 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 November 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 November 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 December 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"26"}}