{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T22:30:06Z","timestamp":1779921006337,"version":"3.53.1"},"reference-count":42,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,3,17]],"date-time":"2025-03-17T00:00:00Z","timestamp":1742169600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,3,17]],"date-time":"2025-03-17T00:00:00Z","timestamp":1742169600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"published-print":{"date-parts":[[2025,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Human-centered dynamic scene understanding plays a pivotal role in enhancing the capability of robotic and autonomous systems, where video-based human-object interaction (V-HOI) detection is a crucial task in semantic scene understanding, which aims to comprehensively understand HOI relationships within a video to benefit the behavioral decisions of mobile robots and autonomous driving systems. Although previous V-HOI detection models have made significant advances in accurate detection on specific datasets, they still lack the general reasoning ability of humans to effectively induce HOI relationships. In this study, we propose V-HOI multi-LLMs collaborated reasoning (V-HOI MLCR), a novel framework consisting of a series of plug-and-play modules that could facilitate the performance of current V-HOI detection models by leveraging the strong reasoning ability of different off-the-shelf pre-trained large language models (LLMs). We design a two-stage collaboration system of different LLMs for the V-HOI task. Specifically, in the first stage, we design a cross-agents reasoning scheme to leverage the LLM to perform reasoning from different aspects. In the second stage, we perform multi-LLMs debate to get the final reasoning answer based on the different knowledge in different LLMs. Additionally, we develop an auxiliary training strategy using CLIP, a large vision-language model to enhance the base V-HOI models\u2019 discriminative ability to better cooperate with LLMs. We validate the superiority of our design by demonstrating its effectiveness in improving the predictive accuracy of the base V-HOI model through reasoning from multiple perspectives.<\/jats:p>","DOI":"10.1007\/s44267-025-00074-1","type":"journal-article","created":{"date-parts":[[2025,3,17]],"date-time":"2025-03-17T03:14:46Z","timestamp":1742181286000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Enhancing human-centered dynamic scene understanding via multiple LLMs collaborated reasoning"],"prefix":"10.1007","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-1203-3363","authenticated-orcid":false,"given":"Hang","family":"Zhang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8680-6010","authenticated-orcid":false,"given":"Wenxiao","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5054-3394","authenticated-orcid":false,"given":"Haoxuan","family":"Qu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4365-4165","authenticated-orcid":false,"given":"Jun","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,3,17]]},"reference":[{"key":"74_CR1","first-page":"8231","volume-title":"Proceedings of the 2023 IEEE international conference on robotics and automation","author":"J. Wang","year":"2023","unstructured":"Wang, J., Huang, J., Zhang, C., & Deng, Z. (2023). Cross-modality time-variant relation learning for generating dynamic scene graphs. In Proceedings of the 2023 IEEE international conference on robotics and automation (pp. 8231\u20138238). Piscataway: IEEE."},{"key":"74_CR2","first-page":"1585","volume-title":"Proceedings of the conference on robot learning","author":"X. Li","year":"2021","unstructured":"Li, X., Guo, D., Liu, H., & Sun, F. (2021). Embodied semantic scene graph generation. In Proceedings of the conference on robot learning (pp. 1585\u20131594). PMLR."},{"issue":"1","key":"74_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TPAMI.2021.3137605","volume":"45","author":"X. Chang","year":"2021","unstructured":"Chang, X., Ren, P., Xu, P., Li, Z., Chen, X., & Hauptmann, A. (2021). A comprehensive survey of scene graphs: generation and application. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1), 1\u201326.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"74_CR4","unstructured":"Li, X., Bai, Y., Cai, P., Wen, L., Fu, D., Zhang, B., Yang, X., Cai, X., Ma, T., Guo, J., et\u00a0al. (2023). Towards knowledge-driven autonomous driving. arXiv preprint. arXiv:2312.04316."},{"key":"74_CR5","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2023.103741","volume":"233","author":"Z. Ni","year":"2023","unstructured":"Ni, Z., Mascar\u00f3, E.V., Ahn, H., & Lee, D. (2023). Human\u2013object interaction prediction in videos through gaze following. Computer Vision and Image Understanding, 233, 103741.","journal-title":"Computer Vision and Image Understanding"},{"key":"74_CR6","first-page":"718","volume-title":"Proceedings of the 16th European conference on computer vision","author":"D.-J. Kim","year":"2020","unstructured":"Kim, D.-J., Sun, X., Choi, J., Lin, S., & Kweon, I. S. (2020). Detecting human-object interactions with action co-occurrence priors. In A. Vedaldi, H. Bischof, T. Brox, & J.-M. Frahm (Eds.), Proceedings of the 16th European conference on computer vision (pp. 718\u2013736). Cham: Springer."},{"key":"74_CR7","first-page":"10163","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y.-L. Li","year":"2020","unstructured":"Li, Y.-L., Liu, X., Lu, H., Wang, S., Liu, J., Li, J., & Lu, C. (2020). Detailed 2D-3D joint representation for human-object interaction. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10163\u201310172). Piscataway: IEEE."},{"key":"74_CR8","first-page":"5011","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"Y.-L. Li","year":"2020","unstructured":"Li, Y.-L., Liu, X., Wu, X., Li, Y., & Lu, C. (2020). HOI analysis: integrating and decomposing human-object interaction. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, & H. Lin (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 5011\u20135022). Red Hook: Curran Associates."},{"key":"74_CR9","first-page":"13617","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"O. Ulutan","year":"2020","unstructured":"Ulutan, O., Iftekhar, A. S. M., & Vsgnet, B. S. M. (2020). Spatial attention network for detecting human object interactions using graph convolutions. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 13617\u201313626). Piscataway: IEEE."},{"key":"74_CR10","first-page":"374","volume-title":"Proceedings of the 17th European conference on computer vision","author":"L. Xu","year":"2022","unstructured":"Xu, L., Qu, H., Kuen, J., Gu, J., & Liu, J. (2022). Meta spatio-temporal debiasing for video scene graph generation. In S. Avidan, G. Brostow, M. Ciss\u00e9, G. M. Farinella, & T. Hassner (Eds.), Proceedings of the 17th European conference on computer vision (pp. 374\u2013390). Cham: Springer."},{"key":"74_CR11","first-page":"69","volume-title":"Proceedings of the 16th European conference on computer vision","author":"X. Zhong","year":"2020","unstructured":"Zhong, X., Ding, C., Qu, X., & Tao, D. (2020). Polysemy deciphering network for human-object interaction detection. In A. Vedaldi, H. Bischof, T. Brox, & J.-M. Frahm (Eds.), Proceedings of the 16th European conference on computer vision (pp. 69\u201385). Cham: Springer."},{"key":"74_CR12","first-page":"482","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Liao","year":"2020","unstructured":"Liao, Y., Liu, S., Wang, F., Chen, Y., Qian, C., & Feng, J. (2020). PPDM: parallel point detection and matching for real-time human-object interaction detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 482\u2013490). Piscataway: IEEE."},{"key":"74_CR13","first-page":"4116","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"T. Wang","year":"2020","unstructured":"Wang, T., Yang, T., Danelljan, M., Khan, F.S., Zhang, X., & Sun, J. (2020). Learning human-object interaction detection using interaction points. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 4116\u20134125). Piscataway: IEEE."},{"key":"74_CR14","first-page":"498","volume-title":"Proceedings of the 16th European conference on computer vision","author":"B. Kim","year":"2020","unstructured":"Kim, B., Choi, T., Kang, J., & Kim, H. J. (2020). UnionDet: union-level detector towards real-time human-object interaction detection. In A. Vedaldi, H. Bischof, T. Brox, & J.-M. Frahm (Eds.), Proceedings of the 16th European conference on computer vision (pp. 498\u2013514). Cham: Springer."},{"key":"74_CR15","first-page":"1291","volume-title":"Proceedings of the AAAI conference on artificial intelligence","author":"H.-S. Fang","year":"2021","unstructured":"Fang, H.-S., Xie, Y., Shao, D., & Lu, C. (2021). Dirv: dense interaction region voting for end-to-end human-object interaction detection. In Proceedings of the AAAI conference on artificial intelligence (pp. 1291\u20131299). Palo Alto: AAAI Press."},{"key":"74_CR16","first-page":"5308","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"A. Jain","year":"2016","unstructured":"Jain, A., Zamir, A. R., Savarese, S., & Saxena, A. (2016). Structural-RNN: deep learning on spatio-temporal graphs. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5308\u20135317). Piscataway: IEEE."},{"key":"74_CR17","doi-asserted-by":"publisher","first-page":"691","DOI":"10.1145\/3394171.3413778","volume-title":"Proceedings of the 28th ACM international conference on multimedia","author":"S.P.R. Sunkesula","year":"2020","unstructured":"Sunkesula, S.P.R., Dabral, R., & Ramakrishnan, G. (2020). Lighten: learning interactions with graph and hierarchical temporal networks for hoi in videos. In Proceedings of the 28th ACM international conference on multimedia (pp. 691\u2013699). New York: ACM."},{"key":"74_CR18","doi-asserted-by":"publisher","first-page":"4985","DOI":"10.1145\/3474085.3475636","volume-title":"Proceedings of the 29th ACM international conference on multimedia","author":"N. Wang","year":"2021","unstructured":"Wang, N., Zhu, G., Zhang, L., Shen, P., Li, H., & Hua, C. (2021). Spatio-temporal interaction graph parsing networks for human-object interaction recognition. In Proceedings of the 29th ACM international conference on multimedia (pp. 4985\u20134993). New York: ACM."},{"key":"74_CR19","first-page":"1","volume-title":"Proceedings of the 2021 IEEE international conference on multimedia and expo","author":"X. Sun","year":"2021","unstructured":"Sun, X., He, Y., Ren, T., & Wu, G. (2021). Spatial-temporal human-object interaction detection. In Proceedings of the 2021 IEEE international conference on multimedia and expo (pp. 1\u20136). Piscataway: IEEE."},{"key":"74_CR20","first-page":"9","volume-title":"Proceedings of the 2021 ACM workshop on intelligent cross-data analysis and retrieval","author":"M.-J. Chiou","year":"2021","unstructured":"Chiou, M.-J., Liao, C.-Y., Wang, L.-W., Zimmermann, R., & Feng, J. (2021). ST-HOI: a spatial-temporal baseline for human-object interaction detection in videos. In Proceedings of the 2021 ACM workshop on intelligent cross-data analysis and retrieval (pp. 9\u201317). New York: ACM."},{"key":"74_CR21","first-page":"8106","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"J. Ji","year":"2021","unstructured":"Ji, J., Desai, R., & Niebles, J. C. (2021). Detecting human-object relationships in videos. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 8106\u20138116). Piscataway: IEEE."},{"key":"74_CR22","first-page":"16372","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Y. Cong","year":"2021","unstructured":"Cong, Y., Liao, W., Ackermann, H., Rosenhahn, B., & Yang, M.Y. (2021). Spatial-temporal transformer for dynamic scene graph generation. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 16372\u201316382). Piscataway: IEEE."},{"key":"74_CR23","first-page":"23345","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"D. Tu","year":"2022","unstructured":"Tu, D., Sun, W., Min, X., Zhai, G., & Shen, W. (2022). Video-based human-object interaction detection from tubelet tokens. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Proceedings of the 36th international conference on neural information processing systems (pp. 23345\u201323357). Red Hook: Curran Associates."},{"key":"74_CR24","doi-asserted-by":"crossref","unstructured":"Guo, Z., Tang, Y., Zhang, R., Wang, D., Wang, Z., Zhao, B., & Li, X. (2023). Viewrefer: Grasp the multi-view knowledge for 3d visual grounding with gpt and prototype guidance. arXiv preprint. arXiv:2303.16894.","DOI":"10.1109\/ICCV51070.2023.01410"},{"key":"74_CR25","first-page":"1","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"Z. Zhao","year":"2024","unstructured":"Zhao, Z., Lee, W.S., & Hsu, D. (2024). Large language models as commonsense knowledge for large-scale task planning. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 1\u201321). Red Hook: Curran Associates."},{"key":"74_CR26","first-page":"8748","volume-title":"Proceedings of the 38th international conference on machine learning","author":"A. Radford","year":"2021","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In M. Meila & T. Zhang (Eds.), Proceedings of the 38th international conference on machine learning (pp. 8748\u20138763). PMLR."},{"key":"74_CR27","first-page":"23507","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Ning","year":"2023","unstructured":"Ning, S., Qiu, L., Liu, Y., & Hoiclip, X. He. (2023). Efficient knowledge transfer for hoi detection with vision-language models. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 23507\u201323517). Piscataway: IEEE."},{"key":"74_CR28","first-page":"1","volume-title":"Proceedings of the 2017 14th IEEE international conference on advanced video and signal based surveillance","author":"A. M. Truong","year":"2017","unstructured":"Truong, A. M., & Yoshitaka, A. (2017). Structured lstm for human-object interaction detection and anticipation. In Proceedings of the 2017 14th IEEE international conference on advanced video and signal based surveillance (pp. 1\u20136). Piscataway: IEEE."},{"key":"74_CR29","first-page":"401","volume-title":"Proceedings of the European conference on computer vision","author":"S. Qi","year":"2018","unstructured":"Qi, S., Wang, W., Jia, B., Shen, J., & Zhu, S.-C. (2018). Learning human-object interactions by graph parsing neural networks. In M. Hebert, V. Ferrari, C. Sminchisescu, & Y. Weiss (Eds.), Proceedings of the European conference on computer vision (pp. 401\u2013417). Cham: Springer."},{"key":"74_CR30","doi-asserted-by":"publisher","first-page":"1049","DOI":"10.18653\/v1\/2023.findings-acl.67","volume-title":"Proceedings of the findings of the association for computational linguistics: ACL 2023","author":"J. Huang","year":"2023","unstructured":"Huang, J., & Chang, K. C.-C. (2023). Towards reasoning in large language models: a survey. In Proceedings of the findings of the association for computational linguistics: ACL 2023 (pp. 1049\u20131065). Stroudsburg: ACL."},{"key":"74_CR31","unstructured":"Li, B., Corona, R., Mangalam, K., Chen, C., Flaherty, D., Belongie, S., Weinberger, K. Q., Malik, J., Darrell, T., & Klein, D. (2022). A vision-free baseline for multimodal grammar induction. arXiv preprint. arXiv:2212.10564."},{"key":"74_CR32","first-page":"24824","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"J. Wei","year":"2022","unstructured":"Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 24824\u201324837). Red Hook: Curran Associates."},{"key":"74_CR33","first-page":"19730","volume-title":"Proceedings of the 40th international conference on machine learning","author":"J. Li","year":"2023","unstructured":"Li, J., Li, D., Savarese, S., & Hoi, S. (2023). BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, & J. Scarlett (Eds.), Proceedings of the 40th international conference on machine learning (pp. 19730\u201319742). PMLR."},{"key":"74_CR34","first-page":"43447","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"P. Lu","year":"2023","unstructured":"Lu, P., Peng, B., Cheng, H., Galley, M., Chang, K.-W., Wu, Y.N., Zhu, S.-C., & Gao, J. (2023). Chameleon: plug-and-play compositional reasoning with large language models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Proceedings of the 36th international conference on neural information processing systems (pp. 43447\u201343478). Red Hook: Curran Associates."},{"key":"74_CR35","first-page":"23716","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"J.-B. Alayrac","year":"2022","unstructured":"Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. (2022). Flamingo: a visual language model for few-shot learning. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 23716\u201323736). Red Hook: Curran Associates."},{"key":"74_CR36","unstructured":"Rozanova, J., Ferreira, D., Dubba, K., Cheng, W., Zhang, D., & Freitas, A. (2021). Grounding natural language instructions: Can large language models capture spatial information? arXiv preprint. arXiv:2109.08634."},{"key":"74_CR37","doi-asserted-by":"publisher","first-page":"71","DOI":"10.18653\/v1\/W19-1608","volume-title":"Proceedings of the combined workshop on spatial language understanding and grounded communication for robotics","author":"M. Ghanimifard","year":"2019","unstructured":"Ghanimifard, M., & Dobnik, S. (2019). What a neural language model tells us about spatial relations. In A. Bhatia, Y. Bisk, P. Kordjamshidi, & J. Thomason (Eds.), Proceedings of the combined workshop on spatial language understanding and grounded communication for robotics (pp. 71\u201381). Stroudsburg: ACL."},{"key":"74_CR38","first-page":"18225","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"W. Feng","year":"2023","unstructured":"Feng, W., Zhu, W., Fu, T.-J., Jampani, V., Akula, A., He, X., Basu, S., Wang, X. E., & Wang, W. Y. (2023). LayoutGPT: compositional visual planning and generation with large language models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Proceedings of the 36th international conference on neural information processing systems (pp. 18225\u201318250). Red Hook: Curran Associates."},{"key":"74_CR39","first-page":"8359","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"G. Gkioxari","year":"2018","unstructured":"Gkioxari, G., Girshick, R., Doll\u00e1r, P., & He, K. (2018). Detecting and recognizing human-object interactions. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 8359\u20138367). Piscataway: IEEE."},{"key":"74_CR40","doi-asserted-by":"publisher","first-page":"17889","DOI":"10.18653\/v1\/2024.emnlp-main.992","volume-title":"Proceedings of the 2024 conference on empirical methods in natural language processing","author":"T. Liang","year":"2024","unstructured":"Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Shi, S., & Tu, Z. (2024). Encouraging divergent thinking in large language models through multi-agent debate. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings of the 2024 conference on empirical methods in natural language processing (pp. 17889\u201317904). Stroudsburg: ACL."},{"key":"74_CR41","volume-title":"Proceedings of the 12th international conference on learning representations","author":"Z. Gou","year":"2024","unstructured":"Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Duan, N., & Chen, W. (2024). CRITIC: large language models can self-correct with tool-interactive critiquing. In Proceedings of the 12th international conference on learning representations. OpenReview.net."},{"key":"74_CR42","first-page":"10236","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. Ji","year":"2020","unstructured":"Ji, J., Krishna, R., Fei-Fei, L., & Niebles, J. C. (2020). Action genome: actions as compositions of spatio-temporal scene graphs. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10236\u201310247). Piscataway: IEEE."}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00074-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-025-00074-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00074-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,17]],"date-time":"2025-03-17T03:15:03Z","timestamp":1742181303000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-025-00074-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,17]]},"references-count":42,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["74"],"URL":"https:\/\/doi.org\/10.1007\/s44267-025-00074-1","relation":{},"ISSN":["2097-3330","2731-9008"],"issn-type":[{"value":"2097-3330","type":"print"},{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,17]]},"assertion":[{"value":"25 September 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 February 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 February 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 March 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Jun Liu is an Associate Editor at Visual Intelligence and was not involved in the editorial review of this article or the decision to publish it. The authors declare that they have no other competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"3"}}