{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T18:29:21Z","timestamp":1772908161845,"version":"3.50.1"},"reference-count":71,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2025,5,7]],"date-time":"2025-05-07T00:00:00Z","timestamp":1746576000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,5,7]],"date-time":"2025-05-07T00:00:00Z","timestamp":1746576000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005950","name":"Hong Kong University of Science and Technology","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005950","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Recent efforts to use natural language for interpretable driving focus mainly on planning, neglecting perception tasks. In this paper, we address this gap by introducing ROLISP (Risk Object Localization and Intention and Suggestion Prediction), which towards interpretable risk object detection and suggestion for ego car motions. Accurate ROLISP implementation requires extensive reasoning to identify critical traffic objects and infer their intentions, prompting us to explore the capabilities of multimodal large language models (MLLMs). However, the limited perception performance of CLIP-ViT vision encoders in existing MLLMs struggles with capturing essential visual perception information,\u00a0e.g., high-resolution, multi-scale and visual-related inductive biases, which are important for autonomous driving. Addressing these challenges, we introduce HiLM-D, a resource-efficient framework that enhances visual information processing in MLLMs for ROLISP. Our method is motivated by the fact that the primary variations in autonomous driving scenarios are the motion trajectories rather than the semantic or appearance information (e.g., the shapes and colors) of objects. Hence, the visual process of HiLM-D\u00a0is a two-stream framework: (i) a temporal reasoning stream, receiving low-resolution dynamic video content, to capture temporal semantics, and (ii) a spatial perception stream, receiving a single high-resolution frame, to capture holistic visual perception-related information. The spatial perception stream can be made very lightweight by a well-designed P-Adapter, which is lightweight, training-efficient, and easily integrated into existing MLLMs. Experiments on the DRAMA-ROLISP dataset show HiLM-D\u2019s significant improvements over current MLLMs, with a <jats:inline-formula>\n              <jats:alternatives>\n                <jats:tex-math>$$3.7\\%$$<\/jats:tex-math>\n                <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mrow>\n                    <mml:mn>3.7<\/mml:mn>\n                    <mml:mo>%<\/mml:mo>\n                  <\/mml:mrow>\n                <\/mml:math>\n              <\/jats:alternatives>\n            <\/jats:inline-formula> in BLEU-4 for captioning and <jats:inline-formula>\n              <jats:alternatives>\n                <jats:tex-math>$$8.7\\%$$<\/jats:tex-math>\n                <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mrow>\n                    <mml:mn>8.7<\/mml:mn>\n                    <mml:mo>%<\/mml:mo>\n                  <\/mml:mrow>\n                <\/mml:math>\n              <\/jats:alternatives>\n            <\/jats:inline-formula> in mIoU for detection. Further tests on the Shikra-RD dataset confirm our method\u2019s generalization capabilities. The DRAMA-ROLISP is available at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/xmed-lab\/HiLM-D\" ext-link-type=\"uri\">https:\/\/github.com\/xmed-lab\/HiLM-D<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s11263-025-02433-3","type":"journal-article","created":{"date-parts":[[2025,5,7]],"date-time":"2025-05-07T13:48:48Z","timestamp":1746625728000},"page":"5379-5395","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["HiLM-D: Enhancing MLLMs with Multi-scale High-Resolution Details for Autonomous Driving"],"prefix":"10.1007","volume":"133","author":[{"given":"Xinpeng","family":"Ding","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianhua","family":"Han","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hang","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1105-8083","authenticated-orcid":false,"given":"Xiaomeng","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,5,7]]},"reference":[{"key":"2433_CR1","doi-asserted-by":"crossref","unstructured":"Alletto, S., Palazzi, A., Solera, F., Calderara, S., & Cucchiara, R. (2016). Dr (eye) ve: A dataset for attention-based tasks with applications to autonomous and assisted driving. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops (pp. 54\u201360).","DOI":"10.1109\/CVPRW.2016.14"},{"key":"2433_CR2","doi-asserted-by":"crossref","unstructured":"Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lu\u010di\u0107, M., & Schmid, C. (2021). Vivit: A video vision transformer. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 6836\u20136846).","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"2433_CR3","unstructured":"Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., & Zhou, J. (2023). Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966"},{"key":"2433_CR4","unstructured":"Bavishi, R., Elsen, E., Hawthorne, C., Nye, M., Odena, A., Somani, A., & Ta\u015f\u0131rlar, S. (2023). Introducing our multimodal models. https:\/\/www.adept.ai\/blog\/fuyu-8b"},{"key":"2433_CR5","doi-asserted-by":"crossref","unstructured":"Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., & Beijbom, O. (2020). nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11621\u201311631).","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"2433_CR6","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. In European conference on computer vision (pp. 213\u2013229). Springer.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"2433_CR7","doi-asserted-by":"crossref","unstructured":"Carreira, J., & Zisserman, A. (2017). Quo vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6299\u20136308).","DOI":"10.1109\/CVPR.2017.502"},{"key":"2433_CR8","unstructured":"Casas, S., Luo, W., & Urtasun, R. (2018). Intentnet: Learning to predict intention from raw sensor data. In Conference on robot learning (pp. 947\u2013956). PMLR."},{"key":"2433_CR9","unstructured":"Chen, Z., Duan, Y., Wang, W., He, J., Lu, T., Dai, J., & Qiao, Y. (2022). Vision transformer adapter for dense predictions."},{"key":"2433_CR10","doi-asserted-by":"crossref","unstructured":"Chen, X., Ma, H., Wan, J., Li, B., & Xia, T. (2017). Multi-view 3d object detection network for autonomous driving. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1907\u20131915).","DOI":"10.1109\/CVPR.2017.691"},{"key":"2433_CR11","unstructured":"Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., & Zhao, R. (2023). Shikra: Unleashing multimodal llm\u2019s referential dialogue magic. arXiv preprint arXiv:2306.15195"},{"key":"2433_CR12","unstructured":"Cheng, Z., Leng, S., Zhang, H., Xin, Y., Li, X., Chen, G., Zhu, Y., Zhang, W., Luo, Z., Zhao, D., & Bing, L. (2024). Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476"},{"key":"2433_CR13","unstructured":"Chiang, W. -L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al. (2023). Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https:\/\/vicuna.lmsys.org. Accessed 14 April 2023."},{"key":"2433_CR14","unstructured":"Dai, W., Li, J., Li, D., Huat, A., Zhao, J., Wang, W., Li, B., Fung, P., & Hoi, S. (2023). Instructblip: Towards general-purpose vision-language models with instruction tuning."},{"key":"2433_CR15","doi-asserted-by":"crossref","unstructured":"Deruyttere, T., Vandenhende, S., Grujicic, D., Van\u00a0Gool, L., & Moens, M. F. (2019). Talk2car: Taking control of your self-driving car. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) (pp. 2088\u20132098).","DOI":"10.18653\/v1\/D19-1215"},{"key":"2433_CR16","doi-asserted-by":"crossref","unstructured":"Dewangan, V., Choudhary, T., Chandhok, S., Priyadarshan, S., Jain, A., Singh, A.K., Srivastava, S., Jatavallabhula, K. M., & Krishna, K. M. (2023). Talk2bev: Language-enhanced bird\u2019s-eye view maps for autonomous driving. arXiv preprint arXiv:2310.02251","DOI":"10.1109\/ICRA57147.2024.10611485"},{"key":"2433_CR17","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C. (2020). X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 203\u2013213).","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"2433_CR18","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., & He, K. (2019). Slowfast networks for video recognition. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 6202\u20136211).","DOI":"10.1109\/ICCV.2019.00630"},{"key":"2433_CR19","doi-asserted-by":"crossref","unstructured":"Gao, M., Tawari, A., & Martin, S. (2019). Goal-oriented object importance estimation in on-road driving videos. In 2019 international conference on robotics and automation (ICRA) (pp. 5509\u20135515). IEEE.","DOI":"10.1109\/ICRA.2019.8793970"},{"key":"2433_CR20","doi-asserted-by":"publisher","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In 2016 IEEE conference on computer vision and pattern recognition (CVPR). https:\/\/doi.org\/10.1109\/cvpr.2016.90","DOI":"10.1109\/cvpr.2016.90"},{"key":"2433_CR21","unstructured":"Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685"},{"key":"2433_CR22","doi-asserted-by":"crossref","unstructured":"Hu, Y., Yang, J., Chen, L., Li, K., Sima, C., Zhu, X., Chai, S., Du, S., Lin, T., & Wang, W. (2023). Planning-oriented autonomous driving. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 17853\u201317862).","DOI":"10.1109\/CVPR52729.2023.01712"},{"key":"2433_CR23","doi-asserted-by":"crossref","unstructured":"Jin, B., Liu, X., Zheng, Y., Li, P., Zhao, H., Zhang, T., Zheng, Y., Zhou, G., & Liu, J. (2023). Adapt: Action-aware driving caption transformer. arXiv preprint arXiv:2302.00673","DOI":"10.1109\/ICRA48891.2023.10160326"},{"key":"2433_CR24","doi-asserted-by":"crossref","unstructured":"Kim, J., & Canny, J. (2017). Interpretable learning for self-driving cars by visualizing causal attention. In Proceedings of the IEEE international conference on computer vision (pp. 2942\u20132950).","DOI":"10.1109\/ICCV.2017.320"},{"key":"2433_CR25","doi-asserted-by":"publisher","unstructured":"Kim, J., Misu, T., Chen, Y. -T., Tawari, A., & Canny, J. (2019). Grounding human-to-vehicle advice for self-driving vehicles. In 2019 IEEE\/CVF conference on computer vision and pattern recognition (CVPR). https:\/\/doi.org\/10.1109\/cvpr.2019.01084","DOI":"10.1109\/cvpr.2019.01084"},{"key":"2433_CR26","doi-asserted-by":"crossref","unstructured":"Kim, J., Rohrbach, A., Darrell, T., Canny, J., & Akata, Z. (2018). Textual explanations for self-driving vehicles. In Proceedings of the European conference on computer vision (ECCV) (pp. 563\u2013578).","DOI":"10.1007\/978-3-030-01216-8_35"},{"key":"2433_CR27","doi-asserted-by":"crossref","unstructured":"Li, C., Chan, S. H., & Chen, Y. -T. (2020). Who make drivers stop? towards driver-centric risk assessment: Risk object identification via causal inference. In 2020 IEEE\/RSJ international conference on intelligent robots and systems (IROS) (pp. 10711\u201310718). IEEE.","DOI":"10.1109\/IROS45743.2020.9341072"},{"key":"2433_CR28","unstructured":"Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., & Qiao, Y. (2023). Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355"},{"key":"2433_CR29","unstructured":"Li, J., Li, D., Savarese, S., & Hoi, S. (2023). Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597"},{"key":"2433_CR30","unstructured":"Li, J., Pan, K., Ge, Z., Gao, M., Ji, W., Zhang, W., Chua, T. -S., Tang, S., Zhang, H., & Zhuang, Y. (2023). Fine-tuning multimodal llms to follow zero-shot demonstrative instructions. In The 12th international conference on learning representations."},{"key":"2433_CR31","doi-asserted-by":"crossref","unstructured":"Li, Z., Yang, B., Liu, Q., Ma, Z., Zhang, S., Yang, J., Sun, Y., Liu, Y., & Bai, X. (2023). Monkey: Image resolution and text label are important things for large multi-modal models. arXiv preprint arXiv:2311.06607","DOI":"10.1109\/CVPR52733.2024.02527"},{"key":"2433_CR32","unstructured":"Li, B., Zhang, P., Yang, J., Zhang, Y., Pu, F., & Liu, Z. (2023). Otterhd: A high-resolution multi-modality model. arXiv preprint arXiv:2311.04219"},{"key":"2433_CR33","doi-asserted-by":"crossref","unstructured":"Liu, H., Li, C., Li, Y., & Lee, Y. J. (2023). Improved Baselines with Visual Instruction Tuning. arXiv:2310.03744","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"2433_CR34","unstructured":"Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., & Lee, Y. J. (2024). LLaVA-NeXT: Improved reasoning, OCR, and world knowledge. https:\/\/llava-vl.github.io\/blog\/2024-01-30-llava-next\/"},{"key":"2433_CR35","unstructured":"Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2023). Visual instruction tuning. arXiv preprint arXiv:2304.08485"},{"key":"2433_CR36","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 10012\u201310022).","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"2433_CR37","doi-asserted-by":"crossref","unstructured":"Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., et al. (2023). Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499","DOI":"10.1007\/978-3-031-72970-6_3"},{"key":"2433_CR38","unstructured":"Loshchilov, I., & Hutter, F. (2016). Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983"},{"key":"2433_CR39","unstructured":"Loshchilov, I., & Hutter, F. (2017). Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101"},{"key":"2433_CR40","doi-asserted-by":"crossref","unstructured":"Luo, W., Yang, B., & Urtasun, R. (2018). Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3569\u20133577).","DOI":"10.1109\/CVPR.2018.00376"},{"key":"2433_CR41","doi-asserted-by":"crossref","unstructured":"Malla, S., Choi, C., Dwivedi, I., Choi, J. H., & Li, J. (2023). Drama: Joint risk localization and captioning in driving. In Proceedings of the IEEE\/CVF winter conference on applications of computer vision (pp. 1043\u20131052).","DOI":"10.1109\/WACV56688.2023.00110"},{"key":"2433_CR42","doi-asserted-by":"crossref","unstructured":"Malla, S., Dariush, B., & Choi, C. (2020). Titan: Future forecast using action priors. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11186\u201311196).","DOI":"10.1109\/CVPR42600.2020.01120"},{"key":"2433_CR43","unstructured":"Ngiam, J., Caine, B., Vasudevan, V., Zhang, Z., Chiang, H. -T. L., Ling, J., Roelofs, R., Bewley, A., Liu, C., Venugopal, A., et al. (2021). Scene transformer: A unified multi-task model for behavior prediction and planning (vol. 2, no. 7). arXiv preprint arXiv:2106.08417"},{"key":"2433_CR44","unstructured":"OpenAI, O. (2023). Gpt-4 technical report"},{"key":"2433_CR45","unstructured":"OpenAI: (2024). Hello gpt-4o. https:\/\/openai.com\/index\/hello-gpt-4o\/"},{"key":"2433_CR46","unstructured":"Oquab, M., Darcet, T., Moutakanni, T., Vo, H. V., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Howes, R., Huang, P. -Y., Xu, H., Sharma, V., Li, S. -W., Galuba, W., Rabbat, M., Assran, M., Ballas, N., Synnaeve, G., Misra, I., Jegou, H., Mairal, J., Labatut, P., Joulin, A., & Bojanowski, P. (2023). DINOv2: Learning robust visual features without supervision."},{"key":"2433_CR47","unstructured":"Pan, J., Lin, Z., Zhu, X., Shao, J., ST-Adapter, H. L. (2022). Parameter-efficient image-to-video transfer learning for action recognition. Preprint at arxiv.org\/abs\/2206.13559"},{"key":"2433_CR48","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019). Pytorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems (vol. 32)."},{"key":"2433_CR49","unstructured":"Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., & Wei, F. (2023). Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824"},{"key":"2433_CR50","doi-asserted-by":"crossref","unstructured":"Petrovskaya, A., & Thrun, S. (2008). Model based vehicle tracking for autonomous driving in urban environments. In Proceedings of robotics: Science and systems IV, Zurich, Switzerland (vol. 34).","DOI":"10.15607\/RSS.2008.IV.023"},{"key":"2433_CR51","doi-asserted-by":"crossref","unstructured":"Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., & Lazebnik, S. (2015). Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision (pp. 2641\u20132649).","DOI":"10.1109\/ICCV.2015.303"},{"key":"2433_CR52","doi-asserted-by":"crossref","unstructured":"Qian, T., Chen, J., Zhuo, L., Jiao, Y., & Jiang, Y. -G. (2023). Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario. arXiv preprint arXiv:2305.14836","DOI":"10.1609\/aaai.v38i5.28253"},{"key":"2433_CR53","unstructured":"Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., & Clark, J. (2021). Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748\u20138763). PMLR."},{"issue":"8","key":"2433_CR54","first-page":"9","volume":"1","author":"A Radford","year":"2019","unstructured":"Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Blog, 1(8), 9.","journal-title":"OpenAI Blog"},{"issue":"1","key":"2433_CR55","first-page":"5485","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1), 5485\u20135551.","journal-title":"The Journal of Machine Learning Research"},{"key":"2433_CR56","doi-asserted-by":"crossref","unstructured":"Shukor, M., Dancette, C., & Cord, M. (2023). ep-alm: Efficient perceptual augmentation of language models. arXiv preprint arXiv:2303.11403","DOI":"10.1109\/ICCV51070.2023.02016"},{"key":"2433_CR57","unstructured":"Simonyan, K., & Zisserman, A. (2014). Two-stream convolutional networks for action recognition in videos. In Advances in neural information processing systems (vol. 27)."},{"key":"2433_CR58","doi-asserted-by":"publisher","DOI":"10.1088\/1757-899x\/1022\/1\/012028","volume":"1022","author":"S Singh","year":"2021","unstructured":"Singh, S., & Saini, B. S. (2021). Autonomous cars: Recent developments, challenges, and possible solutions. IOP Conference Series: Materials Science and Engineering, 1022, Article 012028. https:\/\/doi.org\/10.1088\/1757-899x\/1022\/1\/012028","journal-title":"IOP Conference Series: Materials Science and Engineering"},{"key":"2433_CR59","doi-asserted-by":"crossref","unstructured":"Tawari, A., Mallela, P., & Martin, S. (2018). Learning to attend to salient targets in driving videos using fully convolutional RNN. In 2018 21st international conference on intelligent transportation systems (ITSC) (pp. 3225\u20133232). IEEE.","DOI":"10.1109\/ITSC.2018.8569438"},{"key":"2433_CR60","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. -A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023). Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971"},{"key":"2433_CR61","unstructured":"Wang, W., Chen, Z., Chen, X., Wu, J., Zhu, X., Zeng, G., Luo, P., Lu, T., Zhou, J., Qiao, Y., et al. (2023). Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. arXiv preprint arXiv:2305.11175"},{"key":"2433_CR62","doi-asserted-by":"crossref","unstructured":"Wang, D., Devin, C., Cai, Q. -Z., Yu, F., & Darrell, T. (2019). Deep object-centric policies for autonomous driving. In 2019 international conference on robotics and automation (ICRA) (pp. 8853\u20138859). IEEE.","DOI":"10.1109\/ICRA.2019.8794224"},{"issue":"3","key":"2433_CR63","doi-asserted-by":"publisher","first-page":"415","DOI":"10.1007\/s41095-022-0274-8","volume":"8","author":"W Wang","year":"2022","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., & Shao, L. (2022). Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3), 415\u2013424.","journal-title":"Computational Visual Media"},{"key":"2433_CR64","unstructured":"Xu, R., Yao, Y., Guo, Z., Cui, J., Ni, Z., Ge, C., Chua, T. -S., Liu, Z., Sun, M., & Huang, G. (2024). Llava-uhd: An lmm perceiving any aspect ratio and high-resolution images. arXiv preprint arXiv:2403.11703"},{"key":"2433_CR65","doi-asserted-by":"crossref","unstructured":"Xu, Z., Zhang, Y., Xie, E., Zhao, Z., Guo, Y., Wong, K.K., Li, Z., & Zhao, H. (2023). Drivegpt4: Interpretable end-to-end autonomous driving via large language model. arXiv preprint arXiv:2310.01412","DOI":"10.1109\/LRA.2024.3440097"},{"key":"2433_CR66","unstructured":"Zang, Y., Li, W., Han, J., Zhou, K., Loy, C. C. (2023). Contextual object detection with multimodal large language models. arXiv preprint arXiv:2305.18279"},{"key":"2433_CR67","doi-asserted-by":"crossref","unstructured":"Zeng, K. -H., Chou, S. -H., Chan, F. -H., & Carlos\u00a0Niebles, J., Sun, M. (2017). Agent-centric risk assessment: Accident anticipation and risky region localization. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2222\u20132230).","DOI":"10.1109\/CVPR.2017.146"},{"key":"2433_CR68","unstructured":"Zhang, R., Han, J., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., Gao, P., & Qiao, Y. (2023). Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199"},{"key":"2433_CR69","doi-asserted-by":"crossref","unstructured":"Zhang, H., Li, X., Bing, L. (2023). Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858","DOI":"10.18653\/v1\/2023.emnlp-demo.49"},{"key":"2433_CR70","unstructured":"Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., & Lin, X. V., et al. (2022). Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068"},{"key":"2433_CR71","unstructured":"Zhu, D., Chen, J., Shen, X., Li, X., & Elhoseiny, M. (2023). Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02433-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-025-02433-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02433-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T13:32:27Z","timestamp":1757165547000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-025-02433-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,7]]},"references-count":71,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["2433"],"URL":"https:\/\/doi.org\/10.1007\/s11263-025-02433-3","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,7]]},"assertion":[{"value":"23 June 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 March 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 May 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}