{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T15:15:52Z","timestamp":1780413352597,"version":"3.54.1"},"reference-count":38,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2025,12,21]],"date-time":"2025-12-21T00:00:00Z","timestamp":1766275200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100018542","name":"Natural Science Foundation of Sichuan Province","doi-asserted-by":"publisher","award":["2023NSFSC0243"],"award-info":[{"award-number":["2023NSFSC0243"]}],"id":[{"id":"10.13039\/501100018542","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Talent Introduction Program of Chengdu University of Information Technology","award":["KYTZ202261"],"award-info":[{"award-number":["KYTZ202261"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJGI"],"abstract":"<jats:p>Traditional road traffic accident analysis has long relied on structured data, making it difficult to integrate high-dimensional heterogeneous information such as remote sensing imagery and leading to an incomplete understanding of accident scene environments. This study proposes a road traffic accident analysis framework based on Multimodal Large Language Models. The approach integrates high-resolution remote sensing imagery with structured accident data through a three-stage progressive training pipeline. Specifically, we fine-tune three open-source vision\u2013language models using Low-Rank Adaptation (LoRA) to sequentially optimize the model\u2019s capabilities in visual environmental description, multi-task accident classification, and Chain-of-Thought (CoT) driven causal reasoning. A multimodal dataset was constructed containing remote sensing image descriptions, accident classification labels, and interpretable reasoning chains. Experimental results show that the fine-tuned model achieved a maximum improvement in the CIDEr score for image description tasks. In the joint classification task of accident severity and duration, the model achieved an accuracy of 71.61% and an F1-score of 0.8473. In the CoT reasoning task, both METEOR and CIDEr scores improved significantly. These results validate the effectiveness of structured reasoning mechanisms in multimodal fusion for transportation applications, providing a feasible path toward interpretable and intelligent analysis for real-world traffic management.<\/jats:p>","DOI":"10.3390\/ijgi15010008","type":"journal-article","created":{"date-parts":[[2025,12,24]],"date-time":"2025-12-24T14:27:51Z","timestamp":1766586471000},"page":"8","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Domain-Adapted MLLMs for Interpretable Road Traffic Accident Analysis Using Remote Sensing Imagery"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-2510-7069","authenticated-orcid":false,"given":"Bing","family":"He","sequence":"first","affiliation":[{"name":"School of Computer Science, Chengdu University of Information Technology, Chengdu 610225, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"He","sequence":"additional","affiliation":[{"name":"School of Computer Science, Chengdu University of Information Technology, Chengdu 610225, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qing","family":"Chang","sequence":"additional","affiliation":[{"name":"Sain Associates, Inc., 5021 Technology Drive Northwest, Suite B2, Huntsville, AL 35805, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1409-4738","authenticated-orcid":false,"given":"Wen","family":"Luo","sequence":"additional","affiliation":[{"name":"School of Computer Science, Chengdu University of Information Technology, Chengdu 610225, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lingli","family":"Xiao","sequence":"additional","affiliation":[{"name":"School of Computer Science, Chengdu University of Information Technology, Chengdu 610225, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,12,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"170635","DOI":"10.1155\/2015\/170635","article-title":"The application of data mining technology to build a forecasting model for classification of road traffic accidents","volume":"2015","author":"Shiau","year":"2015","journal-title":"Math. Probl. Eng."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Manawadu, M., and Wijenayake, U. (2023). Predictive Analysis of Accidents Based on US Accident Data. Proceedings of the 2023 14th International Conference on Information and Communication Technology Convergence (ICTC), IEEE.","DOI":"10.1109\/ICTC58733.2023.10393583"},{"key":"ref_3","unstructured":"Yan, Y., Liao, Y., Xu, G., Yao, R., Fan, H., Sun, J., Wang, X., Sprinkle, J., An, Z., and Ma, M. (2025). Large language models for traffic and transportation research: Methodologies, state of the art, and future opportunities. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zhen, H., and Yang, J.J. (2025). CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis. arXiv.","DOI":"10.1016\/j.ait.2025.100030"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"633","DOI":"10.1016\/S0001-4575(99)00094-9","article-title":"Modeling traffic accident occurrence and involvement","volume":"32","author":"Radwan","year":"2000","journal-title":"Accid. Anal. Prev."},{"key":"ref_6","unstructured":"Behboudi, N., Moosavi, S., and Ramnath, R. (2024). Recent advances in traffic accident analysis and prediction: A comprehensive review of machine learning techniques. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"He, S., Sadeghi, M.A., Chawla, S., Alizadeh, M., Balakrishnan, H., and Madden, S. (2021). Inferring high-resolution traffic accident risk maps based on satellite imagery and GPS trajectories. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV48922.2021.01176"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Liu, L., Gao, Z., Luo, P., Duan, W., Hu, M., Mohd Arif Zainol, M.R.R., and Zawawi, M.H. (2023). The influence of visual landscapes on road traffic safety: An assessment using remote sensing and deep learning. Remote Sens., 15.","DOI":"10.3390\/rs15184437"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"672","DOI":"10.1080\/17457300.2024.2409638","article-title":"Mapping the relationship between traffic accidents, road network configuration, and urban land use","volume":"31","author":"Iranmanesh","year":"2024","journal-title":"Int. J. Inj. Control. Saf. Promot."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yan, R., Hu, L., Li, J., and Lin, N. (2024). Accident severity analysis of traffic accident hot spot areas in Changsha City considering built environment. Sustainability, 16.","DOI":"10.3390\/su16073054"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Cheng, G., Cheng, R., Pei, Y., Xu, L., and Qi, W. (2020). Severity assessment of accidents involving roadside trees based on occupant injury analysis. PLoS ONE, 15.","DOI":"10.1371\/journal.pone.0231030"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"108077","DOI":"10.1016\/j.aap.2025.108077","article-title":"When language and vision meet road safety: Leveraging multimodal large language models for video-based traffic accident analysis","volume":"219","author":"Zhang","year":"2025","journal-title":"Accid. Anal. Prev."},{"key":"ref_13","unstructured":"Cai, J., Meng, K., Yang, B., and Shao, G. (2024). Multimodal remote sensing scene classification using VLMs and dual-cross attention networks. arXiv."},{"key":"ref_14","unstructured":"Fan, Z., Wang, P., Zhao, Y., Zhao, Y., Ivanovic, B., Wang, Z., Pavone, M., and Yang, H.F. (2024). Learning traffic crashes as language: Datasets, benchmarks, and what-if causal analyses. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhen, H., Shi, Y., Huang, Y., Yang, J.J., and Liu, N. (2024). Leveraging large language models with chain-of-thought and prompt engineering for traffic crash severity analysis and inference. Computers, 13.","DOI":"10.3390\/computers13090232"},{"key":"ref_16","unstructured":"Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Cui, E., Zhu, J., Ye, S., Tian, H., and Liu, Z. (2024). Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv."},{"key":"ref_17","unstructured":"Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., and Tang, J. (2025). Qwen2.5-vl technical report. arXiv."},{"key":"ref_18","unstructured":"Lu, H., Liu, W., Zhang, B., Wang, B., Dong, K., Liu, B., Sun, J., Ren, T., Li, Z., and Yang, H. (2024). Deepseek-vl: Towards real-world vision-language understanding. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"729","DOI":"10.1016\/S0001-4575(01)00073-2","article-title":"Using logistic regression to estimate the influence of accident factors on accident severity","volume":"34","year":"2002","journal-title":"Accid. Anal. Prev."},{"key":"ref_20","unstructured":"Moosavi, S., Samavatian, M.H., Parthasarathy, S., and Ramnath, R. (2019). A countrywide traffic accident dataset. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Oad, R., Sayani, A.I., and Banitaan, S. (2024). Predicting Severity of US Traffic Accidents: A Machine Learning Approach. Proceedings of the 2024 IEEE International Conference on Electro Information Technology (eIT), IEEE.","DOI":"10.1109\/eIT60633.2024.10609840"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zohra, E.F., Maryam, K., and Hasna, E.E. (2023). Accident Severity Prediction using Machine Learning: A case study on the US Accidents Dataset. Proceedings of the 2023 17th International Conference on Signal-Image Technology & Internet-Based Systems (SITIS), IEEE.","DOI":"10.1109\/SITIS61268.2023.00044"},{"key":"ref_23","first-page":"5582","article-title":"Quantitative study of traffic accident prediction models: A case study of virginia accidents","volume":"14","author":"Almanie","year":"2023","journal-title":"Int. J. Adv. Netw. Appl."},{"key":"ref_24","first-page":"104687","article-title":"Integrating InSAR coherence and air pollution detection satellites to study the impact of war on air quality","volume":"142","author":"Mohamadi","year":"2025","journal-title":"Int. J. Appl. Earth Obs. Geoinf."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"128","DOI":"10.24940\/10.24940\/ijird\/2021\/v10\/i4\/APR21065","article-title":"Mapping and Analysis of the Temporal Changes of Road Traffic Accident Hotspots Using GIS and Remote Sensing in Owerri, Imo State, Nigeria","volume":"10","author":"Ngozi","year":"2021","journal-title":"Int. J. Innov. Res. Dev."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"535","DOI":"10.1109\/JSTARS.2023.3331438","article-title":"Unveiling roadway hazards: Enhancing fatal crash risk estimation through multiscale satellite imagery and self-supervised cross-matching","volume":"17","author":"Liang","year":"2023","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_27","first-page":"345","article-title":"Evaluation of remote sensing technologies for collecting roadside feature data to support highway safety manual implementation","volume":"7","author":"Jalayer","year":"2015","journal-title":"J. Transp. Saf. Secur."},{"key":"ref_28","unstructured":"Mumtarin, M., Chowdhury, M.S., and Wood, J. (2023). Large Language Models in Analyzing Crash Narratives\u2014A Comparative Study of ChatGPT, BARD and GPT-4. arXiv."},{"key":"ref_29","unstructured":"Zheng, O., Abdel-Aty, M., Wang, D., Wang, C., and Ding, S. (2023). Trafficsafetygpt: Tuning a pre-trained large language model to a domain-specific expert in transportation safety. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"4962","DOI":"10.1109\/TIV.2024.3508471","article-title":"Chatsumo: Large language model for automating traffic scenario generation in simulation of urban mobility","volume":"10","author":"Li","year":"2024","journal-title":"IEEE Trans. Intell. Veh."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"9701","DOI":"10.1109\/TMC.2025.3564163","article-title":"Agentscomerge: Large language model empowered collaborative decision making for ramp merging","volume":"24","author":"Hu","year":"2025","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Guo, A., Zhou, Y., Tian, H., Fang, C., Sun, Y., and Sun, W. (2024). Sovar: Build generalizable scenarios from accident reports for autonomous driving testing. Proceedings of the 39th IEEE\/ACM International Conference on Automated Software Engineering, IEEE.","DOI":"10.1145\/3691620.3695037"},{"key":"ref_33","unstructured":"Grigorev, A., Saleh, A.S.M.K., and Ou, Y. (2024). IncidentResponseGPT: Generating traffic incident response plans with generative artificial intelligence. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Lohner, A., Compagno, F., Francis, J., and Oltramari, A. (2024). Enhancing vision-language models with scene graphs for traffic accident understanding. Proceedings of the 2024 IEEE International Automated Vehicle Validation Conference (IAVVC), IEEE.","DOI":"10.1109\/IAVVC63304.2024.10786395"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wu, K., Li, W., and Xiao, X. (2024). AccidentGPT: Large multi-modal foundation model for traffic accident analysis. arXiv.","DOI":"10.5220\/0012422100003636"},{"key":"ref_36","unstructured":"Zhou, X., and Knoll, A.C. (2024). GPT-4V as traffic assistant: An in-depth look at vision language model on complex traffic events. arXiv."},{"key":"ref_37","unstructured":"Chowdhury, M.T.B.Z., Islam, M.R., and Hossain, M. (2024). Durghotona GPT: A Web Scraping and Large Language Model Based Framework to Generate Road Accident Dataset Automatically in Bangladesh. Proceedings of the 2024 27th International Conference on Computer and Information Technology (ICCIT), IEEE."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Kuckreja, K., Danish, M.S., Naseer, M., Das, A., Khan, S., and Khan, F.S. (2024). Geochat: Grounded large vision-language model for remote sensing. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52733.2024.02629"}],"container-title":["ISPRS International Journal of Geo-Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2220-9964\/15\/1\/8\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,26]],"date-time":"2025-12-26T03:28:04Z","timestamp":1766719684000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2220-9964\/15\/1\/8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,21]]},"references-count":38,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["ijgi15010008"],"URL":"https:\/\/doi.org\/10.3390\/ijgi15010008","relation":{},"ISSN":["2220-9964"],"issn-type":[{"value":"2220-9964","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,21]]}}}