{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T16:57:17Z","timestamp":1773421037683,"version":"3.50.1"},"reference-count":49,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T00:00:00Z","timestamp":1773360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["42571292"],"award-info":[{"award-number":["42571292"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100013139","name":"Humanities and Social Science Fund of Ministry of Education of China","doi-asserted-by":"publisher","award":["22YJAZH108"],"award-info":[{"award-number":["22YJAZH108"]}],"id":[{"id":"10.13039\/501100013139","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["B220201021"],"award-info":[{"award-number":["B220201021"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJGI"],"abstract":"<jats:p>Judicial documents have become a significant data source for crime geography research, offering advantages in accessibility and scale compared to highly restricted police-recorded crime data. However, extracting crime addresses from these texts is challenging due to sparse, inconsistent, and incomplete address information. Without proper classification, errors in geocoding and spatial analysis can arise, compromising data quality. To address these limitations, we employed large language models (LLMs) and a structured prompt engineering strategy tailored for this task. Specifically, we propose a fine-tuned LLM, named CAEC_LLM, to extract addresses from judicial documents and classify these crime addresses at various categories with different spatial scales. Experimental results demonstrate that the model achieved an F1-score of 0.79 for address extraction and a classification accuracy of up to 0.74 for the best-performing category, significantly outperforming other LLMs. This study makes two primary contributions: (1) designing an address classification scheme specifically for crime addresses, and (2) developing a fine-tuned LLM for extracting and classifying crime addresses from Chinese judicial documents, enabling LLMs to be used to classify crime addresses into different categories on a spatial scale. These advancements facilitate more accurate crime pattern analysis and data-driven urban planning.<\/jats:p>","DOI":"10.3390\/ijgi15030124","type":"journal-article","created":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T15:06:22Z","timestamp":1773414382000},"page":"124","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Theft Address Extraction and Classification from Chinese Judicial Documents Based on Large Language Model"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1447-6689","authenticated-orcid":false,"given":"Zengli","family":"Wang","sequence":"first","affiliation":[{"name":"Jiangsu Key Laboratory of Soil and Water Processes in Watershed, College of Geography and Remote Sensing, Hohai University, Nanjing 211100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2267-9186","authenticated-orcid":false,"given":"Xiang","family":"Li","sequence":"additional","affiliation":[{"name":"School of Earth Sciences and Engineering, Hohai University, Nanjing 211100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7764-4272","authenticated-orcid":false,"given":"Xiaoping","family":"Rui","sequence":"additional","affiliation":[{"name":"School of Earth Sciences and Engineering, Hohai University, Nanjing 211100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Linfang","family":"Ding","sequence":"additional","affiliation":[{"name":"Jiangsu Key Laboratory of Soil and Water Processes in Watershed, College of Geography and Remote Sensing, Hohai University, Nanjing 211100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingjing","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Land Resources Management, School of Public Administration, China University of Geosciences, Wuhan 430074, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,3,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1337","DOI":"10.1007\/s10888-024-09662-5","article-title":"Local inequality and crime: New evidence from South Africa","volume":"23","year":"2025","journal-title":"J. Econ. Inequal."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Dai, M., and Liu, C. (2020). Multi-label classification of chinese judicial documents based on BERT. Proceedings of the 8th IEEE International Conference on Big Data, Atlanta, GA, USA, 10\u201313 December 2020, IEEE. Electr Network.","DOI":"10.1109\/BigData50022.2020.9377969"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Liu, C., and Hsieh, C. (2006). Exploring phrase-based classification of judicial documents for criminal charges in chinese. Proceedings of the 16th International Symposium on Methodologies for Intelligent Systems, Bari, Italy, 27\u201329 September 2006, Spring.","DOI":"10.1007\/11875604_75"},{"key":"ref_4","first-page":"2193","article-title":"A case similarity calculation model in case pushing of judicial documents","volume":"41","author":"Wang","year":"2019","journal-title":"Comput. Eng. Sci."},{"key":"ref_5","unstructured":"Yin, J. (2025). Research on Automatic Summarization Technology for Law-Related Texts. [Master\u2019s Thesis, People\u2019s Public Security University of China]."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"635","DOI":"10.1007\/s10707-012-0173-8","article-title":"An algorithm for local geoparsing of microtext","volume":"17","author":"Gelernter","year":"2013","journal-title":"GeoInformatica"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1111\/tgis.12510","article-title":"GeoTxt: A scalable geoparsing system for unstructured text geolocation","volume":"23","author":"Karimzadeh","year":"2019","journal-title":"Trans. GIS"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1393","DOI":"10.1111\/tgis.12579","article-title":"Enhancing spatial and textual analysis with EUPEG: An extensible and unified platform for evaluating geoparsers","volume":"23","author":"Wang","year":"2019","journal-title":"Trans. GIS"},{"key":"ref_9","first-page":"1277","article-title":"Research on classification of commodity ultra-short text based on deep random forest","volume":"41","author":"Niu","year":"2022","journal-title":"Trans. Beijing Inst. Technol."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"103395","DOI":"10.1109\/ACCESS.2020.2994187","article-title":"Tracking flooding phase transitions and establishing a passive hotline with AI-enabled social media data","volume":"8","author":"Wang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_11","unstructured":"Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., and Bi, X. (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv."},{"key":"ref_12","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_13","unstructured":"Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., and Ray, A. (2022). Training language models to follow instructions with human feedback. Proceedings of the 36th Conference on Neural Information Processing Systems, New Orleans, LA, USA, 28 November\u20139 December 2022, IEEE. Electr Network."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Razeghi, Y., Logan, R.L., Gardner, M., and Singh, S. (2022). Impact of pretraining term frequencies on few-shot reasoning. arXiv.","DOI":"10.18653\/v1\/2022.findings-emnlp.59"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Ren, Y., Han, J., Lin, Y., Mei, X., and Zhang, L. (2022). An ontology-based and deep learning-driven method for extracting legal facts from Chinese legal texts. Electronics, 11.","DOI":"10.3390\/electronics11121821"},{"key":"ref_16","first-page":"36","article-title":"Knowledge prompt fine-tuning for event extraction","volume":"7","author":"Li","year":"2024","journal-title":"Comput. Mod."},{"key":"ref_17","first-page":"1","article-title":"Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing","volume":"55","author":"Liu","year":"2023","journal-title":"ACM Comput. Surv."},{"key":"ref_18","first-page":"175","article-title":"Information extraction from chinese wheat varieties journal based on large language model","volume":"7","author":"Wei","year":"2025","journal-title":"Front. Data Comput."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Hammerton, J. (June, January 31). Named entity recognition with long short-term memory. Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, Edmonton, AB, Canada.","DOI":"10.3115\/1119176.1119202"},{"key":"ref_20","unstructured":"Sang, E.F., and De Meulder, F. (2003). Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. arXiv."},{"key":"ref_21","unstructured":"Borthwick, A.E. (1999). A Maximum Entropy Approach to Named Entity Recognition. [Ph.D. Thesis, New York University]."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Grishman, R., and Sundheim, B.M. (1996). Message understanding conference-6: A brief history. COLING 1996 Volume 1: The 16th International Conference on Computational Linguistics, Copenhagen, Denmark, 5\u20139 August 1996, Association for Computational Linguistics.","DOI":"10.3115\/992628.992709"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Mikheev, A., Moens, M., and Grover, C. (1999, January 8\u201312). Named entity recognition without gazetteers. Proceedings of the 9th Conference of the European Chapter of the Association-for-Computational-Linguistics, Bergen, Norway.","DOI":"10.3115\/977035.977037"},{"key":"ref_24","unstructured":"Gao, P., Zhang, X., and Qi, G. (2019). Discovering hypernymy relationships in Chinese traffic legal texts. Proceedings of the Joint International Semantic Technology Conference 2019, Hangzhou, China, 25\u201327 November 2019, Springer."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Qi, Y., Zhai, R., Wu, F., Yin, J., Gong, X., Zhu, L., and Yu, H. (2024). CSMNER: A toponym entity recognition model for chinese social media. ISPRS Int. J. Geo-Inf., 13.","DOI":"10.3390\/ijgi13090311"},{"key":"ref_26","first-page":"110","article-title":"Address entity recognition based on multi-layer knowledge perception","volume":"39","author":"Shao","year":"2025","journal-title":"J. Chin. Inf. Process."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1016\/j.procir.2023.05.001","article-title":"Opportunities and challenges of ChatGPT for design knowledge management","volume":"119","author":"Hu","year":"2023","journal-title":"Procedia CIRP"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"4668","DOI":"10.1080\/17538947.2023.2278895","article-title":"Autonomous GIS: The next-generation AI-powered GIS","volume":"16","author":"Li","year":"2023","journal-title":"Int. J. Digit. Earth"},{"key":"ref_29","first-page":"17","article-title":"Use chat gpt to solve programming bugs","volume":"3","author":"Surameery","year":"2023","journal-title":"Int. J. Inf. Technol. Comput. Eng."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1186\/s40561-023-00237-x","article-title":"What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in education","volume":"10","author":"Tlili","year":"2023","journal-title":"Smart Learn. Environ."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Yin, Z., Li, D., and Goldberg, D.W. (2023, January 13). Is ChatGPT a game changer for geocoding-a benchmark for geocoding address parsing techniques. Proceedings of the 2nd ACM SIGSPATIAL International Workshop on Searching and Mining Large Collections of Geospatial Data, Hamburg, Germany.","DOI":"10.1145\/3615890.3628538"},{"key":"ref_32","unstructured":"Hu, Y., and Wang, J. (2020). How do people describe locations during a natural disaster: An analysis of tweets from Hurricane Harvey. arXiv."},{"key":"ref_33","unstructured":"Huang, Q., Tao, M., Zhang, C., An, Z., Jiang, C., Chen, Z., Wu, Z., and Feng, Y. (2023). Lawyer llama technical report. arXiv."},{"key":"ref_34","unstructured":"Cui, J., Ning, M., Li, Z., Chen, B., Yan, Y., Li, H., Ling, B., Tian, Y., and Yuan, L. (2024). Chatlaw: A multi-agent collaborative legal assistant with knowledge graph enhanced mixture-of-experts large language model. arXiv."},{"key":"ref_35","first-page":"12","article-title":"Fine-tuning and application of large language model in law domain","volume":"10","author":"Shen","year":"2024","journal-title":"Big Data Res."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"467","DOI":"10.1038\/s41597-025-04757-8","article-title":"An LLM driven dataset on the spatiotemporal distributions of street and neighborhood crime in China","volume":"12","author":"Zhang","year":"2025","journal-title":"Sci. Data"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"2289","DOI":"10.1080\/13658816.2023.2266495","article-title":"Geo-knowledge-guided GPT models improve the extraction of location descriptions from disaster-related social media messages","volume":"37","author":"Hu","year":"2023","journal-title":"Int. J. Geogr. Inf. Sci."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"214","DOI":"10.1016\/j.compenvurbsys.2007.11.006","article-title":"A comparison of address point, parcel and street geocoding techniques","volume":"32","author":"Zandbergen","year":"2008","journal-title":"Comput. Environ. Urban Syst."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"647","DOI":"10.1111\/j.1749-8198.2008.00205.x","article-title":"Geocoding quality and implications for spatial analysis","volume":"3","author":"Zandbergen","year":"2009","journal-title":"Geogr. Compass"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1111\/j.1467-9671.2011.01270.x","article-title":"Dasymetric mapping using high resolution address point datasets","volume":"15","author":"Zandbergen","year":"2011","journal-title":"Trans. GIS"},{"key":"ref_41","first-page":"99","article-title":"A new method of Chinese address extraction based on address tree model","volume":"44","author":"Kang","year":"2015","journal-title":"Acta Geod. Cartogr. Sin."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"2005","DOI":"10.1007\/s11280-020-00782-2","article-title":"Geographical address representation learning for address matching","volume":"23","author":"Shan","year":"2020","journal-title":"World Wide Web"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"101819","DOI":"10.1016\/j.compenvurbsys.2022.101819","article-title":"W-TextCNN: A TextCNN model with weighted word embeddings for Chinese address pattern classification","volume":"95","author":"Zhang","year":"2022","journal-title":"Comput. Environ. Urban Syst."},{"key":"ref_44","unstructured":"Zeng, A., Xu, B., Wang, B., Zhang, C., Yin, D., Zhang, D., Rojas, D., Feng, G., Zhao, H., and Lai, H. (2024). Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Beltagy, I., Lo, K., and Cohan, A. (2019). SciBERT: A pretrained language model for scientific text. arXiv.","DOI":"10.18653\/v1\/D19-1371"},{"key":"ref_46","unstructured":"Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021). Lora: Low-rank adaptation of large language models. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Zheng, Y., Zhang, R., Zhang, J., Ye, Y., Luo, Z., Feng, Z., and Ma, Y. (2024). Llamafactory: Unified efficient fine-tuning of 100+ language models. arXiv.","DOI":"10.18653\/v1\/2024.acl-demos.38"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Gritta, M., Pilehvar, M.T., and Collier, N. (2018). Which melbourne? Augmenting geocoding with maps. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Melbourne, Australia, 15\u201320 July 2018, Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-1119"},{"key":"ref_49","first-page":"781","article-title":"Geocoding the past world: Unearthing coordinates of early China from texts using generative AI","volume":"141","author":"Chen","year":"2025","journal-title":"Int. J. Geogr. Inf. Sci."}],"container-title":["ISPRS International Journal of Geo-Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2220-9964\/15\/3\/124\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T15:21:33Z","timestamp":1773415293000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2220-9964\/15\/3\/124"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,13]]},"references-count":49,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2026,3]]}},"alternative-id":["ijgi15030124"],"URL":"https:\/\/doi.org\/10.3390\/ijgi15030124","relation":{},"ISSN":["2220-9964"],"issn-type":[{"value":"2220-9964","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,13]]}}}