{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,24]],"date-time":"2025-11-24T16:46:46Z","timestamp":1764002806989,"version":"3.44.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"9","funder":[{"name":"National Key Research and Development Program of China","award":["2022YFC3301801"],"award-info":[{"award-number":["2022YFC3301801"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>Named Entity Recognition (NER) in specialized domains for low-resource languages remains a significant challenge due to data scarcity and the complexity of domain-specific terminology. Existing cross-lingual approaches\u2014spanning model-transfer and data-transfer paradigms\u2014often suffer from semantic drift and inadequate domain adaptation. To address these limitations, we introduce BiLegalNERD, the first bilingual Chinese\u2013Uyghur legal NER dataset, constructed via a semantics-aware label transfer strategy. We further propose CUTLM, a cross-lingual annotation method that combines dual translation with Levenshtein-based alignment to ensure high-fidelity preservation of entity boundaries across languages. In addition, we present BiLegalNER, a domain-adapted multilingual NER model incorporating vocabulary expansion and bilingual fine-tuning, significantly enhancing performance on Uyghur legal texts. Experiments demonstrate that BiLegalNER achieves state-of-the-art results, with F1-scores of 86.65% on automatically generated training data and 89.11% on fully human-annotated data\u2014outperforming the strongest multilingual baseline by 4.18% and 4.64%, respectively. Moreover, CUTLM surpasses prior cross-lingual transfer methods by up to 9.89%, confirming its effectiveness in preserving entity integrity during label projection. These findings establish a new benchmark for Uyghur legal NER and provide a scalable framework for cross-lingual NER in low-resource settings.<\/jats:p>","DOI":"10.1145\/3748325","type":"journal-article","created":{"date-parts":[[2025,7,14]],"date-time":"2025-07-14T11:43:31Z","timestamp":1752493411000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["A Bilingual Legal NER Dataset and Semantics-Aware Cross-Lingual Label Transfer Method for Low-Resource Languages"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-0432-692X","authenticated-orcid":false,"given":"Paerhati","family":"Tulajiang","sequence":"first","affiliation":[{"name":"School of computer science and technology, Dalian University of Technology","place":["Dalian, China"]},{"name":"College of computer science and technology, Xinjiang Normal University","place":["Dalian, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6515-134X","authenticated-orcid":false,"given":"Yuanyuan","family":"Sun","sequence":"additional","affiliation":[{"name":"School of computer science and technology, Dalian University of Technology","place":["Dalian, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6740-1701","authenticated-orcid":false,"given":"Yuanyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of computer science and technology, Dalian University of Technology","place":["Dalian, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-0606-4098","authenticated-orcid":false,"given":"Yingying","family":"Le","sequence":"additional","affiliation":[{"name":"School of computer science and technology, Dalian University of Technology","place":["Dalian, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6223-3642","authenticated-orcid":false,"given":"Kelaiti","family":"Xiao","sequence":"additional","affiliation":[{"name":"School of computer science and technology, Dalian University of Technology","place":["Dalian, China"]},{"name":"College of computer science and technology, Xinjiang Normal University","place":["Dalian, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0872-7688","authenticated-orcid":false,"given":"Hongfei","family":"Lin","sequence":"additional","affiliation":[{"name":"School of computer science and technology, Dalian University of Technology","place":["Dalian, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,11]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"198\u2013207","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Fang Z.","year":"2021","unstructured":"Z. Fang, Y. Cao, T. Li, R. Jia, F. Fang, Y. Shang, and Y. Lu. 2021. TEBNER: Domain specific named entity recognition with type expanded boundary-aware network[C]. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. (2021), 198\u2013207. Retrieved from https:\/\/aclanthology.org\/2021.emnlp-main.18\/"},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"B. Y. Lin W. Gao J. Yan et al. 2021. RockNER: A simple method to create adversarial examples for evaluating the robustness of named entity recognition models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 3728--3737. https:\/\/aclanthology.org\/2021.emnlp-main.302\/","DOI":"10.18653\/v1\/2021.emnlp-main.302"},{"key":"e_1_3_2_4_2","first-page":"5303\u20135308","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Wang R.","year":"2021","unstructured":"R. Wang and R. Henao. 2021. Unsupervised paraphrasing consistency training for low resource named entity recognition[C]. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. (2021), 5303\u20135308. Retrieved from https:\/\/aclanthology.org\/2021.emnlp-main.430\/"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","unstructured":"Z. Nasar S. W. Jaffry and M. K Malik. 2021. Named entity recognition and relation extraction: State-of-the-art. ACM Computing Surveys (CSUR) 54 1 (2021) 1--39. 10.1145\/3445965","DOI":"10.1145\/3445965"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","unstructured":"L. Zhong J. Wu Q. Li et al. 2023. A comprehensive survey on automatic knowledge graph construction. ACM Computing Surveys 56 4 (2023) 1--62. 10.1145\/3618295","DOI":"10.1145\/3618295"},{"key":"e_1_3_2_7_2","article-title":"Convolutional recurrent neural networks for rare sound event detection[J]","volume":"12","author":"Cak\u0131r E.","year":"2019","unstructured":"E. Cak\u0131r and T. Virtanen. 2019. Convolutional recurrent neural networks for rare sound event detection[J]. Deep Neural Networks for Sound Event Detection, 2019, 12. Retrieved from https:\/\/trepo.tuni.fi\/bitstream\/handle\/10024\/114151\/cakir_12.pdf?sequence=1#page=141","journal-title":"Deep Neural Networks for Sound Event Detection"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"M. E. Peters W. Ammar C. Bhagavatula et al. 2017. Semi-supervised sequence tagging with bidirectional language models. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1756--1765. https:\/\/aclanthology.org\/P17-1161\/","DOI":"10.18653\/v1\/P17-1161"},{"issue":"1","key":"e_1_3_2_9_2","article-title":"Empower sequence labeling with task-aware neural language model[C]","volume":"32","author":"Liu L.","year":"2018","unstructured":"L. Liu, J. Shang, X. Ren, F. F. Xu, H. Gui, J. Peng, and J. Han. 2018. Empower sequence labeling with task-aware neural language model[C]. In Proceedings of the AAAI Conference on Artificial Intelligence 32, 1 (2018), 5253--5260. Retrieved from https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/view\/12006","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.psychres.2021.114135"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3638762"},{"key":"e_1_3_2_12_2","first-page":"4996\u20135001","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Pires T.","unstructured":"T. Pires, E. Schlinger, and D. Garrette. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 4996\u20135001. Retrieved from https:\/\/aclanthology.org\/P19-1493\/"},{"volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","key":"e_1_3_2_13_2","unstructured":"A. Conneau, K. Khandelwal, N. Goyal, et al. 2020. Unsupervised cross-lingual representation learning at scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 8440--8451. https:\/\/aclanthology.org\/2020.acl-main.747\/"},{"key":"e_1_3_2_14_2","unstructured":"Z. Yang Z. Xu Y. Cui et al. 2022. CINO: A chinese minority pre-trained language model. Proceedings of the 29th International Conference on Computational Linguistics. 3937--3949. https:\/\/aclanthology.org\/2022.coling-1.346\/"},{"key":"e_1_3_2_15_2","unstructured":"J. Devlin M. W. Chang K. Lee et al. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Vol. 1 (Long and Short Papers). 4171--4186. https:\/\/aclanthology.org\/N19-1423\/"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3675397"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454921"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"H. Huang Y. Liang N. Duan et al. 2019. Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2485--2494. https:\/\/aclanthology.org\/D19-1252\/","DOI":"10.18653\/v1\/D19-1252"},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","unstructured":"Z. Chi L. Dong F. Wei et al. 2021. INFOXLM: An information-theoretic framework for cross-lingual language model Pre-training. 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies NAACL-HLT 2021. Association for Computational Linguistics (ACL'21). 3576--3588. https:\/\/aclanthology.org\/anthology-files\/pdf\/naacl\/2021.naacl-main.280.pdf","DOI":"10.18653\/v1\/2021.naacl-main.280"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3617831#core-collateral-purchase-access"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649456"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Lin Pan Chung-Wei Hang Haode Qi Abhishek Shah Saloni Potdar and Mo Yu. 2021. Multilingual BERT Post-Pretraining Alignment. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Online. Association for Computational Linguistics. 210--219. https:\/\/aclanthology.org\/2021.naacl-main.20\/","DOI":"10.18653\/v1\/2021.naacl-main.20"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"T. Chang Z. Tu and B. Bergen. 2022. The geometry of multilingual language model representations. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 119--136. https:\/\/aclanthology.org\/2022.emnlp-main.9\/","DOI":"10.18653\/v1\/2022.emnlp-main.9"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3653299"},{"key":"e_1_3_2_25_2","first-page":"2956\u20132961","article-title":"Tibert: Tibetan pre-trained language model[C]","author":"Liu S.","year":"2022","unstructured":"S. Liu, J. Deng, Y. Sun, and X. Zhao. 2022. Tibert: Tibetan pre-trained language model[C]. 2022 IEEE International Conference on Systems, Man, and Cybernetics (SMC). IEEE, (2022), 2956\u20132961. Retrieved from https:\/\/ieeexplore.ieee.org\/abstract\/document\/9945074","journal-title":"2022 IEEE International Conference on Systems"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"J. Xie Z. Yang G. Neubig et al. 2018. Neural cross-lingual named entity recognition with minimal resources. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 369--379. https:\/\/aclanthology.org\/D18-1034\/","DOI":"10.18653\/v1\/D18-1034"},{"key":"e_1_3_2_27_2","doi-asserted-by":"crossref","unstructured":"J. Ni G. Dinu and R. Floria. 2017. Weakly supervised cross-lingual named entity recognition via effective annotation and representation projection. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1470--1480. https:\/\/aclanthology.org\/P17-1135\/","DOI":"10.18653\/v1\/P17-1135"},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","unstructured":"S. Wu M. Dredze and Bentz Beto. 2019. Becas: The surprising cross-lingual effectiveness of BERT. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 833--844. https:\/\/aclanthology.org\/D19-1077\/","DOI":"10.18653\/v1\/D19-1077"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","unstructured":"X. Li L. Bing W. Zhang Z. Li and W. Lam. 2020. Unsupervised cross-lingual adaptation for sequence tagging and beyond[J]. arXiv:2010.12405. Retrieved from https:\/\/arxiv.org\/abs\/2010.12405 2020. DOI:10.48550\/arXiv.2010.12405","DOI":"10.48550\/arXiv.2010.12405"},{"key":"e_1_3_2_30_2","first-page":"6505\u20136514","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Wu Q.","year":"2020","unstructured":"Q. Wu, Z. Lin, B. Karlsson, J. G. Lou, and B. Huang. 2020. Single-\/multi-source cross-lingual NER via teacher-student learning on unlabeled data in target language[C]. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. (2020), 6505\u20136514. Retrieved from https:\/\/aclanthology.org\/2020.acl-main.581\/"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/543"},{"key":"e_1_3_2_32_2","first-page":"5834\u20135846","article-title":"MulDA: A multilingual data augmentation framework for low-resource cross-lingual NER[C]","author":"Liu L.","year":"2021","unstructured":"L. Liu, B. Ding, L. Bing, S. Joty, L. Si, and C. Miao. 2021. MulDA: A multilingual data augmentation framework for low-resource cross-lingual NER[C]. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (2021), 5834\u20135846. Retrieved from https:\/\/aclanthology.org\/2021.acl-long.453\/","journal-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)"},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"J. Yang S. Huang S. Ma Y. Yin L. Dong D. Zhang H. Guo Z. Li and F. Wei. 2022. CROP: Zero-shot cross-lingual named entity recognition with multilingual labeled sequence translation. Findings of the Association for Compu tational Linguistics: EMNLP 2022. 486--496. https:\/\/aclanthology.org\/2022.findings-emnlp.34\/","DOI":"10.18653\/v1\/2022.findings-emnlp.34"},{"key":"e_1_3_2_34_2","first-page":"5775\u20135796","article-title":"Frustratingly easy label projection for cross-lingual transfer[C]","author":"Chen Y.","year":"2023","unstructured":"Y. Chen, C. Jiang, A. Ritter, and W. Xu. 2023. Frustratingly easy label projection for cross-lingual transfer[C]. Findings of the Association for Computational Linguistics: ACL (2023), 5775\u20135796. Retrieved from https:\/\/aclanthology.org\/2023.findings-acl.357\/","journal-title":"Findings of the Association for Computational Linguistics: ACL"},{"key":"e_1_3_2_35_2","first-page":"5995\u20136009","article-title":"CoLaDa: A collaborative label denoising framework for cross-lingual named entity recognition[C]","author":"Ma T.","year":"2023","unstructured":"T. Ma, Q. Wu, H. Jiang, B. F. Karlsson, T. Zhao, and C. Y. Lin. 2023. CoLaDa: A collaborative label denoising framework for cross-lingual named entity recognition[C]. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (2023), 5995\u20136009. Retrieved from https:\/\/aclanthology.org\/2023.acl-long.330\/","journal-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)"},{"key":"e_1_3_2_36_2","first-page":"8438\u20138449","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Zhou R.","year":"2022","unstructured":"R. Zhou, X. Li, L. Bing, E. Cambria, L. Si, and C. Miao. 2022. ConNER: Consistency training for cross-lingual named entity recognition[C]. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 8438\u20138449. Retrieved from https:\/\/aclanthology.org\/2022.emnlp-main.577\/"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1504\/IJICT.2019.102996"},{"key":"e_1_3_2_38_2","first-page":"273\u2013278","article-title":"Named entity recognition of uyghur organizations based on conditional random fields","volume":"40","author":"Maimaiti Maihemuti","year":"2019","unstructured":"Maihemuti Maimaiti, Lulu Wang, Tuerge Eniubulayin, Aishan Wumair, and Kaharjiang Abide Rexiti. 2019. Named entity recognition of uyghur organizations based on conditional random fields. Computer Engineering and Design 40 (2019), 273\u2013278","journal-title":"Computer Engineering and Design"},{"key":"e_1_3_2_39_2","unstructured":"Lulu Wang Wumaier Aishan Maimaiti Maihemuti Abiderexiti Kahaerjiang and Yibulayin Tuerge. 2018. A Semi-supervised approach to uyghur named entity recognition based on CRF. Journal of Chinese Information Processing 32 11 (2018) 16--26."},{"key":"e_1_3_2_40_2","unstructured":"Lulu Wang Aishan Wumaier Tuergen Yibulayin Maihemuti Maimaiti and Kahaerjiang Abiderexiti. 2019. Uyghur Named Entity Recognition Based on Deep Neural Network. Journal of Chinese Information Processing 33 3 (2019) 64--70."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3665244"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.16383\/j.aas.2017.c150769"},{"issue":"2","key":"e_1_3_2_43_2","first-page":"58\u201365","article-title":"Uyghur named entity recognition based on transfer learning","volume":"52","author":"Kong Xiangpeng","year":"2020","unstructured":"Xiangpeng Kong, Wu Shouer, Slamu Qimeng Yang, and Zhe Li. 2020. Uyghur named entity recognition based on transfer learning. Journal of Northeast Normal University: Natural Science Edition 52, 2 (2020), 58\u201365.","journal-title":"Journal of Northeast Normal University: Natural Science Edition"},{"key":"e_1_3_2_44_2","first-page":"417\u2013426","volume-title":"Proceedings of the 13th Language Resources and Evaluation Conference","author":"Yeshpanov R.","year":"2022","unstructured":"R. Yeshpanov, Y. Khassanov, and H. A. Varol. 2022. KazNERD: Kazakh named entity recognition dataset[C]. In Proceedings of the 13th Language Resources and Evaluation Conference. 417\u2013426. Retrieved from https:\/\/aclanthology.org\/2022.lrec-1.44\/"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.5555\/645530.655813"},{"key":"e_1_3_2_46_2","first-page":"1064\u20131074","article-title":"End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF[C]","author":"Ma X.","year":"2016","unstructured":"X. Ma and E. Hovy. 2016. End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF[C]. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1064\u20131074. Retrieved from https:\/\/aclanthology.org\/P16-1101\/","journal-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3748325","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T13:36:40Z","timestamp":1757597800000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3748325"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,11]]},"references-count":45,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3748325"],"URL":"https:\/\/doi.org\/10.1145\/3748325","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2025,9,11]]},"assertion":[{"value":"2024-10-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}