{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:19:59Z","timestamp":1750220399008,"version":"3.41.0"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2021,12,13]],"date-time":"2021-12-13T00:00:00Z","timestamp":1639353600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2022,5,31]]},"abstract":"<jats:p>In this article, we propose a new encoding scheme for named entity recognition (NER) called Joined Type-Length encoding (JoinedTL). Unlike most existing named entity encoding schemes, which focus on flat entities, JoinedTL can label nested named entities in a single sequence. JoinedTL uses a packed encoding to represent both type and span of a named entity, which not only results in less tagged tokens compared to existing encoding schemes, but also enables it to support nested NER. We evaluate the effectiveness of JoinedTL for nested NER on three nested NER datasets: GENIA in English, GermEval in German, and PerNest, our newly created nested NER dataset in Persian. We apply CharLSTM+WordLSTM+CRF, a three-layer sequence tagging model on three datasets encoded using JoinedTL and two existing nested NE encoding schemes, i.e., JoinedBIO and JoinedBILOU. Our experiment results show that CharLSTM+WordLSTM+CRF trained with JoinedTL encoded datasets can achieve competitive F1 scores as the ones trained with datasets encoded by two other encodings, but with 27%\u201348% less tagged tokens. To leverage the power of three different encodings, i.e., JoinedTL, JoinedBIO, and JoinedBILOU, we propose an encoding-based ensemble method for nested NER. Evaluation results show that the ensemble method achieves higher F1 scores on all datasets than the three models each trained using one of the three encodings. By using nested NE encodings including JoinedTL with CharLSTM+WordLSTM+CRF, we establish new state-of-the-art performance with an F1 score of 83.7 on PerNest, 74.9 on GENIA, and 70.5 on GermEval, surpassing two recent neural models specially designed for nested NER.<\/jats:p>","DOI":"10.1145\/3487057","type":"journal-article","created":{"date-parts":[[2021,12,13]],"date-time":"2021-12-13T14:36:25Z","timestamp":1639406185000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Joined Type Length Encoding for Nested Named Entity Recognition"],"prefix":"10.1145","volume":"21","author":[{"given":"Mohammad Sadegh","family":"Sheikhaei","sequence":"first","affiliation":[{"name":"School of Computing, Queen\u2019s University, Kingston, Ontario, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hasan","family":"Zafari","sequence":"additional","affiliation":[{"name":"School of Computing, Queen\u2019s University, Kingston, Ontario, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuan","family":"Tian","sequence":"additional","affiliation":[{"name":"School of Computing, Queen\u2019s University, Kingston, Ontario, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,12,13]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383306"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.5555\/1572392.1572404"},{"key":"e_1_3_2_4_2","unstructured":"Darina Benikova Chris Biemann Max Kisselew and Sebastian Pado. 2014. GermEval 2014 named entity recognition shared task: Companion paper. Retrieved 25 Nov. 2021. from https:\/\/hildok.bsz-bw.de\/frontdoor\/index\/index\/docId\/283."},{"key":"e_1_3_2_5_2","first-page":"2524","volume-title":"Proceedings of the 9th International Conference on Language Resources and Evaluation","author":"Benikova Darina","year":"2014","unstructured":"Darina Benikova, Chris Biemann, and Marc Reznicek. 2014. NoSta-D named entity annotation for German: Guidelines and dataset. In Proceedings of the 9th International Conference on Language Resources and Evaluation. European Language Resources Association, 2524\u20132531. Retrieved from http:\/\/www.lrec-conf.org\/proceedings\/lrec2014\/pdf\/276_Paper.pdf."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-010-9132-x"},{"key":"e_1_3_2_7_2","volume-title":"A Maximum Entropy Approach to Named Entity Recognition","author":"Borthwick Andrew Eliot","year":"1999","unstructured":"Andrew Eliot Borthwick. 1999. A Maximum Entropy Approach to Named Entity Recognition. Ph.D. Dissertation. New York University, New York, NY. UMI Order Number:AAI 9945252."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00104"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3015467"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.5555\/1699510.1699529"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1585"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.29252\/jsdp.14.3.127"},{"key":"e_1_3_2_13_2","unstructured":"Zhiheng Huang Wei Xu and Kai Yu. 2015. Bidirectional LSTM-CRF Models for Sequence Tagging. CoRR abs\/1508.01991. http:\/\/arxiv.org\/abs\/1508.01991."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1131"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3329710"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1079"},{"issue":"1","key":"e_1_3_2_17_2","first-page":"i180\u2013i182","article-title":"GENIA corpus\u2019a semantically annotated corpus for bio-textmining","volume":"19","author":"Kim J.-D.","year":"2003","unstructured":"J.-D. Kim, T. Ohta, Y. Tateisi, and J. Tsujii. 2003. GENIA corpus\u2019a semantically annotated corpus for bio-textmining. Bioinformatics 19, suppl_1 (July 2003), i180\u2013i182. DOI:https:\/\/doi.org\/10.1093\/bioinformatics\/btg1023arXiv:https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/19\/suppl_1\/i180\/614820\/btg1023.pdf.","journal-title":"Bioinformatics"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.3115\/1073336.1073361"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1030"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.3038670"},{"key":"e_1_3_2_21_2","first-page":"5849","article-title":"A unified MRC framework for named entity recognition","author":"Li Xiaoya","year":"2019","unstructured":"Xiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han, Fei Wu, and Jiwei Li. 2019. A unified MRC framework for named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 5849\u20135859.","journal-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.5555\/3504035.3504679"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1102"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.571"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3129290"},{"key":"e_1_3_2_26_2","article-title":"Ancora: Multilingual and multilevel annotated corpora","author":"Mart\u00ed Maria Ant\u00f2nia","year":"2007","unstructured":"Maria Ant\u00f2nia Mart\u00ed, Mariona Taul\u00e9, Manu Bertran, and Llu\u00eds M\u00e0rquez. 2007. Ancora: Multilingual and multilevel annotated corpora. MS, Universitat de Barcelona (2007).","journal-title":"MS, Universitat de Barcelona"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.11613\/BM.2012.031"},{"key":"e_1_3_2_28_2","first-page":"51","volume-title":"Proceedings of the Australasian Language Technology Workshop 2006","author":"Moll\u00e1 Diego","year":"2006","unstructured":"Diego Moll\u00e1, Menno van Zaanen, and Daniel Smith. 2006. Named entity recognition for question answering. In Proceedings of the Australasian Language Technology Workshop 2006, 51\u201358. Retrieved from https:\/\/www.aclweb.org\/anthology\/U06-1009."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1276"},{"key":"e_1_3_2_30_2","unstructured":"Binh An Nguyen Kiet Van Nguyen and Ngan Luu-Thuy Nguyen. 2019. Error Analysis for Vietnamese Named Entity Recognition on Deep Neural Network Models. CoRR abs\/1911.07228. http:\/\/arxiv.org\/abs\/1911.07228."},{"key":"e_1_3_2_31_2","volume-title":"Proceedings of the 11th International Conference on Language Resources and Evaluation","author":"Poostchi Hanieh","year":"2018","unstructured":"Hanieh Poostchi, Ehsan Zare Borzeshi, and Massimo Piccardi. 2018. BiLSTM-CRF for persian named-entity recognition ArmanPersoNERCorpus: The first entity-annotated persian dataset. In Proceedings of the 11th International Conference on Language Resources and Evaluation. European Language Resources Association. Retrieved from https:\/\/www.aclweb.org\/anthology\/L18-1701."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-25007-6_27"},{"key":"e_1_3_2_33_2","volume-title":"Proceedings of the 3rd Workshop on Very Large Corpora","author":"Ramshaw Lance","year":"1995","unstructured":"Lance Ramshaw and Mitch Marcus. 1995. Text chunking using transformation-based learning. In Proceedings of the 3rd Workshop on Very Large Corpora. Retrieved from https:\/\/www.aclweb.org\/anthology\/W95-0107."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.5555\/1596374.1596399"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.5555\/927230"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","unstructured":"Marek Rei. 2017. Semi-supervised multitask learning for sequence labeling. CoRR abs\/1704.07156.","DOI":"10.18653\/v1\/P17-1194"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.3115\/977035.977059"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Mahsa Sadat Shahshahani Mahdi Mohseni Azadeh Shakery and Heshaam Faili. 2018. PEYMA: A tagged corpus for Persian named entities. CoRR.","DOI":"10.29252\/jsdp.16.1.91"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/11424918_40"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00334"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1309"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.5555\/2380921.2380942"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1527"},{"key":"e_1_3_2_44_2","article-title":"Recognizing and encoding disorder concepts in clinical text using machine learning and vector space model","volume":"1179","author":"Tang Buzhou","year":"2013","unstructured":"Buzhou Tang, Yonghui Wu, M. Jiang, Joshua Denny, and H. Xu. 2013. Recognizing and encoding disorder concepts in clinical text using machine learning and vector space model. CEUR Workshop Proceedings 1179 (Jan. 2013).","journal-title":"CEUR Workshop Proceedings"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.3115\/1119176.1119195"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1019"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1124"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.525"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.486"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1021\/acs.jcim.9b00470"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S18-2021"},{"key":"e_1_3_2_52_2","first-page":"3879","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Yang Jie","year":"2018","unstructured":"Jie Yang, Shuailong Liang, and Yue Zhang. 2018. Design challenges and misconceptions in neural sequence labeling. In Proceedings of the 27th International Conference on Computational Linguistics. Association for Computational Linguistics, 3879\u20133889. Retrieved from https:\/\/www.aclweb.org\/anthology\/C18-1327."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-4013"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.577"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1034"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487057","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3487057","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:18:47Z","timestamp":1750191527000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487057"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,13]]},"references-count":54,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,5,31]]}},"alternative-id":["10.1145\/3487057"],"URL":"https:\/\/doi.org\/10.1145\/3487057","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2021,12,13]]},"assertion":[{"value":"2020-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-12-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}