{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,13]],"date-time":"2026-02-13T22:52:30Z","timestamp":1771023150921,"version":"3.50.1"},"reference-count":26,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2016,10,14]],"date-time":"2016-10-14T00:00:00Z","timestamp":1476403200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2017,6,30]]},"abstract":"<jats:p>Manually constructing an annotated Named Entity (NE) in a bilingual corpus is a time-consuming, labor--intensive, and expensive process, but this is necessary for natural language processing (NLP) tasks such as cross-lingual information retrieval, cross-lingual information extraction, machine translation, etc. In this article, we present an automatic approach to construct an annotated NE in English-Vietnamese bilingual corpus from a bilingual parallel corpus by proposing an aligned NE method. Basing this corpus on a bilingual corpus in which the initial NEs are extracted from its own language separately, the approach tries to correct unrecognized NEs or incorrectly recognized NEs before aligning the NEs by using a variety of bilingual constraints. The generated corpus not only improves the NE recognition results but also creates alignments between English NEs and Vietnamese NEs, which are necessary for training NE translation models. The experimental results show that the approach outperforms the baseline methods effectively. In the English-Vietnamese NE alignment task, the F-measure increases from 68.58% to 79.77%. Thanks to the improvement of the NE recognition quality, the proposed method also increases significantly: the F-measure goes from 84.85% to 88.66% for the English side and from 75.71% to 85.55% for the Vietnamese side. By providing the additional semantic information for the machine translation systems, the BLEU score increases from 33.04% to 45.11%.<\/jats:p>","DOI":"10.1145\/2990191","type":"journal-article","created":{"date-parts":[[2016,10,14]],"date-time":"2016-10-14T13:34:24Z","timestamp":1476452064000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["An Approach to Construct a Named Entity Annotated English-Vietnamese Bilingual Corpus"],"prefix":"10.1145","volume":"16","author":[{"given":"Long H. B.","family":"Nguyen","sequence":"first","affiliation":[{"name":"University of Science, HCM City, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dien","family":"Dinh","sequence":"additional","affiliation":[{"name":"University of Science, HCM City, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Phuoc","family":"Tran","sequence":"additional","affiliation":[{"name":"Ton Duc Thang University, HCM City, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,10,14]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073150"},{"key":"e_1_2_1_2_1","unstructured":"P. F. Brown S. A. Della Pietra V. J. Della Pietra and R. L. Mercer. 1993. The mathematics of statistical machine translation: Parameter estimation. Computational Linguistics.   P. F. Brown S. A. Della Pietra V. J. Della Pietra and R. L. Mercer. 1993. The mathematics of statistical machine translation: Parameter estimation. Computational Linguistics."},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL).","author":"Che Wanxiang","year":"2013"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics (ACL), 631--639","author":"Chen Yufeng"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1162\/COLI_a_00122"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.3115\/1075096.1075108"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-5010"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of 4th IEEE International Conference on Computer Science - Research, Innovation and Vision of the Future 2006 (RIVF\u201906)","author":"Dien D."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing.","author":"Feng Donghui","year":"2004"},{"key":"e_1_2_1_10_1","doi-asserted-by":"crossref","volume-title":"FASTUS: A Cascaded Finite-State Transducer for Extracting Information from Natural Language Text","author":"Hobbs J.","DOI":"10.3115\/1075671.1075701"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.3115\/1119384.1119386"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.3115\/1075218.1075268"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the 5th Workshop on South and Southeast Asian Natural Language Processing.","author":"Hung Ngo Quoc","year":"2014"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/972695.972699"},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing (EMNLP). 388--395","author":"Koehn Philipp","year":"2004"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.3115\/1118905.1118922"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1165255.1165257"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1162\/089120100561683"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.3115\/1067807.1067842"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.3115\/1075218.1075274"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1162\/089120103321337421"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-014-3127-5"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.3115\/993268.993313"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the 27th AAAI Conference on Artificial Intelligence (AAAI).","author":"Wang Mengqiu"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.3115\/1072228.1072238"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2990191","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2990191","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:39:34Z","timestamp":1750217974000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2990191"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,10,14]]},"references-count":26,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2017,6,30]]}},"alternative-id":["10.1145\/2990191"],"URL":"https:\/\/doi.org\/10.1145\/2990191","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,10,14]]},"assertion":[{"value":"2015-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-10-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}