{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:25:40Z","timestamp":1750307140805,"version":"3.41.0"},"reference-count":16,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2011,12,1]],"date-time":"2011-12-01T00:00:00Z","timestamp":1322697600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Transactions on Asian Language Information Processing"],"published-print":{"date-parts":[[2011,12]]},"abstract":"<jats:p>\n            Bilingual Named Entity (NE) pairs are valuable resources for many NLP applications. Since comparable corpora are more accessible, abundant and up-to-date, recent researches have concentrated on mining bilingual lexicons using comparable corpora. Leveraging comparable corpora, this research presents a novel approach to mining English-Chinese NE translations by combining multi-dimension features from various information sources for every possible NE pair, which include the transliteration model, English-Chinese matching, Chinese-English matching, translation model, length, and context vector. These features are integrated into one model with linear combination and minimum sample risk (MSR) algorithm. As for the high type-dependence of NE translation, we integrate different features according to different NE types. We experiment with the above individual feature or integrated features to mine person NE (PN) pairs, location NE (LN) pairs and organization NE (ON) pairs. When using transliteration and length to mine PN pairs, we achieve the best performance of 84.9% (\n            <jats:italic>F<\/jats:italic>\n            -score). The LN pairs can be mined with the features of transliteration model, length, translation model, English-Chinese matching and Chinese-English matching. And the best performance is 83.4% (\n            <jats:italic>F<\/jats:italic>\n            -score). The ON pairs can be mined with the features of English-Chinese matching and Chinese-English matching. It reaches the best performance with 84.1% (\n            <jats:italic>F<\/jats:italic>\n            -score).\n          <\/jats:p>","DOI":"10.1145\/2025384.2025387","type":"journal-article","created":{"date-parts":[[2011,12,6]],"date-time":"2011-12-06T19:05:23Z","timestamp":1323198323000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Mining English-Chinese Named Entity Pairs from Comparable Corpora"],"prefix":"10.1145","volume":"10","author":[{"given":"Lishuang","family":"Li","sequence":"first","affiliation":[{"name":"Dalian University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peng","family":"Wang","sequence":"additional","affiliation":[{"name":"Dalian University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Degen","family":"Huang","sequence":"additional","affiliation":[{"name":"Dalian University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lian","family":"Zhao","sequence":"additional","affiliation":[{"name":"Dalian University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2011,12]]},"reference":[{"volume-title":"Proceedings of the 2nd European Conference on Research and Advanced Technology for Digital Libraries (ECDL\u201998)","author":"Braschler M.","key":"e_1_2_1_1_1"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1330291.1330292"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.3115\/1219840.1219885"},{"volume-title":"Proceedings of the 3rd Annual Workshop on Very Large Corpora (VLC\u201995)","year":"1995","author":"Pascale F.","key":"e_1_2_1_4_1"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220575.1220602"},{"volume-title":"Proceedings of Machine Translation Summit XI (MT\u201907)","author":"Gao G. H.","key":"e_1_2_1_6_1"},{"volume-title":"Proceedings of the 23rd International Conference on Computational Linguistics (CL\u201910)","author":"Huang D. G.","key":"e_1_2_1_7_1"},{"key":"e_1_2_1_8_1","first-page":"23","article-title":"Named entity translation with Web mining and transliteration","volume":"21","author":"Jiang L.","year":"2007","journal-title":"J. Chinese Inf. Proc."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220835.1220846"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1198296.1198298"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220355.1220444"},{"volume-title":"Proceedings of the 20th Pacific Asia Conference on Language, Information and Computation (PACL\u201906)","author":"Lu M.","key":"e_1_2_1_12_1"},{"key":"e_1_2_1_13_1","first-page":"3","article-title":"Chinese word segmentation based on the marginal probabilities generated by CRFs","volume":"23","author":"Luo Y. Y.","year":"2009","journal-title":"J. Chinese Inf. Proc."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220175.1220185"},{"volume-title":"Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201906)","author":"Tao T.","key":"e_1_2_1_15_1"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1360\/jos180196"}],"container-title":["ACM Transactions on Asian Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2025384.2025387","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2025384.2025387","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T09:48:52Z","timestamp":1750240132000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2025384.2025387"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,12]]},"references-count":16,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2011,12]]}},"alternative-id":["10.1145\/2025384.2025387"],"URL":"https:\/\/doi.org\/10.1145\/2025384.2025387","relation":{},"ISSN":["1530-0226","1558-3430"],"issn-type":[{"type":"print","value":"1530-0226"},{"type":"electronic","value":"1558-3430"}],"subject":[],"published":{"date-parts":[[2011,12]]},"assertion":[{"value":"2010-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-12-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}