{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:09:37Z","timestamp":1750306177736,"version":"3.41.0"},"reference-count":22,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2017,3,17]],"date-time":"2017-03-17T00:00:00Z","timestamp":1489708800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2017,9,30]]},"abstract":"<jats:p>This article proposes a technique for mining bilingual lexicons from pairs of parallel short word sequences. The technique builds a generative model from a corpus of training data consisting of such pairs. The model is a hierarchical nonparametric Bayesian model that directly induces a bilingual lexicon while training. The model learns in an unsupervised manner and is designed to exploit characteristics of the language pairs being mined. The proposed model is capable of utilizing commonly used word-pair frequency information and additionally can employ the internal character alignments within the words themselves. It is thereby capable of mining transliterations and can use reliably aligned transliteration pairs to support the mining of other words in their context. The model is also capable of performing word reordering and word deletion during the alignment process, and it is furthermore capable of operating in the absence of full segmentation information. In this work, we study two mining tasks based on English-Japanese and English-Chinese language pairs, and compare the proposed approach to baselines based on a simpler models that use only word-pair frequency information. Our results show that the proposed method is able to mine bilingual word pairs at higher levels of precision and recall than the baselines.<\/jats:p>","DOI":"10.1145\/3003726","type":"journal-article","created":{"date-parts":[[2017,3,20]],"date-time":"2017-03-20T12:29:43Z","timestamp":1490012983000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Inducing a Bilingual Lexicon from Short Parallel Multiword Sequences"],"prefix":"10.1145","volume":"16","author":[{"given":"Andrew","family":"Finch","sequence":"first","affiliation":[{"name":"National Institute of Information and Communications Technology, Kyoto, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Taisuke","family":"Harada","sequence":"additional","affiliation":[{"name":"Kyushu University, Fukuoka City, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kumiko","family":"Tanaka-Ishii","sequence":"additional","affiliation":[{"name":"University of Tokyo, Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eiichiro","family":"Sumita","sequence":"additional","affiliation":[{"name":"National Institute of Information and Communications Technology, Kyoto, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,3,17]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/1690219.1690256"},{"key":"e_1_2_1_2_1","first-page":"10","volume-title":"Proceedings of the 2010 Named Entities Workshop. 53--56","author":"Darwish Kareem","year":"2010","unstructured":"Kareem Darwish . 2010 . Transliteration mining with phonetic conflation and iterative training . In Proceedings of the 2010 Named Entities Workshop. 53--56 . http:\/\/www.aclweb.org\/anthology\/W 10 - 2407 Kareem Darwish. 2010. Transliteration mining with phonetic conflation and iterative training. In Proceedings of the 2010 Named Entities Workshop. 53--56. http:\/\/www.aclweb.org\/anthology\/W10-2407"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the 7th International Workshop on Spoken Language Translation (IWSLT\u201910)","author":"Finch Andrew","year":"2010","unstructured":"Andrew Finch and Eiichiro Sumita . 2010 . A Bayesian model of bilingual segmentation for transliteration . In Proceedings of the 7th International Workshop on Spoken Language Translation (IWSLT\u201910) . 259--266. Andrew Finch and Eiichiro Sumita. 2010. A Bayesian model of bilingual segmentation for transliteration. In Proceedings of the 7th International Workshop on Spoken Language Translation (IWSLT\u201910). 259--266."},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the 2012 International Workshop on Spoken Language Translation (IWSLT\u201912)","author":"Finch Andrew M.","year":"2012","unstructured":"Andrew M. Finch , Ohnmar Htun , and Eiichiro Sumita . 2012 . The NICT translation system for IWSLT 2012 . In Proceedings of the 2012 International Workshop on Spoken Language Translation (IWSLT\u201912) . 121--125. http:\/\/www.isca-speech.org\/archive\/iwslt_12\/sltc_121.html. Andrew M. Finch, Ohnmar Htun, and Eiichiro Sumita. 2012. The NICT translation system for IWSLT 2012. In Proceedings of the 2012 International Workshop on Spoken Language Translation (IWSLT\u201912). 121--125. http:\/\/www.isca-speech.org\/archive\/iwslt_12\/sltc_121.html."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499955.2499957"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220175.1220260"},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Ohnmar Htun Andrew Finch Eiichiro Sumita and Yoshiki Mikami. 2012. Improving transliteration mining by integrating expert knowledge with statistical approaches. International Journal of Computer Applications 58 Article No. 17.  Ohnmar Htun Andrew Finch Eiichiro Sumita and Yoshiki Mikami. 2012. Improving transliteration mining by integrating expert knowledge with statistical approaches. International Journal of Computer Applications 58 Article No. 17.","DOI":"10.5120\/9373-3821"},{"key":"e_1_2_1_8_1","first-page":"1211","article-title":"Generalized weighted Chinese restaurant processes for species sampling mixture models","volume":"13","author":"Ishwaran Hemant","year":"2003","unstructured":"Hemant Ishwaran and Lancelot F. James . 2003 . Generalized weighted Chinese restaurant processes for species sampling mixture models . Statistica Sinica 13 , 4, 1211 -- 1235 . Hemant Ishwaran and Lancelot F. James. 2003. Generalized weighted Chinese restaurant processes for species sampling mixture models. Statistica Sinica 13, 4, 1211--1235.","journal-title":"Statistica Sinica"},{"key":"e_1_2_1_9_1","first-page":"10","volume-title":"Proceedings of the 2010 Named Entities Workshop. 39--47","author":"Jiampojamarn Sittichai","year":"2010","unstructured":"Sittichai Jiampojamarn , Kenneth Dwyer , Shane Bergsma , Aditya Bhargava , Qing Dou , Mi-Young Kim , and Grzegorz Kondrak . 2010 . Transliteration generation and mining with limited training resources . In Proceedings of the 2010 Named Entities Workshop. 39--47 . http:\/\/www.aclweb.org\/anthology\/W 10 - 2405 Sittichai Jiampojamarn, Kenneth Dwyer, Shane Bergsma, Aditya Bhargava, Qing Dou, Mi-Young Kim, and Grzegorz Kondrak. 2010. Transliteration generation and mining with limited training resources. In Proceedings of the 2010 Named Entities Workshop. 39--47. http:\/\/www.aclweb.org\/anthology\/W10-2405"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/972764.972767"},{"key":"e_1_2_1_11_1","first-page":"10","volume-title":"Report of NEWS 2010 transliteration mining shared task. In Proceedings of the 2010 Named Entities Workshop. 21--28","author":"Kumaran A.","year":"2010","unstructured":"A. Kumaran , Mitesh M. Khapra , and Haizhou Li . 2010 . Report of NEWS 2010 transliteration mining shared task. In Proceedings of the 2010 Named Entities Workshop. 21--28 . http:\/\/www.aclweb.org\/anthology\/W 10 - 2403 A. Kumaran, Mitesh M. Khapra, and Haizhou Li. 2010. Report of NEWS 2010 transliteration mining shared task. In Proceedings of the 2010 Named Entities Workshop. 21--28. http:\/\/www.aclweb.org\/anthology\/W10-2403"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/2390948.2390976"},{"key":"e_1_2_1_13_1","first-page":"13","volume-title":"Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 393--398","author":"Li Tingting","year":"2013","unstructured":"Tingting Li , Tiejun Zhao , Andrew Finch , and Chunyue Zhang . 2013 . A tightly-coupled unsupervised clustering and bilingual alignment model for transliteration . In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 393--398 . http:\/\/www.aclweb.org\/anthology\/P 13 - 2070 Tingting Li, Tiejun Zhao, Andrew Finch, and Chunyue Zhang. 2013. A tightly-coupled unsupervised clustering and bilingual alignment model for transliteration. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 393--398. http:\/\/www.aclweb.org\/anthology\/P13-2070"},{"volume-title":"Proceedings of the 5th International Workshop on Frontiers in Handwriting Recognition. 233--238","author":"Lopresti D.","key":"e_1_2_1_14_1","unstructured":"D. Lopresti , A. Tomkins , and J. Zhou . 1997. Algorithms for matching hand-drawn sketches . In Proceedings of the 5th International Workshop on Frontiers in Handwriting Recognition. 233--238 . D. Lopresti, A. Tomkins, and J. Zhou. 1997. Algorithms for matching hand-drawn sketches. In Proceedings of the 5th International Workshop on Frontiers in Handwriting Recognition. 233--238."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/1687878.1687894"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002553"},{"key":"e_1_2_1_17_1","first-page":"10","volume-title":"Proceedings of the 2010 Named Entities Workshop. 57--61","author":"Noeman Sara","year":"2010","unstructured":"Sara Noeman and Amgad Madkour . 2010 . Language independent transliteration mining system using finite state automata framework . In Proceedings of the 2010 Named Entities Workshop. 57--61 . http:\/\/www.aclweb.org\/anthology\/W 10 - 2408 Sara Noeman and Amgad Madkour. 2010. Language independent transliteration mining system using finite state automata framework. In Proceedings of the 2010 Named Entities Workshop. 57--61. http:\/\/www.aclweb.org\/anthology\/W10-2408"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1162\/089120103321337421"},{"key":"e_1_2_1_19_1","volume-title":"Retrieved","author":"Pitman Jim","year":"1995","unstructured":"Jim Pitman and Marc Yor . 1995 . The Two-Parameter Poisson-Dirichlet Distribution Derived from a Stable Subordinator . Retrieved November 19, 2016, from http:\/\/digitalassets.lib.berkeley.edu\/sdtr\/ucb\/text\/433.pdf. Jim Pitman and Marc Yor. 1995. The Two-Parameter Poisson-Dirichlet Distribution Derived from a Stable Subordinator. Retrieved November 19, 2016, from http:\/\/digitalassets.lib.berkeley.edu\/sdtr\/ucb\/text\/433.pdf."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.682181"},{"key":"e_1_2_1_21_1","first-page":"12","volume-title":"Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 469--477","author":"Sajjad Hassan","year":"2012","unstructured":"Hassan Sajjad , Alexander Fraser , and Helmut Schmid . 2012 . A statistical model for unsupervised and semi-supervised transliteration mining . In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 469--477 . http:\/\/www.aclweb.org\/anthology\/P 12 - 1049 Hassan Sajjad, Alexander Fraser, and Helmut Schmid. 2012. A statistical model for unsupervised and semi-supervised transliteration mining. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 469--477. http:\/\/www.aclweb.org\/anthology\/P12-1049"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1198\/016214502753479464"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3003726","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3003726","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:49:55Z","timestamp":1750218595000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3003726"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,3,17]]},"references-count":22,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2017,9,30]]}},"alternative-id":["10.1145\/3003726"],"URL":"https:\/\/doi.org\/10.1145\/3003726","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2017,3,17]]},"assertion":[{"value":"2015-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-03-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}