{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,31]],"date-time":"2025-12-31T12:18:46Z","timestamp":1767183526184,"version":"3.41.0"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2015,6,12]],"date-time":"2015-06-12T00:00:00Z","timestamp":1434067200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004543","name":"China Scholarship Council","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004543","id-type":"DOI","asserted-by":"crossref"}]},{"name":"JSPS KAKENHI","award":["26730121"],"award-info":[{"award-number":["26730121"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2015,6,12]]},"abstract":"<jats:p>\n            A machine-readable bilingual dictionary plays a crucial role in many natural language processing tasks, such as statistical machine translation and cross-language information retrieval. In this article, we propose a framework for extracting a bilingual dictionary from comparable corpora by exploiting a novel combination of topic modeling and word aligners such as the IBM models. Using a multilingual topic model, we first convert a comparable\n            <jats:italic>document<\/jats:italic>\n            -aligned corpus into a parallel\n            <jats:italic>topic<\/jats:italic>\n            -aligned corpus. This novel topic-aligned corpus is similar in structure to the\n            <jats:italic>sentence<\/jats:italic>\n            -aligned corpus frequently employed in statistical machine translation and allows us to extract a bilingual dictionary using a word alignment model.\n          <\/jats:p>\n          <jats:p>The main advantages of our framework is that (1) no seed dictionary is necessary for bootstrapping the process, and (2) multilingual comparable corpora in more than two languages can also be exploited. In our experiments on a large-scale Wikipedia dataset, we demonstrate that our approach can extract higher precision dictionaries compared to previous approaches and that our method improves further as we add more languages to the dataset.<\/jats:p>","DOI":"10.1145\/2699939","type":"journal-article","created":{"date-parts":[[2015,6,18]],"date-time":"2015-06-18T18:14:05Z","timestamp":1434651245000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Multilingual Topic Models for Bilingual Dictionary Extraction"],"prefix":"10.1145","volume":"14","author":[{"given":"Xiaodong","family":"Liu","sequence":"first","affiliation":[{"name":"Nara Institute of Science and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kevin","family":"Duh","sequence":"additional","affiliation":[{"name":"Nara Institute of Science and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuji","family":"Matsumoto","sequence":"additional","affiliation":[{"name":"Nara Institute of Science and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,6,12]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC\u201914)","author":"Aker Ahmet","year":"2014","unstructured":"Ahmet Aker , Monica Lestari Paramita , Marcis Pinnis , and Robert Gaizauskas . 2014 . Bilingual dictionaries for all EU languages . In Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC\u201914) . Ahmet Aker, Monica Lestari Paramita, Marcis Pinnis, and Robert Gaizauskas. 2014. Bilingual dictionaries for all EU languages. In Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC\u201914)."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/1964750.1964758"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553378"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the Wikimania Conference.","author":"Arai Yoshiaki","year":"2008","unstructured":"Yoshiaki Arai , Tomohiro Fukuhara , Hidetaka Masuda , and Hiroshi Nakagawa . 2008 . Analyzing interlanguage links of Wikipedias . In Proceedings of the Wikimania Conference. Yoshiaki Arai, Tomohiro Fukuhara, Hidetaka Masuda, and Hiroshi Nakagawa. 2008.Analyzing interlanguage links of Wikipedias. In Proceedings of the Wikimania Conference."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the Workshop on Multiword Expressions: From Parsing and Generation to the Real World (MWE\u201911)","author":"Baldwin Timothy","year":"2011","unstructured":"Timothy Baldwin . 2011 . MWEs and topic modelling: Enhancing machine learning with linguistics . In Proceedings of the Workshop on Multiword Expressions: From Parsing and Generation to the Real World (MWE\u201911) . Association for Computational Linguistics, Stroudsburg, PA, 1--1. http:\/\/dl.acm.org\/citation.cfm?id= 2021121.2021123. Timothy Baldwin. 2011. MWEs and topic modelling: Enhancing machine learning with linguistics. In Proceedings of the Workshop on Multiword Expressions: From Parsing and Generation to the Real World (MWE\u201911). Association for Computational Linguistics, Stroudsburg, PA, 1--1. http:\/\/dl.acm.org\/citation.cfm?id=2021121.2021123."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1214\/06-BA104"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944937"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the 6th International Conference on Language Resources and Evaluation (LREC\u201908)","author":"Bond Francis","year":"2008","unstructured":"Francis Bond , Hitoshi Isahara , Kyoko Kanzaki , and Kiyotaka Uchimoto . 2008 . Boot-strapping a WordNet using multiple existing WordNets . In Proceedings of the 6th International Conference on Language Resources and Evaluation (LREC\u201908) . Francis Bond, Hitoshi Isahara, Kyoko Kanzaki, and Kiyotaka Uchimoto. 2008. Boot-strapping a WordNet using multiple existing WordNets. In Proceedings of the 6th International Conference on Language Resources and Evaluation (LREC\u201908)."},{"volume-title":"Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence (UAI\u201909)","author":"Boyd-Graber Jordan","key":"e_1_2_1_9_1","unstructured":"Jordan Boyd-Graber and David M. Blei . 2009. Multilingual topic models for unaligned text . In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence (UAI\u201909) . Jordan Boyd-Graber and David M. Blei. 2009. Multilingual topic models for unaligned text. In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence (UAI\u201909)."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/972470.972474"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS\u201914)","author":"Stanislas Lauly Sarath Chandar A. P.","year":"2014","unstructured":"Sarath Chandar A. P. , Stanislas Lauly , Hugo Larochelle , Mitesh M. Khapra , Balaraman Ravindran , Vikas C. Raykar , and Amrita Saha . 2014 . An autoencoder approach to learning bilingual word representations . In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS\u201914) . Sarath Chandar A. P., Stanislas Lauly, Hugo Larochelle, Mitesh M. Khapra, Balaraman Ravindran, Vikas C. Raykar, and Amrita Saha. 2014.An autoencoder approach to learning bilingual word representations. In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS\u201914)."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-HLT\u201911)","author":"Jagadeesh Jagarlamudi Hal","year":"2011","unstructured":"Hal Daume III and Jagadeesh Jagarlamudi . 2011 . Domain adaptation for machine translation by mining unseen words . In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-HLT\u201911) . 407--412. Hal Daume III and Jagadeesh Jagarlamudi. 2011. Domain adaptation for machine translation by mining unseen words. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-HLT\u201911). 407--412."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1645953.1646020"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.3115\/1072228.1072394"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2442076.2442077"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201904)","author":"Fung Pascale","year":"2004","unstructured":"Pascale Fung and Percy Cheung . 2004 . Mining very-non-parallel corpora: Parallel sentence and lexicon extraction via bootstrapping and EM . In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201904) . Pascale Fung and Percy Cheung. 2004. Mining very-non-parallel corpora: Parallel sentence and lexicon extraction via bootstrapping and EM. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201904)."},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics (ACL-COLING\u201998)","author":"Fung Pascale","year":"1998","unstructured":"Pascale Fung and Yuen Yee Lo . 1998 . Translating unknown words using nonparallel, comparable texts . In Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics (ACL-COLING\u201998) . Pascale Fung and Yuen Yee Lo. 1998. Translating unknown words using nonparallel, comparable texts. In Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics (ACL-COLING\u201998)."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.3115\/1218955.1219022"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/383952.383965"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the 46th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-HLT\u201908)","author":"Haghighi Aria","year":"2008","unstructured":"Aria Haghighi , Percy Liang , Taylor Berg-Kirkpatrick , and Dan Klein . 2008 . Learning bilingual lexicons from monolingual corpora . In Proceedings of the 46th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-HLT\u201908) . 771--779. Aria Haghighi, Percy Liang, Taylor Berg-Kirkpatrick, and Dan Klein. 2008. Learning bilingual lexicons from monolingual corpora. In Proceedings of the 46th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-HLT\u201908). 771--779."},{"volume-title":"Proceedings of the International Conference on Neural Information Processing Systems (NIPS\u201910)","author":"Hoffman Matthew","key":"e_1_2_1_22_1","unstructured":"Matthew Hoffman , Francis R. Bach , and David M. Blei . 2010. Online learning for latent dirichlet allocation . In Proceedings of the International Conference on Neural Information Processing Systems (NIPS\u201910) . 856--864. Matthew Hoffman, Francis R. Bach, and David M. Blei. 2010. Online learning for latent dirichlet allocation. In Proceedings of the International Conference on Neural Information Processing Systems (NIPS\u201910). 856--864."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1110"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-12275-0_39"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the International Conference on Computational Linguistics (COLING\u201912)","author":"Klementiev Alexandre","year":"2012","unstructured":"Alexandre Klementiev , Ivan Titov , and Binod Bhattarai . 2012 . Inducing crosslingual distributed representations of words . In Proceedings of the International Conference on Computational Linguistics (COLING\u201912) . 1459--1474. Alexandre Klementiev, Ivan Titov, and Binod Bhattarai. 2012. Inducing crosslingual distributed representations of words. In Proceedings of the International Conference on Computational Linguistics (COLING\u201912). 1459--1474."},{"key":"e_1_2_1_26_1","volume-title":"Statistical Machine Translation","author":"Koehn Philipp","unstructured":"Philipp Koehn . 2010. Statistical Machine Translation ( 1 st Ed.). Cambridge University Press , New York, NY . Philipp Koehn. 2010. Statistical Machine Translation (1st Ed.). Cambridge University Press, New York, NY.","edition":"1"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.3115\/1118627.1118629"},{"volume-title":"Proceedings of the 6th Workshop on Building and Using Comparable Corpora. 11--15","year":"2013","key":"e_1_2_1_28_1","unstructured":"Hong-seok Kwon, Hyeong-won Seo, and Jae-hoon Kim. 2013 . Bilingual lexicon extraction via pivot language and word alignment tool . In Proceedings of the 6th Workshop on Building and Using Comparable Corpora. 11--15 . Hong-seok Kwon, Hyeong-won Seo, and Jae-hoon Kim. 2013. Bilingual lexicon extraction via pivot language and word alignment tool. In Proceedings of the 6th Workshop on Building and Using Comparable Corpora. 11--15."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/1873781.1873851"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220835.1220849"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 17th Conference on Computational Natural Language Learning (CoNLL\u201913)","author":"Liu Xiaodong","year":"2013","unstructured":"Xiaodong Liu , Kevin Duh , and Yuji Matsumoto . 2013 . Topic models + word alignment = a flexible framework for extracting bilingual dictionary from comparable corpus . In Proceedings of the 17th Conference on Computational Natural Language Learning (CoNLL\u201913) . 212. Xiaodong Liu, Kevin Duh, and Yuji Matsumoto. 2013. Topic models + word alignment = a flexible framework for extracting bilingual dictionary from comparable corpus. In Proceedings of the 17th Conference on Computational Natural Language Learning (CoNLL\u201913). 212."},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the International Workshop on the Future of the Dictionary.","author":"Magnini Bernardo","year":"1994","unstructured":"Bernardo Magnini , Carlo Strapparava , Fabio Ciravegna , and Emanuele Pianta . 1994 . Multilingual lexical knowledge bases: Applied WordNet prospects . In Proceedings of the International Workshop on the Future of the Dictionary. Bernardo Magnini, Carlo Strapparava, Fabio Ciravegna, and Emanuele Pianta. 1994. Multilingual lexical knowledge bases: Applied WordNet prospects. In Proceedings of the International Workshop on the Future of the Dictionary."},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the 47th Meeting of the Association for Computational Linguistics (ACL\u201909)","author":"Soderland Stephen","year":"2009","unstructured":"Mausam, Stephen Soderland , Oren Etzioni , Daniel S. Weld , Michael Skinner , and Jeff Bilmes . 2009 . Compiling a massive, multilingual dictionary via probabilistic inference . In Proceedings of the 47th Meeting of the Association for Computational Linguistics (ACL\u201909) . Mausam, Stephen Soderland, Oren Etzioni, Daniel S. Weld, Michael Skinner, and Jeff Bilmes. 2009.Compiling a massive, multilingual dictionary via probabilistic inference. In Proceedings of the 47th Meeting of the Association for Computational Linguistics (ACL\u201909)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.5555\/1699571.1699627"},{"key":"e_1_2_1_35_1","unstructured":"Thomas Minka. 2000. Estimating a Dirichlet distribution. Microsoft Research.  Thomas Minka. 2000. Estimating a Dirichlet distribution. Microsoft Research."},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-HLT\u201911)","author":"Neubig Graham","year":"2011","unstructured":"Graham Neubig , Yosuke Nakata , and Shinsuke Mori . 2011 . Pointwise prediction for robust, adaptable Japanese morphological analysis . In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-HLT\u201911) . 529--533. Graham Neubig, Yosuke Nakata, and Shinsuke Mori. 2011. Pointwise prediction for robust, adaptable Japanese morphological analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-HLT\u201911). 529--533."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526904"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1162\/089120103321337421"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/1620853.1620914"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.3115\/981658.981709"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.3115\/1072133.1072185"},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the Seminaire Papillon.","author":"Sadat Fatiha","year":"2002","unstructured":"Fatiha Sadat , Herve Dejean , and Eric Gaussier . 2002 . A combination of models for bilingual lexicon extraction from comparable corpora . In Proceedings of the Seminaire Papillon. Fatiha Sadat, Herve Dejean, and Eric Gaussier. 2002. A combination of models for bilingual lexicon extraction from comparable corpora. In Proceedings of the Seminaire Papillon."},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL\u201912)","author":"Tamura Akihiro","year":"2012","unstructured":"Akihiro Tamura , Taro Watanabe , and Eiichiro Sumita . 2012 . Bilingual lexicon extraction from comparable corpora using label propagation . In Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL\u201912) . 24--36. Akihiro Tamura, Taro Watanabe, and Eiichiro Sumita. 2012. Bilingual lexicon extraction from comparable corpora using label propagation. In Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL\u201912). 24--36."},{"key":"e_1_2_1_45_1","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems (NIPS\u201906)","author":"Teh Yee W.","year":"2006","unstructured":"Yee W. Teh , David Newman , and Max Welling . 2006 . A collapsed variational Bayesian inference algorithm for latent Dirichlet allocation . In Proceedings of the International Conference on Neural Information Processing Systems (NIPS\u201906) . 1353--1360. Yee W. Teh, David Newman, and Max Welling. 2006. A collapsed variational Bayesian inference algorithm for latent Dirichlet allocation. In Proceedings of the International Conference on Neural Information Processing Systems (NIPS\u201906). 1353--1360."},{"key":"e_1_2_1_46_1","volume-title":"Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (ACL\u201912)","author":"Vaswani Ashish","year":"2012","unstructured":"Ashish Vaswani , Liang Huang , and David Chiang . 2012 . Smaller alignment models for better translations: Unsupervised word alignment with the l 0-norm . In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (ACL\u201912) . 311--319. Ashish Vaswani, Liang Huang, and David Chiang. 2012. Smaller alignment models for better translations: Unsupervised word alignment with the l 0-norm. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (ACL\u201912). 311--319."},{"key":"e_1_2_1_47_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201913)","author":"Volkova Svitlana","year":"2013","unstructured":"Svitlana Volkova , Theresa Wilson , and David Yarowsky . 2013 . Exploring demographic language variations to improve multilingual sentiment analysis in social media . In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201913) . 18--21. Svitlana Volkova, Theresa Wilson, and David Yarowsky. 2013. Exploring demographic language variations to improve multilingual sentiment analysis in social media. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201913). 18--21."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.5555\/2002736.2002832"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.5555\/1687878.1687902"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.5555\/1858681.1858796"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2699939","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2699939","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T06:16:59Z","timestamp":1750227419000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2699939"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,6,12]]},"references-count":48,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2015,6,12]]}},"alternative-id":["10.1145\/2699939"],"URL":"https:\/\/doi.org\/10.1145\/2699939","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2015,6,12]]},"assertion":[{"value":"2014-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-06-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}