{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,1]],"date-time":"2026-02-01T00:03:15Z","timestamp":1769904195388,"version":"3.49.0"},"reference-count":48,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2020,10,6]],"date-time":"2020-10-06T00:00:00Z","timestamp":1601942400000},"content-version":"vor","delay-in-days":554,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>We propose a novel geometric approach for learning bilingual mappings given monolingual embeddings and a bilingual dictionary. Our approach decouples the source-to-target language transformation into (a) language-specific rotations on the original embeddings to align them in a common, latent space, and (b) a language-independent similarity metric in this common space to better model the similarity between the embeddings. Overall, we pose the bilingual mapping problem as a classification problem on smooth Riemannian manifolds. Empirically, our approach outperforms previous approaches on the bilingual lexicon induction and cross-lingual word similarity tasks.<\/jats:p>\n               <jats:p>We next generalize our framework to represent multiple languages in a common latent space. Language-specific rotations for all the languages and a common similarity metric in the latent space are learned jointly from bilingual dictionaries for multiple language pairs. We illustrate the effectiveness of joint learning for multiple languages in an indirect word translation setting.<\/jats:p>","DOI":"10.1162\/tacl_a_00257","type":"journal-article","created":{"date-parts":[[2019,4,16]],"date-time":"2019-04-16T19:33:41Z","timestamp":1555443221000},"page":"107-120","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":18,"title":["Learning Multilingual Word Embeddings in Latent Metric Space: A Geometric Approach"],"prefix":"10.1162","volume":"7","author":[{"given":"Pratik","family":"Jawanpuria","sequence":"first","affiliation":[{"name":"Microsoft, India. pratik.jawanpuria@microsoft.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arjun","family":"Balgovind","sequence":"additional","affiliation":[{"name":"IIT Madras, India. barjun@cse.iitm.ac.in"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anoop","family":"Kunchukuttan","sequence":"additional","affiliation":[{"name":"Microsoft, India. ankunchu@microsoft.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bamdev","family":"Mishra","sequence":"additional","affiliation":[{"name":"Microsoft, India. bamdevm@microsoft.com"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2019,4,1]]},"reference":[{"key":"2021060823283543600_bib1","unstructured":"Pierre-Antoine\n              Absil\n            , RobertMahony, and RodolpheSepulchre. \n          2008. Optimization Algorithms on Matrix Manifolds. Princeton University Press, Princeton, NJ."},{"key":"2021060823283543600_bib2","unstructured":"Waleed\n              Ammar\n            , GeorgeMulcaire, YuliaTsvetkov, GuillaumeLample, ChrisDyer, and Noah A.Smith. \n          2016. Massively multilingual word embeddings. Technical report, arXiv preprint arXiv:1602.01925."},{"key":"2021060823283543600_bib3","doi-asserted-by":"crossref","unstructured":"Mikel\n              Artetxe\n            , GorkaLabaka, and EnekoAgirre. \n          2016. Learning principled bilingual mappings of word embeddings while preserving monolingual invariance. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 2289\u20132294.","DOI":"10.18653\/v1\/D16-1250"},{"key":"2021060823283543600_bib4","doi-asserted-by":"crossref","unstructured":"Mikel\n              Artetxe\n            , GorkaLabaka, and EnekoAgirre. \n          2017. Learning bilingual word embeddings with (almost) no bilingual data. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 451\u2013462.","DOI":"10.18653\/v1\/P17-1042"},{"key":"2021060823283543600_bib5","unstructured":"Mikel\n              Artetxe\n            , GorkaLabaka, and EnekoAgirre. \n          2018a. Generalizing and improving bilingual word embedding mappings with a multi-step framework of linear transformations. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 5012\u20135019."},{"key":"2021060823283543600_bib6","doi-asserted-by":"crossref","unstructured":"Mikel\n              Artetxe\n            , GorkaLabaka, and EnekoAgirre. \n          2018b. A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 789\u2013798. https:\/\/github.com.artetxem\/vecmap.","DOI":"10.18653\/v1\/P18-1073"},{"key":"2021060823283543600_bib7","doi-asserted-by":"crossref","unstructured":"Antonio Valerio Miceli\n              Barone\n            \n          . 2016. Towards cross-lingual distributed representations without parallel text trained with adversarial autoencoders. In Proceedings of the 1st Workshop on Representation Learning for NLP.","DOI":"10.18653\/v1\/W16-1614"},{"key":"2021060823283543600_bib8","doi-asserted-by":"crossref","unstructured":"Silv\u00e8re\n              Bonnabel\n             and RodolpheSepulchre. \n          2010. Riemannian metric and geometric mean for positive semidefinite matrices of fixed rank. SIAM Journal on Matrix Analysis and Applications, 31(3):1055\u20131070.","DOI":"10.1137\/080731347"},{"key":"2021060823283543600_bib9","unstructured":"Nicolas\n              Boumal\n            , BamdevMishra, Pierre-AntoineAbsil, and RodolpheSepulchre. \n          2014. Manopt, a Matlab toolbox for optimization on manifolds. Journal of Machine Learning Research, 15(Apr):1455\u20131459."},{"key":"2021060823283543600_bib10","doi-asserted-by":"crossref","unstructured":"Jos\u00e9\n              Camacho-Collados\n            , Mohammad TaherPilehvar, NigelCollier, and RobertoNavigli. \n          2017. SemEval-2017 Task 2: Multilingual and Cross-lingual Semantic Word Similarity. In Proceedings of the 11th International Workshop on Semantic Evaluation.","DOI":"10.18653\/v1\/S17-2002"},{"key":"2021060823283543600_bib11","doi-asserted-by":"crossref","unstructured":"Jos\u00e9\n              Camacho-Collados\n            , Mohammad TaherPilehvar, and RobertoNavigli. \n          2016. Nasari: Integrating explicit knowledge and corpus statistics for a multilingual representation of concepts and entities. Artificial Intelligence, 240:36\u201364.","DOI":"10.1016\/j.artint.2016.07.005"},{"key":"2021060823283543600_bib12","unstructured":"Sarath\n              Chandar\n            , StanislasLauly, HugoLarochelle, MiteshKhapra, BalaramanRavindran, Vikas CRaykar, and AmritaSaha. \n          2014. An autoencoder approach to learning bilingual word representations. In Proceedings of the Advances in Neural Information Processing Systems, pages 1853\u20131861."},{"key":"2021060823283543600_bib13","doi-asserted-by":"crossref","unstructured":"Xilun\n              Chen\n             and ClaireCardie. \n          2018. Unsupervise multilingual word embeddings. In Proceedings of the Conference on Empirical Methods in Natural Language Processing.","DOI":"10.18653\/v1\/D18-1024"},{"key":"2021060823283543600_bib14","unstructured":"Alexis\n              Conneau\n            , GuillaumeLample, Marc\u2019AurelioRanzato, LudovicDenoyer, and Herv\u00e9J\u00e9gou. \n          2018. Word translation without parallel data. In Proceedings of the International Conference on Learning Representations. https:\/\/github.com\/facebookresearch\/MUSE."},{"key":"2021060823283543600_bib15","unstructured":"Georgiana\n              Dinu\n             and MarcoBaroni. \n          2015. Improving zero-shot learning by mitigating the hubness problem. In Workshop track of International Conference on Learning Representations."},{"key":"2021060823283543600_bib16","doi-asserted-by":"crossref","unstructured":"Yerai\n              Doval\n            , JoseCamacho-Collados, LuisEspinosa-Anke, and StevenSchockaert. \n          2018. Improving cross-lingual word embeddings by meeting in the middle. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. https:\/\/github.com\/yeraidam\/meemi","DOI":"10.18653\/v1\/D18-1027"},{"key":"2021060823283543600_bib17","doi-asserted-by":"crossref","unstructured":"Long\n              Duong\n            , HiroshiKanayama, TengfeiMa, StevenBird, and TrevorCohn. \n          2017. Multilingual training of crosslingual word embeddings. In Proceedings of the Conference of the European Chapter of the Association for Computational Linguistics, pages 894\u2013904.","DOI":"10.18653\/v1\/E17-1084"},{"key":"2021060823283543600_bib18","doi-asserted-by":"crossref","unstructured":"Alan\n              Edelman, \n            \n              Tom\u00e1s A.\n              Arias\n            , and Steven T.Smith. \n          1998. The geometry of algorithms with orthogonality constraints. SIAM Journal on Matrix Analysis and Applications, 20(2): 303\u2013353.","DOI":"10.1137\/S0895479895290954"},{"key":"2021060823283543600_bib19","doi-asserted-by":"crossref","unstructured":"Manaal\n              Faruqui and \n            \n              Chris\n              Dyer. \n          \n          2014. Improving vector space word representations using multilingual correlation. In Proceedings of the Conference of the European Chapter of the Association for Computational Linguistics, pages 462\u2013471.","DOI":"10.3115\/v1\/E14-1049"},{"key":"2021060823283543600_bib20","unstructured":"Stephan\n              Gouws, \n            \n              Yoshua\n              Bengio\n            , and GregCorrado. \n          2015. Bilbowa: Fast bilingual distributed representations without word alignments. In Proceedings of the International Conference on Machine Learning, pages 748\u2013756."},{"key":"2021060823283543600_bib21","doi-asserted-by":"crossref","unstructured":"John C.\n              Gower\n            \n          . 1975. Generalized procrustes analysis. Psychometrika, 40(1):33\u201351.","DOI":"10.1007\/BF02291478"},{"key":"2021060823283543600_bib22","unstructured":"Edouard\n              Grave, \n            \n              Armand\n              Joulin, and \n            \n              Quentin\n              Berthet. \n          \n          2018. Unsupervised alignment of embeddings with Wasserstein Procrustes. Technical report, arXiv preprint arXiv:1805.11222."},{"key":"2021060823283543600_bib23","unstructured":"Jiatao\n              Gu\n            , HanyHassan, JacobDevlin, and Victor OKLi. \n          2018. Universal neural machine translation for extremely low resource languages. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies."},{"key":"2021060823283543600_bib24","unstructured":"Mehrtash\n              Harandi, \n            \n              Mathieu\n              Salzmann\n            , and RichardHartley. \n          2017. Joint dimensionality reduction and metric learning: A geometric take. In Proceedings of the International Conference on Machine Learning."},{"key":"2021060823283543600_bib25","doi-asserted-by":"crossref","unstructured":"Karl Moritz\n              Hermann\n             and PhilBlunsom. \n          2014. Multilingual models for compositional distributed semantics. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 58\u201368.","DOI":"10.3115\/v1\/P14-1006"},{"key":"2021060823283543600_bib26","doi-asserted-by":"crossref","unstructured":"Yedid\n              Hoshen\n             and LiorWolf. \n          2018. Non-adversarial unsupervised word translation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 469\u2013478.","DOI":"10.18653\/v1\/D18-1043"},{"key":"2021060823283543600_bib27","doi-asserted-by":"crossref","unstructured":"Kejun\n              Huang, \n            \n              Matt\n              Gardner, \n            \n              Evangelos E.\n              Papalexakis, \n            \n              Christos\n              Faloutsos, \n            \n              Nikos D.\n              Sidiropoulos, \n            \n              Tom M.\n              Mitchell, \n            \n              Partha Pratim\n              Talukdar\n            , and XiaoFu. \n          2015. Translation invariant word embeddings. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1084\u20131088.","DOI":"10.18653\/v1\/D15-1127"},{"key":"2021060823283543600_bib28","unstructured":"Wen\n              Huang, \n            \n              Pierre-Antoine\n              Absil, \n            \n              Kyle A.\n              Gallivan\n            , and PaulHand. \n          2016. ROPTLIB: An object-oriented C++ library for optimization on Riemannian manifolds. Technical report, FSU16-14.v2, Florida State University."},{"key":"2021060823283543600_bib29","doi-asserted-by":"crossref","unstructured":"Armand\n              Joulin, \n            \n              Piotr\n              Bojanowski, \n            \n              Tomas\n              Mikolov, \n            \n              Edouard\n              Grave\n            , and Herv\u00e8J\u00e8gou. \n          2018. Loss in translation: Learning bilingual word mapping with a retrieval criterion. In Proceedings of the Conference on Empirical Methods in Natural Language Processing.","DOI":"10.18653\/v1\/D18-1330"},{"key":"2021060823283543600_bib30","doi-asserted-by":"crossref","unstructured":"Y.\n              Kementchedjhieva, \n            \n              S.\n              Ruder, \n            \n              R.\n              Cotterell\n            , and A.S\u00f8gaard. \n          2018. Generalizing Procrustes analysis for better bilingual dictionary induction. In Proceedings of the SIGNLL Conference on Computational Natural Language Learning.","DOI":"10.18653\/v1\/K18-1021"},{"key":"2021060823283543600_bib31","unstructured":"Alexandre\n              Klementiev, \n            \n              Ivan\n              Titov\n            , and BinodBhattarai. \n          2012. Inducing crosslingual distributed representations of words. In Proceedings of the International Conference on Computational Linguistics: Technical Papers, pages 1459\u20131474."},{"key":"2021060823283543600_bib32","doi-asserted-by":"crossref","unstructured":"Angeliki\n              Lazaridou, \n            \n              Georgiana\n              Dinu\n            , and MarcoBaroni. \n          2015. Hubness and pollution: Delving into cross-space mapping for zero-shot learning. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 1959\u20131970.","DOI":"10.3115\/v1\/P15-1027"},{"key":"2021060823283543600_bib33","unstructured":"John M.\n              Lee\n            \n          . 2003. Introduction to Smooth Manifolds, second edition, Springer-Verlag, New York."},{"key":"2021060823283543600_bib34","unstructured":"Gilles\n              Meyer, \n            \n              Silv\u00e8re\n              Bonnabel\n            , and RodolpheSepulchre. \n          2011. Linear regression under fixed-rank constraints: A Riemannian approach. In Proceedings of the International Conference on Machine Learning, pages 545\u2013552."},{"key":"2021060823283543600_bib35","unstructured":"Tomas\n              Mikolov\n            , KaiChen, GregCorrado, and JeffreyDean. \n          2013a. Efficient estimation of word representations in vector space. Technical report, arXiv preprint arXiv:1301.3781."},{"key":"2021060823283543600_bib36","unstructured":"Tomas\n              Mikolov\n            , Quoc VLe, and IlyaSutskever. \n          2013b. Exploiting similarities among languages for machine translation. Technical report, arXiv preprint arXiv:1309.4168."},{"key":"2021060823283543600_bib37","doi-asserted-by":"crossref","unstructured":"Bamdev\n              Mishra\n            , GillesMeyer, Silv\u00e8reBonnabel, and RodolpheSepulchre. \n          2014. Fixed-rank matrix factorizations and Riemannian low-rank optimization. Computational Statistics, 29(3):591\u2013621.","DOI":"10.1007\/s00180-013-0464-z"},{"key":"2021060823283543600_bib38","doi-asserted-by":"crossref","unstructured":"Hiroyuki\n              Sato\n             and ToshihiroIwai. \n          2013. A new, globally convergent Riemannian conjugate gradient method. Optimization: A Journal of Mathematical Programming and Operations Research, 64(4):1011\u20131031.","DOI":"10.1080\/02331934.2013.836650"},{"key":"2021060823283543600_bib39","doi-asserted-by":"crossref","unstructured":"Peter H.\n              Sch\u00f6nemann\n            \n          . 1966. A generalized solution of the orthogonal Procrustes problem. Psychometrika, 31(1):1\u201310.","DOI":"10.1007\/BF02289451"},{"key":"2021060823283543600_bib40","unstructured":"Samuel L.\n              Smith\n            , David H. P.Turban, StevenHamblin, and Nils Y.Hammerla. \n          2017a. Aligning the fastText vectors of 78 languages. URL: https:\/\/github.com\/Babylonpartners\/fastText_multilingual."},{"key":"2021060823283543600_bib41","unstructured":"Samuel L.\n              Smith\n            , David H. P.Turban, StevenHamblin, and Nils Y.Hammerla. \n          2017b. Offline bilingual word vectors, orthogonal transformations and the inverted softmax. In Proceedings of the International Conference on Learning Representations."},{"key":"2021060823283543600_bib42","doi-asserted-by":"crossref","unstructured":"Anders\n              S\u00f8gaard\n            , SebastianRuder, and IvanVuli\u0107. \n          2018. On the limitations of unsupervised bilingual dictionary induction. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, pages 778\u2013788.","DOI":"10.18653\/v1\/P18-1072"},{"key":"2021060823283543600_bib43","doi-asserted-by":"crossref","unstructured":"Robert\n              Speer\n             and JoannaLowry-Duda. \n          2017. ConceptNet at SemEval-2017 Task 2: Extending word embeddings with multilingual relational knowledge. In Proceedings of the 11th International Workshop on Semantic Evaluations.","DOI":"10.18653\/v1\/S17-2008"},{"key":"2021060823283543600_bib44","unstructured":"James\n              Townsend\n            , NiklasKoep, and SebastianWeichwald. \n          2016. Pymanopt: A python toolbox for optimization on manifolds using automatic differentiation. Journal of Machine Learning Research, 17(137):1\u20135. URL: https:\/\/pymanopt.github.io."},{"key":"2021060823283543600_bib45","doi-asserted-by":"crossref","unstructured":"Chao\n              Xing\n            , DongWang, ChaoLiu, and YiyeLin. \n          2015. Normalized word embedding and orthogonal transform for bilingual word translation. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1006\u20131011.","DOI":"10.3115\/v1\/N15-1104"},{"key":"2021060823283543600_bib46","doi-asserted-by":"crossref","unstructured":"Meng\n              Zhang\n            , YangLiu, HuanboLuan, and MaosongSun. \n          2017a. Adversarial training for unsupervised bilingual lexicon induction. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 1959\u20131970.","DOI":"10.18653\/v1\/P17-1179"},{"key":"2021060823283543600_bib47","doi-asserted-by":"crossref","unstructured":"Meng\n              Zhang\n            , YangLiu, HuanboLuan, and MaosongSun. \n          2017b. Earth mover\u2019s distance minimization for unsupervised bilingual lexicon induction. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1934\u20131945.","DOI":"10.18653\/v1\/D17-1207"},{"key":"2021060823283543600_bib48","doi-asserted-by":"crossref","unstructured":"Huiwei\n              Zhou\n            , LongChen, FulinShi, and DegenHuang. \n          2015. Learning bilingual sentiment word embeddings for cross-language sentiment classification. In Proceedings of the Annual Meeting of the Association for Computational Linguistics and the International Joint Conference on Natural Language Processing, pages 430\u2013440.","DOI":"10.3115\/v1\/P15-1042"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00257\/1923013\/tacl_a_00257.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"http:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00257\/1923013\/tacl_a_00257.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,6,9]],"date-time":"2021-06-09T03:23:39Z","timestamp":1623209019000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00257\/43509\/Learning-Multilingual-Word-Embeddings-in-Latent"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,4,1]]},"references-count":48,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00257","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2019,4]]},"published":{"date-parts":[[2019,4,1]]}}}