{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T09:38:14Z","timestamp":1783503494350,"version":"3.55.0"},"reference-count":62,"publisher":"MIT Press","issue":"3","license":[{"start":{"date-parts":[[2023,5,25]],"date-time":"2023-05-25T00:00:00Z","timestamp":1684972800000},"content-version":"vor","delay-in-days":144,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,9,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Large multilingual language models typically share their parameters across all languages, which enables cross-lingual task transfer, but learning can also be hindered when training updates from different languages are in conflict. In this article, we propose novel methods for using language-specific subnetworks, which control cross-lingual parameter sharing, to reduce conflicts and increase positive transfer during fine-tuning. We introduce dynamic subnetworks, which are jointly updated with the model, and we combine our methods with meta-learning, an established, but complementary, technique for improving cross-lingual transfer. Finally, we provide extensive analyses of how each of our methods affects the models.<\/jats:p>","DOI":"10.1162\/coli_a_00482","type":"journal-article","created":{"date-parts":[[2023,5,25]],"date-time":"2023-05-25T19:40:35Z","timestamp":1685043635000},"page":"613-641","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":9,"title":["Cross-Lingual Transfer with Language-Specific Subnetworks for Low-Resource Dependency Parsing"],"prefix":"10.1162","volume":"49","author":[{"given":"Rochelle","family":"Choenni","sequence":"first","affiliation":[{"name":"University of Amsterdam, The Institute for Logic, Language and Computation (ILLC). r.m.v.k.choenni@uva.nl"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dan","family":"Garrette","sequence":"additional","affiliation":[{"name":"Google Research. dhgarrette@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ekaterina","family":"Shutova","sequence":"additional","affiliation":[{"name":"University of Amsterdam, The Institute for Logic, Language and Computation (ILLC). e.shutova@uva.nl"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2023,9,1]]},"reference":[{"key":"2023111518560186500_bib1","doi-asserted-by":"publisher","first-page":"4762","DOI":"10.18653\/v1\/2021.findings-emnlp.410","article-title":"MAD-G: Multilingual adapter generation for efficient cross-lingual transfer","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"Ansell","year":"2021"},{"key":"2023111518560186500_bib2","article-title":"How to train your MAML","volume-title":"Seventh International Conference on Learning Representations","author":"Antoniou","year":"2019"},{"key":"2023111518560186500_bib3","first-page":"3874","article-title":"Massively multilingual neural machine translation in the wild: Findings and challenges","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Arivazhagan","year":"2019"},{"key":"2023111518560186500_bib4","first-page":"20852","article-title":"The generalization-stability tradeoff in neural network pruning","volume-title":"Advances in Neural Information Processing Systems","author":"Bartoldson","year":"2020"},{"key":"2023111518560186500_bib5","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1308.3432","article-title":"Estimating or propagating gradients through stochastic neurons for conditional computation","author":"Bengio","year":"2013","journal-title":"arXiv preprint arXiv:1308.3432"},{"key":"2023111518560186500_bib6","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2205.11758","article-title":"Analyzing the mono-and cross-lingual pretraining dynamics of multilingual language models","author":"Blevins","year":"2022","journal-title":"arXiv preprint arXiv:2205.11758"},{"key":"2023111518560186500_bib7","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2109.12683","article-title":"On the prunability of attention heads in multilingual BERT","author":"Budhraja","year":"2021","journal-title":"arXiv preprint arXiv:2109.12683"},{"key":"2023111518560186500_bib8","first-page":"15834","article-title":"The lottery ticket hypothesis for pre-trained BERT networks","volume-title":"Advances in Neural Information Processing Systems","author":"Chen","year":"2020"},{"key":"2023111518560186500_bib9","doi-asserted-by":"publisher","first-page":"5564","DOI":"10.18653\/v1\/2020.acl-main.493","article-title":"Finding universal grammatical relations in multilingual BERT","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Chi","year":"2020"},{"issue":"3","key":"2023111518560186500_bib10","doi-asserted-by":"publisher","first-page":"635","DOI":"10.1162\/coli_a_00444","article-title":"Investigating language relationships in multilingual sentence encoders through the lens of linguistic typology","volume":"48","author":"Choenni","year":"2022","journal-title":"Computational Linguistics"},{"key":"2023111518560186500_bib11","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2021.sigtyp-1.5","article-title":"Improving the performance of UDify with linguistic typology knowledge","volume-title":"Proceedings of the Third Workshop on Computational Typology and Multilingual NLP","author":"Choudhary","year":"2021"},{"key":"2023111518560186500_bib12","first-page":"1396","article-title":"On the shortest arborescence of a directed graph","volume":"14","author":"Chu","year":"1965","journal-title":"Scientica Sinica"},{"key":"2023111518560186500_bib13","doi-asserted-by":"publisher","first-page":"276","DOI":"10.18653\/v1\/W19-4828","article-title":"What does BERT Look at? An analysis of BERT\u2019s attention","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Clark","year":"2019"},{"key":"2023111518560186500_bib14","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of NAACL-HLT","author":"Devlin","year":"2019"},{"key":"2023111518560186500_bib15","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1611.01734","article-title":"Deep biaffine attention for neural dependency parsing","author":"Dozat","year":"2016","journal-title":"arXiv preprint arXiv:1611.01734"},{"key":"2023111518560186500_bib16","first-page":"1126","article-title":"Model-agnostic meta-learning for fast adaptation of deep networks","volume-title":"International Conference on Machine Learning","author":"Finn","year":"2017"},{"key":"2023111518560186500_bib17","article-title":"Discovering language-neutral sub-networks in multilingual language models","author":"Foroutan","year":"2022","journal-title":"arXiv preprint arXiv:2205.12672"},{"key":"2023111518560186500_bib18","article-title":"The lottery ticket hypothesis: Finding sparse, trainable neural networks","volume-title":"International Conference on Learning Representations","author":"Frankle","year":"2018"},{"key":"2023111518560186500_bib19","doi-asserted-by":"publisher","first-page":"4878","DOI":"10.18653\/v1\/2021.findings-acl.431","article-title":"Climbing the tower of treebanks: Improving low-resource dependency parsing via hierarchical source selection","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Glava\u0161","year":"2021"},{"key":"2023111518560186500_bib20","doi-asserted-by":"publisher","first-page":"3622","DOI":"10.18653\/v1\/D18-1398","article-title":"Meta-learning for low-resource neural machine translation","volume-title":"2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018","author":"Gu","year":"2020"},{"key":"2023111518560186500_bib21","first-page":"1135","article-title":"Learning both weights and connections for efficient neural network","volume-title":"Advances in Neural Information Processing Systems","author":"Han","year":"2015"},{"key":"2023111518560186500_bib22","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2210.05709","article-title":"Shapley head pruning: Identifying and removing interference in multilingual transformers","author":"Held","year":"2022","journal-title":"arXiv preprint arXiv:2210.05709"},{"key":"2023111518560186500_bib23","first-page":"351","article-title":"Domain specific sub-network for multi-domain neural machine translation","volume-title":"Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing","author":"Hendy","year":"2022"},{"key":"2023111518560186500_bib24","first-page":"2790","article-title":"Parameter-efficient transfer learning for NLP","volume-title":"International Conference on Machine Learning","author":"Houlsby","year":"2019"},{"key":"2023111518560186500_bib25","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1911.12246","article-title":"Do attention heads in BERT track syntactic dependencies?","author":"Htut","year":"2019","journal-title":"arXiv preprint arXiv:1911.12246"},{"key":"2023111518560186500_bib26","doi-asserted-by":"publisher","first-page":"4163","DOI":"10.18653\/v1\/2020.findings-emnlp.372","article-title":"TinyBERT: Distilling BERT for natural language understanding","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Jiao","year":"2020"},{"key":"2023111518560186500_bib27","article-title":"Adam: A method for stochastic optimization","volume-title":"International Conference on Learning Representations","author":"Kingma","year":"2015"},{"key":"2023111518560186500_bib28","doi-asserted-by":"publisher","first-page":"2779","DOI":"10.18653\/v1\/D19-1279","article-title":"75 languages, 1 model: Parsing universal dependencies universally","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Kondratyuk","year":"2019"},{"key":"2023111518560186500_bib29","article-title":"ALBERT: A lite BERT for self-supervised learning of language representations","volume-title":"International Conference on Learning Representations","author":"Lan","year":"2019"},{"key":"2023111518560186500_bib30","doi-asserted-by":"publisher","first-page":"8503","DOI":"10.18653\/v1\/2022.acl-long.582","article-title":"Meta-learning for fast cross-lingual adaptation in dependency parsing","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Langedijk","year":"2022"},{"key":"2023111518560186500_bib31","doi-asserted-by":"publisher","first-page":"4483","DOI":"10.18653\/v1\/2020.emnlp-main.363","article-title":"From zero to hero: On the limitations of zero-shot language transfer with multilingual transformers","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Lauscher","year":"2020"},{"key":"2023111518560186500_bib32","doi-asserted-by":"publisher","first-page":"817","DOI":"10.18653\/v1\/2021.acl-short.103","article-title":"Lightweight adapter tuning for multilingual speech translation","volume-title":"The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP 2021)","author":"Le","year":"2021"},{"key":"2023111518560186500_bib33","article-title":"Sequential reptile: Inter-task gradient alignment for multilingual learning","volume-title":"International Conference on Learning Representations","author":"Lee","year":"2021"},{"key":"2023111518560186500_bib34","article-title":"Pruning filters for efficient convNets","author":"Li","year":"2016","journal-title":"arXiv preprint arXiv:1608.08710"},{"key":"2023111518560186500_bib35","doi-asserted-by":"publisher","first-page":"1852","DOI":"10.18653\/v1\/2022.acl-long.130","article-title":"Probing structured pruning on multilingual pre-trained models: Settings, algorithms, and efficiency","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Li","year":"2022"},{"key":"2023111518560186500_bib36","doi-asserted-by":"publisher","first-page":"293","DOI":"10.18653\/v1\/2021.acl-long.25","article-title":"Learning language specific sub-network for multilingual machine translation","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Lin","year":"2021"},{"key":"2023111518560186500_bib37","doi-asserted-by":"publisher","first-page":"8","DOI":"10.18653\/v1\/E17-2002","article-title":"URIEL and Lang2Vec: Representing languages as typological, geographical, and phylogenetic vectors","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers","author":"Littell","year":"2017"},{"key":"2023111518560186500_bib38","doi-asserted-by":"publisher","first-page":"6882","DOI":"10.1109\/ICASSP43922.2022.9747671","article-title":"Language adaptive cross-lingual speech representation learning with sparse sharing sub-networks","volume-title":"ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Lu","year":"2022"},{"key":"2023111518560186500_bib39","first-page":"14037","article-title":"Are sixteen heads really better than one?","volume-title":"Advances in Neural Information Processing Systems","author":"Michel","year":"2019"},{"key":"2023111518560186500_bib40","first-page":"629","article-title":"Selective sharing for multilingual dependency parsing","volume-title":"Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Long Papers-Volume 1","author":"Naseem","year":"2012"},{"key":"2023111518560186500_bib41","first-page":"1659","article-title":"Universal dependencies v1: A multilingual treebank collection","volume-title":"Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916)","author":"Nivre","year":"2016"},{"key":"2023111518560186500_bib42","doi-asserted-by":"publisher","first-page":"4547","DOI":"10.18653\/v1\/2020.emnlp-main.368","article-title":"Zero-shot cross-lingual transfer with meta learning","volume-title":"the 2020 Conference on Empirical Methods in Natural Language Processing","author":"Nooralahzadeh","year":"2020"},{"key":"2023111518560186500_bib43","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2209.02982","article-title":"Improving the cross-lingual generalisation in visual question answering","author":"Nooralahzadeh","year":"2022","journal-title":"arXiv preprint arXiv:2209.02982"},{"key":"2023111518560186500_bib44","doi-asserted-by":"publisher","first-page":"7654","DOI":"10.18653\/v1\/2020.emnlp-main.617","article-title":"MAD-X: An adapter-based framework for multi-task cross-lingual transfer","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Pfeiffer","year":"2020"},{"key":"2023111518560186500_bib45","doi-asserted-by":"publisher","first-page":"4996","DOI":"10.18653\/v1\/P19-1493","article-title":"How multilingual is multilingual BERT?","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Pires","year":"2019"},{"key":"2023111518560186500_bib46","doi-asserted-by":"publisher","first-page":"3208","DOI":"10.18653\/v1\/2020.emnlp-main.259","article-title":"When BERT plays the lottery, all tickets are winning","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Prasanna","year":"2020"},{"key":"2023111518560186500_bib47","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1706.05098","article-title":"An overview of multi-task learning in deep neural networks","author":"Ruder","year":"2017","journal-title":"arXiv preprint arXiv:1706.05098"},{"key":"2023111518560186500_bib48","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1910.01108","article-title":"DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter","author":"Sanh","year":"2019","journal-title":"arXiv preprint arXiv:1910.01108"},{"key":"2023111518560186500_bib49","doi-asserted-by":"publisher","first-page":"8936","DOI":"10.1609\/aaai.v34i05.6424","article-title":"Learning sparse sharing architectures for multiple tasks","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Sun","year":"2020"},{"key":"2023111518560186500_bib50","doi-asserted-by":"publisher","first-page":"281","DOI":"10.18653\/v1\/D19-6132","article-title":"Zero-shot dependency parsing with pre-trained multilingual sentence representations","volume-title":"Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019)","author":"Tran","year":"2019"},{"key":"2023111518560186500_bib51","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2106.16171","article-title":"Revisiting the primacy of English in zero-shot cross-lingual transfer","author":"Turc","year":"2021","journal-title":"arXiv preprint arXiv:2106.16171"},{"key":"2023111518560186500_bib52","doi-asserted-by":"publisher","first-page":"2302","DOI":"10.18653\/v1\/2020.emnlp-main.180","article-title":"UDapter: Language adaptation for truly universal dependency parsing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"\u00dcst\u00fcn","year":"2020"},{"key":"2023111518560186500_bib53","doi-asserted-by":"publisher","first-page":"176","DOI":"10.18653\/v1\/2021.eacl-demos.22","article-title":"Massive choice, ample tasks (MaChAmp): A toolkit for multi-task learning in NLP","volume-title":"16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations (EACL)","author":"van der Goot","year":"2021"},{"key":"2023111518560186500_bib54","first-page":"6000","article-title":"Attention is all you need","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani","year":"2017"},{"key":"2023111518560186500_bib55","doi-asserted-by":"publisher","first-page":"5797","DOI":"10.18653\/v1\/P19-1580","article-title":"Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned","volume-title":"57th Annual Meeting of the Association for Computational Linguistics","author":"Voita","year":"2019"},{"key":"2023111518560186500_bib56","doi-asserted-by":"publisher","first-page":"4438","DOI":"10.18653\/v1\/2020.emnlp-main.359","article-title":"On negative interference in multilingual models: Findings and a meta-learning treatment","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Wang","year":"2020"},{"key":"2023111518560186500_bib57","doi-asserted-by":"publisher","first-page":"9274","DOI":"10.1609\/aaai.v34i05.6466","article-title":"Enhanced meta-learning for cross-lingual named entity recognition with minimal resources","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Wu","year":"2020"},{"key":"2023111518560186500_bib58","doi-asserted-by":"publisher","first-page":"530","DOI":"10.18653\/v1\/2022.acl-short.58","article-title":"S4-Tuning: A simple cross-lingual sub-network tuning method","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Xu","year":"2022"},{"key":"2023111518560186500_bib59","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2209.05735","article-title":"Learning ASR pathways: A Sparse multilingual ASR model","author":"Yang","year":"2022","journal-title":"arXiv preprint arXiv:2209.05735"},{"key":"2023111518560186500_bib60","first-page":"5824","article-title":"Gradient surgery for multi-task learning","author":"Yu","year":"2020"},{"key":"2023111518560186500_bib61","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/K18-2001","article-title":"CoNLL 2018 shared task: Multilingual parsing from raw text to universal dependencies","volume-title":"Proceedings of the CoNLL 2018 Shared Task: Multilingual parsing from raw text to universal dependencies","author":"Zeman","year":"2018"},{"key":"2023111518560186500_bib62","doi-asserted-by":"publisher","first-page":"36","DOI":"10.1016\/j.aiopen.2021.05.003","article-title":"Know what you don\u2019t need: Single-shot meta-pruning for attention heads","author":"Zhang","year":"2021"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/49\/3\/613\/2177420\/coli_a_00482.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/49\/3\/613\/2177420\/coli_a_00482.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,15]],"date-time":"2023-11-15T18:57:18Z","timestamp":1700074638000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/49\/3\/613\/116157\/Cross-Lingual-Transfer-with-Language-Specific"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"references-count":62,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,9,1]]},"published-print":{"date-parts":[[2023,9,1]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00482","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023]]},"published":{"date-parts":[[2023]]}}}