{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:24:40Z","timestamp":1760239480272,"version":"build-2065373602"},"reference-count":34,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2020,11,27]],"date-time":"2020-11-27T00:00:00Z","timestamp":1606435200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Recently, the pretraining of models has been successfully applied to unsupervised and semi-supervised neural machine translation. A cross-lingual language model uses a pretrained masked language model to initialize the encoder and decoder of the translation model, which greatly improves the translation quality. However, because of a mismatch in the number of layers, the pretrained model can only initialize part of the decoder\u2019s parameters. In this paper, we use a layer-wise coordination transformer and a consistent pretraining translation transformer instead of a vanilla transformer as the translation model. The former has only an encoder, and the latter has an encoder and a decoder, but the encoder and decoder have exactly the same parameters. Both models can guarantee that all parameters in the translation model can be initialized by the pretrained model. Experiments on the Chinese\u2013English and English\u2013German datasets show that compared with the vanilla transformer baseline, our models achieve better performance with fewer parameters when the parallel corpus is small.<\/jats:p>","DOI":"10.3390\/fi12120215","type":"journal-article","created":{"date-parts":[[2020,11,28]],"date-time":"2020-11-28T03:51:16Z","timestamp":1606535476000},"page":"215","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Keeping Models Consistent between Pretraining and Translation for Low-Resource Neural Machine Translation"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8932-923X","authenticated-orcid":false,"given":"Wenbo","family":"Zhang","sequence":"first","affiliation":[{"name":"Xinjiang Technical Institute of Physics &amp; Chemistry, Chinese Academy of Sciences, Urumqi 830011, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"Xinjiang Laboratory of Minority Speech and Language Information Processing, Urumqi 830011, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiao","family":"Li","sequence":"additional","affiliation":[{"name":"Xinjiang Technical Institute of Physics &amp; Chemistry, Chinese Academy of Sciences, Urumqi 830011, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"Xinjiang Laboratory of Minority Speech and Language Information Processing, Urumqi 830011, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yating","family":"Yang","sequence":"additional","affiliation":[{"name":"Xinjiang Technical Institute of Physics &amp; Chemistry, Chinese Academy of Sciences, Urumqi 830011, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"Xinjiang Laboratory of Minority Speech and Language Information Processing, Urumqi 830011, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rui","family":"Dong","sequence":"additional","affiliation":[{"name":"Xinjiang Technical Institute of Physics &amp; Chemistry, Chinese Academy of Sciences, Urumqi 830011, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"Xinjiang Laboratory of Minority Speech and Language Information Processing, Urumqi 830011, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gongxu","family":"Luo","sequence":"additional","affiliation":[{"name":"Xinjiang Technical Institute of Physics &amp; Chemistry, Chinese Academy of Sciences, Urumqi 830011, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"Xinjiang Laboratory of Minority Speech and Language Information Processing, Urumqi 830011, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,11,27]]},"reference":[{"key":"ref_1","unstructured":"Sutskever, I., Vinyals, O., and Le, Q.V. (2014, January 8\u201313). Sequence to sequence learning with neural networks. Proceedings of the Advances in Neural Information Processing Systems, Montr\u00e9al, QC, Canada."},{"key":"ref_2","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural Machine Translation by Jointly Learning to Align and Translate. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Luong, M.T., Pham, H., and Manning, C.D. (2015, January 17\u201321). Effective Approaches to Attention-based Neural Machine Translation. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lisbon, Portugal.","DOI":"10.18653\/v1\/D15-1166"},{"key":"ref_4","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., and Kaiser, L. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_5","unstructured":"Hassan, H., Aue, A., Chen, C., Chowdhary, V., Clark, J., Federmann, C., Huang, X., Junczys-Dowmunt, M., Lewis, W., and Li, M. (2018). Achieving human parity on automatic chinese to english news translation. arXiv."},{"key":"ref_6","unstructured":"Wu, Y., Schuster, M., Chen, Z., Le, Q.V., Macherey, M.N.W., Krikun, J.M., Cao, Y., Cao, Q., Macherey, K., and Klinger, J. (2016). Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. arXiv."},{"key":"ref_7","unstructured":"Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. (2017, January 6\u201311). Convolutional sequence to sequence learning. Proceedings of the 34th International Conference on Machine Learning-Volume 70, Sydney, Australia."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Koehn, P., and Knowles, R. (2017). Six challenges for neural machine translation. arXiv.","DOI":"10.18653\/v1\/W17-3204"},{"key":"ref_9","unstructured":"Nguyen, T.Q., and Chiang, D. (2017). Transfer Learning across Low-Resource, Related Languages for Neural Machine Translation. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zoph, B., Yuret, D., May, J., and Knight, K. (2016). Transfer learning for low-resource neural machine translation. arXiv.","DOI":"10.18653\/v1\/D16-1163"},{"key":"ref_11","unstructured":"Cheng, Y., Xu, W., He, Z., He, W., Wu, H., Sun, M., and Liu, Y. (2020, November 01). Semi-Supervised Learning for Neural Machine Translation. Available online: https:\/\/link.springer.com\/chapter\/10.1007\/978-981-32-9748-7_3."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Gu, J., Wang, Y., Chen, Y., Cho, K., and Li, V.O.K. (November, January 31). Meta-Learning for Low-Resource Neural Machine Translation. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.","DOI":"10.18653\/v1\/D18-1398"},{"key":"ref_13","unstructured":"Artetxe, M., Labaka, G., and Agirre, E. (August, January 30). Learning bilingual word embeddings with (almost) no bilingual data. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vancouver, BC, Canada."},{"key":"ref_14","first-page":"1137","article-title":"A neural probabilistic language model","volume":"3","author":"Bengio","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_15","unstructured":"Sennrich, R., Haddow, B., and Birch, A. (2020, November 01). Improving Neural Machine Translation Models with Monolingual Data. Available online: https:\/\/arxiv.org\/abs\/1511.06709."},{"key":"ref_16","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_17","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"key":"ref_18","unstructured":"Lample, G., and Conneau, A. (2019). Crosslingual language model pretraining. arXiv."},{"key":"ref_19","unstructured":"Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.Y. (2019). Mass: Masked sequence to sequence pre-training for language generation. arXiv."},{"key":"ref_20","unstructured":"He, T., Tan, X., Xia, Y., He, D., Qin, T., Chen, Z., and Liu, T.Y. (2018, January 3\u20138). Layer-wise coordination between encoder and decoder for neural machine translation. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_21","unstructured":"Koehn, P., Hoang, H., Birch, A., Burch, C.C., Federico, M., Bertoldi, N., Corwan, B., Shen, W., Moran, C., and Zens, R. (2020, November 01). Moses: Open Source Toolkit for Statistical Machine Translation. Available online: https:\/\/www.aclweb.org\/anthology\/P07-2.pdf."},{"key":"ref_22","unstructured":"Gulcehre, C., Firat, O., Xu, K., Cho, K., Barrault, L., Lin, H., Bougares, F., Schwenk, H., and Bengio, Y. (2015). On using monolingual corpora in neural machine translation. arXiv."},{"key":"ref_23","unstructured":"Skorokhodov, I., Rykachevskiy, A., Emelyanenko, D., Slotin, S., and Ponkratov, A. (2020, November 01). Semi-Supervised Neural Machine Translation with Language Models. Available online: https:\/\/www.aclweb.org\/anthology\/W18-2205.pdf."},{"key":"ref_24","unstructured":"Xia, Y., He, D., Qin, T., Wang, L., Yu, N., Liu, T.Y., and Ma, W.Y. (2018, January 6\u201312). Dual learning for machine translation. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_25","unstructured":"Wang, Y., Xia, Y., Zhao, L., Bian, J., Qin, T., and Liu, G. (2020, November 01). Dual Transfer Learning for Neural Machine Translation with Marginal Distribution Regularization. Available online: https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2017\/11\/17041-72820-1-SM.pdf."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Liu, S., Li, M., and Chen, E. (2018). Joint Training for Neural Machine Translation Models with Monolingual Data. arXiv.","DOI":"10.1609\/aaai.v32i1.11248"},{"key":"ref_27","unstructured":"Fadaee, M., Bisazza, A., and Monz, C. (2020, November 01). Data Augmentation for Low-Resource Neural Machine Translation. Available online: https:\/\/arxiv.org\/abs\/1705.00440."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zhang, J., and Zong, C. (2016, January 1\u20135). Exploiting source-side monolingual data in neural machine translation. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, TX, USA.","DOI":"10.18653\/v1\/D16-1160"},{"key":"ref_29","unstructured":"Edunov, S., Ott, M., Auli, M., and Grangier, D. (2020, November 01). Understanding Back-Translation at Scale. Available online: https:\/\/arxiv.org\/abs\/1808.09381."},{"key":"ref_30","unstructured":"Poncelas, A., Shterionov, D., Way, A., Wenniger, G.M.B., and Passban, P. (2018). Investigating backtranslation in neural machine translation. arXiv."},{"key":"ref_31","unstructured":"Sennrich, R., Haddow, B., and Birch, A. (2020, November 01). Neural Machine Translation of Rare Words with Subword Units. Available online: https:\/\/arxiv.org\/abs\/1508.07909."},{"key":"ref_32","unstructured":"Diederik, K., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_33","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Nitish","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_34","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2020, November 01). BLEU: A Method for Automatic Evaluation of Machine Translation. Available online: https:\/\/www.aclweb.org\/anthology\/P02-1040.pdf."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/12\/12\/215\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:38:40Z","timestamp":1760179120000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/12\/12\/215"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,27]]},"references-count":34,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2020,12]]}},"alternative-id":["fi12120215"],"URL":"https:\/\/doi.org\/10.3390\/fi12120215","relation":{},"ISSN":["1999-5903"],"issn-type":[{"type":"electronic","value":"1999-5903"}],"subject":[],"published":{"date-parts":[[2020,11,27]]}}}