{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T02:00:47Z","timestamp":1781316047325,"version":"3.54.1"},"reference-count":45,"publisher":"China Science Publishing & Media Ltd.","issue":"2","content-domain":{"domain":["engine.scichina.com"],"crossmark-restriction":false},"short-container-title":["DI"],"published-print":{"date-parts":[[2026,6,1]]},"DOI":"10.3724\/2096-7004.di.2025.0109","type":"journal-article","created":{"date-parts":[[2025,9,22]],"date-time":"2025-09-22T07:05:53Z","timestamp":1758524753000},"page":"20250109","update-policy":"https:\/\/doi.org\/10.1360\/scp-crossmark-policy-page","source":"Crossref","is-referenced-by-count":0,"title":["CoRe-KD: Contrastive Multi-Teacher and Retention-Aware Knowledge Distillation for Continual Multilingual Neural Machine Translation"],"prefix":"10.3724","volume":"8","author":[{"given":"Lin","family":"Zhou","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Degen","family":"Huang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2026","published-online":{"date-parts":[[2025,9,22]]},"reference":[{"key":"null","unstructured":"J. Liu, K. Huang, J. Li, H. Liu, J. Su, and D. Huang, \u201cAdaptive token-level cross-lingual feature mixing for multilingual neural machine translation,\u201d in Proc. EMNLP, 2022, pp. 10097\u201310113."},{"key":"null","unstructured":"K. Huang, P. Li, J. Ma, T. Yao, and Y. Liu, \u201cKnowledge transfer in incremental learning for multilingual neural machine translation,\u201d in Proc. ACL, 2023, pp. 15286\u201315304."},{"key":"null","unstructured":"A. Ebrahimi and K. Kann, \u201cHow to adapt your pretrained multilingual model to 1600 languages,\u201d in Proc. ACL-IJCNLP, 2021, pp. 4555\u20134567."},{"key":"null","unstructured":"O. Feyisetan, B. Balle, T. Drake, and T. Diethe, \u201cPrivacy-and utility-preserving textual analysis via calibrated multivariate perturbations,\u201d in Proc. Int. Conf. Web Search Data Mining, 2020, pp. 178\u2013186."},{"key":"null","unstructured":"M. B. Ring, Continual Learning in Reinforcement Environments. Univ. Texas Austin, 1994."},{"key":"null","unstructured":"S. Gu, B. Hu, and Y. Feng, \u201cContinual learning of neural machine translation within low forgetting risk regions,\u201d in Proc. EMNLP, 2022, pp. 1707\u20131718."},{"key":"null","unstructured":"K. Huang, P. Li, J. Liu, M. Sun, and Y. Liu, \u201cLearn and consolidate: Continual adaptation for zero-shot and multilingual neural machine translation,\u201d in Proc. EMNLP, 2023, pp. 13938\u201313951."},{"key":"null","unstructured":"J. Wu, Y. Liu, and C. Zong, \u201cF-MALLOC: Feed-forward memory allocation for continual learning in neural machine translation,\u201d in Proc. NAACL-HLT, 2024, pp. 7180\u20137192."},{"key":"null","unstructured":"R. French, \u201cCatastrophic interference in connectionist networks: Can it be predicted, can it be prevented?\u201d in Adv. Neural Inf. Process. Syst., vol. 6, 1993, pp. 1176\u20131177."},{"key":"null","unstructured":"X. Garcia, N. Constant, A. Parikh, and O. Firat, \u201cTowards continual learning for multilingual machine translation via vocabulary substitution,\u201d in Proc. NAACL-HLT, 2021, pp. 1184\u20131192."},{"key":"null","unstructured":"Z. Liu, G. I. Winata, and P. Fung, \u201cContinual mixed-language pre-training for extremely low-resource neural machine translation,\u201d in Proc. ACL-IJCNLP Findings, 2021, pp. 2706\u20132718."},{"key":"null","unstructured":"H. Khayrallah, B. Thompson, K. Duh, and P. Koehn, \u201cRegularized training objective for continued training for domain adaptation in neural machine translation,\u201d in Proc. Workshop Neural Mach. Transl. Gener., 2018, pp. 36\u201344."},{"key":"null","unstructured":"B. Thompson, J. Gwinnup, H. Khayrallah, K. Duh, and P. Koehn, \u201cOvercoming catastrophic forgetting during domain adaptation of neural machine translation,\u201d in Proc. NAACL-HLT, 2019, pp. 2062\u20132068."},{"key":"null","unstructured":"A. Bapna and O. Firat, \u201cSimple, scalable adaptation for neural machine translation,\u201d in Proc. EMNLP-IJCNLP, 2019, pp. 1538\u20131548."},{"key":"null","unstructured":"Y. Cao, H.-R. Wei, B. Chen, and X. Wan, \u201cContinual learning for neural machine translation,\u201d in Proc. NAACL-HLT, 2021, pp. 3964\u20133974."},{"key":"null","unstructured":"P. Dakwale and C. Monz, \u201cFine-tuning for neural machine translation with limited degradation across in- and out-of-domain data,\u201d in Proc. Mach. Transl. Summit, 2017, pp. 156\u2013169."},{"key":"null","unstructured":"K. Kanwatchara, T. Horsuwan, P. Lertvittayakumjorn, B. Kijsirikul, and P. Vateekul, \u201cRational lamol: A rational-based lifelong learning framework,\u201d in Proc. ACL-IJCNLP, 2021, pp. 2942\u20132953."},{"key":"null","unstructured":"J. Liang, C. Zhao, M. Wang, X. Qiu, and L. Li, \u201cFinding sparse structures for domain specific neural machine translation,\u201d in Proc. AAAI Conf. Artif. Intell., vol. 35, 2021, pp. 13333\u201313342."},{"key":"null","unstructured":"W. Peng, C. Huang, T. Li, Y. Chen, and Q. Liu, \u201cDictionary-based data augmentation for cross-domain neural machine translation,\u201d arXiv preprint arXiv:2004.02577, 2020."},{"key":"null","unstructured":"C. de Masson D\u2019Autume, S. Ruder, L. Kong, and D. Yogatama, \u201cEpisodic memory in lifelong language learning,\u201d in Adv. Neural Inf. Process. Syst., vol. 32, 2019."},{"key":"null","unstructured":"F.-K. Sun, C.-H. Ho, and H.-Y. Lee, \u201cLAMOL: Language modeling for lifelong language learning,\u201d arXiv preprint arXiv:1909.03329, 2019."},{"key":"null","unstructured":"G. Hinton, O. Vinyals, and J. Dean, \u201cDistilling the knowledge in a neural network,\u201d arXiv preprint arXiv:1503.02531, 2015."},{"key":"null","unstructured":"T. Fukuda, M. Suzuki, G. Kurata, S. Thomas, J. Cui, and B. Ramabhadran, \u201cEfficient knowledge distillation from an ensemble of teachers,\u201d in Proc. Interspeech, 2017."},{"key":"null","unstructured":"J. Guo, Y. Liang, and J. Xu, \u201cContinual learning with confidence-based multi-teacher knowledge distillation for neural machine translation,\u201d in Proc. Int. Conf. Nat. Lang. Process., 2024, pp. 336\u2013343."},{"key":"null","unstructured":"K. Kwon, H. Na, H. Lee, and N. S. Kim, \u201cAdaptive knowledge distillation based on entropy,\u201d in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process., 2020, pp. 7409\u20137413."},{"key":"null","unstructured":"S. Du, S. You, X. Li, J. Wu, F. Wang, C. Qian, and C. Zhang, \u201cAgree to disagree: Adaptive ensemble knowledge distillation in gradient space,\u201d in Adv. Neural Inf. Process. Syst., vol. 33, 2020, pp. 12345\u201312355."},{"key":"null","unstructured":"R. Hadsell, S. Chopra, and Y. LeCun, \u201cDimensionality reduction by learning an invariant mapping,\u201d in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., vol. 2, 2006, pp. 1735\u20131742."},{"key":"null","unstructured":"T. Gao, X. Yao, and D. Chen, \u201cSimCSE: Simple contrastive learning of sentence embeddings,\u201d in Proc. EMNLP, 2021, pp. 6894\u20136910."},{"key":"null","unstructured":"Y. Liang, F. Meng, J. Wang, J. Xu, Y. Chen, and J. Zhou, \u201cD2tv: Dual knowledge distillation and target-oriented vision modeling for many-to-many multimodal summarization,\u201d in Proc. EMNLP Findings, 2023, pp. 14910\u201314922."},{"key":"null","unstructured":"Y. Liu and P. Liu, \u201cSimCLS: A simple framework for contrastive learning of abstractive summarization,\u201d in Proc. ACL-IJCNLP, 2021, pp. 1065\u20131072."},{"key":"null","unstructured":"Y. Liang, F. Meng, J. Xu, Y. Chen, and J. Zhou, \u201cScheduled multi-task learning for neural chat translation,\u201d in Proc. ACL, 2022, pp. 4375\u20134388."},{"key":"null","unstructured":"Y. Liang, C. Zhou, F. Meng, J. Xu, Y. Chen, J. Su, and J. Zhou, \u201cTowards making the most of dialogue characteristics for neural chat translation,\u201d in Proc. EMNLP, 2021, pp. 67\u201379."},{"key":"null","unstructured":"X. Cheng, S. Gao, L. Liu, D. Zhao, and R. Yan, \u201cNeural machine translation with contrastive translation memories,\u201d in Proc. EMNLP, 2022, pp. 3591\u20133601."},{"key":"null","unstructured":"X. Pan, M. Wang, L. Wu, and L. Li, \u201cContrastive learning for many-to-many multilingual neural machine translation,\u201d in Proc. ACL-IJCNLP, 2021, pp. 244\u2013258."},{"key":"null","unstructured":"A. F. et al., \u201cBeyond english-centric multilingual machine translation,\u201d J. Mach. Learn. Res., vol. 22, no. 107, pp. 1\u201348, 2021."},{"key":"null","unstructured":"M. R. C.-J. et al., \u201cNo language left behind: Scaling human-centered machine translation,\u201d arXiv preprint arXiv:2207.04672, 2022."},{"key":"null","unstructured":"M.-T. Luong and C. D. Manning, \u201cStanford neural machine translation systems for spoken language domains,\u201d in Proc. 12 th  Int. Workshop Spoken Lang. Transl., 2015, pp. 76\u201379."},{"key":"null","unstructured":"J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, and A. G.-B. et al., \u201cOvercoming catastrophic forgetting in neural networks,\u201d Proc. Nat. Acad. Sci. USA, vol. 114, no. 13, pp. 3521\u20133526, 2017."},{"key":"null","unstructured":"S. Gu, Y. Feng, and W. Xie, \u201cPruning-then-expanding model for domain adaptation of neural machine translation,\u201d in Proc. NAACL-HLT, 2021, pp. 3942\u20133952."},{"key":"null","unstructured":"M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, \u201cFairseq: A fast, extensible toolkit for sequence modeling,\u201d in Proc. NAACL-HLT (Demos), 2019, pp. 48\u201353."},{"key":"null","unstructured":"A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, \u0141. Kaiser, and I. Polosukhin, \u201cAttention is all you need,\u201d in Adv. Neural Inf. Process. Syst., vol. 30, 2017."},{"key":"null","unstructured":"D. Kinga and J. Ba, \u201cA method for stochastic optimization,\u201d in Int. Conf. Learn. Represent., 2015."},{"key":"null","unstructured":"N. A. et al., \u201cMassively multilingual neural machine translation in the wild: Findings and challenges,\u201d arXiv preprint arXiv:1907.05019, 2019."},{"key":"null","unstructured":"O. B. et al., \u201cProceedings of the third conference on machine translation,\u201d in Proc. Conf. Mach. Transl., 2018."},{"key":"null","unstructured":"A. G. et al., \u201cThe LLaMA 3\u00a0herd of models,\u201d arXiv preprint arXiv:2407.21783, 2024."}],"container-title":["Data Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.sciengine.com\/sci-open\/api\/v1\/open\/file\/pdf\/FB567064D25C40129F817FFBEBC9C383","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.sciengine.com\/doi\/10.3724\/2096-7004.di.2025.0109","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.sciengine.com\/sci-open\/api\/v1\/open\/file\/pdf\/FB567064D25C40129F817FFBEBC9C383","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T01:26:02Z","timestamp":1781313962000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.sciengine.com\/doi\/10.3724\/2096-7004.di.2025.0109"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,22]]},"references-count":45,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,9,22]]},"published-print":{"date-parts":[[2026,6,1]]}},"URL":"https:\/\/doi.org\/10.3724\/2096-7004.di.2025.0109","relation":{},"ISSN":["2096-7004"],"issn-type":[{"value":"2096-7004","type":"print"}],"subject":[],"published":{"date-parts":[[2025,9,22]]}}}