{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:06:13Z","timestamp":1750309573387,"version":"3.41.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"6","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62166042"],"award-info":[{"award-number":["62166042"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Natural Science Foundation of Xinjiang, China","award":["2021D01C076"],"award-info":[{"award-number":["2021D01C076"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>Transfer learning plays a crucial role in low-resource machine translation by addressing the challenge of poor model performance due to limited data in low-resource languages, thereby improving translation accuracy. Current research methods not only utilize pre-trained parent models for parameter initialization and fine-tuning but also use the soft labels output by these parent models to enhance the consistency between parent and child models. However, even if the parent model performs well, there are still instances where certain token predictions are unstable. During training, if the child model incorporates these unstable token predictions, it can hinder its learning effectiveness; the child model might not fully comprehend the parent model\u2019s prediction strategy, potentially affecting overall translation performance. To address this, we propose a training strategy called Token-Level Filter Training, designed to effectively filter out unstable token predictions from the parent model, thereby transferring the parent model\u2019s positive knowledge to the child model. Additionally, we introduce a hierarchical ranking loss method to help the child model better learn the parent model\u2019s prediction strategies and sequence order, thus enhancing translation accuracy and fluency. Experimental results show that our method outperforms baseline methods on the public datasets Global Voices (Id, Ca, Hu, Pl) and WMT17 (Turkish\u2013English), with BLEU score improvements of 1.47, 0.91, 0.50, 0.54, and 0.55, respectively. These results demonstrate the effectiveness and superiority of the proposed method.<\/jats:p>","DOI":"10.1145\/3732779","type":"journal-article","created":{"date-parts":[[2025,4,28]],"date-time":"2025-04-28T10:54:26Z","timestamp":1745837666000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["TFT-TL: Token-Level Filter Training Transfer Learning for Low-Resource Neural Machine Translation"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-8144-0519","authenticated-orcid":false,"given":"Wei","family":"Dai","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Xinjiang University","place":["Urumqi, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3026-5010","authenticated-orcid":false,"given":"Dongfang","family":"Han","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xinjiang University","place":["Urumqi, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1639-8899","authenticated-orcid":false,"given":"Turdi","family":"Tohti","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xinjiang University","place":["Urumqi, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6481-0692","authenticated-orcid":false,"given":"Yi","family":"Liang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xinjiang University","place":["Urumqi, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8055-0437","authenticated-orcid":false,"given":"Zicheng","family":"Zuo","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xinjiang University","place":["Urumqi, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-9901-2613","authenticated-orcid":false,"given":"Yuanyuan","family":"Liao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xinjiang University","place":["Urumqi, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-9380-1066","authenticated-orcid":false,"given":"Qingwen","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xinjiang University","place":["Urumqi, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,18]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_3_1_2_2","DOI":"10.18653\/v1\/2020.acl-main.688"},{"unstructured":"Mikel Artetxe Gorka Labaka Eneko Agirre and Kyunghyun Cho. 2017. Unsupervised neural machine translation. arXiv preprint arXiv:1710.11041 (2017). Retrieved from https:\/\/arxiv.org\/abs\/1710.11041","key":"e_1_3_1_3_2"},{"unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 (2014). Retrieved from https:\/\/arxiv.org\/abs\/1409.0473","key":"e_1_3_1_4_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_5_2","DOI":"10.18653\/v1\/2020.emnlp-main.615"},{"key":"e_1_3_1_6_2","volume-title":"Proceedings of the Conference on Natural Language Processing","author":"Benaicha Moncef","year":"2023","unstructured":"Moncef Benaicha, David Thulke, and Metin Turan. 2023. Leveraging cross-lingual transfer learning in spoken named entity recognition systems. In Proceedings of the Conference on Natural Language Processing. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:259342064"},{"key":"e_1_3_1_7_2","volume-title":"Advances in Neural Information Processing Systems","author":"Berthelot David","year":"2019","unstructured":"David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A. Raffel. 2019. MixMatch: A holistic approach to semi-supervised learning. In Advances in Neural Information Processing Systems. H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32, Curran Associates, Inc. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2019\/file\/1cd138d0499a68f4bb72bee04bbec2d7-Paper.pdf"},{"unstructured":"Saughmon Boujkian. 2024. Cross-linguistic examination of machine translation transfer learning. arXiv:2501.00045. Retrieved from https:\/\/arxiv.org\/abs\/2501.00045","key":"e_1_3_1_8_2"},{"key":"e_1_3_1_9_2","first-page":"109980","volume-title":"Advances in Neural Information Processing Systems","volume":"37","author":"Chen Junru","year":"2024","unstructured":"Junru Chen, Tianyu Cao, Jing Xu, Jiahe Li, Zhilong Chen, Tao Xiao, and YANG YANG. 2024. Con4m: Context-aware consistency learning framework for segmented time series classification. In Advances in Neural Information Processing Systems. A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37, Curran Associates, Inc., 109980\u2013110009. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2024\/file\/c6b96bf6380bd81bbfbce5db0e054a41-Paper-Conference.pdf"},{"doi-asserted-by":"publisher","key":"e_1_3_1_10_2","DOI":"10.18653\/v1\/P17-1176"},{"key":"e_1_3_1_11_2","first-page":"1180","volume-title":"Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 (ICML\u201915)","author":"Ganin Yaroslav","year":"2015","unstructured":"Yaroslav Ganin and Victor Lempitsky. 2015. Unsupervised domain adaptation by backpropagation. In Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 (ICML\u201915). JMLR.org, 1180\u20131189."},{"doi-asserted-by":"publisher","key":"e_1_3_1_12_2","DOI":"10.18653\/v1\/2024.naacl-long.14"},{"doi-asserted-by":"publisher","key":"e_1_3_1_13_2","DOI":"10.18653\/v1\/2024.findings-naacl.203"},{"unstructured":"Mozhdeh Gheini and Jonathan May. 2019. A universal parent model for low-resource neural machine translation transfer. arXiv:1909.06516. Retrieved from https:\/\/arxiv.org\/abs\/1909.06516","key":"e_1_3_1_14_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_15_2","DOI":"10.18653\/v1\/N18-1032"},{"key":"e_1_3_1_16_2","volume-title":"Advances in Neural Information Processing Systems","author":"Huang Jiayuan","year":"2006","unstructured":"Jiayuan Huang, Arthur Gretton, Karsten Borgwardt, Bernhard Sch\u00f6lkopf, and Alex Smola. 2006. Correcting sample selection bias by unlabeled data. In Advances in Neural Information Processing Systems. B. Sch\u00f6lkopf, J. Platt, and T. Hoffman (Eds.), Vol. 19, MIT Press. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2006\/file\/a2186aa7c086b46ad4e8bf81e2a3a19b-Paper.pdf"},{"doi-asserted-by":"publisher","key":"e_1_3_1_17_2","DOI":"10.18653\/v1\/N19-1383"},{"doi-asserted-by":"publisher","key":"e_1_3_1_18_2","DOI":"10.1109\/ICCV.2017.167"},{"doi-asserted-by":"publisher","key":"e_1_3_1_19_2","DOI":"10.18653\/v1\/2023.acl-long.211"},{"key":"e_1_3_1_20_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Inan Hakan","year":"2017","unstructured":"Hakan Inan, Khashayar Khosravi, and Richard Socher. 2017. Tying word vectors and word classifiers: A loss framework for language modeling. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=r1aPbsFle"},{"doi-asserted-by":"publisher","key":"e_1_3_1_21_2","DOI":"10.18653\/v1\/2022.acl-long.17"},{"doi-asserted-by":"publisher","key":"e_1_3_1_22_2","DOI":"10.18653\/v1\/2020.emnlp-main.7"},{"doi-asserted-by":"publisher","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations (ICLR). San Diego CA USA. Retrieved from 10.48550\/arxiv.1412.6980","key":"e_1_3_1_23_2","DOI":"10.48550\/arxiv.1412.6980"},{"doi-asserted-by":"publisher","key":"e_1_3_1_24_2","DOI":"10.18653\/v1\/W18-6325"},{"key":"e_1_3_1_25_2","first-page":"79","volume-title":"Proceedings of Machine Translation Summit X: Papers","author":"Koehn Philipp","year":"2005","unstructured":"Philipp Koehn. 2005. Europarl: A parallel corpus for statistical machine translation. In Proceedings of Machine Translation Summit X: Papers. 79\u201386. Retrieved from https:\/\/aclanthology.org\/2005.mtsummit-papers.11"},{"doi-asserted-by":"publisher","key":"e_1_3_1_26_2","DOI":"10.3115\/1557769.1557821"},{"doi-asserted-by":"publisher","key":"e_1_3_1_27_2","DOI":"10.1214\/aoms\/1177729694"},{"doi-asserted-by":"publisher","key":"e_1_3_1_28_2","DOI":"10.18653\/v1\/2022.emnlp-main.574"},{"key":"e_1_3_1_29_2","first-page":"10890","volume-title":"Advances in Neural Information Processing Systems","author":"liang xiaobo","year":"2021","unstructured":"xiaobo liang, Lijun Wu, Juntao Li, Yue Wang, Qi Meng, Tao Qin, Wei Chen, Min Zhang, and Tie-Yan Liu. 2021. R-Drop: Regularized dropout for neural networks. In Advances in Neural Information Processing Systems. M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34, Curran Associates, Inc., 10890\u201310905. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/5a66b9200f29ac3fa0ae244cc2a51b39-Paper.pdf"},{"doi-asserted-by":"publisher","key":"e_1_3_1_30_2","DOI":"10.1109\/18.61115"},{"doi-asserted-by":"publisher","key":"e_1_3_1_31_2","DOI":"10.18653\/v1\/2024.emnlp-main.768"},{"doi-asserted-by":"publisher","key":"e_1_3_1_32_2","DOI":"10.1109\/TASLP.2019.2941587"},{"doi-asserted-by":"publisher","key":"e_1_3_1_33_2","DOI":"10.18653\/v1\/2021.emnlp-main.408"},{"doi-asserted-by":"publisher","key":"e_1_3_1_34_2","DOI":"10.18653\/v1\/D15-1166"},{"doi-asserted-by":"publisher","key":"e_1_3_1_35_2","DOI":"10.1109\/TPAMI.2018.2858821"},{"key":"e_1_3_1_36_2","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation","author":"Negi Gaurav","year":"2024","unstructured":"Gaurav Negi, Rajdeep Sarkar, Omnia Zayed, and Paul Buitelaar. 2024. A hybrid approach to aspect based sentiment analysis using transfer learning. In Proceedings of the International Conference on Language Resources and Evaluation. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:268692009"},{"doi-asserted-by":"publisher","key":"e_1_3_1_37_2","DOI":"10.18653\/v1\/D18-1103"},{"key":"e_1_3_1_38_2","first-page":"296","volume-title":"Proceedings of the 8th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Nguyen Toan Q.","year":"2017","unstructured":"Toan Q. Nguyen and David Chiang. 2017. Transfer learning across low-resource, related languages for neural machine translation. In Proceedings of the 8th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). Greg Kondrak and Taro Watanabe (Eds.), Asian Federation of Natural Language Processing, Taipei, Taiwan, 296\u2013301. Retrieved from https:\/\/aclanthology.org\/I17-2050"},{"doi-asserted-by":"publisher","key":"e_1_3_1_39_2","DOI":"10.18653\/v1\/N19-4009"},{"doi-asserted-by":"publisher","key":"e_1_3_1_40_2","DOI":"10.1109\/TKDE.2009.191"},{"doi-asserted-by":"publisher","key":"e_1_3_1_41_2","DOI":"10.18653\/v1\/W18-6319"},{"doi-asserted-by":"publisher","key":"e_1_3_1_42_2","DOI":"10.18653\/v1\/E17-2025"},{"doi-asserted-by":"publisher","key":"e_1_3_1_43_2","DOI":"10.1145\/3567592"},{"doi-asserted-by":"publisher","key":"e_1_3_1_44_2","DOI":"10.3758\/s13423-015-0889-1"},{"doi-asserted-by":"publisher","key":"e_1_3_1_45_2","DOI":"10.18653\/v1\/P16-1162"},{"key":"e_1_3_1_46_2","first-page":"223","volume-title":"Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers","author":"Snover Matthew","year":"2006","unstructured":"Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006. A study of translation edit rate with targeted human annotation. In Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers. Association for Machine Translation in the Americas, Cambridge, Massachusetts, USA, 223\u2013231. Retrieved from https:\/\/aclanthology.org\/2006.amta-papers.25"},{"unstructured":"Xu Tan Yichong Leng Jiale Chen Yi Ren Tao Qin and Tie-Yan Liu. 2019. A study of multilingual neural machine translation. arXiv:1912.11625. Retrieved from https:\/\/arxiv.org\/abs\/1912.11625","key":"e_1_3_1_47_2"},{"key":"e_1_3_1_48_2","first-page":"2214","volume-title":"Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC\u201912)","author":"Tiedemann J\u00f6rg","year":"2012","unstructured":"J\u00f6rg Tiedemann. 2012. Parallel data, tools and interfaces in OPUS. In Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC\u201912). Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet U\u011fur Do\u011fan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (Eds.), European Language Resources Association (ELRA), Istanbul, Turkey, 2214\u20132218. Retrieved from http:\/\/www.lrec-conf.org\/proceedings\/lrec2012\/pdf\/463_Paper.pdf"},{"key":"e_1_3_1_49_2","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30, Curran Associates, Inc. Retrieved fromhttps:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf"},{"doi-asserted-by":"publisher","key":"e_1_3_1_50_2","DOI":"10.18653\/v1\/2020.acl-main.278"},{"doi-asserted-by":"publisher","key":"e_1_3_1_51_2","DOI":"10.18653\/v1\/2024.wmt-1.69"},{"doi-asserted-by":"publisher","unstructured":"Xiangyu Wu Qing-Yuan Jiang Yang Yang Yi-Feng Wu Qing-Guo Chen and Jianfeng Lu. 2024. TAI++: Text as image for multi-label image classification by co-learning transferable prompt. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI\u201924). Jeju Korea. DOI:10.24963\/ijcai.2024\/578","key":"e_1_3_1_52_2","DOI":"10.24963\/ijcai.2024\/578"},{"key":"e_1_3_1_53_2","first-page":"6256","volume-title":"Advances in Neural Information Processing Systems","author":"Xie Qizhe","year":"2020","unstructured":"Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020. Unsupervised data augmentation for consistency training. In Advances in Neural Information Processing Systems. H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33, Curran Associates, Inc., 6256\u20136268. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/44feb0096faa8326192570788b38c1d1-Paper.pdf"},{"key":"e_1_3_1_54_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Xie Shufang","year":"2022","unstructured":"Shufang Xie, Ang Lv, Yingce Xia, Lijun Wu, Tao Qin, Tie-Yan Liu, and Rui Yan. 2022. Target-side input augmentation for sequence to sequence generation. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=pz1euXohm4H"},{"doi-asserted-by":"publisher","key":"e_1_3_1_55_2","DOI":"10.18653\/v1\/2022.acl-short.68"},{"key":"e_1_3_1_56_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhang* Tianyi","year":"2020","unstructured":"Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=SkeHuCVFDr"},{"doi-asserted-by":"publisher","key":"e_1_3_1_57_2","DOI":"10.18653\/v1\/2024.emnlp-main.111"},{"doi-asserted-by":"publisher","key":"e_1_3_1_58_2","DOI":"10.1007\/s11263-023-01911-w"},{"doi-asserted-by":"publisher","key":"e_1_3_1_59_2","DOI":"10.18653\/v1\/N16-1004"},{"doi-asserted-by":"publisher","key":"e_1_3_1_60_2","DOI":"10.18653\/v1\/D16-1163"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3732779","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:19:24Z","timestamp":1750295964000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3732779"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,18]]},"references-count":59,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3732779"],"URL":"https:\/\/doi.org\/10.1145\/3732779","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2025,6,18]]},"assertion":[{"value":"2024-12-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}